moonsize

    A Wasm size profiler for MoonBit — section breakdown, function ranking, call graph and retained-size analysis

    wasm
    webassembly
    size
    profiling
    cli
    Download zip
    Version
    0.3.0
    License
    MIT
    Last updated
    6 hours ago
    Downloads
    10

    Dependencies

    #moonsize

    moon add BigSaltyMan/moonsize

    A size analyzer for WebAssembly, built for MoonBit's output. It reads a .wasm file, measures every section, attributes the code section to named functions, follows the call graph, and reports what could be deleted.

    The analysis model follows Twiggy: sizes are attributed to individual functions, functions are grouped by package, and reachability from the module's roots decides what is really used.

    Chinese version: README.zh.md.

    #Building

    moon build --release ./_build/native/release/build/cmd/main/main.exe app.wasm

    The module declares supported_targets = "+native", so a project that depends on it builds with --target native as well.

    #Usage

    usage: moonsize <file.wasm> [--top <n>] [--retained] [--dead-code] [--compress] [--call-graph <path>] [--html <path>] [--max-size <size>] [--baseline <path.wasm>] Report the section table of a WebAssembly binary, rank its heaviest functions, roll the bytes up by module, and follow the call graph to find what can be deleted. --top <n> how many functions to rank (default 10) --retained what deleting each function would free --dead-code functions no root can reach --compress what the file and each section cost gzip-compressed --call-graph <path> write the call graph as .dot or .json --html <path> write the charts as a self-contained HTML report --max-size <size> fail with exit code 3 above this size, where size is bytes or a KB/MB/GB count such as 10KB or 1.5MB --baseline <path> compare against another module and print the change; --max-size then bounds the growth, not the file Exit codes: 0 ok, 1 unreadable input, 2 bad command line, 3 over budget.

    Without flags it prints the section table, the heaviest functions and the per-package roll-up. --retained and --dead-code add their sections and can be combined. --call-graph and --html each write a file instead of printing a report, so they are used on their own:

    moonsize app.wasm --retained --dead-code # the full text report moonsize app.wasm --html report.html # charts, opens in a browser moonsize app.wasm --call-graph graph.dot # for Graphviz

    #Using the library

    The parsing, the analysis and every report are ordinary packages in this module, so a program can take the numbers instead of the text. moon addBigSaltyMan/moonsize and import it; Options::default is the way in, because Options is read-only and no other package may write one.

    fn main {
    let bytes : Bytes = read("app.wasm")
    match @moonsize.parse_module(bytes) {
    Ok(parsed) =>
    match @moonsize.analyze(bytes, parsed) {
    Ok(analysis) => {
    let options = @moonsize.Options::default("app.wasm")
    // What the five heaviest functions cost, as text or as numbers:
    // `analysis.stats` and `analysis.retained` are there to be read.
    println(@moonsize.render_top_functions(analysis.stats, analysis.data.length(), 5))
    println(@moonsize.render_html(analysis, options.path))
    }
    Err(error) => println(@moonsize.describe_error(error))
    }
    Err(error) => println(@moonsize.describe_error(error))
    }
    }

    The library does no I/O, so reading the file and choosing an exit code stay with the caller. supported_targets = "+native" applies to a program that depends on this module too: build it with --target native.

    #Analysis

    Everything below is built on one idea: a function's bytes only matter if something can call it.

    #Symbol names

    Attribution starts with the names in the binary's name section, and MoonBit does not publish how it mangles them, so the rules here were read back from the name section of real builds (moonc 0.1.20260920, the wasm backend):

    • a symbol is _M0<kind> and then a path: F a function, M a method, I a trait implementation;
    • the package is B for the builtin package, C for moonbitlang/core, or P and an index — which runs into the first component's own length, so P55probe is index 5 followed by 5probe;
    • a component is its length and then its name, and the length counts the mangled name;
    • anything that cannot appear in an identifier is escaped as _ and the byte in hex, and a literal underscore is doubled: to__string_2einner is to_string.inner, _24default__impl is $default_impl, and a non-ASCII character is escaped one byte at a time, so 中 is _e4_b8_ad;
    • a generic instantiation carries G...E with one code per argument, where a builtin is a single letter, a tuple is U...E, and a named type is R and a path;
    • a closure is either an environment named __moonbit_<fn> or a body carrying C<id>l<line>, the id of the closure and the line it was written on;
    • a trait implementation names two packages, its type's and its trait's.

    The reports print the path through display_name. demangle_full is the fuller decoder, for callers that want escapes resolved, arguments spelled out and closures marked. Both are display aids rather than a codec: when one cannot decode a symbol exactly it returns it untouched, because a report that invents a name is worse than one that shows a raw symbol. The known limits are where that happens — a generic argument whose type code is not one of the letters a build produced (Int, Double, String, Bool, Char, Byte, Float, Int64, UInt64, Unit) is left undecoded rather than guessed, a named argument is only decoded as the last one in its list because its path has no terminator to separate it from what follows, and a symbol over 4096 bytes is refused outright so a forged name cannot drive the walk into a deep recursion.

    #The call graph

    The decoder walks every function body instruction by instruction — the full MVP instruction set, the 0xFC saturating and bulk-memory opcodes, the exception handling and tail-call opcodes MoonBit emits, and the 0xFB garbage-collection opcodes its wasm-gc backend uses. Each body contributes:

    • call x — a precise edge to x;
    • ref.func x — a reference, which counts as reachable because the function may run later as a closure;
    • call_indirect — a conservative edge to every function any element segment places in that table. The table's contents are not known statically, so the candidates are all of them; a missing edge would mean a missing function and a function wrongly reported dead, which is the one mistake worth avoiding here.

    A function is a root if the outside world can reach it: it is exported, it is the start function, or it sits in a table. Everything reachable from those roots is live; everything else is dead.

    #Retained size

    SIZE is a function on its own. RETAINED is what deleting it would actually free: its own bytes plus every function that becomes unreachable with it, with DIES counting how many that is.

    That set is exactly the functions the deleted one dominates — those every path from a root has to pass through. Taking the module's graph, adding a virtual root above the real ones, and computing the dominator tree gives every function's retained size in one pass. Cycles fall out of the fixpoint rather than needing a special case, and the answer is exact rather than an approximation from local predecessor counts: if two branches both call a function, deleting one branch does not free it, but deleting the point where both branches meet does.

    Reading the table:

    • RETAINED == SIZE — the function is load-bearing only for itself. Deleting it saves its own bytes and nothing more.
    • RETAINED ≫ SIZE with a high DIES — a good candidate. Removing it takes a whole subtree with it.
    • IND — the function can be reached through a call_indirect, so removing it is only safe if no table entry and no indirect call expects it.

    Both numbers are printed on purpose. SIZE says how much the compiler could win by making the function smaller; RETAINED says how much it wins by removing it, which is usually the cheaper change.

    #Dead code

    Anything unreachable from the roots cannot run: no export reaches it, no start function, no table entry, and no chain of calls from any of those. Its bytes are already paid for and never used, so the list is a deletion list.

    MoonBit's compiler does dead-code elimination well, so a small program usually reports none — that is a finding too.

    #Compressed size

    A WebAssembly module is not served raw. It is served gzipped, and how much that saves depends on what the bytes are: machine code compresses a little, symbol names and strings compress a lot, and already-packed data barely moves at all. A section table read raw therefore misleads in both directions, and --compress answers the question a network asks instead:

    COMPRESSED SIZE raw 10675 B gzip 5207 B (48.7% of raw, ratio 2.05x) BY SECTION SECTION RAW GZIP RATIO SHARE(GZIP) code 5166 2712 1.90x 52.0% custom (name) 4975 1983 2.50x 38.0% data 252 174 1.44x 3.3% custom (producers) 71 95 0.74x 1.8% type 59 66 0.89x 1.2% function 50 63 0.79x 1.2% import 37 61 0.60x 1.1% export 21 45 0.46x 0.8% global 13 34 0.38x 0.6% element 8 32 0.25x 0.6% table 7 31 0.22x 0.5% memory 5 29 0.17x 0.5% datacount 3 27 0.11x 0.5%

    SHARE(GZIP) is what the section costs as a share of the compressed file rather than of the raw one, and it is the column to read. Here the code section is 48.3% of the file raw but 52.0% of it compressed, while the name section goes the other way — 46.6% raw, 38.0% compressed. The bytes the compiler emitted are the ones worth attacking; the symbol names are close to free on the wire, which is worth knowing before spending an afternoon on --strip.

    Two things to keep in mind reading it. It is an estimate: the compressor runs at its default level, which is the level servers use, but a server may be configured differently, and it compresses the whole response rather than each section on its own. And every section is compressed as its own stream, so each pays its own gzip container of about twenty bytes — which is invisible for a section of kilobytes and dominant for one of a handful of bytes, and is why the smallest sections above read as growing when compressed. The shares barely notice, because twenty bytes is nothing against a table of kilobytes.

    #Judging what can be deleted

    1. Start with --dead-code. Every row there is free to remove; check that nothing outside the module calls it by name (an exported name would have made it a root, so this is already accounted for).
    2. Then look at --retained for large RETAINED values with DIES above zero. Those are the functions whose removal cascades.
    3. Treat IND rows with care: an indirect candidate can only be removed when the table entry and the call sites go with it.

    #Example output

    A 10,675-byte MoonBit program (--target wasm, debug), reduced to the analysis sections:

    Retained size # INDEX SIZE RETAINED DIES SHARE IND FUNCTION 1 47 158 5162 46 48.3% ____moonbit__main 2 39 301 1773 10 16.6% int::Int::to__string_2einner 3 37 24 887 8 8.3% println 4 34 9 819 5 7.6% moonbit.println 5 33 206 810 4 7.5% moonbit.fprintln 6 28 49 717 7 6.7% moonbit.decref 7 29 409 668 6 6.2% moonbit.gc.free 8 43 633 633 0 5.9% int__to__string__dec DIES counts the other functions that become unreachable with this one. IND marks a function a call_indirect could reach. Dead code 0 of 47 functions are unreachable from the roots 0 bytes (0.0% of file) every function is reachable

    ____moonbit__main retains 5,162 bytes — 48.3% of the file, 46 functions — which is what a program with a single entry point looks like: everything hangs off it. int__to__string__dec is the largest single function at 633 bytes but retains only itself, so shrinking it is a compiler problem rather than a deletion.

    --call-graph graph.dot writes the same graph for Graphviz, with dead functions dashed and indirect edges dotted; --call-graph graph.json writes it with the roots, every edge and every indirect site's candidate set.

    #HTML report

    --html <path> writes the same analysis as one self-contained page: no server, no build step, no network. Open the file and the four charts are there, with the compressed sizes the text report only shows under --compress.

    moonsize HTML report

    A worked example is checked in at examples/report.html, generated from examples/fib.wasm — a 10,675-byte MoonBit program. It is one file: the chart library is inlined, so it can be moved, emailed or opened from anywhere. The source it was built from is examples/fib.mbt, which is also a package here: building --target wasm produces the example alongside the tool.

    Four headline numbers. File size, the code section, what the file costs gzipped, and how much of it is unreachable. The gzip card carries the share of raw and the ratio underneath, and the dead-code card turns red when there is anything to delete, which is the one number most people open the report for.

    Sections — a horizontal bar per section, widest first, with the share of the file in the tooltip. This is the stage-one section table, drawn. Both bar charts are sized to their row count, so every row keeps its label; a fixed-height box makes ECharts drop most of them, and the rows that survive look like headings for the unlabelled ones below.

    The sections chart has two views, switched by the raw / gzip legend in its corner: the same bars measured before and after compression. They share one axis and overlap exactly, so switching shows how much shorter the compressed bars are instead of rescaling the axis to hide it — the code section barely moves, the name section loses half its length. The tooltip gives both numbers and the ratio in either view.

    Top 20 functions by body size — what the compiler could shrink. Body size is used rather than the encoded size so that the ranking answers "where is the code", not "where are the size prefixes".

    Modules — a donut of how the code section divides between packages, as a share of the code section rather than of the file, so the slices mean "which package is responsible for the code". Slices below 0.5% are rolled into a single other slice, named on hover: a slice that thin cannot carry a label, and leaving it in draws a sliver nobody can read or aim at. Labels sit inside the ring, so they never collide with the legend.

    Treemap — module → function, where area is bytes. This is the one chart that shows the whole binary at once: the big cells are the functions worth looking at, and cells are drawn red when the function is unreachable. The area stays the raw size — a treemap drawn by compressed bytes would be a different chart, and a less useful one — so the compressed size and ratio are reported in the tooltip instead of drawn.

    Hovering any bar, slice or cell shows the exact bytes and percentage. Every number in the page comes from the same Analysis the text report uses, so the two cannot disagree — including the compressed ones, which the page computes with the same compressor and the same level as --compress.

    The library is vendored in assets/echarts.min.js and inlined into the page. When that file is not next to the working directory — running a installed binary from elsewhere, say — the page falls back to a CDN <scriptsrc> tag instead and the command says so; that report needs a network connection to draw.

    #CI integration

    --max-size turns the report into a gate, and --baseline turns it into a comparison, so a build can fail on growth instead of on a number somebody has to keep updating by hand.

    moonsize app.wasm --max-size 100KB # a ceiling on the file moonsize app.wasm --baseline main.wasm --max-size 5KB # at most 5 KB of growth

    With a baseline, --max-size bounds the change, not the file: the question a pull request asks is what it costs, not how big the program already was. A run with a baseline prints the comparison after the report, heaviest section first, with the change as bytes and as a share of what it was:

    SIZE COMPARISON baseline current delta total 10675 11039 +364 (+3.4%) code 5166 5299 +133 (+2.5%) custom (name) 4975 5205 +230 (+4.6%) data 252 252 0 (0.0%) custom (producers) 71 71 0 (0.0%) ...

    The exit codes are 0 ok, 1 unreadable input, 2 bad command line and 3 over budget, so a job can fail on the budget alone or tell a broken input apart from a broken build.

    This repository gates itself in .github/workflows/ci.yml: formatting, moon check --target native --deny-warn, moon test, and the checked-in example held to a budget.

    - uses: moonbit-community/setup-moonbit@v1 - name: The example stays inside its budget run: moon run cmd/main -- --max-size 100KB examples/fib.wasm

    The same three lines gate any project that produces a .wasm:

    - uses: moonbit-community/setup-moonbit@v1 - run: moon build --target wasm --release - name: Size budget run: moonsize _build/wasm/release/build/app/app.wasm --max-size 250KB

    To catch growth rather than an absolute size, build the base branch too and pass it as the baseline; --max-size then reads as the allowance:

    - name: Build the base branch run: | git worktree add "$RUNNER_TEMP/base" "origin/${{ github.base_ref }}" (cd "$RUNNER_TEMP/base" && moon build --target wasm --release) - name: Size budget run: | moonsize _build/wasm/release/build/app/app.wasm \ --baseline "$RUNNER_TEMP/base/_build/wasm/release/build/app/app.wasm" \ --max-size 5KB

    .github/workflows/size-check.yml does that on every pull request. examples/ is a package of this module, so both sides build the example the same way and the comparison is between two builds of the same source:

    moon build --target wasm # -> examples wasm moon run cmd/main -- that.wasm --baseline base.wasm --max-size 5KB

    #Development

    moon check --target native --deny-warn moon test

    The test suite covers the instruction table one opcode at a time — every entry carries a canonical encoding and the length the specification gives it — and pins three real MoonBit function bodies, including one whose locals use two-byte garbage-collection reference types, so a wrong immediate width fails a test instead of quietly producing a wrong call graph.

    #Regenerating the example

    The numbers quoted in examples/report.html and in this README — the file size, the section shares, the retained sizes — are read out of a real build rather than written by hand. Changing examples/fib.mbt or moving to a new toolchain makes them stale, and all three steps have to be run again:

    moon build --target wasm moon run cmd/main -- --html examples/report.html examples/fib.wasm # then re-capture examples/report.png from that page

    The first step builds examples/ and leaves the artifact at _build/wasm/debug/build/examples/examples.wasm; examples/fib.wasm is a copy of it, the report is generated from that, and the screenshot is a capture of the report at 1400px wide and full height. Nothing under examples/ is edited by hand, so a stale number there means a stale build rather than a typo.

    .moonignore keeps the rendered page out of the published package — it is 1.2 MB of what would otherwise be a 2.4 MB module, and nobody moon adds this for a screenshot — while the source, the example binary and the chart library all ship.

    Analysis

    pub struct Analysis {
    data : Bytes
    parsed : WasmModule
    stats : Array[FunctionStat]
    graph : CallGraph
    roots :
    Set
    [Int]
    reachable :
    Set
    [Int]
    retained : Array[RetainedSize]
    indirect :
    Set
    [Int]
    }

    A parsed module with its sizes, call graph and reachability resolved.

    CallEdge

    pub struct CallEdge {
    caller : Int
    callee : Int
    kind : EdgeKind
    offset : Int
    }

    One edge of the call graph.

    CallGraph

    pub struct CallGraph {
    calls : Map[Int,
    Set
    [Int]]
    called_by : Map[Int,
    Set
    [Int]]
    indirect_sites : Array[IndirectSite]
    edges : Array[CallEdge]
    }

    The reference graph between functions.

    CallKind

    pub enum CallKind {
    Direct(Int)
    Indirect(Int)
    }

    How a call site reaches its target.

    CallSite

    pub struct CallSite {
    kind : CallKind
    type_index : Int?
    offset : Int
    }

    One call found in a function body.

    EdgeKind

    pub enum EdgeKind {
    Direct
    Indirect
    Reference
    }

    How one function refers to another.

    ElementSegment

    pub struct ElementSegment {
    table_index : Int
    passive : Bool
    functions : Array[Int]
    }

    One entry of the element section.

    Export

    pub struct Export {
    name : String
    kind : Int
    index : Int
    }

    One entry of the export section.

    FunctionBody

    pub struct FunctionBody {
    index : Int
    body_size : Int
    total_size : Int
    offset : Int
    }

    One function body decoded out of the code section.

    The spec stores a body as a size-prefixed blob, so an entry costs its own length prefix on top of the body itself. Both numbers are kept: the body is what a compiler can shrink, the total is what the file actually spends.

    FunctionStat

    pub struct FunctionStat {
    index : Int
    name : String
    module_name : String
    body_size : Int
    total_size : Int
    offset : Int
    }

    One function's byte budget, with the name and package the binary gives it.

    IndirectSite

    pub struct IndirectSite {
    caller : Int
    table_index : Int
    type_index : Int
    offset : Int
    candidates : Array[Int]
    }

    A call_indirect, with the targets its table could supply.

    Instruction

    pub struct Instruction {
    opcode : Opcode
    offset : Int
    size : Int
    operands : Array[Int]
    }

    A decoded instruction: what it is, where it starts and what it carries.

    Opcode

    pub enum Opcode {
    Unreachable
    Nop
    Block
    Loop
    If
    Else
    Try
    Catch
    Throw
    Rethrow
    ThrowRef
    End
    Br
    BrIf
    BrTable
    Return
    Call
    CallIndirect
    ReturnCall
    ReturnCallIndirect
    CallRef
    ReturnCallRef
    Drop
    Select
    Delegate
    CatchAll
    SelectT
    LocalGet
    LocalSet
    LocalTee
    GlobalGet
    GlobalSet
    TableGet
    TableSet
    I32Load
    I64Load
    F32Load
    F64Load
    I32Load8S
    I32Load8U
    I32Load16S
    I32Load16U
    I64Load8S
    I64Load8U
    I64Load16S
    I64Load16U
    I64Load32S
    I64Load32U
    I32Store
    I64Store
    F32Store
    F64Store
    I32Store8
    I32Store16
    I64Store8
    I64Store16
    I64Store32
    MemorySize
    MemoryGrow
    I32Const
    I64Const
    F32Const
    F64Const
    I32Eqz
    I32Eq
    I32Ne
    I32LtS
    I32LtU
    I32GtS
    I32GtU
    I32LeS
    I32LeU
    I32GeS
    I32GeU
    I64Eqz
    I64Eq
    I64Ne
    I64LtS
    I64LtU
    I64GtS
    I64GtU
    I64LeS
    I64LeU
    I64GeS
    I64GeU
    F32Eq
    F32Ne
    F32Lt
    F32Gt
    F32Le
    F32Ge
    F64Eq
    F64Ne
    F64Lt
    F64Gt
    F64Le
    F64Ge
    I32Clz
    I32Ctz
    I32Popcnt
    I32Add
    I32Sub
    I32Mul
    I32DivS
    I32DivU
    I32RemS
    I32RemU
    I32And
    I32Or
    I32Xor
    I32Shl
    I32ShrS
    I32ShrU
    I32Rotl
    I32Rotr
    I64Clz
    I64Ctz
    I64Popcnt
    I64Add
    I64Sub
    I64Mul
    I64DivS
    I64DivU
    I64RemS
    I64RemU
    I64And
    I64Or
    I64Xor
    I64Shl
    I64ShrS
    I64ShrU
    I64Rotl
    I64Rotr
    F32Abs
    F32Neg
    F32Ceil
    F32Floor
    F32Trunc
    F32Nearest
    F32Sqrt
    F32Add
    F32Sub
    F32Mul
    F32Div
    F32Min
    F32Max
    F32Copysign
    F64Abs
    F64Neg
    F64Ceil
    F64Floor
    F64Trunc
    F64Nearest
    F64Sqrt
    F64Add
    F64Sub
    F64Mul
    F64Div
    F64Min
    F64Max
    F64Copysign
    I32WrapI64
    I32TruncF32S
    I32TruncF32U
    I32TruncF64S
    I32TruncF64U
    I64ExtendI32S
    I64ExtendI32U
    I64TruncF32S
    I64TruncF32U
    I64TruncF64S
    I64TruncF64U
    F32ConvertI32S
    F32ConvertI32U
    F32ConvertI64S
    F32ConvertI64U
    F32DemoteF64
    F64ConvertI32S
    F64ConvertI32U
    F64ConvertI64S
    F64ConvertI64U
    F64PromoteF32
    I32ReinterpretF32
    I64ReinterpretF64
    F32ReinterpretI32
    F64ReinterpretI64
    I32Extend8S
    I32Extend16S
    I64Extend8S
    I64Extend16S
    I64Extend32S
    RefNull
    RefIsNull
    RefFunc
    RefEq
    RefAsNonNull
    BrOnNull
    BrOnNonNull
    I32TruncSatF32S
    I32TruncSatF32U
    I32TruncSatF64S
    I32TruncSatF64U
    I64TruncSatF32S
    I64TruncSatF32U
    I64TruncSatF64S
    I64TruncSatF64U
    MemoryInit
    DataDrop
    MemoryCopy
    MemoryFill
    TableInit
    ElemDrop
    TableCopy
    TableGrow
    TableSize
    TableFill
    StructNew
    StructNewDefault
    StructGet
    StructGetS
    StructGetU
    StructSet
    ArrayNew
    ArrayNewDefault
    ArrayNewFixed
    ArrayNewData
    ArrayNewElem
    ArrayGet
    ArrayGetS
    ArrayGetU
    ArraySet
    ArrayLen
    ArrayFill
    ArrayCopy
    ArrayInitData
    ArrayInitElem
    RefTest
    RefTestNull
    RefCast
    RefCastNull
    BrOnCast
    BrOnCastFail
    AnyConvertExtern
    ExternConvertAny
    RefI31
    I31GetS
    I31GetU
    }

    Every instruction the decoder recognises, one variant per mnemonic.

    The set covers the MVP, the sign-extension, saturating-truncation, bulk-memory and reference-type extensions MoonBit's linear-memory backend emits, and the garbage-collection instructions its wasm-gc backend emits.

    Opcode::name

    fn Opcode::name(self : Opcode) -> String

    The spec mnemonic for an opcode, spelled as the text format writes it.

    Opcode::operands

    fn Opcode::operands(self : Opcode) -> Operands

    The immediate operands each opcode carries.

    Operands

    pub enum Operands {
    None_
    BlockType
    Label
    LabelTable
    FunctionIndex
    CallIndirect
    LocalIndex
    GlobalIndex
    TableIndex
    TagIndex
    MemoryIndex
    MemoryArg
    I32Const
    I64Const
    F32Const
    F64Const
    HeapType
    RefType
    TypeIndex
    TypeFieldPair
    TypePair
    TypeCountPair
    TypeDataPair
    TypeElemPair
    DataIndex
    ElemIndex
    DataMemoryPair
    MemoryPair
    ElemTablePair
    TablePair
    CastPair
    TypeVector
    }

    How the bytes after an opcode are laid out.

    The shapes named <a><b>Pair read a then b, in encoding order. Naming them here rather than inlining the reads keeps every immediate width in one table that can be reviewed — and tested — against the specification.

    Options

    pub struct Options {
    path : String
    top : Int
    retained : Bool
    dead_code : Bool
    call_graph : String?
    html : String?
    max_size : Int?
    baseline : String?
    compress : Bool
    }

    What to analyze, how much of it to show, and which reports to produce.

    Options::default

    fn Options::default(path : String) -> Options

    The defaults parse_options starts from, with one flat report for the module at path. For callers that link this module as a library instead of running the CLI: every flag the command line can set is left at its absent value.

    Reader

    pub struct Reader {
    data : Bytes
    pos : Int
    }

    A forward-only cursor over a byte buffer.

    The shape follows jtenner/starshine's src/binary decoder, where a decode step takes (bytes, offset) and reports how far it advanced, so callers can walk a stream without hidden global state.

    Reader::at_end

    fn Reader::at_end(self : Reader) -> Bool

    Reader::new

    fn Reader::new(data : Bytes) -> Reader

    Reader::position

    fn Reader::position(self : Reader) -> Int

    Reader::read_bits

    fn Reader::read_bits(self : Reader, count : Int) -> Result[Int, WasmError]

    Read a little-endian bit pattern of count bytes as a single integer.

    Floating point constants have no structural meaning for a size analyzer, so they are kept as the bits the file holds rather than converted and back.

    Reader::read_byte

    fn Reader::read_byte(self : Reader) -> Result[Int, WasmError]

    Read a single byte, or report the offset that ran out of input.

    Reader::read_bytes

    fn Reader::read_bytes(self : Reader, count : Int) -> Result[BytesView, WasmError]

    Read count raw bytes and advance past them.

    A view is enough for every caller here: names are decoded straight from it and section payloads are skipped by size rather than copied.

    Reader::read_reftype

    fn Reader::read_reftype(self : Reader) -> Result[Int, WasmError]

    Consume a reference type.

    A reference type is either a one-byte shorthand such as funcref (0x70), or one of the two constructors 0x63/0x64 followed by a heap type. The shorthand byte is returned, or 0 when the two-byte form was used.

    Reader::read_signed_leb

    fn Reader::read_signed_leb(self : Reader, bits : Int) -> Result[Int, WasmError]

    Read a signed LEB128 integer of at most bits significant bits.

    Block types, heap types and the i32.const/i64.const immediates are all signed, and the sign lives in the last group's bit 6, so leaving it out would turn a small negative constant into a large positive one. The bits bound is what makes a non-terminating encoding malformed: a group may only continue while fewer than bits bits have been consumed.

    Reader::read_u32_leb

    fn Reader::read_u32_leb(self : Reader) -> Result[Int, WasmError]

    Read an unsigned LEB128 integer of at most 32 significant bits.

    WebAssembly encodes section payload lengths, vector counts and custom section name lengths this way. Five groups of seven bits cover u32, so a sixth continuation bit means the value is malformed.

    Reader::read_valtype

    fn Reader::read_valtype(self : Reader) -> Result[Unit, WasmError]

    Consume a value type, whose width is not always one byte.

    Numeric and vector types are a single byte, but a garbage-collection reference type is the two-byte 0x63/0x64 form followed by a heap type. Reading one byte per type desynchronises the stream on the first struct-typed local, so every value type goes through here.

    Reader::remaining

    fn Reader::remaining(self : Reader) -> Int

    Reader::seek

    fn Reader::seek(self : Reader, offset : Int) -> Unit

    Move the cursor to an absolute offset.

    RetainedSize

    pub struct RetainedSize {
    index : Int
    own_size : Int
    retained : Int
    retained_functions : Int
    }

    What deleting one function would free.

    Section

    pub struct Section {
    id : Int
    kind : SectionId
    offset : Int
    payload_size : Int
    total_size : Int
    custom_name : String?
    }

    One decoded section header, with the byte budget it accounts for.

    SectionCompression

    pub struct SectionCompression {
    label : String
    raw : Int
    gzip : Int
    }

    A section measured before and after compression.

    SectionId

    pub enum SectionId {
    Custom
    Type
    Import
    Function
    Table
    Memory
    Global
    Export
    Start
    Element
    Code
    Data
    DataCount
    Tag
    }

    The section identifiers defined by the WebAssembly core specification.

    SectionId::from_id

    fn SectionId::from_id(id : Int) -> SectionId?

    Map a raw section id byte to its meaning, or None when unassigned.

    SectionId::is_custom

    fn SectionId::is_custom(self : SectionId) -> Bool

    Custom sections are the only ones that carry a name, and they are worth telling apart from the numbered sections.

    SectionId::name

    fn SectionId::name(self : SectionId) -> String

    The lower-case name used by wasm-objdump and friends.

    Sections

    pub struct Sections {
    code : Section?
    imports : Section?
    names : Section?
    exports : Section?
    elements : Section?
    tables : Section?
    start : Section?
    }

    The sections the analysis passes need, picked out of the section table once.

    A well-formed module holds at most one of each numbered section, so the last one seen wins. Several custom sections are legal, and name is the only one this tool reads.

    TableType

    pub struct TableType {
    element_type : Int
    minimum : Int
    maximum : Int?
    }

    A table declared by the table section.

    WasmError

    pub enum WasmError {
    UnexpectedEnd(Int, Int)
    BadMagic
    InvalidSectionId(Int)
    LebOverflow(Int)
    InvalidImportKind(Int)
    UnknownOpcode(Int, Int)
    UnknownPrefixedOpcode(Int, Int, Int)
    }

    Every failure mode the decoder can report, with the offset that explains it.

    WasmModule

    pub struct WasmModule {
    version : Int
    file_size : Int
    sections : Array[Section]
    }

    A decoded module preamble plus its section table.

    DEFAULT_TOP

    let DEFAULT_TOP : Int

    How many functions --top ranks when the flag is absent.

    EXIT_INPUT

    let EXIT_INPUT : Int

    The input could not be read, or is not a WebAssembly module.

    EXIT_OK

    let EXIT_OK : Int

    The analysis ran and nothing was over budget.

    EXIT_OVER_SIZE

    let EXIT_OVER_SIZE : Int

    The module is over the size budget the invocation set.

    EXIT_USAGE

    let EXIT_USAGE : Int

    The command line did not make sense.

    analyze

    fn analyze(data : Bytes, parsed : WasmModule) -> Result[Analysis, WasmError]

    Analyse a module: names and sizes, then the call graph and what it implies.

    build

    fn build(data : Bytes, bodies : Array[FunctionBody], elements : Array[ElementSegment]) -> Result[CallGraph, WasmError]

    Build the call graph of a module.

    Every body is decoded once: call gives a precise edge, ref.func gives a reference edge, and call_indirect gives one conservative edge per function its table can hold plus a record of the site itself.

    collect_function_stats

    fn collect_function_stats(parsed : WasmModule, data : Bytes) -> Result[Array[FunctionStat], WasmError]

    Read a module's function names and sizes.

    A convenience over analyze for callers that only want the size table; the two share function_stats, so they cannot disagree about attribution.

    compress_gzip

    fn compress_gzip(data : Bytes) -> Result[Bytes, String]

    Compress data as a gzip stream.

    The compressor runs at its default level — six, which is what servers use — so the result is what a transfer would pay rather than the best a compressor could manage. It stays an estimate in two ways worth knowing about: a server may be configured for a different level, and it compresses the response as a whole rather than each section on its own.

    compress_sections

    fn compress_sections(data : Bytes, parsed : WasmModule) -> Array[SectionCompression]

    Compress each section of parsed on its own, out of the file it was read from, heaviest raw section first.

    Each section becomes its own gzip stream, so each pays its own container — roughly twenty bytes of magic, header, CRC and length. That is invisible for a section of kilobytes and dominant for one of a handful of bytes, which is why a tiny section can come back larger than it went in.

    This is what render_compression is handed instead of bare Sections: a Section records where its bytes are, not what they are, and a section cannot be compressed without them.

    compressed_size

    fn compressed_size(data : Bytes) -> Int

    The compressed size of data in bytes, or its raw size when it cannot be compressed at all.

    Falling back rather than failing keeps a report renderable: one stream the compressor refuses should not take the run down, and claiming it shrank would be worse than saying it did not.

    count_imported_functions

    fn count_imported_functions(data : Bytes, section : Section) -> Result[Int, WasmError]

    Count the function imports of a module.

    Imported functions occupy the first slots of the function index space, so this count is what turns a code-section position into the index that the name section and every call instruction actually use.

    decode_instruction

    fn decode_instruction(bytes : Bytes, offset : Int) -> Result[Instruction, WasmError]

    Decode the instruction that starts at offset.

    An unassigned opcode is an error rather than a skip: silently stepping over one byte would turn a decoding bug into a wrong call graph, and every number downstream of this function is a size.

    decode_instructions

    fn decode_instructions(bytes : Bytes, body_offset : Int, body_size : Int) -> Result[Array[Instruction], WasmError]

    Skip a function body's local declaration and decode the expression.

    body_offset is the first byte after the body's size prefix, so the locals vector comes first. The declared body_size is a hard bound: an instruction that would run past it is reported rather than read out of the next body, which is what keeps a wrong immediate width from quietly becoming a wrong call graph.

    demangle_full

    fn demangle_full(symbol : String) -> String

    The full decoded name for a symbol: escapes resolved, generic arguments spelled out, lifted closures marked.

    This is a display transform, not a codec. Every step is allowed to give up, and when one does the original symbol is returned unchanged, because a size report that invents a name is worse than one that shows a raw symbol.

    Known limits, all of which fall back to the raw symbol: a type code outside the builtin letters, a named type argument that is not the last one in a list (its path has no terminator, so it cannot be told apart from what follows), and any symbol that is not mangled by MoonBit in the first place.

    describe_error

    fn describe_error(error : WasmError) -> String

    A one-line explanation for a failed parse.

    describe_size_failure

    fn describe_size_failure(options : Options, current_size : Int, baseline_size : Int?) -> String

    Why the size gate failed, in one line, and an empty string when it passed.

    It asks the gate rather than re-deriving the answer, so the message and the exit code can never disagree about whether the run was over budget.

    display_name

    fn display_name(name : String) -> String

    The name to print for a symbol: demangled when the binary mangles it, otherwise left exactly as the producer wrote it.

    find_sections

    fn find_sections(parsed : WasmModule) -> Sections

    Pick the interesting sections out of a parsed section table.

    format_report

    fn format_report(parsed : WasmModule, path : String) -> String

    The user-facing report: one line per section, each with its byte count.

    The result carries no trailing newline, so the caller decides how to end it.

    function_stats

    fn function_stats(bodies : Array[FunctionBody], names : Map[Int, String]?) -> Array[FunctionStat]

    Attribute a body to a name and a package.

    Attribution is best effort by design: a release binary keeps no function names, and the sizes are still worth reporting, so unnamed functions come back labelled (unnamed) rather than failing the run.

    indirect_targets

    fn indirect_targets(graph : CallGraph) ->
    Set
    [Int]

    The functions any call_indirect could reach, for labelling report rows.

    module_of

    fn module_of(name : String) -> String

    The package a symbol belongs to.

    Everything in the name except the final segment is the package, so demo::sample::src::fib lands in demo::sample::src and tlsf/searchBlock in tlsf. Names with no package at all are grouped under (top-level), which keeps the roll-up honest instead of inventing one bucket per function.

    parse_code_section

    fn parse_code_section(data : Bytes, section : Section, base_index? : Int) -> Result[Array[FunctionBody], WasmError]

    Decode the code section into one record per function body.

    base_index is the number of imported functions, so the returned index values line up with the function index space. Bodies are measured and skipped, never interpreted, which keeps this a structural pass like parse_module.

    parse_count

    fn parse_count(text : String) -> Int?

    Parse a non-negative decimal integer, or None for anything else.

    This toolchain's core library ships no string-to-Int parser, and the only number the CLI accepts is --top, so the grammar stays deliberately narrow: one or more ASCII digits, saturating into a rejection rather than an overflow.

    parse_element_section

    fn parse_element_section(data : Bytes, section : Section) -> Result[Array[ElementSegment], WasmError]

    Decode the element section.

    The eight encodings differ only in whether a table index, an offset expression, an element kind or a reference type is present, so the shared tail — a vector of function indices or of constant expressions — is read once after the mode byte has been handled.

    parse_export_section

    fn parse_export_section(data : Bytes, section : Section) -> Result[Array[Export], WasmError]

    Decode the export section.

    parse_module

    fn parse_module(data : Bytes) -> Result[WasmModule, WasmError]

    Decode the module preamble and walk every section header.

    Payloads are measured but not interpreted, which keeps this step a pure structural pass over the file.

    parse_name_section

    fn parse_name_section(data : Bytes, section : Section) -> Map[Int, String]?

    Decode the function names out of a name custom section.

    The result is keyed by function index. None means the section is not a name section at all; Some of an empty map is a real answer, because names are optional producer metadata and a stripped binary simply has none. Subsection ids other than 1 (locals, labels, types) are skipped whole, and a malformed subsection ends the walk without discarding earlier names.

    parse_options

    fn parse_options(args : Array[String]) -> Result[Options, String]

    Parse a full argv, including the program name in slot zero.

    The grammar is moonsize <file.wasm> [--top <n>] [--retained] [--dead-code][--call-graph <path>] [--html <path>] [--max-size <size>] [--baseline<path.wasm>] [--compress]: exactly one positional path, and a repeat of a flag keeps the last value. Anything else is reported as an error so a typo never silently analyses the wrong thing.

    --call-graph and --html each write a file rather than a report, so combining either with another output mode is rejected instead of quietly ignoring one of them.

    parse_size

    fn parse_size(text : String) -> Int?

    Parse a byte size: a decimal count with an optional KB, MB or GB suffix, so --max-size 10KB and --max-size 1.5MB both work.

    The suffixes are binary — a KB is 1024 bytes — which is what a size report means by the word. A fraction has to be followed by a suffix: --max-size1.5 is rejected, because half a byte is not a size.

    parse_start_section

    fn parse_start_section(data : Bytes, section : Section) -> Result[Int, WasmError]

    The function the module runs at instantiation, from the start section.

    parse_table_section

    fn parse_table_section(data : Bytes, section : Section) -> Result[Array[TableType], WasmError]

    Decode the table section.

    reachable_from_roots

    fn reachable_from_roots(graph : CallGraph, roots :
    Set
    [Int]) ->
    Set
    [Int]

    Everything reachable from roots, the roots included.

    An explicit stack rather than recursion: real modules nest thousands of calls deep, and the depth of the graph should not be the depth of the C stack.

    render_compression

    fn render_compression(raw_size : Int, gzip_size : Int, sections : Array[SectionCompression]) -> String

    The compression report: what the file costs as a transfer, and which sections that cost is made of.

    SHARE(GZIP) is a section's compressed size as a share of the whole file's compressed size, not of its raw size, because that is the distribution a network pays for. A section that is large raw and small compressed is not the expensive one, and only this column says so.

    The totals and the per-section numbers come from compressing each piece on its own, so the sections do not add up to the total. Two effects are at work: gzip finds redundancy between sections that compressing one in isolation cannot see, and every stream pays its own header and trailer — about 20 bytes — which a section smaller than that shows up as a ratio below 1. The shares barely move for the same reason: 20 bytes against a table of kilobytes.

    render_dead_code

    fn render_dead_code(stats : Array[FunctionStat], reachable :
    Set
    [Int], file_size : Int, top : Int) -> String

    The dead-code section: functions no root can reach.

    This is the report that pays for the analysis. Anything listed here cannot run: no export, no start function, no table entry and no path of calls leads to it, so its bytes are already paid for and never used.

    render_dot

    fn render_dot(stats : Array[FunctionStat], graph : CallGraph, reachable :
    Set
    [Int]) -> String

    Render the call graph as Graphviz DOT.

    Dead functions are drawn dashed, which is what makes the picture useful: the solid part is the program, the dashed part is what could be deleted. Indirect edges are dotted, because they are candidates rather than certainties.

    render_graph

    fn render_graph(path : String, analysis : Analysis, target : String) -> String

    Render a graph export, choosing the format from the target file's extension.

    The extension is the only signal available and the one a user already thinks in: .json means JSON, anything else means DOT, which is what Graphviz tooling expects from a file called graph.dot.

    render_html

    fn render_html(analysis : Analysis, file_path : String, echarts? : String) -> String

    Render the whole analysis as one HTML document.

    echarts is the chart library's source. Passing it inlines the library, so the report opens from any directory with no network and no sibling file; leaving it empty emits a CDN <script src> instead, which keeps the command working for someone who has the binary but not the repository around it. The library does no I/O, so reading assets/echarts.min.js is the caller's job.

    render_json

    fn render_json(path : String, stats : Array[FunctionStat], graph : CallGraph, roots :
    Set
    [Int], reachable :
    Set
    [Int], retained : Array[RetainedSize]) -> String

    Render the call graph, the roots and the retained sizes as JSON.

    render_module_summary

    fn render_module_summary(stats : Array[FunctionStat], file_size : Int) -> String

    The per-package roll-up of every function in the binary.

    render_retained

    fn render_retained(stats : Array[FunctionStat], retained : Array[RetainedSize], indirect :
    Set
    [Int], file_size : Int, top : Int) -> String

    The retained-size table: what deleting each function would free.

    SIZE is the function on its own and RETAINED is that plus every function that becomes unreachable with it, with DIES counting how many others that is. A row where the two sizes match is load-bearing only for itself; a row where RETAINED dwarfs SIZE is a good place to look when shrinking a binary, because removing it takes everything underneath with it.

    render_size_comparison

    fn render_size_comparison(baseline : WasmModule, current : WasmModule) -> String

    The size comparison table: the whole file and every section, against the module it is being compared with.

    Sections are ordered by their current size, so the table reads like the section report above it, and the totals come first because that is the number a budget is spent against.

    render_top_functions

    fn render_top_functions(stats : Array[FunctionStat], file_size : Int, top : Int) -> String

    The Top-N function table, heaviest first.

    Shares are of the whole file, so a row here can be read against the section table above it. top is clamped to the number of functions available.

    retained_sizes

    fn retained_sizes(graph : CallGraph, sizes : Map[Int, Int], roots :
    Set
    [Int]) -> Array[RetainedSize]

    The retained size of every reachable function.

    Deleting a function makes exactly the functions it dominates unreachable: those every path from a root has to pass through. Computing that with an explicit "remove it and see what dies" walk would be correct but quadratic, and it would still need cycle handling; the dominator tree answers the same question for every function at once, and cycles fall out of the fixpoint without a special case.

    roots_of

    fn roots_of(exports : Array[Export], elements : Array[ElementSegment], start : Int?) ->
    Set
    [Int]

    The roots of a module: what the outside world can reach without going through code the module already has.

    Exported functions are the module's public surface, the start function runs at instantiation, and anything an element section puts in a table can be called from outside or through call_indirect.

    same_kind

    fn same_kind(a : SectionId, b : SectionId) -> Bool

    True when two section ids are the same variant.

    Section ids carry no payload and this toolchain derives no Eq for them, so the comparison is written out once here rather than spelled at each call site that has to pick a section out of the table.

    scan_calls

    fn scan_calls(bytes : Bytes, body_offset : Int, body_size : Int) -> Result[Array[CallSite], WasmError]

    Collect every call a function body makes.

    Only call and call_indirect count here. ref.func is a reference rather than a call, and the call graph picks those up separately.

    size_exit_code

    fn size_exit_code(options : Options, current_size : Int, baseline_size : Int?) -> Int

    The exit code a module of current_size bytes earns, measured against baseline_size when the invocation named a baseline.

    Without one, --max-size is a ceiling on the file. With one it is the growth the file may show over the baseline, because that is the question a CI job is asking when it compares a branch against the branch it merges into — an absolute ceiling there would be answering something nobody asked.

    usage

    fn usage() -> String

    Printed when the tool is invoked without a file, or with arguments it cannot make sense of.