amazon-ion-moonbit-core

    A bounded MoonBit Amazon Ion 1.0 value-stream codec with text and binary decoding, deterministic encoding, ordered structs, annotations, symbols, and resource limits.

    amazon-ion
    ion
    serialization
    codec
    data-format
    Download zip
    Version
    0.1.1
    License
    Apache-2.0
    Last updated
    12 hours ago
    Downloads
    3

    #Amazon Ion for MoonBit

    这是一个面向 MoonBit 的 Amazon Ion 1.0 数据格式库。它把同一组值映射到 Ion 文本和 Ion 二进制两种表示,适合配置交换、跨语言 fixture、消息边界和需要长期保留类型信息的应用。

    库的核心取舍是:公共值模型保持数据的结构与顺序,解码过程默认有资源上限,写出过程保持输入顺序并产生稳定结果。这样可以在不引入文件系统、网络或 C 依赖的情况下,将 codec 嵌入 native、WebAssembly 和 JavaScript 程序。

    #包布局

    包作用
    format_modelIon 值、symbol、decimal、timestamp、限制和错误模型
    decode文本扫描解析与二进制读取
    encode文本写出与二进制写出
    根包对常用 parse/encode 操作的简短转发 API
    fixtures小型有效样例与覆盖说明
    maintenance规范、互操作性和工具链维护记录

    #支持范围

    当前实现覆盖 Ion 1.0 的 null、bool、整数、Float、decimal、timestamp、string、symbol、blob、clob、list、sexp、struct、类型注解和局部符号表。文本端也支持长字符串、长 clob、数值分隔符和常用 Ion 转义。

    IonValue 会保留未知 symbol SID、annotation 顺序、struct 字段顺序和重复字段。文本与二进制写出采用保序确定性策略,但不宣称实现外部 Canonical Ion 标准。

    解析器和写出器都支持深度、值数量、容器元素数量、文本字节数、blob/clob 字节数、总输入字节数和 symbol 数量限制。格式错误包含稳定错误类别;二进制错误带字节偏移,文本错误带行列位置。

    以下能力明确留给后续版本:Ion 1.1、shared symbol table catalog、Ion Schema、Ion Hash、压缩、文件系统 I/O、零拷贝和完整流式 API。遇到未实现的扩展时会返回错误,不静默降级。

    #最小用法

    import { "LL124-Arch/amazon-ion-moonbit-core" @ion }

    let values = @ion.parse_text("{answer: 42, enabled: true}")
    let binary = @ion.encode_binary(values[:])
    let decoded = @ion.parse_binary(binary)
    let text = @ion.encode_text(decoded[:])

    需要更细的包边界时,可以直接使用 decode、encode 和 format_model。examples/tiny_read 展示文本读取后再写出,examples/tiny_write 展示从模型构造 struct 并生成文本。

    模块标识为 LL124-Arch/amazon-ion-moonbit-core,与公开 GitHub 仓库的归属保持一致。

    #安全标量访问

    format_model 提供 IonValue::string、bool、int、decimal 和 timestamp 构造函数,以及对应的 as_* 投影。投影会忽略值外层的 Ion annotation;类型不匹配(也包括 typed null)返回 None,不会将值强制转换。整数始终以 BigInt 保存,decimal 的 coefficient/exponent 和 timestamp 的精度、时区偏移及小数秒均原样保留。

    import {
    "LL124-Arch/amazon-ion-moonbit-core/format_model" @ion,
    "moonbitlang/core/bigint",
    }

    let value = @ion.IonValue::int(
    @bigint.BigInt::from_string("123456789012345678901234567890"),
    ).with_annotations([@ion.IonSymbol::from_text("id")])

    match value.as_int() {
    Some(identifier) => println(identifier.to_string())
    None => println("expected an Ion int")
    }

    #文档序列处理

    根包还提供面向完整、有界文档的序列操作:slice_values 选择零基范围,filter_values_by_kind 按 Ion 类型筛选,merge_documents 按输入顺序合并多个文档并检查配置的值、深度、符号、载荷和容器限制。reencode_text_range 与 reencode_binary_range 则复用解析器和写出器来选择一段文档值并生成新的完整文档;二进制输出的版本标记由输出选项控制。它们不是流式 API。

    #写出前预检与诊断

    调用方如果先构造了 Ion 值,可以在写出前用 preflight_values 对可从值模型测得的预算做一次检查。成功时返回 IonStats;失败时仍抛出带精确 LimitKind 的 IonError::Limit,便于按资源类别恢复或记录。最终文本或二进制大小依旧由编码器的总字节限制检查,因为逃逸和二进制长度编码会影响实际输出大小。

    import {
    "LL124-Arch/amazon-ion-moonbit-core" @ion,
    "LL124-Arch/amazon-ion-moonbit-core/format_model" @model,
    }

    let values = @ion.parse_text("{name: \"ion\"}")
    let stats = @ion.preflight_values(
    values[:],
    limits=@model.Limits::new(max_depth=8, max_text_bytes=1024),
    )
    println("values: \{stats.value_count()}")

    let limit_error = try @ion.preflight_values(
    values[:],
    limits=@model.Limits::new(max_values=0),
    ) catch {
    error => error
    } noraise {
    _ => fail("expected configured limit")
    }
    match limit_error.limit_kind() {
    Some(kind) => println("budget exceeded: \{to_repr(kind)}")
    None => println(limit_error.summary())
    }

    #互操作诊断

    encode_hex 和 decode_hex 提供稳定的小写十六进制转换;解码器允许 ASCII 空白与 # 行注释,因而可以直接读取仓库中的可审查二进制 fixture。decode_hex 会拒绝非十六进制字符和不完整字节对。IonError::location() 将错误的偏移、行和列封装为 IonLocation,其 format() 固定为 line:column (offset N);IonError::summary() 则提供适合日志的稳定摘要。IonError::limit_kind() 让调用方在不解析摘要文本的情况下识别资源预算。需要核对同一文档的两种表示时,diagnose_text_binary 返回两端的顶层值数量和精确匹配结果;比较不会重排字段、丢弃重复字段、注解或未解析 symbol SID。

    #验证与样例

    仓库中的 fixtures/valid 保存文本样例和最小二进制样例;ion_test.mbt 覆盖空容器、嵌套值、注解、重复字段、符号表、decimal、timestamp、blob/clob、截断输入和资源限制。

    moon fmt --check moon check moon test moon test --target native moon test --target wasm moon test --target wasm-gc moon test --target js moon build --target native moon build --target wasm moon build --target js moon doc --quiet

    面向评审的逐步演示命令和预期输出见 roundtrip_lab/README.md。项目架构取舍、AI 工具使用和规范来源见开发复盘。

    官方 Ion 实现互操作检查:

    cd roundtrip_lab npm ci node ion-js-diff.mjs

    roundtrip_lab 使用锁定的 amazon-ion/ion-js 5.2.1,仅作为开发时的对照实现,不是运行时依赖。

    #规范参考

    实现边界以 Amazon Ion 1.0 规范、Ion 文本文法 和 Ion 二进制规范 为准。跨语言实现可从 Amazon Ion 官方实现列表 继续核对。

    #许可证

    Apache-2.0,见 LICENSE。

    InteropReport

    pub(all) struct InteropReport {
    text_values : Int
    binary_values : Int
    matches : Bool
    } derive(Eq,
    Debug
    )

    Result of comparing text and binary Ion documents without normalizing them.

    InteropReport::binary_value_count

    fn InteropReport::binary_value_count(self : InteropReport) -> Int

    Return the number of top-level values found in the binary input.

    InteropReport::matches

    fn InteropReport::matches(self : InteropReport) -> Bool

    Report whether parsed values are exactly equal, including annotations, symbols, duplicate fields, and field order.

    InteropReport::text_value_count

    fn InteropReport::text_value_count(self : InteropReport) -> Int

    Return the number of top-level values found in the text input.

    IonStats

    pub(all) struct IonStats {
    value_count : Int
    max_depth : Int
    max_container_items : Int
    symbol_count : Int
    text_bytes : Int
    blob_bytes : Int
    } derive(Eq,
    Debug
    )

    Summary information for a collection of Ion values.

    IonStats::first_exceeded_limit

    Return the first measured resource budget that this document exceeds.

    TotalBytes is intentionally omitted: the final encoded size depends on text escaping, binary length encoding, and output options, and is checked by the encoder after it has produced the representation.

    IonStats::fits

    Check the measured value, depth, container, symbol, text, and binary budgets.

    IonStats::max_container_items

    fn IonStats::max_container_items(self : IonStats) -> Int

    Return the largest element, field, or annotation count in one container.

    IonStats::max_depth

    fn IonStats::max_depth(self : IonStats) -> Int

    Return the deepest structural level encountered by the analysis.

    IonStats::payload_bytes

    fn IonStats::payload_bytes(self : IonStats) -> Int

    Return the combined UTF-8 and blob/clob payload size in bytes.

    IonStats::symbol_count

    fn IonStats::symbol_count(self : IonStats) -> Int

    Return the number of symbol occurrences, including field names and annotations.

    IonStats::value_count

    fn IonStats::value_count(self : IonStats) -> Int

    Return the number of Ion values, excluding annotation wrappers.

    PathSegment

    pub(all) enum PathSegment {
    Field(String)
    Index(Int)
    } derive(Eq,
    Debug
    )

    An unambiguous segment in a read-only Ion path query.

    Field selects every matching ordered struct field, including a field whose name is numeric. Index selects one element from a list or s-expression.

    analyze

    Analyze value count, nesting, symbols, and payload sizes before encoding.

    binary_version

    fn binary_version(input : Bytes) -> (Int, Int)?

    Return the Ion version marker when the input begins with a binary version marker.

    decode_hex

    fn decode_hex(input : String) -> Bytes raise
    IonError

    Decode hexadecimal fixture text.

    ASCII whitespace and # line comments are ignored; every remaining character must be a hexadecimal digit and digits must occur in pairs.

    diagnose_text_binary

    Parse text and binary Ion and compare their preserved value models.

    encode_hex

    fn encode_hex(input : Bytes) -> String

    Encode bytes as lowercase hexadecimal for portable binary Ion fixtures.

    filter_values_by_kind

    Copy document values whose Ion type matches kind, preserving their order.

    Type annotations do not affect the match, consistent with IonValue::kind.

    is_binary

    fn is_binary(input : Bytes) -> Bool

    Report whether the input begins with an Ion binary version marker.

    list_append

    Append an element to a list, preserving any annotations on the list.

    Returns None when value is not a list; s-expressions are intentionally not edited by this list-specific operation.

    list_replace

    Replace a zero-based list element without changing its length or annotations.

    Returns None for a non-list or an out-of-bounds (including negative) index.

    merge_documents

    Concatenate documents in their supplied order after enforcing sequence limits.

    The returned document has no synthetic Ion version marker: version markers are document framing and are emitted by encode_binary when configured. Nested values are retained as supplied, while the top-level sequence is newly owned.

    parse_bytes

    Parse either UTF-8 Ion text or a binary document based on its prefix.

    parse_one_bytes

    Parse exactly one value from either UTF-8 text or binary input.

    preflight_values

    Analyze a document and reject the first measured resource budget it exceeds.

    This is suitable immediately before writing an application-owned value document. It preserves the IonError::Limit category and its LimitKind; encoded byte size remains checked by encode_text or encode_binary.

    Example

    test {
    let values = [@model.IonValue::string("ion")]
    let stats = preflight_values(
    values[:],
    limits=@model.Limits::new(max_text_bytes=3),
    )
    assert_eq(stats.value_count(), 1)
    }

    reencode_binary_range

    Parse a binary document, select a bounded range of its values, and encode it.

    A selected binary range is written as one new Ion document, so output binary version-marker emission is controlled by options rather than copied from the input framing.

    reencode_text_range

    Parse a text document, select a bounded range of its values, and encode it.

    Parsing enforces input_limits; encoding enforces the limits carried by options. This deliberately works on complete, bounded documents rather than presenting itself as a streaming API.

    select_document_typed_path

    Apply an explicit path to every root value in an Ion document.

    Results retain document order and repeated-field order. An empty path copies the document roots into the returned result, rather than exposing the input array for mutation.

    select_field

    Return every field with the requested name from a struct value.

    select_first_path

    Return the first value selected by a path, if the path has a result.

    select_first_typed_path

    Return the first value selected by an explicit path, if any.

    select_index

    Return an element by index from a list or s-expression.

    select_path

    Select values by alternating struct field names and non-negative indexes.

    select_typed_path

    Select values with explicit field and index path segments.

    Unlike select_path, a Field("0") always means the struct field named "0", never list element zero. Repeated fields and document order are preserved, while a missing field, invalid container kind, or out-of-range index returns an empty result without changing the input.

    slice_values

    Copy at most length document values beginning at the zero-based start.

    Negative starts and non-positive lengths select an empty sequence. The input values themselves are not modified.

    struct_append_field

    Append a text-named field to a struct, keeping its field order and annotations.

    Returns None when value is not a struct (after unwrapping annotations).

    struct_remove_field

    Remove the zero-based matching occurrence of a text-named struct field.

    Other duplicate fields retain their original relative order. Returns None when the requested occurrence is absent, negative, or value is not a struct.

    struct_replace_field

    Replace the zero-based matching occurrence of a text-named struct field.

    The existing field symbol is retained, so symbol IDs and field order stay intact. Returns None for a non-struct, a negative occurrence, or no matching field.