docx2html

    Native MoonBit DOCX to HTML/Markdown converter ported from Mammoth

    docx
    html
    markdown
    mammoth
    Download zip
    Version
    0.6.1
    License
    Apache-2.0
    Last updated
    10 days ago
    Downloads
    43

    #moonbitlang/docx2html

    Native MoonBit DOCX reader and converter, ported from Mammoth. The current focus is the common docx -> html, docx -> markdown, and raw-text paths with Mammoth-compatible diagnostics where the JavaScript library exposes them.

    #Install

    moon add moonbitlang/docx2html

    #Usage

    ///|
    test "convert document model to html" {
    let doc = @docx2html.document([
    @docx2html.paragraph([@docx2html.text("Hello.")]),
    ])
    let result = @docx2html.convert_document_to_html(doc)
    inspect(result.value, content="<p>Hello.</p>")
    }

    Use convert or convert_document with output_format=Markdown when the format is chosen dynamically:

    ///|
    test "choose markdown output" {
    let doc = @docx2html.document([
    @docx2html.paragraph([@docx2html.text("Hello.")]),
    ])
    let result = @docx2html.convert_document(doc, output_format=Markdown)
    inspect(result.value, content="Hello\\.\n\n")
    }

    For DOCX bytes, use convert, convert_to_html, convert_to_markdown, or extract_raw_text. The native API accepts BytesView, so callers can decide how to load files. External linked images stay disabled by default, matching Mammoth; pass external_file_access=true and a read_external_file callback when the source document is allowed to read sibling files. Use read_docx_with_messages when you need the parsed document model together with DOCX reader diagnostics; read_docx is the document-only convenience wrapper.

    #Style Maps

    Pass explicit style-map lines with style_map. If you already have a Mammoth-style multi-line string, normalize it with read_style_map_string so comments, blank lines, and trimming follow Mammoth's option parser:

    ///|
    test "parse mammoth style map strings" {
    let style_map_text =
    #|# ignored
    #|
    #|p[style-name='Heading 1'] => h2
    let style_map = @docx2html.read_style_map_string(style_map_text)
    let doc = @docx2html.document([
    @docx2html.paragraph(
    [@docx2html.text("Title")],
    properties=@docx2html.paragraph_properties(style_name=Some("Heading 1")),
    ),
    ])
    let result = @docx2html.convert_document_to_html(doc, style_map~)
    inspect(result.value, content="<h2>Title</h2>")
    }

    Use parse_style_map or parse_style_map_string when you want to validate a style map before conversion while preserving the valid mappings:

    ///|
    test "preflight style map diagnostics" {
    let parsed = @docx2html.parse_style_map_string(
    "p.SectionTitle => h2\np => span#",
    )
    inspect(parsed.mappings.length(), content="1")
    debug_inspect(
    parsed.messages,
    content=(
    #|[Warning("Did not understand this style mapping, so ignored it: p => span#\nError was at character number 10: Expected end but got unrecognisedCharacter \"#\"")]
    ),
    )
    }

    #Document Model

    The document model exposes enum constructors for pattern matching and lower-snake helpers for building nodes and property records:

    ///|
    test "build document model with helpers" {
    let doc = @docx2html.document([
    @docx2html.paragraph(
    [
    @docx2html.hyperlink(
    [@docx2html.text("MoonBit")],
    href=Some("https://www.moonbitlang.com"),
    ),
    ],
    properties=@docx2html.paragraph_properties(style_name=Some("Heading 1")),
    ),
    @docx2html.table([
    @docx2html.table_row([
    @docx2html.table_cell([@docx2html.paragraph([@docx2html.text("Cell")])]),
    ]),
    ]),
    ])
    let result = @docx2html.convert_document_to_html(doc, style_map=[
    "p[style-name='Heading 1'] => h2",
    ])
    inspect(
    result.value,
    content="<h2><a href=\"https://www.moonbitlang.com\">MoonBit</a></h2><table><tr><td><p>Cell</p></td></tr></table>",
    )
    }

    Embedded style maps can be inspected or rewritten without a JavaScript-style mutable ZIP object. Use read_embedded_style_map(docx) to read the raw mammoth/style-map part, and embed_style_map(docx, "p => h1") to return a new DOCX archive with the style map, relationship entry, and content-type override updated.

    #Transforms

    Document transforms run before conversion and recurse through child-bearing nodes. Use the typed helpers instead of Mammoth's JavaScript string type names:

    ///|
    test "transform runs before conversion" {
    let doc = @docx2html.document([
    @docx2html.paragraph([@docx2html.run([@docx2html.text("Hello.")])]),
    ])
    let transform_document = @docx2html.transform_runs(fn(_run) {
    @docx2html.run([@docx2html.text("Goodbye.")])
    })
    let result = @docx2html.convert_document_to_html(doc, transform_document~)
    inspect(result.value, content="<p>Goodbye.</p>")
    let runs = @docx2html.document_descendants_where(doc, fn(element) {
    element is Run(..)
    })
    inspect(runs.length(), content="1")
    }

    #Native CLI

    From this repository checkout, the native executable mirrors the common Mammoth CLI paths:

    moon run --target native cmd/docx2html -- input.docx moon run --target native cmd/docx2html -- --output-format=markdown input.docx moon run --target native cmd/docx2html -- --style-map style-map input.docx output.html moon run --target native cmd/docx2html -- --output-dir out input.docx

    The executable examples in tests/cram/cli.md are verified by:

    moon cram test tests/cram

    #Image Conversion

    Images are emitted as data URIs by default. Pass convert_image to override that behavior. data_uri_image is the default converter, and inline_image is the MoonBit name for Mammoth's images.inline/images.imgElement helper:

    ///|
    test "custom image conversion" {
    let image = @docx2html.Image::{
    content_type: "image/png",
    alt_text: Some("chart"),
    data: b"abc",
    }
    let doc = @docx2html.document([@docx2html.paragraph([Image(image)])])
    let convert_image = @docx2html.inline_image(fn(_image) {
    { "src": "/assets/chart.png" }
    })
    let result = @docx2html.convert_document_to_html(doc, convert_image~)
    inspect(
    result.value,
    content="<p><img alt=\"chart\" src=\"/assets/chart.png\" /></p>",
    )
    }

    Image converters return diagnostics explicitly instead of throwing:

    ///|
    test "custom image conversion diagnostics" {
    let image = @docx2html.image("image/png", b"abc", alt_text=Some("chart"))
    let convert_image = fn(_image : @docx2html.Image) {
    @docx2html.image_conversion([], messages=[@docx2html.error("image omitted")])
    }
    let result = @docx2html.convert_document_to_html(image, convert_image~)
    inspect(result.value, content="")
    debug_inspect(result.messages, content="[Error(\"image omitted\")]")
    }

    #Results

    Conversion results carry rendered text plus diagnostics. Use combine_results when stitching several conversion fragments together:

    ///|
    test "combine conversion results" {
    let combined = @docx2html.combine_results([
    @docx2html.success("One"),
    { value: "Two", messages: [Warning("same")], },
    { value: "Three", messages: [Warning("same")], },
    ])
    inspect(combined.value, content="OneTwoThree")
    debug_inspect(combined.messages, content="[Warning(\"same\")]")
    }

    #Status

    This is an active native-first port. The verified scope covers the core docx -> html, docx -> markdown, and raw-text paths, plus embedded style maps, external-image loader callbacks, notes, comments, tables, hyperlinks, complex fields, XML/ZIP helpers, and the native CLI. See MammothParity.md for the current parity ledger, StressTesting.md for large-file comparison notes, and the intentional API differences from Mammoth's JavaScript surface.

    BreakType

    Kind of break represented in the document model.

    Comment

    Word comment metadata and body content.

    ConversionResult

    Converted value together with accumulated diagnostics.

    DocumentElement

    Node in the Mammoth-style document tree.

    DocumentMatcher

    Matcher for document elements in style-map rules.

    DocxPackageResult

    Package-level read result: body document plus sections and header/footer stories. Returned by read_docx_package.

    DocxReadResult

    Document tree plus diagnostics produced while reading a DOCX package.

    DocxXmlResourceLimit

    Machine-readable XML work or allocation guard raised by DOCX processing.

    HtmlNode

    Node in the intermediate HTML tree.

    HtmlPathElement

    One element in a parsed style-map HTML path.

    Image

    Image payload and metadata read from a DOCX package.

    ImageConversion

    Rendered image nodes plus any diagnostics from image conversion.

    Indent

    Paragraph indentation values from WordprocessingML.

    Message

    Diagnostic emitted while reading or converting a document.

    Note

    Footnote or endnote content in the document model.

    Numbering

    List numbering metadata attached to a paragraph.

    OutputFormat

    Selects whether conversion renders HTML or Markdown.

    ParagraphProperties

    Formatting metadata attached to a paragraph.

    RunProperties

    Formatting metadata attached to a run.

    StyleMapParseResult

    Parsed style mappings together with style-map diagnostics.

    StyleMapping

    Parsed mapping from a document matcher to an HTML path.

    StyleNameMatcher

    Matcher for style names in style-map rules.

    TableProperties

    Formatting metadata attached to a table.

    VerticalAlignment

    Vertical positioning for text within a run.

    XmlElement

    XML element with a name, attributes, and child nodes.

    XmlNode

    Node in the lightweight XML tree.

    ZipArchive

    In-memory ZIP archive used to read DOCX parts.

    DocxError

    pub(all) suberror DocxError {
    InvalidZip(message~ : String)
    InvalidXml(message~ : String)
    MissingPart(message~ : String)
    Unsupported(message~ : String)
    ResourceLimit(limit~ :
    DocxXmlResourceLimit
    , message~ : String)
    WriteResourceLimit(kind~ : String, limit~ : Int64, actual~ : Int64, message~ : String)
    } derive(Eq,
    Debug
    )

    Errors raised for invalid archives, invalid XML, missing parts, or unsupported input.

    DocxError::equal

    #deprecated("implicit trait-method promotion is being removed; call via the trait")
    fn DocxError::equal(DocxError, DocxError) -> Bool

    DocxError::not_equal

    #deprecated("implicit trait-method promotion is being removed; call via the trait")
    fn DocxError::not_equal(x : DocxError, y : DocxError) -> Bool

    DocxError::to_repr

    #deprecated("implicit trait-method promotion is being removed; call via the trait")
    fn DocxError::to_repr(DocxError) ->
    Repr

    bookmark_start

    Builds a bookmark-start element.

    checkbox

    Builds a checkbox element.

    column_break

    Builds a column-break element.

    combine_results

    Combines conversion results, preserving values and diagnostics.

    comment

    fn comment(comment_id : String, body : Array[
    DocumentElement
    ], author_name? : String, author_initials? : String) ->
    Comment

    Builds a document comment.

    comment_reference

    fn comment_reference(comment_id : String) ->
    DocumentElement

    Builds a comment-reference element.

    comment_reply

    fn comment_reply(author~ : String, initials? : String, date? : String, reply_to~ : Int, done? : Bool, body : Array[
    DocumentElement
    ]) ->
    CommentSpec
    raise DocxError

    Builds an anchorless REPLY to an earlier comment spec (by its index in the array passed to write_docx_with_comments); the thread linkage lands in word/commentsExtended.xml. See @docx.comment_reply.

    comment_spec

    fn comment_spec(author~ : String, initials? : String, date? : String, from~ : Int, to~ : Int, done? : Bool, body : Array[
    DocumentElement
    ]) ->
    CommentSpec
    raise DocxError

    Validates and builds one comment for write_docx_with_comments: author/initials/date metadata, the inclusive 0-based range of body block indexes it anchors (from..to, both endpoints top-level paragraphs), and its paragraph-only plain-content body. See @docx.comment_spec for the fail-closed rules (non-empty author, lexical xsd:dateTime date, ordered range, non-empty body).

    convert

    fn convert(docx : BytesView, output_format? :
    OutputFormat
    , style_map? : Array[String], include_default_style_map? : Bool, include_embedded_style_map? : Bool, ignore_empty_paragraphs? : Bool, id_prefix? : String, pretty_print? : Bool, convert_image? : (
    Image
    ) ->
    ImageConversion
    , transform_document? : (
    DocumentElement
    ) ->
    DocumentElement
    , external_file_access? : Bool, read_external_file? : (String) -> Bytes?) ->
    ConversionResult
    raise DocxError

    Converts DOCX bytes using the requested output format.

    convert_document

    Converts a document tree using the requested output format.

    convert_document_to_html

    Converts a document tree to HTML.

    convert_document_to_markdown

    Converts a document tree to Markdown.

    convert_to_html

    fn convert_to_html(docx : BytesView, style_map? : Array[String], include_default_style_map? : Bool, include_embedded_style_map? : Bool, ignore_empty_paragraphs? : Bool, id_prefix? : String, pretty_print? : Bool, convert_image? : (
    Image
    ) ->
    ImageConversion
    , transform_document? : (
    DocumentElement
    ) ->
    DocumentElement
    , external_file_access? : Bool, read_external_file? : (String) -> Bytes?) ->
    ConversionResult
    raise DocxError

    Converts DOCX bytes to HTML.

    convert_to_markdown

    fn convert_to_markdown(docx : BytesView, style_map? : Array[String], include_default_style_map? : Bool, include_embedded_style_map? : Bool, ignore_empty_paragraphs? : Bool, id_prefix? : String, convert_image? : (
    Image
    ) ->
    ImageConversion
    , transform_document? : (
    DocumentElement
    ) ->
    DocumentElement
    , external_file_access? : Bool, read_external_file? : (String) -> Bytes?) ->
    ConversionResult
    raise DocxError

    Converts DOCX bytes to Markdown.

    data_uri_image

    Converts one image to an inline data URI image result.

    data_uri_image_converter

    Converts images to inline data URI image nodes.

    default_style_map_lines

    fn default_style_map_lines() -> Array[String]

    Returns Mammoth-compatible default style-map lines.

    document_descendants

    Returns all descendants of a document element.

    document_descendants_where

    Returns descendants that satisfy a predicate.

    embed_style_map

    fn embed_style_map(docx : BytesView, style_map : String) -> Bytes raise DocxError

    Embeds or replaces the DOCX style map part.

    error

    fn error(message : String) ->
    Message

    Creates an error diagnostic.

    escape_html

    fn escape_html(value : String) -> String

    Escapes text for HTML content.

    escape_html_attribute

    fn escape_html_attribute(value : String) -> String

    Escapes text for an HTML attribute value.

    escape_markdown_text

    fn escape_markdown_text(value : String) -> String

    Escapes text for Markdown output.

    extract_raw_text

    fn extract_raw_text(docx : BytesView, external_file_access? : Bool, read_external_file? : (String) -> Bytes?) ->
    ConversionResult
    raise DocxError

    Extracts raw text from DOCX bytes.

    extract_raw_text_from_document

    Extracts raw text from a document tree.

    fresh_html_element

    fn fresh_html_element(tag : String, attributes? : Map[String, String], children? : Array[
    HtmlNode
    ], separator? : String) ->
    HtmlNode

    Builds an HTML element with a fresh generated id.

    Validates and builds one header or footer story for write_docx_with_annotations: variant is the OOXML reference type (default, first, even) and the body is block-level plain content (paragraphs and tables; fields allowed, hyperlinks/images/notes not). See @docx.header_footer_spec.

    html_element

    fn html_element(tag : String, attributes? : Map[String, String], children? : Array[
    HtmlNode
    ], fresh? : Bool, separator? : String) ->
    HtmlNode

    Builds an HTML element node.

    html_text

    fn html_text(value : String) ->
    HtmlNode

    Builds an HTML text node.
    fn hyperlink(children : Array[
    DocumentElement
    ], href? : String?, anchor? : String?, target_frame? : String?) ->
    DocumentElement

    Builds a hyperlink element.

    identity_document_transform

    Returns a document element unchanged.

    image

    fn image(content_type : String, data : Bytes, alt_text? : String?) ->
    DocumentElement

    Builds an image element.

    image_conversion

    Builds an image conversion result from rendered nodes and diagnostics.

    img_element

    Creates an image converter that emits an img element.

    indent

    fn indent(start? : String?, end? : String?, first_line? : String?, hanging? : String?) ->
    Indent

    Builds paragraph indentation metadata.

    inline_image

    Alias for img_element matching Mammoth image API style.

    join_zip_path

    fn join_zip_path(parts : Array[String]) -> String

    Joins ZIP path components with normalized separators.

    line_break

    Builds a line-break element.

    new_blank_docx

    fn new_blank_docx() -> Bytes

    Builds a minimal, schema-valid blank docx (one empty paragraph, Normal style, Letter page). See @docx.new_blank_docx.

    note

    Builds a footnote or endnote.

    note_reference

    fn note_reference(note_type : String, note_id : String) ->
    DocumentElement

    Builds a note-reference element.

    note_spec

    Validates and builds one footnote/endnote body for write_docx_with_annotations (which array it goes in decides the kind). Plain-content, paragraph-only, no nested notes. See @docx.note_spec.

    numbering

    fn numbering(is_ordered : Bool, level : Int) ->
    Numbering

    Builds paragraph numbering metadata.

    open_zip

    fn open_zip(data : BytesView) ->
    ZipArchive
    raise DocxError

    Opens ZIP bytes into an in-memory archive.

    page_break

    Builds a page-break element.

    paragraph_properties

    fn paragraph_properties(style_id? : String?, style_name? : String?, numbering? :
    Numbering
    ?, alignment? : String?, indent? :
    Indent
    ) ->
    ParagraphProperties

    Builds paragraph formatting metadata.

    parse_style_map

    Parses style-map lines and preserves diagnostics.

    parse_style_map_string

    fn parse_style_map_string(style_map : String) ->
    StyleMapParseResult

    Parses a style-map string and preserves diagnostics.

    read_docx

    fn read_docx(docx : BytesView, external_file_access? : Bool, read_external_file? : (String) -> Bytes?) ->
    DocumentElement
    raise DocxError

    Reads DOCX bytes into a document tree.

    read_docx_annotated

    fn read_docx_annotated(docx : BytesView, external_file_access? : Bool, read_external_file? : (String) -> Bytes?) ->
    DocxAnnotatedResult
    raise DocxError

    Reads DOCX bytes into the package representation PLUS the annotation index (comments with threading and anchors, note references). See @docx.read_docx_annotated; the package half is identical to read_docx_package.

    read_docx_package

    fn read_docx_package(docx : BytesView, external_file_access? : Bool, read_external_file? : (String) -> Bytes?) ->
    DocxPackageResult
    raise DocxError

    Reads DOCX bytes into the package-level representation: the body document (identical to read_docx_with_messages) plus sections and header/footer stories. See @docx.read_docx_package.

    read_docx_with_messages

    fn read_docx_with_messages(docx : BytesView, external_file_access? : Bool, read_external_file? : (String) -> Bytes?) ->
    DocxReadResult
    raise DocxError

    Reads DOCX bytes into a document tree and diagnostics.

    read_embedded_style_map

    fn read_embedded_style_map(docx : BytesView) -> String? raise DocxError

    Reads the embedded style map from DOCX bytes, if present.

    read_style_map

    Parses style-map lines and returns valid mappings.

    read_style_map_string

    fn read_style_map_string(style_map : String) -> Array[String]

    Splits a style-map string into meaningful lines.

    read_style_mapping

    fn read_style_mapping(line : String) ->
    StyleMapping
    ?

    Parses one style-map line into a mapping, if valid.

    read_xml_string

    fn read_xml_string(value : String, namespace_map? : Map[String, String]) ->
    XmlElement
    raise DocxError

    Parses XML text into an element.

    run_properties

    fn run_properties(style_id? : String?, style_name? : String?, is_bold? : Bool, is_underline? : Bool, is_italic? : Bool, is_strikethrough? : Bool, is_all_caps? : Bool, is_small_caps? : Bool, vertical_alignment? :
    VerticalAlignment
    , font? : String?, font_size? : Int?, highlight? : String?) ->
    RunProperties

    Builds run formatting metadata.

    simplify_html

    Simplifies adjacent and redundant HTML nodes.

    split_zip_path

    fn split_zip_path(path : String) -> (String, String)

    Splits a ZIP path into directory and filename.

    success

    Creates a successful conversion result with no diagnostics.

    table_cell

    Builds a table-cell element.

    table_properties

    fn table_properties(style_id? : String?, style_name? : String?) ->
    TableProperties

    Builds table formatting metadata.

    text

    Builds a text element.

    warning

    fn warning(message : String) ->
    Message

    Creates a warning diagnostic.

    write_docx

    Serializes semantic body content (paragraphs and runs, F1 surface) into a schema-valid docx. Fail-closed: unsupported elements raise Unsupported rather than dropping content. See @docx.write_docx.

    write_docx_with_annotations

    The full annotation writer: write_docx plus comments, footnotes, and endnotes. Body runs reference notes with note_reference(kind, index) (0-based into the matching array; exactly one reference per note). See @docx.write_docx_with_annotations.

    write_docx_with_comments

    write_docx plus comments: anchors are emitted into the anchored paragraphs in the canonical shape and the definitions land in word/comments.xml as a main-part relationship, with dense ids (a spec's array index is its w:id). See @docx.write_docx_with_comments.

    write_html

    fn write_html(nodes : Array[
    HtmlNode
    ], pretty_print? : Bool) -> String

    Serializes HTML nodes to a string.

    write_markdown

    fn write_markdown(nodes : Array[
    HtmlNode
    ]) -> String

    Serializes HTML nodes to Markdown.

    write_xml_string

    fn write_xml_string(element :
    XmlElement
    , namespaces? : Map[String, String]) -> String

    Serializes an XML element to text.

    xml_element

    fn xml_element(name : String, attributes? : Map[String, String], children? : Array[
    XmlNode
    ]) ->
    XmlElement

    Builds an XML element.

    xml_text

    fn xml_text(value : String) ->
    XmlNode

    Builds an XML text node.

    zip_archive

    fn zip_archive(entries : Map[String, BytesView]) ->
    ZipArchive
    raise DocxError

    Builds a ZIP view from an entry map.