krueger

    Parser and parsing utilities for Elm and Elm-like dialects (e.g. Morphir) in MoonBit

    elm
    elm-like
    morphir
    parser
    scanner
    ast
    Download zip
    Author
    Version
    0.5.0
    License
    Apache-2.0
    Last updated
    4 hours ago
    Downloads
    276

    #moonrockz/krueger

    Parser and parsing utilities for Elm and Elm-like dialects (such as Morphir) in MoonBit.

    krueger parses Elm 0.19.1 source into an exact mirror of the stil4m/elm-syntax 7.3.9 AST. On a corpus of 363 files from published packages, its JSON output is byte for byte the same as elm-syntax's.

    • Scanner: lossless tokens. The trivia (whitespace and comments) and the lexemes of the tokens give back the source text.
    • Parser: tokens to AST, a concrete syntax tree (CST) and diagnostics. A declaration that does not parse is left out and reported; the rest of the file is parsed.
    • AST: elm-syntax 7.3.9 types, with a JSON encoder and decoder that match elm-syntax.
    • Diagnostics: elm make-style error reports (terminal, plain text, and the elm make --report=json shape).
    • Syntax tree: a read-only node model with parent links, position lookup, and six traversal styles: walk, fold, Visitor, pull events, push events and a cursor.
    • Dialects: the rules of elm make 0.19.1 or of elm-syntax 7.3.9, and hooks for Elm-like languages.
    • Doc-comment attributes: @name value data in doc comments, read into the parse result.
    • Printer and formatter: Elm source from an AST in the elm-format layout, and format, which writes what elm-format 0.8.7 writes and keeps every comment.

    #Installation

    moon add moonrockz/krueger

    Then import it in your package's moon.pkg:

    import { "moonrockz/krueger", }

    #Usage

    Parse a module and read the names of its declarations:

    test "parse a module" {
    let source = @krueger.SourceText::new(
    "module Main exposing (main)\n\nmain =\n greet \"world\"\n",
    )
    let result = @krueger.parse_module(source)
    guard result.diagnostics.is_empty() else { fail("syntax errors") }
    guard result.ast is Some(file) else { fail("no AST") }
    for declaration in file.declarations {
    match declaration.value {
    FunctionDeclaration(f) => println(f.declaration.value.name.value)
    _ => ()
    }
    }
    }

    Show a syntax error the way elm make does:

    test "report an error" {
    let source = @krueger.SourceText::new("module Main exposing (..)\n\nx = (1, \n")
    let result = @krueger.parse_module(source)
    for diagnostic in result.diagnostics {
    println(@krueger.render_plain(diagnostic, source, "src/Main.elm"))
    }
    }

    Visit every expression with the syntax tree:

    test "count expressions" {
    let result = @krueger.parse_module(
    @krueger.SourceText::new("module Main exposing (..)\n\nx = f 1 2\n"),
    )
    guard @krueger.NodeRef::of_result(result) is Some(root) else { return }
    let count = @krueger.fold(root, 0, (n, node) => {
    (if node.category() == "expression" { n + 1 } else { n }, Continue)
    })
    println(count)
    }

    let result = @krueger.parse_module(@krueger.SourceText::new(source))
    guard result.ast is Some(file) else { return }
    let text = @krueger.print_file(file) // elm-format layout, width 120

    print_file raises PrintError when the AST cannot print as valid Elm. The error's path leads to the bad node.

    See Generate Elm code for ASTs built in code.

    #Format Elm source

    test "format a module" {
    let source = "module A exposing (a)\n\na = f x -- why\n y\n"
    println(@krueger.format(source))
    }

    This prints what elm-format 0.8.7 writes:

    module A exposing (a) a = f x -- why y

    format keeps every comment, and raises ParseFailed with the diagnostics when the source has a syntax error. layout=Width(80) breaks lines to fit 80 columns instead of where the source breaks them. See Format Elm source.

    #Documentation

    • Cookbook: task articles, such as building an AST explorer, converting to a unist tree, writing a lint rule and editor features. Their code is tested.
    • API reference: the doc comments on mooncakes.io. Most entry points have a tested example.

    #Development

    See AGENTS.md for the architecture, the parity and rejection checks, and the mise tasks. Install the git hooks with:

    mise run hooks:install

    #License

    Apache-2.0

    AttributeGroup

    The parse result and doc-comment attributes.

    AttributeSyntax

    Dialects: rejection rules, operator table and other extension data.

    AttributeTarget

    The parse result and doc-comment attributes.

    Block

    Scanner types: source text, tokens, trivia, positions, spans and diagnostics (with their report blocks).

    Case

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    CaseBlock

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    Cases

    Elm.Syntax.Expression.Cases

    Chunk

    Scanner types: source text, tokens, trivia, positions, spans and diagnostics (with their report blocks).

    Color

    Scanner types: source text, tokens, trivia, positions, spans and diagnostics (with their report blocks).

    Comment

    Scanner types: source text, tokens, trivia, positions, spans and diagnostics (with their report blocks).

    CommentKind

    Scanner types: source text, tokens, trivia, positions, spans and diagnostics (with their report blocks).

    Control

    Read-only node model over a parse result, and its traversals (walk, fold, accept, events, TreeCursor).

    Declaration

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    DeclarationCst

    Concrete syntax tree: every token and top-level declaration, with trivia.

    DecodeError

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    DefaultModuleData

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    Diagnostic

    Scanner types: source text, tokens, trivia, positions, spans and diagnostics (with their report blocks).

    Dialect

    Dialects: rejection rules, operator table and other extension data.

    DocAttribute

    The parse result and doc-comment attributes.

    Documentation

    type Documentation = String

    Elm.Syntax.Documentation.Documentation

    EffectModuleData

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    EnterEvent

    Read-only node model over a parse result, and its traversals (walk, fold, accept, events, TreeCursor).

    Event

    Read-only node model over a parse result, and its traversals (walk, fold, accept, events, TreeCursor).

    EventReader

    Read-only node model over a parse result, and its traversals (walk, fold, accept, events, TreeCursor).

    EventSource

    Read-only node model over a parse result, and its traversals (walk, fold, accept, events, TreeCursor).

    ExposedType

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    Exposing

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    Expression

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    File

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    FormatError

    Elm source from an AST, in the elm-format layout and fitted to a line width (default 120), and Elm source formatted as elm-format 0.8.7 does, with its comments (format, format_parsed).

    Function

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    FunctionCst

    Concrete syntax tree: every token and top-level declaration, with trivia.

    FunctionImplementation

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    Handler

    Read-only node model over a parse result, and its traversals (walk, fold, accept, events, TreeCursor).

    Import

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    ImportCst

    Concrete syntax tree: every token and top-level declaration, with trivia.

    Infix

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    InfixDirection

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    KeywordKind

    Scanner types: source text, tokens, trivia, positions, spans and diagnostics (with their report blocks).

    Lambda

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    Layout

    Elm source from an AST, in the elm-format layout and fitted to a line width (default 120), and Elm source formatted as elm-format 0.8.7 does, with its comments (format, format_parsed).

    LeaveEvent

    Read-only node model over a parse result, and its traversals (walk, fold, accept, events, TreeCursor).

    LetBlock

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    LetDeclaration

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    Location

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    Module

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    ModuleCst

    Concrete syntax tree: every token and top-level declaration, with trivia.

    ModuleHeaderCst

    Concrete syntax tree: every token and top-level declaration, with trivia.

    ModuleName

    type ModuleName = ArrayView[String]

    Elm.Syntax.ModuleName.ModuleName

    NameKind

    Elm source from an AST, in the elm-format layout and fitted to a line width (default 120), and Elm source formatted as elm-format 0.8.7 does, with its comments (format, format_parsed).

    Node

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    NodePath

    Read-only node model over a parse result, and its traversals (walk, fold, accept, events, TreeCursor).

    NodeRef

    Read-only node model over a parse result, and its traversals (walk, fold, accept, events, TreeCursor).

    OperatorDef

    Dialects: rejection rules, operator table and other extension data.

    ParseResult

    The parse result and doc-comment attributes.

    PathStep

    Read-only node model over a parse result, and its traversals (walk, fold, accept, events, TreeCursor).

    Pattern

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    Position

    Scanner types: source text, tokens, trivia, positions, spans and diagnostics (with their report blocks).

    PrintError

    Elm source from an AST, in the elm-format layout and fitted to a line width (default 120), and Elm source formatted as elm-format 0.8.7 does, with its comments (format, format_parsed).

    PrintProblem

    Elm source from an AST, in the elm-format layout and fitted to a line width (default 120), and Elm source formatted as elm-format 0.8.7 does, with its comments (format, format_parsed).

    QualifiedNameRef

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    Range

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    RecordDefinition

    Elm.Syntax.TypeAnnotation.RecordDefinition

    RecordField

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    RecordSetter

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    Rule

    Dialects: rejection rules, operator table and other extension data.

    ScanErrorList

    Scanner types: source text, tokens, trivia, positions, spans and diagnostics (with their report blocks).

    Scanner

    Scanner types: source text, tokens, trivia, positions, spans and diagnostics (with their report blocks).

    Severity

    Scanner types: source text, tokens, trivia, positions, spans and diagnostics (with their report blocks).

    Signature

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    SourceText

    Scanner types: source text, tokens, trivia, positions, spans and diagnostics (with their report blocks).

    Span

    Scanner types: source text, tokens, trivia, positions, spans and diagnostics (with their report blocks).

    Token

    Scanner types: source text, tokens, trivia, positions, spans and diagnostics (with their report blocks).

    TokenKind

    Scanner types: source text, tokens, trivia, positions, spans and diagnostics (with their report blocks).

    TokenStream

    Scanner types: source text, tokens, trivia, positions, spans and diagnostics (with their report blocks).

    TopLevelExpose

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    Tree

    Read-only node model over a parse result, and its traversals (walk, fold, accept, events, TreeCursor).

    TreeCursor

    Read-only node model over a parse result, and its traversals (walk, fold, accept, events, TreeCursor).

    Trivia

    Scanner types: source text, tokens, trivia, positions, spans and diagnostics (with their report blocks).

    Type

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    TypeAlias

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    TypeAliasCst

    Concrete syntax tree: every token and top-level declaration, with trivia.

    TypeAnnotation

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    UnionTypeCst

    Concrete syntax tree: every token and top-level declaration, with trivia.

    ValueConstructor

    elm-syntax 7.3.9 model. Elm.Syntax.Comments.Comment is not re-exported because the scanner's Comment type uses the same name.

    Visitor

    Read-only node model over a parse result, and its traversals (walk, fold, accept, events, TreeCursor).

    accept

    Walk root (see walk) and call one visitor method per node: functions (top-level or let) → visit_function; other declarations → visit_declaration; expressions, patterns, types → visit_expression, visit_pattern, visit_type; case branches → visit_case; imports, comments, doc attributes → visit_import, visit_comment, visit_attribute; every other node (doc comments included) → visit_other. The method's Control steers the walk; leave runs after each node's children.

    A visitor needs its own type, so this example is not a doc test (a doc test holds only test blocks). The cookbook article "Choose a traversal" has a tested visitor: https://github.com/moonrockz/krueger/blob/main/docs/cookbook/traversal.mbt.md

    struct Names {
    functions : Array[String]
    mut patterns : Int
    }

    impl @syntax.Visitor for Names with fn visit_function(self, n) {
    match n {
    Declaration({ value: FunctionDeclaration(f), .. }, _) =>
    self.functions.push(f.declaration.value.name.value)
    _ => ()
    }
    Continue
    }

    impl @syntax.Visitor for Names with fn visit_pattern(self, _) {
    self.patterns 1
    Continue
    }

    test {
    let src = "module Main exposing (..)\n\nadd x y =\n x + y\n\nz = 0\n"
    let result = @parser.parse_module(
    @scanner.SourceText::new(src),
    @scanner.DefaultScanner::new(),
    )
    let root = @syntax.NodeRef::of_result(result).unwrap()
    let names = { functions: [], patterns: 0, }
    @syntax.accept(root, names)
    debug_inspect(names.functions, content="[\"add\", \"z\"]")
    inspect(names.patterns, content="2")
    }

    diagnostic_description

    fn diagnostic_description(code : String) -> String?

    The fixed meaning of a diagnostic code, or None for an unknown code.

    test {
    inspect(
    @krueger.diagnostic_description("KR-SCAN-004").unwrap_or("?"),
    content="Unterminated string, char or GLSL literal",
    )
    debug_inspect(@krueger.diagnostic_description("KR-NONE"), content="None")
    }

    encode_attributes

    Attribute groups as JSON, for tools.

    Each group is {"target": ..., "attributes": [...]}. The target is "module" or {"declaration": name, "range": [...]}. An attribute is {"name": ..., "arguments": [...], "range": [...]}, with names and arguments as elm-syntax nodes; @docs is {"docs": [...], "range":[...]}. Ranges are elm-syntax ranges: [startRow, startColumn, endRow,endColumn]. Int arguments are written as elm-syntax writes them (see encode_attributes_with).

    test {
    let text =
    #|module Main exposing (x)
    #|
    #|{-| Answers.
    #|
    #|@since 2
    #|-}
    #|x = 42
    #|
    let source = @scanner.SourceText::new(text)
    let result = @parser.parse_module(source, @scanner.DefaultScanner::new())
    // The first doc comment documents the module.
    let json = @parser.encode_attributes(result.attributes)
    inspect(
    json.stringify(),
    content=(
    #|[{"target":"module","attributes":[{"name":{"range":[5,2,5,7],"value":"since"},"arguments":[{"range":[5,8,5,9],"value":{"type":"integer","integer":2}}],"range":[5,1,5,9]}]}]
    ),
    )
    }

    encode_attributes_with

    fn encode_attributes_with(groups : Array[
    AttributeGroup
    ], exact_ints~ : Bool) -> Json

    Attribute groups as JSON. With exact_ints, Int arguments are written with their exact digits (see @ast.encode_file_with). Without it, an Int above 2^53 is written as the nearest Double, as elm-syntax does.

    test {
    let text =
    #|module Main exposing (x)
    #|
    #|{-| Ids.
    #|
    #|@id 9007199254740993
    #|-}
    #|x = 1
    #|
    let source = @scanner.SourceText::new(text)
    let groups = @parser.parse_module(source, @scanner.DefaultScanner::new()).attributes
    let exact = @parser.encode_attributes_with(groups, exact_ints=true).stringify()
    inspect(exact.contains("9007199254740993"), content="true")
    let nearest = @parser.encode_attributes(groups).stringify()
    inspect(nearest.contains("9007199254740993"), content="false")
    }

    encode_diagnostics

    All diagnostics (warnings too) with krueger's codes, as JSON.

    Each diagnostic is an object with code, severity (error, warning or info), title, message and span. The long report is not included; use render_plain or render_elm_json for it.

    test {
    let source = @scanner.SourceText::new(
    "module Main exposing (..)\n\nmain =\n (1 + 2\n",
    )
    let result = @parser.parse_module(source, @scanner.DefaultScanner::new())
    let json = @report.encode_diagnostics(result.diagnostics)
    inspect(
    json.stringify(),
    content=(
    #|[{"code":"KR-PARSE-004","severity":"error","title":"UNFINISHED PARENTHESES","message":"Malformed function declaration: expected `,` or `)`, found end of input","span":{"start":{"line":4,"column":11},"end":{"line":4,"column":11}}}]
    ),
    )
    }

    fold

    Walk root (see walk) with an accumulator: enter and leave take the current value and return the next one, in walk's callback order. The result is the value after the last callback. The accumulator belongs to the caller; fold does not copy it.

    test {
    let src = "module Main exposing (..)\n\nadd x y =\n x + y\n"
    let result = @parser.parse_module(
    @scanner.SourceText::new(src),
    @scanner.DefaultScanner::new(),
    )
    let root = @syntax.NodeRef::of_result(result).unwrap()
    // Count the expressions and find the deepest nesting.
    let (count, _, deepest) = @syntax.fold(
    root,
    (0, 0, 0),
    (acc, n) => {
    let (count, depth, deepest) = acc
    let count = count + (if n.category() == "expression" { 1 } else { 0 })
    let depth = depth + 1
    let deepest = if depth > deepest { depth } else { deepest }
    ((count, depth, deepest), Continue)
    },
    leave=(acc, _) => (acc.0, acc.1 - 1, acc.2),
    )
    inspect(count, content="3")
    inspect(deepest, content="5")
    }

    format

    Formats Elm source as elm-format 0.8.7 does (layout=ElmFormat, the default), or with the line-width layout of print_file (layout=Width(n)).

    • The output is the source of normalize_file(ast): exposed items and imports in elm-format's order, elm-format's parentheses and literal forms, and doc comments with their Markdown and Elm code formatted.
    • Every regular comment prints once, where elm-format puts it: a comment leads the node after it, trails the node before it on its line, or goes before the closing token of its container. A comment moves with an exposed item or an import that elm-format reorders. A block comment gets elm-format's form ({-a-} gives {- a -}). No comment is dropped: one that the printer cannot place raises Unprintable with UnplacedComment.
    • Raises ParseFailed when the source has a syntax error in dialect.

    format(format(s)) == format(s), except for the doc comments on which elm-format 0.8.7 is not idempotent (see normalize_file).

    test {
    let source = "module A exposing (a)\n\na = f x -- why\n y\n"
    inspect(
    @printer.format(source),
    content=(
    #|module A exposing (a)
    #|
    #|
    #|a =
    #| f x
    #| -- why
    #| y
    #|
    ),
    )
    }

    format_parsed

    Formats a parse result, as format formats source. Give the dialect that parsed it: the printer uses its operator table. Raises ParseFailed when the result has a diagnostic with severity Error. Use it when the parse result is already there, for example after a check of its diagnostics.

    test {
    let result = @parser.parse_module(
    @scanner.SourceText::new("module A exposing (a)\n\na = [1,2]\n"),
    @scanner.DefaultScanner::new(),
    )
    inspect(
    @printer.format_parsed(result, layout=Width(40)),
    content=(
    #|module A exposing (a)
    #|
    #|
    #|a =
    #| [ 1, 2 ]
    #|
    ),
    )
    }

    is_lower_start

    fn is_lower_start(c : Char) -> Bool

    Whether c can start a lower-case name (a Unicode lower-case letter). Lower-case names are values, functions and type variables.

    test {
    inspect(@scanner.is_lower_start('a'), content="true")
    inspect(@scanner.is_lower_start('é'), content="true")
    inspect(@scanner.is_lower_start('A'), content="false")
    inspect(@scanner.is_lower_start('_'), content="false")
    }

    is_name_part

    fn is_name_part(c : Char) -> Bool

    Whether c can continue a name: a Unicode letter or number, or _.

    test {
    inspect(@scanner.is_name_part('x'), content="true")
    inspect(@scanner.is_name_part('_'), content="true")
    inspect(@scanner.is_name_part('7'), content="true")
    inspect(@scanner.is_name_part('ß'), content="true")
    inspect(@scanner.is_name_part('-'), content="false")
    inspect(@scanner.is_name_part('\''), content="false")
    }

    is_upper_start

    fn is_upper_start(c : Char) -> Bool

    Whether c can start an upper-case name (a Unicode upper-case or title-case letter). Upper-case names are modules, types and constructors. In the elm-syntax-7.3.9 dialect the scanner does not start a name with a title-case letter (rule titlecase-name-start).

    test {
    inspect(@scanner.is_upper_start('M'), content="true")
    inspect(@scanner.is_upper_start('Ä'), content="true")
    inspect(@scanner.is_upper_start('Dž'), content="true")
    inspect(@scanner.is_upper_start('m'), content="false")
    }

    kind_table

    fn kind_table() -> Array[(String, String, Array[String])]

    Every node type: (category, kind, fields), one row per kind. The kinds and field names are the elm-syntax JSON vocabulary.

    test {
    let rows = @syntax.kind_table()
    let row = rows.filter(r => r.0 == "expression" && r.1 == "ifBlock")[0]
    debug_inspect(row.2, content="[\"clause\", \"then\", \"else\"]")
    // `list` is a kind of both expressions and patterns.
    let lists = rows.filter(r => r.1 == "list").map(r => r.0)
    debug_inspect(lists, content="[\"expression\", \"pattern\"]")
    }

    normalize_file

    The file in the order and with the parentheses that print_file gives it (elm-format's, and the ones that the parser needs). For a parsed file, parsing the output of print_file gives this AST (without ranges and regular comments). format gives this AST too, with two differences that come from elm-format: it keeps parentheses that have a comment inside them, and it keeps a chain of operators of the same precedence and different directions flat (a |> f <| g), where this AST has (a |> f) <| g, as print_file writes it (elm make rejects the flat chain). dialect gives the operator table.

    • The module's exposed items are a set: a duplicate is removed (T and T(..) give T(..)). They are sorted: operators, then types, then values, each by name. When the module documentation has @docs lines, the items that they name come first, in the order of those lines.
    • The imports are sorted by module name. The imports of one module are merged into the first: a later alias replaces an earlier one, and the exposing lists are joined ((..) wins). An alias equal to the module name is removed (import A as A). Each exposing list is a sorted set, as above.
    • Parentheses that neither the parser nor elm-format needs go (elm-format formatExpression, syntaxParens): case (f x) of gives case f x of, a - (f b) gives a - f b, f (-x) gives f -x. Parentheses around an operator chain in a chain stay.
    • Parentheses are added where print_file writes them: elm-format's, around a lambda, if, case or let at the end of an operator chain, except after <| (x |> (\y -> y)), and around a constructor pattern with arguments before or after :: and in an as ((Just a) :: rest); and the ones that the parser needs, for example around an operator chain that cannot join its parent's chain ((a + b) * c).

    • Doc comments (documentation and the doc comments in File.comments) are written as elm-format writes them: Markdown through @markdown.format_doc, Elm code in them formatted (see format). As elm-format, this is not idempotent for some doc comments: @docsT(..) (it names no exposed item, then it names T(..) as T), emphasis that starts after a letter (a*b c* gives a_b c_, then text), two bullet lists in a row (one list the next time), [a] [a] (one link), and a fenced Elm code block that starts with a blank line (the line goes the next time).

    Nodes keep their ranges; an expression that replaces its parentheses takes their range. Nothing else changes.

    test {
    let text = "module A exposing (b, a)\n\nimport C\nimport B as B\n\n\na =\n 1\n"
    let file = @parser.parse_module(
    @scanner.SourceText::new(text),
    @scanner.DefaultScanner::new(),
    ).ast.unwrap()
    inspect(
    @printer.print_file(@printer.normalize_file(file)),
    content=(
    #|module A exposing (a, b)
    #|
    #|import B
    #|import C
    #|
    #|
    #|a =
    #| 1
    #|
    ),
    )
    }

    parse_module

    Tokenize and parse source in dialect (default: Elm 0.19.1).

    The result has the elm-syntax AST (ast), the CST (cst), the diagnostics and the doc-comment attributes. The parser does not stop at the first error: a declaration that does not parse is left out of the AST and the CST, and is reported. ast is None only when the module header is missing or malformed, or when the source does not scan. Check diagnostics for an entry with severity Error before you trust the AST.

    test {
    let source = @krueger.SourceText::new(
    "module Main exposing (main)\n\nmain =\n 1 + 2\n",
    )
    let result = @krueger.parse_module(source)
    inspect(result.diagnostics.length(), content="0")
    guard result.ast is Some(file) else { fail("no AST") }
    guard file.declarations[0].value is FunctionDeclaration(f) else {
    fail("not a function")
    }
    inspect(f.declaration.value.name.value, content="main")
    debug_inspect(
    f.declaration.value.name.range,
    content="{ start: { row: 3, column: 1 }, end: { row: 3, column: 5 } }",
    )
    }

    A syntax error is a diagnostic with a code, a title and a message:

    test {
    let source = @krueger.SourceText::new(
    "module Main exposing (..)\n\nmain =\n (1 + 2\n\nother = 3\n",
    )
    let result = @krueger.parse_module(source)
    let d = result.diagnostics[0]
    inspect(d.code, content="KR-PARSE-004")
    inspect(d.title, content="UNFINISHED PARENTHESES")
    debug_inspect(d.severity, content="Error")
    // The declaration after the error still parses.
    guard result.ast is Some(file) else { fail("no AST") }
    inspect(file.declarations.length(), content="1")
    }

    Pass a dialect to accept or reject other syntax. elm make rejects a leading zero; elm-syntax 7.3.9 accepts it:

    test {
    let source = @krueger.SourceText::new(
    "module Main exposing (..)\n\nx =\n 007\n",
    )
    let strict = @krueger.parse_module(source)
    inspect(
    strict.diagnostics[0].message,
    content="Malformed function declaration: numbers cannot start with zeros [rule: leading-zero]",
    )
    let lenient = @krueger.parse_module(
    source,
    dialect=@krueger.Dialect::elm_syntax_7_3_9(),
    )
    inspect(lenient.diagnostics.length(), content="0")
    }

    parse_tokens

    Parse a token stream in dialect (default: Elm 0.19.1).

    Use it when you already have the tokens from tokenize. Tokenize and parse with the same dialect: the dialect's reserved words and operator symbols change the tokens. parse_module does both steps.

    test {
    let dialect = @krueger.Dialect::elm_0_19_1()
    let source = @krueger.SourceText::new("module Main exposing (..)\n\nx = 1\n")
    guard @krueger.tokenize(source, dialect~) is Ok(tokens) else {
    fail("scan error")
    }
    let result = @krueger.parse_tokens(tokens, dialect~)
    inspect(result.diagnostics.length(), content="0")
    inspect(result == @krueger.parse_module(source, dialect~), content="true")
    }

    The Elm source of a declaration, without a trailing line feed. A port's doc comment is not part of its declaration (elm-syntax keeps it in File.comments), so only print_file prints it.

    test {
    let r : @ast.Range = {
    start: { row: 0, column: 0, },
    end: { row: 0, column: 0, },
    }
    let int : @ast.Node[@ast.TypeAnnotation] = {
    range: r,
    value: Typed({ range: r, value: ([][:], "Int"), }, [][:]),
    }
    let decl : @ast.Node[@ast.Declaration] = {
    range: r,
    value: AliasDeclaration({
    documentation: None,
    name: { range: r, value: "Age", },
    generics: [][:],
    type_annotation: int,
    }),
    }
    inspect(
    @printer.print_declaration(decl),
    content=(
    #|type alias Age =
    #| Int
    ),
    )
    }

    The Elm source of an expression, with elm-format's parentheses: parentheses are added where precedence needs them and where elm-format writes them (a lambda, if, case or let at the end of an operator chain, except after <|); parentheses in the AST that neither needs go (see normalize_file). if, case and let are always multi-line (elm-format). Raises PrintError for an expression that cannot print as valid Elm; the error's path starts at e.

    test {
    let r : @ast.Range = {
    start: { row: 0, column: 0, },
    end: { row: 0, column: 0, },
    }
    fn v(name : String) -> @ast.Node[@ast.Expression] {
    { range: r, value: FunctionOrValue([][:], name), }
    }
    let sum : @ast.Node[@ast.Expression] = {
    range: r,
    value: OperatorApplication("+", Left, v("a"), v("b")),
    }
    let product : @ast.Node[@ast.Expression] = {
    range: r,
    value: OperatorApplication("*", Left, sum, v("c")),
    }
    inspect(@printer.print_expression(product), content="(a + b) * c")
    }

    The Elm source of a file in the elm-format layout, ending with one line feed. Fixed shapes (declaration bodies, custom types, if, case, let) are always multi-line; other constructs stay on one line when they fit in width columns. Prints documentation and the doc comments of File.comments (module and port documentation), not regular comments. It prints normalize_file(file): the exposing lists and the imports in elm-format's order, with a module's exposing list grouped by the @docs lines of its documentation. Doc comments are written as elm-format writes them (Markdown and Elm code formatted, LF line ends; see @markdown.format_doc). As elm-format, this is not idempotent for some doc comments (see normalize_file). Text inside GLSL is written as it is, line ends included. Raises PrintError for an AST that cannot print as valid Elm; the error's path starts at the file.

    UnConsPattern(a, AsPattern(b, c)) prints as a :: b as c, which elm make 0.19.1 and elm-format read as (a :: b) as c (see print_pattern).

    test {
    let text = "module Main exposing (main)\n\n\nmain =\n 1 + 2\n"
    let result = @parser.parse_module(
    @scanner.SourceText::new(text),
    @scanner.DefaultScanner::new(),
    )
    inspect(@printer.print_file(result.ast.unwrap()) == text, content="true")
    }

    The Elm source of a pattern. Raises PrintError for a pattern that Elm cannot write (a negative number, a bad name, a tuple with one item); the error's path starts at p.

    UnConsPattern(a, AsPattern(b, c)) prints as a :: b as c, because krueger and elm-syntax read that text back as the same AST. elm make 0.19.1 and elm-format read a :: b as c as (a :: b) as c. To bind only the tail, use ParenthesizedPattern: a :: (b as c).

    test {
    let r : @ast.Range = {
    start: { row: 0, column: 0, },
    end: { row: 0, column: 0, },
    }
    fn pvar(name : String) -> @ast.Node[@ast.Pattern] {
    { range: r, value: VarPattern(name), }
    }
    let inner : @ast.Node[@ast.Pattern] = {
    range: r,
    value: UnConsPattern(pvar("a"), pvar("b")),
    }
    let outer : @ast.Node[@ast.Pattern] = {
    range: r,
    value: UnConsPattern(inner, pvar("rest")),
    }
    // `::` is right-associative, so an `::` on its left gets parentheses.
    inspect(@printer.print_pattern(outer), content="(a :: b) :: rest")
    // elm-format writes a constructor with arguments before `::` in
    // parentheses.
    let just : @ast.Node[@ast.Pattern] = {
    range: r,
    value: NamedPattern({ module_name: [][:], name: "Just", }, [pvar("a")][:]),
    }
    inspect(
    @printer.print_pattern({
    range: r,
    value: UnConsPattern(just, pvar("rest")),
    }),
    content="(Just a) :: rest",
    )
    }

    The Elm source of a type annotation, laid out in width columns where elm-format allows a choice. Parentheses are added where precedence needs them. Raises PrintError when the type cannot print as valid Elm; the error's path starts at t.

    test {
    let r : @ast.Range = {
    start: { row: 0, column: 0, },
    end: { row: 0, column: 0, },
    }
    let a : @ast.Node[@ast.TypeAnnotation] = {
    range: r,
    value: GenericType("a"),
    }
    let maybe : @ast.Node[@ast.TypeAnnotation] = {
    range: r,
    value: Typed({ range: r, value: ([][:], "Maybe"), }, [a][:]),
    }
    let list : @ast.Node[@ast.TypeAnnotation] = {
    range: r,
    value: Typed({ range: r, value: ([][:], "List"), }, [maybe][:]),
    }
    inspect(@printer.print_type_annotation(list), content="List (Maybe a)")
    }

    push_events

    fn[S :
    EventSource
    , H :
    Handler
    ] push_events(source : S, handler : H) -> Unit

    Pull every event from source and push it to handler: on_enter for Enter (its SkipChildren skips the node's children, its Stop ends the run) and on_leave for Leave. The source is usually an EventReader; a streaming parser can implement EventSource too, and handlers do not change.

    A handler needs its own type, so this example is not a doc test (a doc test holds only test blocks). The cookbook article "Choose a traversal" has a tested handler: https://github.com/moonrockz/krueger/blob/main/docs/cookbook/traversal.mbt.md

    struct Paths {
    found : Array[String]
    }

    impl @syntax.Handler for Paths with fn on_enter(self, e) {
    if e.category == "pattern" {
    self.found.push(e.path.to_string())
    }
    Continue
    }

    test {
    let src = "module Main exposing (..)\n\nadd x y =\n x + y\n"
    let result = @parser.parse_module(
    @scanner.SourceText::new(src),
    @scanner.DefaultScanner::new(),
    )
    let root = @syntax.NodeRef::of_result(result).unwrap()
    let paths = { found: [], }
    @syntax.push_events(@syntax.EventReader::new(root), paths)
    inspect(paths.found.length(), content="2")
    inspect(paths.found[1], content="declarations[0].declaration[0].arguments[1]")
    }

    render_elm_json

    Errors in diags in the shape of elm make --report=json output. Warnings are left out: elm make has none.

    The result is {"type": "compile-errors", "errors": [...]}, with one entry for the file when it has errors and none otherwise. The entry's name is source.module_name (default Main). Each problem has the title, the region of the problem and the message as elm make chunks.

    test {
    let source = @scanner.SourceText::new(
    "module Main exposing (..)\n\nmain =\n (1 + 2\n",
    module_name="Main",
    )
    let result = @parser.parse_module(source, @scanner.DefaultScanner::new())
    let json = @report.render_elm_json(result.diagnostics, source, "src/Main.elm")
    guard json is { "errors": [{ "problems": [problem], .. }], .. } else {
    fail("unexpected shape")
    }
    guard problem is { "title": String(title), "region": region, .. } else {
    fail("unexpected problem")
    }
    inspect(title, content="UNFINISHED PARENTHESES")
    inspect(
    region.stringify(),
    content=(
    #|{"start":{"line":4,"column":11},"end":{"line":4,"column":11}}
    ),
    )
    }

    render_plain

    A diagnostic as plain text, laid out like elm make output.

    The text starts with a -- TITLE ---- path header. The report follows, reflowed to 80 columns, with numbered source lines from source and ^ markers under the problem. path is only shown in the header.

    test {
    let source = @scanner.SourceText::new(
    "module Main exposing (..)\n\nmain =\n (1 + 2\n",
    )
    let result = @parser.parse_module(source, @scanner.DefaultScanner::new())
    let text = @report.render_plain(result.diagnostics[0], source, "src/Main.elm")
    inspect(
    text,
    content=(
    #|-- UNFINISHED PARENTHESES ----------------------------------------- src/Main.elm
    #|
    #|I was partway through parsing an expression in parentheses, but I got stuck
    #|here:
    #|
    #|4| (1 + 2
    #| ^
    #|I was expecting to see `,` or `)` next.
    #|
    ),
    )
    }

    render_terminal

    fn render_terminal(diag :
    Diagnostic
    , source :
    SourceText
    , path : String) -> String

    A diagnostic for a terminal, with ANSI colors like elm make.

    The text is render_plain's text with ANSI escape codes around the header, the ^ markers and the styled chunks of the report.

    test {
    let source = @scanner.SourceText::new(
    "module Main exposing (..)\n\nmain =\n (1 + 2\n",
    )
    let result = @parser.parse_module(source, @scanner.DefaultScanner::new())
    let text = @report.render_terminal(
    result.diagnostics[0],
    source,
    "src/Main.elm",
    )
    // The header is cyan: ESC [36m.
    inspect(
    text.has_prefix("\u{1b}[36m-- UNFINISHED PARENTHESES"),
    content="true",
    )
    }

    standard_operators

    The infix operators of Elm 0.19.1 as elm-syntax 7.3.9 knows them, including </> and <?> from elm/url.

    tokenize

    Tokenize source in dialect (default: Elm 0.19.1).

    The stream is lossless: each token's trivia_before, lexeme and trivia_after rebuild the source. Whitespace and comments are trivia, not tokens. A scan error (for example an unterminated string) gives Err with the diagnostics. Use the same dialect for parse_tokens.

    test {
    let source = @krueger.SourceText::new("x = 1 -- one\n")
    guard @krueger.tokenize(source) is Ok(stream) else { fail("scan error") }
    debug_inspect(
    stream.tokens.map(t => t.lexeme),
    content=(
    #|["x", "=", "1"]
    ),
    )
    // The comment is trivia after the last token.
    inspect(stream.tokens[2].trivia_after.length() > 0, content="true")
    let bad = @krueger.SourceText::new("s = \"abc\n")
    guard @krueger.tokenize(bad) is Err(errors) else { fail("no error") }
    inspect(errors.diagnostics[0].code, content="KR-SCAN-004")
    }

    version

    fn version() -> String

    The version of krueger, the same as version in moon.mod.

    test {
    inspect(@krueger.version().split(".").count(), content="3")
    }

    walk

    Visit root and every node below it in pre-order (source order). enter runs when a node is reached and decides what happens next; leave runs after all of a node's children, with the same node object. An explicit stack, not recursion, holds the open nodes, so trees of any depth work on every target. All traversal state belongs to this call.

    test {
    let src = "module Main exposing (..)\n\nf x = x\n"
    let result = @parser.parse_module(
    @scanner.SourceText::new(src),
    @scanner.DefaultScanner::new(),
    )
    let root = @syntax.NodeRef::of_result(result).unwrap()
    // Log enter and leave; skip the module header.
    let log = []
    @syntax.walk(
    root,
    n => {
    log.push("+" + n.kind())
    if n.category() == "module" {
    SkipChildren
    } else {
    Continue
    }
    },
    leave=n => log.push("-" + n.kind()),
    )
    inspect(
    log.join(" "),
    content="+file +normal -normal +function +implementation +name -name +var -var +functionOrValue -functionOrValue -implementation -function -file",
    )
    // Stop at the first pattern: the walk ends there, with no more `leave`.
    let seen = []
    @syntax.walk(
    root,
    n => {
    seen.push("+" + n.kind())
    if n.category() == "pattern" {
    Stop
    } else {
    Continue
    }
    },
    leave=n => seen.push("-" + n.kind()),
    )
    inspect(
    seen.join(" "),
    content="+file +normal +module_name -module_name +all -all -normal +function +implementation +name -name +var",
    )
    }