moonrockz/krueger/scanner does not have a README file

    Scanner

    pub trait Scanner {
    fn tokenize(Self, SourceText) -> Result[TokenStream, ScanErrorList]
    }

    A tokenizer for Elm source. @parser.parse_module takes any Scanner; DefaultScanner is the built-in one.

    Block

    pub(all) enum Block {
    Text(Array[Chunk])
    Excerpt(context~ : Span, highlight~ : Span)
    Hint(Array[Chunk])
    Note(Array[Chunk])
    Example(String)
    } derive(Eq,
    Debug
    )

    One block of a report. Excerpt shows the source lines of context with markers under highlight; renderers read the lines from the source.

    Chunk

    pub(all) enum Chunk {
    Plain(String)
    Styled(text~ : String, color~ : Color?, bold~ : Bool, underline~ : Bool)
    } derive(Eq,
    Debug
    )

    A piece of report text: plain, or styled like elm make styles code and keywords.

    Chunk::code

    fn Chunk::code(text : String) -> Chunk

    Elm code inside a paragraph, in yellow as elm make shows it.

    Chunk::keyword

    fn Chunk::keyword(text : String) -> Chunk

    An Elm keyword inside a paragraph, in bright cyan as elm make shows it.

    Color

    pub(all) enum Color {
    Red
    Yellow
    Green
    Cyan
    Magenta
    Blue
    Black
    White
    VividRed
    VividYellow
    VividGreen
    VividCyan
    VividMagenta
    VividBlue
    VividBlack
    VividWhite
    } derive(Eq,
    Debug
    )

    Colors of elm make reports. Vivid* are the bright variants; in the --report=json output they are upper case ("RED"), the others lower case.

    Color::json_name

    fn Color::json_name(self : Color) -> String

    The color name in elm make --report=json output.

    Comment

    pub(all) struct Comment {
    kind : CommentKind
    text : String
    span : Span
    } derive(Eq,
    Debug
    )

    A comment in the source. text is the comment exactly as written, delimiters included (-- note, {- note -}). A line comment does not include its line break.

    CommentKind

    pub(all) enum CommentKind {
    Line
    Block
    Doc
    } derive(Eq,
    Debug
    )

    The kind of a comment: Line is -- ..., Block is {- ... -} and Doc is {-| ... -}.

    DefaultScanner

    pub struct DefaultScanner {
    dialect :
    Dialect

    }

    The built-in Elm 0.19.1 scanner. Make one with DefaultScanner::new.

    DefaultScanner::new

    A scanner for dialect (reserved words and operator symbols). Without dialect it uses Dialect::elm_0_19_1(). Give the parser the same dialect.

    test {
    let scanner = @scanner.DefaultScanner::new()
    let source = @scanner.SourceText::new(
    "module Main exposing (main)\n\nmain = 1 + 2\n",
    )
    guard scanner.tokenize(source) is Ok(stream) else { fail("scan failed") }
    debug_inspect(
    stream.tokens.map(t => (t.kind, t.lexeme)),
    content=(
    #|[
    #| (Keyword(Module), "module"),
    #| (Identifier, "Main"),
    #| (Keyword(Exposing), "exposing"),
    #| (LParen, "("),
    #| (Identifier, "main"),
    #| (RParen, ")"),
    #| (Identifier, "main"),
    #| (Equals, "="),
    #| (IntLiteral, "1"),
    #| (Operator("+"), "+"),
    #| (IntLiteral, "2"),
    #|]
    ),
    )
    }

    DefaultScanner::tokenize

    fn DefaultScanner::tokenize(self : DefaultScanner, _source : SourceText) -> Result[TokenStream, ScanErrorList]

    Split source into tokens and attach the trivia to them (see Token).

    Returns Err with every scan problem (KR-SCAN-001 to KR-SCAN-004) when there is one; then there is no token stream. The scan is lossless: the tokens and their trivia give source.text back.

    Diagnostic

    pub(all) struct Diagnostic {
    code : String
    severity : Severity
    message : String
    span : Span
    title : String
    report : Array[Block]
    } derive(Eq,
    Debug
    )

    A problem found in the source.

    • code names the problem: KR-SCAN-* from the scanner, KR-PARSE-* and KR-ATTR-* from the parser.
    • severity: an Error makes the source invalid; a Warning or an Info does not.
    • span is where the problem is.

    The scanner reports its problems as a ScanErrorList:

    test {
    let source = @scanner.SourceText::new("x = \"abc\n")
    guard @scanner.DefaultScanner::new().tokenize(source) is Err(errors) else {
    fail("expected a scan error")
    }
    let d = errors.diagnostics[0]
    inspect(d.code, content="KR-SCAN-004")
    debug_inspect(d.severity, content="Error")
    inspect(d.title, content="ENDLESS STRING")
    inspect(d.message, content="Unterminated string literal")
    debug_inspect((d.span.start.column, d.span.end.column), content="(5, 9)")
    }

    KeywordKind

    pub(all) enum KeywordKind {
    Module
    Exposing
    Import
    As
    Type
    If
    Then
    Else
    Let
    In
    Case
    Of
    Port
    Where
    Custom(String)
    } derive(Eq,
    Debug
    )

    Reserved words of Elm 0.19.1 as elm-syntax 7.3.9 treats them. alias, infix and effect are not reserved; they are identifiers.

    Position

    pub(all) struct Position {
    offset : Int
    line : Int
    column : Int
    } derive(Eq,
    Debug
    )

    A place in the source, between two characters.

    • offset counts UTF-16 code units from the start of the text, from 0.
    • line counts lines from 1. LF and CRLF end a line.
    • column counts code points from the start of the line, from 1. A surrogate pair is one column, a lone surrogate one, a lone \r one.

    So outside ASCII the offset and the column differ:

    test {
    let source = @scanner.SourceText::new("x = \"😀😀\" ++ y\n")
    guard @scanner.DefaultScanner::new().tokenize(source) is Ok(stream) else {
    fail("scan failed")
    }
    // The string literal: 6 UTF-16 units (two surrogate pairs and two
    // quotes), but 4 columns.
    let string = stream.tokens[2].span
    debug_inspect((string.start.offset, string.end.offset), content="(4, 10)")
    debug_inspect((string.start.column, string.end.column), content="(5, 9)")
    // The `++` after it.
    let plus = stream.tokens[3].span.start
    debug_inspect((plus.offset, plus.line, plus.column), content="(11, 1, 10)")
    }

    RawToken

    pub(all) struct RawToken {
    kind : RawTokenKind
    lexeme : String
    span : Span
    } derive(Eq,
    Debug
    )

    A lexical unit before trivia is attached: a significant token or one piece of trivia, with its exact text and its span. lex_raw makes them; normalize_tokens turns them into a TokenStream.

    RawTokenKind

    pub(all) enum RawTokenKind {
    Significant(TokenKind)
    Whitespace
    Newline
    LineComment
    BlockComment
    DocComment
    } derive(Eq,
    Debug
    )

    The kind of a RawToken: a significant token, or one kind of trivia.

    ScanErrorList

    pub(all) struct ScanErrorList {
    diagnostics : Array[Diagnostic]
    } derive(Eq,
    Debug
    )

    All the problems that a scan found, in source order. A scan that finds a problem gives no tokens.

    Severity

    pub(all) enum Severity {
    Error
    Warning
    Info
    } derive(Eq,
    Debug
    )

    How serious a Diagnostic is. Only an Error makes a file fail to parse; a Warning (for example KR-PARSE-009) keeps the result.

    SourceText

    pub(all) struct SourceText {
    module_name : String?
    text : String
    } derive(Eq,
    Debug
    )

    The text of one Elm source file. module_name is the optional Elm module name (for example Main). render_elm_json writes it as name; the parser does not use it.

    SourceText::new

    fn SourceText::new(text : String, module_name? : String) -> SourceText

    Make a SourceText from text, with an optional Elm module name (for example Main). render_elm_json writes module_name as name; the parser does not use it.

    test {
    let source = @scanner.SourceText::new("module Main exposing (..)\n")
    debug_inspect(source.module_name, content="None")
    let named = @scanner.SourceText::new("x = 1\n", module_name="Main")
    debug_inspect(named.module_name, content="Some(\"Main\")")
    }

    Span

    pub(all) struct Span {
    start : Position
    end : Position
    } derive(Eq,
    Debug
    )

    The part of the source from start to end. end is the position just after the last character, so an empty span has start == end. See Position for how offsets, lines and columns count.

    Token

    pub(all) struct Token {
    kind : TokenKind
    lexeme : String
    span : Span
    trivia_before : Array[Trivia]
    trivia_after : Array[Trivia]
    } derive(Eq,
    Debug
    )

    A significant token with the trivia around it.

    • lexeme is the exact source text of the token, and span is where it is.
    • trivia_before is all the trivia between the previous token and this one.
    • trivia_after is empty, except on the last token: there it holds the trivia up to the end of the text.

    The scan is lossless. Join trivia_before, lexeme and trivia_after of every token, in order, and you get the source back. A text with no tokens keeps its trivia in TokenStream::trivia.

    test {
    fn trivia_text(t : @scanner.Trivia) -> String {
    match t {
    Whitespace(text, _) | Newline(text, _) => text
    Comment(comment) => comment.text
    }
    }

    let text = "module Main exposing (..)\n\n-- The answer.\nanswer = 42 {- end -}\n"
    guard @scanner.DefaultScanner::new().tokenize(@scanner.SourceText::new(text))
    is Ok(stream) else {
    fail("scan failed")
    }
    let out = StringBuilder()
    for t in stream.trivia {
    out.write_string(trivia_text(t))
    }
    for token in stream.tokens {
    for t in token.trivia_before {
    out.write_string(trivia_text(t))
    }
    out.write_string(token.lexeme)
    for t in token.trivia_after {
    out.write_string(trivia_text(t))
    }
    }
    inspect(out.to_string() == text, content="true")
    // The line comment is trivia before `answer`.
    let answer = stream.tokens[6]
    inspect(answer.lexeme, content="answer")
    debug_inspect(
    answer.trivia_before.map(trivia_text),
    content=(
    #|["\n", "\n", "-- The answer.", "\n"]
    ),
    )
    }

    TokenKind

    pub(all) enum TokenKind {
    Keyword(KeywordKind)
    Identifier
    IntLiteral
    FloatLiteral
    StringLiteral
    CharLiteral
    Glsl
    LParen
    RParen
    LBracket
    RBracket
    LBrace
    RBrace
    Comma
    Equals
    Dot
    DotDot
    Colon
    Pipe
    Arrow
    Backslash
    Underscore
    Operator(String)
    } derive(Eq,
    Debug
    )

    The kind of a significant token. Keywords come from the dialect, so a dialect with extra reserved words gives Keyword(Custom(...)).

    TokenKind::as_kw

    fn TokenKind::as_kw() -> TokenKind

    The token kind of the as keyword.

    TokenKind::case_kw

    fn TokenKind::case_kw() -> TokenKind

    The token kind of the case keyword.

    TokenKind::else_kw

    fn TokenKind::else_kw() -> TokenKind

    The token kind of the else keyword.

    TokenKind::exposing_kw

    fn TokenKind::exposing_kw() -> TokenKind

    The token kind of the exposing keyword.

    TokenKind::if_kw

    fn TokenKind::if_kw() -> TokenKind

    The token kind of the if keyword.

    TokenKind::import_kw

    fn TokenKind::import_kw() -> TokenKind

    The token kind of the import keyword.

    TokenKind::in_kw

    fn TokenKind::in_kw() -> TokenKind

    The token kind of the in keyword.

    TokenKind::keyword

    fn TokenKind::keyword(kind : KeywordKind) -> TokenKind

    The token kind of the keyword kind.

    TokenKind::let_kw

    fn TokenKind::let_kw() -> TokenKind

    The token kind of the let keyword.

    TokenKind::module_kw

    fn TokenKind::module_kw() -> TokenKind

    The token kind of the module keyword.

    TokenKind::of_kw

    fn TokenKind::of_kw() -> TokenKind

    The token kind of the of keyword.

    TokenKind::port_kw

    fn TokenKind::port_kw() -> TokenKind

    The token kind of the port keyword.

    TokenKind::then_kw

    fn TokenKind::then_kw() -> TokenKind

    The token kind of the then keyword.

    TokenKind::type_kw

    fn TokenKind::type_kw() -> TokenKind

    The token kind of the type keyword.

    TokenKind::where_kw

    fn TokenKind::where_kw() -> TokenKind

    The token kind of the where keyword.

    TokenStream

    pub(all) struct TokenStream {
    tokens : Array[Token]
    trivia : Array[Trivia]
    } derive(Eq,
    Debug
    )

    The result of a scan: the significant tokens in source order, each with its trivia.

    Trivia

    pub(all) enum Trivia {
    Whitespace(String, Span)
    Newline(String, Span)
    Comment(Comment)
    } derive(Eq,
    Debug
    )

    Source text that is not a token: a run of spaces (or a lone \r), one line break (\n or \r\n) or a comment. Each case keeps its exact text, so trivia and tokens together give the source back (see Token).

    invalid_sequence

    fn invalid_sequence(span : Span) -> Diagnostic

    A character Elm does not allow here.

    is_lower_start

    fn is_lower_start(c : Char) -> Bool

    Whether c can start a lower-case name (a Unicode lower-case letter). Lower-case names are values, functions and type variables.

    test {
    inspect(@scanner.is_lower_start('a'), content="true")
    inspect(@scanner.is_lower_start('é'), content="true")
    inspect(@scanner.is_lower_start('A'), content="false")
    inspect(@scanner.is_lower_start('_'), content="false")
    }

    is_name_part

    fn is_name_part(c : Char) -> Bool

    Whether c can continue a name: a Unicode letter or number, or _.

    test {
    inspect(@scanner.is_name_part('x'), content="true")
    inspect(@scanner.is_name_part('_'), content="true")
    inspect(@scanner.is_name_part('7'), content="true")
    inspect(@scanner.is_name_part('ß'), content="true")
    inspect(@scanner.is_name_part('-'), content="false")
    inspect(@scanner.is_name_part('\''), content="false")
    }

    is_upper_start

    fn is_upper_start(c : Char) -> Bool

    Whether c can start an upper-case name (a Unicode upper-case or title-case letter). Upper-case names are modules, types and constructors. In the elm-syntax-7.3.9 dialect the scanner does not start a name with a title-case letter (rule titlecase-name-start).

    test {
    inspect(@scanner.is_upper_start('M'), content="true")
    inspect(@scanner.is_upper_start('Ä'), content="true")
    inspect(@scanner.is_upper_start('Ç…'), content="true")
    inspect(@scanner.is_upper_start('m'), content="false")
    }

    lex_raw

    Split Elm source into significant tokens and trivia. Diagnostics: KR-SCAN-001 unterminated block comment, KR-SCAN-002 malformed doc comment, KR-SCAN-003 a character Elm does not allow here, KR-SCAN-004 an unterminated string, char or GLSL literal.

    The result is a flat list: trivia is not attached to tokens yet, and the lexemes in order give the source back. DefaultScanner::tokenize is lex_raw followed by normalize_tokens.

    test {
    let source = @scanner.SourceText::new("x = 1 -- one\n")
    guard @scanner.lex_raw(source) is Ok(raw) else { fail("scan failed") }
    debug_inspect(
    raw.map(r => (r.kind, r.lexeme)),
    content=(
    #|[
    #| (Significant(Identifier), "x"),
    #| (Whitespace, " "),
    #| (Significant(Equals), "="),
    #| (Whitespace, " "),
    #| (Significant(IntLiteral), "1"),
    #| (Whitespace, " "),
    #| (LineComment, "-- one"),
    #| (Newline, "\n"),
    #|]
    ),
    )
    }

    malformed_doc_comment

    fn malformed_doc_comment(span : Span) -> Diagnostic

    A doc comment with no end. span runs from {-| to the end of the input.

    normalize_tokens

    fn normalize_tokens(raw_tokens : Array[RawToken]) -> TokenStream

    Attach trivia to the significant tokens of raw_tokens (the output of lex_raw). Trivia goes to the next token's trivia_before; trivia after the last token goes to its trivia_after. When there is no token, the trivia goes to TokenStream::trivia. Nothing is dropped.

    raw_to_trivia

    fn raw_to_trivia(raw : RawToken) -> Trivia?

    The trivia for raw, or None when raw is a significant token.

    tab_character

    fn tab_character(span : Span) -> Diagnostic

    A tab character.

    unexpected_character

    fn unexpected_character(span : Span) -> Diagnostic

    A character that is not a symbol and is not used in Elm, such as a backtick or a non-ASCII letter outside a string.

    unexpected_semicolon

    fn unexpected_semicolon(span : Span) -> Diagnostic

    A semicolon.

    unterminated_block_comment

    fn unterminated_block_comment(span : Span) -> Diagnostic

    A block comment with no end. span runs from {- to the end of the input.

    unterminated_char

    fn unterminated_char(span : Span) -> Diagnostic

    A character literal with no closing quote.

    unterminated_glsl

    fn unterminated_glsl(span : Span) -> Diagnostic

    A GLSL block with no closing |].

    unterminated_multiline_string

    fn unterminated_multiline_string(span : Span) -> Diagnostic

    A multi-line string with no closing """.

    unterminated_string

    fn unterminated_string(span : Span) -> Diagnostic

    A string with no closing quote. span runs from the opening quote to where the scanner stopped (the end of the line, or of the input).

    Powered by MoonBit

    Site sourceReport issuePackagesBuild queueSkillsStatistics

    © 2026 mooncakes.io