moonocr

    A pure-MoonBit OCR engine: BMP/PGM/PPM decoding, binarization, segmentation, feature extraction and nearest-neighbour classification of digits and letters.

    ocr
    image
    computer-vision
    netpbm
    bmp
    Download zip
    Author
    Version
    0.1.0
    License
    Apache-2.0
    Last updated
    6 hours ago
    Downloads
    1

    #moonocr

    纯 MoonBit 实现的光学字符识别(OCR)引擎。它读取一张图像,找到其中的文本,返回识别出的字符——不依赖 MoonBit 核心库以外的任何第三方包。

    引擎识别 数字 0-9、字母 A-Z / a-z(共 62 类),支持单行与多行,对噪声和不均匀光照有一定鲁棒性。

    #项目目标

    • 用纯 MoonBit 从零实现一条完整的 OCR 管线,补上 MoonBit 生态中「图像 → 文本」的底座;
    • 零外部依赖,可被 moon add 直接组合,也可编译到 wasm 部署;
    • 以可测试、可验证的鲁棒性语义(缩放不变、噪声剔除、部件装配)为目标,而非单纯追求识别类别数量。

    #管线

    图像 ──▶ 解码 ──▶ 二值化 ──▶ 分割 ──▶ 特征 ──▶ 分类 ──▶ 文本 (BMP/PGM/PPM) (Otsu/固定/自适应) (连通域+行分组+部件) (8×8 网格) (最近邻)

    阶段模块说明
    图像模型image.mbtImage、BT.601 rgb_to_gray
    解码bmp.mbt / pgm.mbt / ppm.mbtparse_bmp、parse_pgm(P1/P2/P4/P5)、parse_ppm(P3/P6)
    二值化binarize.mbt固定阈值 / Otsu / 自适应 / Sauvola / Niblack
    矫正deskew.mbt投影直方图、图像旋转、倾斜检测与 deskew
    分割segment.mbt连通域、噪声过滤、行分组、部件装配
    特征feature.mbt保持纵横比的 8×8 网格
    分类classify.mbt汉明距离最近邻 / 加权 / 平移增强
    识别recognize.mbt单行 / 多行

    #编译与运行

    需要 MoonBit 工具链(moon)。

    moon check # 类型检查 moon test # 运行全部测试(79 项) moon bench # 运行基准

    命令行演示(每个参数是一行文本,渲染成图像后再识别回来):

    moon run cmd/main -- Hello World 42 ABCabc 2026

    #示例

    上面的命令会输出(真实终端输出,输入渲染成图像后走完整管线识别回来,大小写与数字均正确):

    input: Hello World 42 ABCabc 2026 output: Hello World 42 ABCabc 2026

    引擎内部先把文本渲染成 8×8 位图字形(# 为墨迹、. 为纸白),再走完整管线识别。例如字母 A 与数字 7 的字形:

    字母 A 数字 7 ...##... .######. ..####.. .....##. .##..##. ....##.. .##..##. ...##... .######. ...##... .##..##. ...##... .##..##. ...##... .##..##. ...##...

    #基准

    moon bench 对渲染出的测试输入逐阶段计时(本机测量,数值随机器而变,仅供数量级参考):

    阶段平均耗时
    classify(单个字形)~10 µs
    binarize_otsu~92 µs
    connected_components~149 µs
    binarize_adaptive~327 µs
    binarize_niblack~496 µs
    binarize_sauvola~529 µs
    recognize_digits(12 位数字)~408 µs
    recognize_text(3 行文本)~469 µs

    分类器缓存了 62 个内置模板(每次识别不再重算),单个字形分类约 10 µs。

    #使用教程

    #解码图像

    let img = parse_pgm(bytes) // 或 parse_bmp / parse_ppm
    let gray = img.at(x, y) // 0..=255,0 为墨迹,255 为纸白

    所有解码器返回灰度 Image,输入非法时抛出 DecodeError。

    #识别文本

    let text = recognize_text(img) // 多行,以 "\n" 连接
    let line = recognize_digits(img) // 单行

    两者内部用 Otsu 二值化、滤除噪声斑点、分割字形(并把 i/j 的点和竖重新装配),再与 62 个内置模板做最近邻分类。

    #底层构建块

    let bin = binarize_otsu(img)
    let adaptive = binarize_adaptive(img, window=15, c=20)
    let sauvola = binarize_sauvola(img, window=15, k=0.3, r=128.0)
    let niblack = binarize_niblack(img, window=15, k=-0.2)
    let straight = deskew(img, max_angle=10.0, step=1.0)

    let components = connected_components(bin)
    let big = filter_small(components, min_area=8)
    let lines = group_lines(big)
    let glyphs = merge_parts(lines[0])

    let grid = glyph_grid(bin, glyphs[0], size=8)
    let m = classify(grid, alphanumeric_references()) // Match?
    let w = classify_weighted(grid, alphanumeric_references(), center_weights())
    let s = classify_shifted(grid, alphanumeric_references(), size=8)

    #支持的图像格式

    • BMP — 8/24/32 位无压缩 BI_RGB,自底向上。
    • PGM — P1(ASCII 位图)、P2(ASCII 灰度)、P4(位图)、P5(二进制灰度),maxval ≤ 255。
    • PPM — P3(ASCII RGB)、P6(二进制 RGB),maxval ≤ 255;用 BT.601 亮度转灰度。

    另提供 write_pgm_p5 / write_ppm_p6 编码器用于往返测试。

    #工作原理

    • 二值化:Otsu 求全局阈值;自适应版用积分图求局部窗口均值;Sauvola/Niblack 用局部均值与标准差应对光照不均。
    • 矫正:deskew 对轻微旋转的文本求水平投影直方图,在候选角度里选方差最大者旋转回正。
    • 分割:8 连通泛洪给墨迹区域打标签;按垂直重叠合并成行;merge_parts 把被细缝拆开的字形(如 i/j 的点)重新拼回;filter_small 滤除噪声。
    • 特征:每个字形的包围盒按保持纵横比的方式下采样到 8×8 网格,每格记录是否有半数以上像素为墨。
    • 分类:网格间汉明距离最近邻,可选加权距离与平移增强(shift 变体),对轻微偏心字形更鲁棒。参考模板走同一条特征管线生成,因此干净渲染必然与自身距离为 0。

    #测试

    moon test # 79 项测试,全部通过

    #限制

    • 分类器基于内置 8×8 字体模板,未做训练;真实字体与重度变形会降低准确率。
    • 仅识别 0-9、A-Z、a-z(不含标点与重音符号)。
    • 核心库不做文件 I/O:解码器接收 Bytes,由调用方提供图像字节。
    • deskew 面向小角度倾斜(默认 ±10°),旋转为最近邻采样且不改变画布尺寸;大角度或透视变形不在范围之内。

    #开源协议

    Apache-2.0

    DecodeError

    pub suberror DecodeError {
    DecodeError(String)
    }

    unsupported feature.

    Component

    pub struct Component {
    x : Int
    y : Int
    w : Int
    h : Int
    area : Int
    }

    pixels.

    Component::bottom

    fn Component::bottom(self : Component) -> Int

    Exclusive bottom edge of the bounding box.

    Component::right

    fn Component::right(self : Component) -> Int

    Exclusive right edge of the bounding box.

    Image

    pub struct Image {
    width : Int
    height : Int
    pixels : Bytes
    }

    Image::at

    fn Image::at(self : Image, x : Int, y : Int) -> Int

    x is the column and y is the row, both zero-based from the top-left.

    Image::byte_length

    fn Image::byte_length(self : Image) -> Int

    Byte length of the underlying pixel buffer.

    Image::from_packed

    fn Image::from_packed(width : Int, height : Int, packed : Bytes, channels : Int) -> Image

    conventional RGB(A) order.

    Image::new

    fn Image::new(width : Int, height : Int, pixels : Bytes) -> Image

    Construct an image from raw grayscale bytes (row-major, one byte per pixel).

    Image::pixel_count

    fn Image::pixel_count(self : Image) -> Int

    Total number of pixels.

    Match

    pub struct Match {
    label : String
    distance : Int
    }

    smaller distance means a closer match).

    Reference

    pub struct Reference {
    label : String
    grid : Array[Int]
    }

    A labelled feature grid used as a template.

    Reference::new

    fn Reference::new(label : String, grid : Array[Int]) -> Reference

    alphanumeric_references

    fn alphanumeric_references() -> Array[Reference]

    The 62 templates for digits 0-9 and letters A-Z / a-z, from the built-in font.

    binarize_adaptive

    fn binarize_adaptive(img : Image, window : Int, c : Int) -> Image

    single global threshold would fail.

    binarize_fixed

    fn binarize_fixed(img : Image, threshold : Int) -> Image

    (0), the rest become paper (255).

    binarize_niblack

    fn binarize_niblack(img : Image, window : Int, k : Double) -> Image

    Niblack local binarization. Each pixel's threshold is the local mean plus a multiple of the local standard deviation:

    T = m + k * s

    k is usually negative (around -0.2), placing the threshold below the mean so that only pixels clearly darker than their neighbourhood become ink. Niblack recovers text under uneven illumination but, unlike Sauvola, can raise noise in low-variance regions.

    binarize_otsu

    fn binarize_otsu(img : Image) -> Image

    Binarize with the globally optimal Otsu threshold.

    binarize_sauvola

    fn binarize_sauvola(img : Image, window : Int, k : Double, r : Double) -> Image

    Sauvola local binarization. Each pixel's threshold adapts to both the local mean m and local standard deviation s of its window-sized neighbourhood:

    T = m * (1 + k * (s / r - 1))

    r is the dynamic range of s (usually 128 for 8-bit gray) and k controls sensitivity (typically 0.2-0.5). Pixels darker than T become ink. Unlike the fixed-offset mean threshold, Sauvola suppresses noise in flat regions while keeping text in high-contrast regions.

    cell_distance

    fn cell_distance(a : Array[Int], b : Array[Int]) -> Int

    Number of positions at which the two feature vectors differ.

    center_weights

    fn center_weights() -> Array[Int]

    A 64-cell weight grid that emphasises the glyph interior (weight 2) over the outer ring (weight 1). The interior holds the stroke structure that stays stable across fonts, while the border is more sensitive to noise.

    char_grid

    fn char_grid(c : Char) -> Array[Int]

    zero even when the raw font bitmap is not horizontally symmetric.

    char_rows

    fn char_rows(c : Char) -> Array[Int]

    lowercase letters A-Z / a-z. Any other character yields all paper.

    classify

    fn classify(features : Array[Int], refs : Array[Reference]) -> Match?

    reference.

    classify_shifted

    fn classify_shifted(features : Array[Int], refs : Array[Reference], size : Int) -> Match?

    Classify against an augmented reference set in which every template is expanded into its shifted variants. A glyph then matches its class as long as it is close to any one shift of that class's template.

    classify_weighted

    fn classify_weighted(features : Array[Int], refs : Array[Reference], weights : Array[Int]) -> Match?

    Classify by weighted Hamming distance, giving each cell the importance in weights. Otherwise identical to classify.

    connected_components

    fn connected_components(img : Image) -> Array[Component]

    Components are returned in scan order, top-to-bottom then left-to-right.

    deskew

    fn deskew(img : Image, max_angle : Double, step : Double) -> Image

    Deskew an image in one step: detect the corrective rotation and apply it, returning a copy whose text lines are as horizontal as possible.

    detect_skew

    fn detect_skew(img : Image, max_angle : Double, step : Double) -> Double

    Estimate the corrective rotation for slightly tilted text. It tries angles from -max_angle to max_angle in steps of step degrees and returns the one that makes the horizontal projection profile sharpest. Rotating the input by the returned angle straightens its lines. A blank or uniformly inked image has no preferred angle, so 0.0 is returned.

    digit_references

    fn digit_references() -> Array[Reference]

    The ten digit templates from the built-in font, labelled "0" through "9".

    filter_small

    fn filter_small(components : Array[Component], min_area : Int) -> Array[Component]

    stray specks of noise before further analysis.

    font_rows

    fn font_rows(digit : Int) -> Array[Int]

    outside 0..=9.

    font_rows_lower

    fn font_rows_lower(c : Char) -> Array[Int]

    so those two are two-component glyphs.

    font_rows_upper

    fn font_rows_upper(c : Char) -> Array[Int]

    c is outside A-Z. Same bit order and ink convention as font_rows.

    glyph_grid

    fn glyph_grid(img : Image, c : Component, size : Int) -> Array[Int]

    stretched out of shape.

    group_lines

    fn group_lines(components : Array[Component]) -> Array[Array[Component]]

    reading order, with each line's components still unsorted horizontally.

    histogram

    fn histogram(img : Image) -> Array[Int]

    Count how many pixels hold each of the 256 gray levels.

    merge_parts

    fn merge_parts(components : Array[Component]) -> Array[Component]

    which makes the rule independent of the rendering scale.

    otsu_threshold

    fn otsu_threshold(img : Image) -> Int

    groups. Pixels at or below the returned threshold are the ink.

    parse_bmp

    fn parse_bmp(data : Bytes) -> Image raise DecodeError

    Raises DecodeError when the input is not a supported uncompressed BMP.

    parse_pgm

    fn parse_pgm(data : Bytes) -> Image raise DecodeError

    Parse a Netpbm image (P1, P2, P4 or P5) into a grayscale Image.

    parse_ppm

    fn parse_ppm(data : Bytes) -> Image raise DecodeError

    uses a feature outside the supported subset.

    projection_profile

    fn projection_profile(img : Image) -> Array[Int]

    The number of ink pixels (value < 128) in each row. A horizontal line of text concentrates ink into a small band of rows, so this profile is the basis of skew detection.

    recognize_digits

    fn recognize_digits(img : Image) -> String

    that match nothing are skipped.

    recognize_text

    fn recognize_text(img : Image) -> String

    and joined with newlines.

    render_rows

    fn render_rows(rows : Array[Int], scale : Int) -> Image

    used to turn font data into images for recognition tests and demos.

    render_text

    fn render_text(text : String, scale : Int, gap : Int) -> Image

    as paper.

    render_text_block

    fn render_text_block(lines : Array[String], scale : Int, gap_x : Int, gap_y : Int) -> Image

    right.

    rgb_to_gray

    fn rgb_to_gray(r : Int, g : Int, b : Int) -> Int

    coefficients (0.299 R + 0.587 G + 0.114 B), computed in integer arithmetic.

    rotate

    fn rotate(img : Image, angle : Double, fill : Int) -> Image

    Rotate an image by angle degrees (positive counter-clockwise) about its centre, using nearest-neighbour sampling. Source pixels that fall outside the image are replaced by fill (use 255 for paper), so the result keeps the original dimensions.

    rows_to_grid

    fn rows_to_grid(rows : Array[Int]) -> Array[Int]

    representation consumed by the classifier.

    shift_variants

    fn shift_variants(grid : Array[Int], size : Int) -> Array[Array[Int]]

    The five shifted variants of a size-by-size grid: the original plus one-cell shifts up, down, left and right. Augmenting references with these makes classification tolerant of glyphs that sit slightly off-centre.

    weighted_distance

    fn weighted_distance(a : Array[Int], b : Array[Int], weights : Array[Int]) -> Int

    Weighted Hamming distance: every cell where the two grids differ adds its weight instead of 1, so a disagreement in an important cell counts more. Cells past the end of weights fall back to weight 1.

    write_pgm_p5

    fn write_pgm_p5(img : Image) -> Bytes

    trips through parse_pgm.

    write_ppm_p6

    fn write_ppm_p6(img : Image) -> Bytes

    through parse_ppm byte-for-byte in grayscale.