yimai_prophecy_moonbit

译脉·先知 2.0 预知记忆网络引擎的 MoonBit 零依赖实现:带预测能力的记忆网络,支持可解释预测与语义召回。

memory
prediction
translation
moonbit
hebbian
explainable
Download zip
Version
0.1.1
License
MIT
Last updated
20 hours ago
Downloads
3

Dependencies

#译脉·先知 2.0 (MoonBit)

一套带预测能力的记忆网络——让本地智能体记住工作流,并在遇到同类任务时预测下一步需求、给出可白盒解释的路径。这是「译脉·先知 2.0 预知记忆网络」引擎的 MoonBit 零依赖实现(仅 core/json + core/math)。

在「预测记忆」内核之外,已落地 #22 翻译记忆(TM)/ 术语库(TB)一等公民:真正的 fuzzy match(含匹配率%)、concordance 检索、TBX 术语库强制对齐与一致性校验——让引擎从「只预测」走向「预测 + 检索 + 术语守门」。

📜 变更历史CHANGELOG.md 记录每次 P 增量的完整 commit 列表与影响范围;本 README 只描述当前状态。

Tests Hit@3 Modern Corpus Service API MoonBit License


#Table of Contents


#Features

  • Predictive memory network — remembers workflows and predicts the most likely next step before the user asks.
  • White-box explainability — every prediction/explain returns the concrete node/edge/transition path, never a black box.
  • Semantic recall — given a new source sentence, activates a spreading network to recall the exact bilingual terms/sentences that keep terminology consistent.
  • Cold-start generalization — role-level abstraction (D8) induces domain-independent process skeletons, so an unseen project still gets a sensible next step.
  • Deterministic & reproducible — a logical clock replaces wall-clock time; identical call sequences yield byte-identical to_json output (no RNG).
  • Zero third-party dependencies — pure MoonBit core (json + math only); nothing to install beyond the moon toolchain.
  • Serializable — full engine state exports/imports as JSON for persistence and cross-session restore.
  • Fast inference (快) — role inverted-index (role_members) + per-source Top-8 pruning keep the hot path off full-graph scans; a pred_cache (fully invalidated on any engine change, not LRU) short-circuits repeated (context, k) queries.
  • Accurate (准) — second-order Markov (trans2, P(w3|w1,w2) blended at λ=0.4), multi-granularity role keys (前二/前四/前后各二), elastic forgetting (recency-aware edge decay), adaptive Hebbian LR, per-domain bias ΔW (LoRA-style), online contrastive learning (cl_step), and attention-gated edge weights in recall.
  • Explainable & bilingual (美)TermNode (mark_term) boosts terminology recall with a +5.0 activation and a "term hit" flag; explain_card returns a white-box activation_path / prediction_path / value_breakdown JSON; align_diff gives a character-level LCS edit script for bilingual alignment.
  • TM / TermBase (检索 + 守门)#22 新增 add_tm / fuzzy_match / concordance / load_tbx / enforce_terms / check_terms:真正的 fuzzy match(匹配率 %)、concordance 检索、TBX(ISO 30042) 术语库解析、术语强制对齐与一致性校验(见 §#22)。
  • Incremental & collaborative (中/长周期) — Write-Ahead Log (wal_*) for event-sourced replay, active-learning candidates by uncertainty + diversity, and federated increment export/import (fed_*) for cross-agent coordination.
  • Pure-MoonBit service layer (纯 MoonBit 全栈)cmd/service 起本地 HTTP server(127.0.0.1:8787),27 个 /api/* 端点 + /mcp MCP Server(13 基础 + retrieve_prompt / bleu / chrf / style_check / style_report / back_align / term_conflicts / fed_export / fed_import / distill_inject / active_learning / metrics / health / mqm_re_annotate;MCP 层共 25 tools)全部实测通过,含 BLEU/chrF++ 评测、风格检查、回译对齐、术语冲突、TMPlm、联邦/蒸馏注入、主动学习推荐;TM 状态经 fs 原子写持久化(tmp+rename),重启恢复闭环——引擎到服务零桥接语言


#How it works (D1–D8)

The engine is a neuro-inspired memory network. Eight modules map directly to constants in engine.mbt:

ModuleAlgorithmWhat it does
D1 Synaptic graphHebbian w ← w + LR·(1−w)Co-occurrence creates edges; weights decay over time.
D2 Activation spreadMulti-hop a·w·decaySeed node → activate similar nodes → recall by spreading.
D3 Forward model1st-order + 2nd-order Markov src→{dst} / (w1,w2)→{w3}Predict next step by blending context-weighted 1st-order transitions with λ·P(w3\|w1,w2) (second-order).
D4 Value pricingV = α·U_past + β·U_pred + γ·C_graph + δ·R − ε·Costβ=0.45 dominates — predicted-hit value ranks highest.
D5 Episode sequenceepisode logRecords sequences for consolidation replay.
D6 Consolidationprune + constraint-contract snapshotMeta-cognitive explore control; contract roll-back via restore.
D7 Uncertaintydistribution entropyEmits confidence / uncertainty.
D8 Concept abstractionmulti-granularity role transitions (cold-start)前二 / 前四 / 前后各二 role keys induce cross-topic rules.

Why deterministic: a self.clock (incremented on every remember/observe) substitutes wall-clock time, so results are reproducible and dependency-free. Because REC_TAU ≫ training steps, the recency term is ≈ 1.


#Project Structure

yimai_prophecy_moonbit/ ├── engine.mbt # Layer0: 零依赖预测记忆引擎内核(D1–D8, #22 TM/TB, MQM) ├── util.mbt # 编码/TF-IDF/对齐/URL解码工具函数(P4 decode_pct 新增) ├── yimai_prophecy_moonbit.mbt # Lib 主入口(routes_meta 单一源 + lib 共享 helper 文档化) ├── tests/ # P5 仓库整理:18 个测试按主题分到 3 个 sub-package │ ├── core/ # 核心/经典测试(69 测试) │ │ ├── moon.pkg # sub-package(独立 wasm-gc 测试目标) │ │ ├── _test_helpers.mbt # 跨子包共享 helper(canon/topics + 6 fn,pub) │ │ ├── yimai_prophecy_moonbit_test.mbt # 主测试(L1–L2 批量 Hit@3) │ │ ├── yimai_prophecy_moonbit_accept_test.mbt # 验收测试台 │ │ ├── yimai_prophecy_moonbit_bench_p1.mbt # 基准 P1(pruned vs full acc) │ │ ├── yimai_prophecy_moonbit_benchmark_test.mbt │ │ ├── yimai_prophecy_moonbit_golden_test.mbt # 黄金集回归 │ │ ├── yimai_prophecy_moonbit_long_text_test.mbt # 长文本 / 分段 / 数字 token │ │ ├── yimai_prophecy_moonbit_tm_test.mbt # TM 专项 │ │ ├── yimai_prophecy_moonbit_v2_test.mbt # V2 引擎 API │ │ └── yimai_prophecy_moonbit_wbtest.mbt # 白盒内部测试 │ ├── corpus/ # 语料/数据集驱动测试(54 测试) │ │ ├── moon.pkg │ │ ├── _test_helpers.mbt # inline 副本(与 core/ 同步) │ │ ├── yimai_prophecy_moonbit_extended_corpus_test.mbt # 6 个领域 11 测试 │ │ ├── yimai_prophecy_moonbit_modern_corpus_test.mbt # Modern Corpus:8 个前沿领域 │ │ ├── yimai_prophecy_moonbit_roadmap_test.mbt # Roadmap 增量语料 │ │ ├── yimai_prophecy_moonbit_frontier_corpus_test.mbt # Frontier Corpus:10 个领域 14 测试 │ │ └── yimai_prophecy_moonbit_business_corpus_test.mbt # 商务领域(P5 新增,ISO 11669 / GB/T 30539) │ └── feature/ # 扩展/新功能测试(52 测试) │ ├── moon.pkg │ ├── yimai_prophecy_moonbit_extension_test.mbt # 扩展能力回归(E1–E18) │ ├── yimai_prophecy_moonbit_quality_test.mbt # P4 质量/安全(T30–T37) │ ├── yimai_prophecy_moonbit_routes_test.mbt # 端点元数据单一源测试 │ └── yimai_prophecy_moonbit_mqm_reannotation_test.mbt # P5 MQM 二次标注(Google 2025-10-28) ├── cmd/ │ ├── main/moon.pkg # demo 程序(训练 → Hit@3 → replay → D8 冷启动 → consolidate → reward → restore) │ └── service/ │ ├── moon.pkg # 服务入口(`moon run cmd/service --target native` → 127.0.0.1:8787) │ ├── mcp.mbt # MCP Server 实现(spec 2025-11-25 Streamable HTTP) │ ├── routes.mbt # 27 个 HTTP 端点路由(24 + metrics + health + mqm_re_annotate) │ ├── tm_store.mbt # 引擎持久化(`save_store` 深度守卫 P4) │ └── web/ # 前端工作台(静态资源,`serve_static` URL 解码 P4 修复) ├── scripts/ │ ├── dev.ps1 # 一键起服务(env check → build → run → seed → smoke) │ ├── push.ps1 # 双 remote 推送(github via ghproxy.net + gitlink) │ ├── smoke.ps1 # 烟雾测试(27 端点 + MCP) │ └── reorganize_repo.py # 仓库整理复现脚本(git mv + add @lib. prefix + sub-package init) ├── .githooks/ │ └── pre-commit # 本地门禁(`moon check` + `moon test --target wasm-gc`) ├── docs/ │ ├── skill/SKILL.md # WorkBuddy 技能编排手册(frontmatter agent_created=true) │ ├── harness-configs/ # 13 个 harness 配置(Claude Code / Cursor / Gemini CLI / ...) │ ├── plans/ # 项目级 plan / note(按日期 YYYY-MM-DD-<topic>.md 命名) │ └── roadmap.md # 项目路线图(中文转英文,仓库国际友好) ├── AGENTS.md # AI agent 集成指南(含 Project layout 段:未来 _test.mbt 必须在子目录) ├── README.md # 项目说明(badge 175/175 + P6 hardening + International Standards) ├── CHANGELOG.md # 版本变更记录 └── LICENSE # MIT License

三层架构
  • Layer0(内核)engine.mbt + util.mbt —— 零依赖,仅依赖 moonbitlang/core/json + core/math
  • Layer1(服务层)cmd/service/* —— 纯 MoonBit HTTP/MCP 服务,27 个端点,原子写持久化
  • Layer2(知识层)docs/* + AGENTS.md + SKILL.md —— 文档 + 技能编排 + 集成指南

测试策略
  • 契约回归:175 个 test 跨 19 个文件(3 sub-package:core 69 / corpus 54 / feature 52)
  • 门禁scripts/dev.ps1 + .githooks/pre-commit + .github/workflows/ci.yml —— 构建/提交前自动运行 moon test --target wasm-gc
  • 目标:175/175 全绿(wasm-gc 目标,可复现)


#Installation

As a MoonBit library, add the dependency:

moon add Across2005/yimai_prophecy_moonbit

Then declare the import in your package's moon.pkg (recommended alias @lib):

import {
"Across2005/yimai_prophecy_moonbit" @lib,
"moonbitlang/core/json" @json,
}

Requires the MoonBit toolchain (moon, v0.1.2026+).

#Out-of-the-box setup (Windows, AI agents welcome)

Clone the repo, then run the one-shot dev workflow (checks env → builds the native service → starts it on 127.0.0.1:8787 → seeds sample TM pairs → smoke-tests all 27 REST endpoints + 25 MCP tools):

powershell -ExecutionPolicy Bypass -File scripts/dev.ps1

Or step by step: scripts/setup.ps1 (env check) → build.ps1 (compile, needs MSVC) → run.ps1 (start) → seed.ps1 (sample data) → smoke.ps1 (verify).

确定性回归门禁(纯本地,零云端依赖): scripts/dev.ps1 在构建前自动运行 moon test --target wasm-gc(175/175 契约回归,P6 hardening 后),任何一项失败即中止。 此外,仓库自带本地 pre-commit hook(.githooks/pre-commitmoon check + moon test --target wasm-gc),已通过 git config core.hooksPath .githooks 接入本仓库——每个 commit 前自动挡住破坏确定性契约的改动。

Windows native prerequisites: cmd/service requires MSVC (link.native.cc in cmd/service/moon.pkg and cmd/main/moon.pkg points at cl.exe — update both if the path differs on your machine); after a MoonBit toolchain upgrade, rebuild the core native bundle once (cd ~/.moon/lib/core && moon clean --target-dir _build/native &&moon bundle --target native --release). AI agents: see AGENTS.md for the full out-of-the-box guide.


#Quick Start

A copy-paste minimal example: train a short workflow, then predict the next need — and (with #22) manage a translation memory + termbase.

pub fn quickstart() -> Unit {
let mut eng = @lib.ProphecyEngine::make()

// 1) observe() records real steps in order; the engine maintains a
// context window and a 1st/2nd-order Markov transition model internally.
let _ = eng.observe("解析源文件结构", "step")
let _ = eng.observe("提取核心术语表并锁定", "step")
let _ = eng.observe("生成双语对照草稿", "step")

// 2) predict Top-3 most likely next steps from current context.
let pred = eng.predict(3)
println(@json.stringify(pred))

// 3) recall: given a query, return related memories (with activation + path).
let hits = eng.recall("术语", 5)
println(@json.stringify(hits))

// 4) persist & restore.
let snap = eng.to_json()
let eng2 = @lib.ProphecyEngine::from_json(snap)
let _ = eng2

// 5) #22 — TM / TermBase: add memory, load a TBX glossary, align & verify.
let _ = eng.add_tm("电池包热管理策略", "Battery pack thermal management strategy")
let tbx =
"<martif><text><body>" +
"<termEntry id=\"1\"><langSet xml:lang=\"en-US\"><ntig><termGrp><term>sensor</term></termGrp></ntig></langSet>" +
"<langSet xml:lang=\"zh-CN\"><ntig><termGrp><term>传感器</term></termGrp></ntig></langSet></termEntry>" +
"</body></text></martif>"
let _ = eng.load_tbx(tbx)
let tmx = eng.fuzzy_match("电池包热管理", 3, 0.70) // Top-K with match_pct
let v = eng.check_terms("install the sensor", "安装设备") // 1 violation (漏译 传感器)
println(@json.stringify(tmx))
println(@json.stringify(v))
}

observe(text, mtype) builds edges/transitions from the current context window and advances the logical clock. To control co-occurrence manually, call remember(text, mtype, ctx) with an explicit ctx array.

Build & run the bundled demo:

moon build moon run cmd/main # training → Hit@3 → replay prediction → D8 cold-start → consolidate → reward → restore

#Start the HTTP service (pure MoonBit, cmd/service)

moon build cmd/service --target native && moon run cmd/service --target native # 译脉引擎服务: http://127.0.0.1:8787

Then smoke-test the API (Windows native build requires MSVC — see Service Layer):

curl -X POST localhost:8787/api/add_tm -d '{"src":"电池包热管理策略","tgt":"Battery pack thermal management strategy"}' curl -X POST localhost:8787/api/fuzzy_match -d '{"query":"电池包热管理方案","k":3,"threshold":0.5}' curl -X POST localhost:8787/api/check_terms -d '{"source":"install the sensor","target":"安装设备"}' curl -X POST localhost:8787/api/qe_auto -d '{"source":"a","target":"b","match_rate":0.8}' curl -X POST localhost:8787/api/predict -d '{"k":3}'


#Modern Corpus Evaluation (2025–2026)

The engine is evaluated end-to-end on a cutting-edge, purely English corpus spanning 8 modern domains sourced from real 2025–2026 research trends. All transcripts are fully reproducible via moon test --target wasm-gc --filter Layer* (yimai_prophecy_moonbit_modern_corpus_test.mbt).

#Corpus domains (Modern + Extended + Frontier)

#DomainSample training content
1AI Safety & AlignmentRLHF reward hacking audits, red-teaming frontier models against CBRN knowledge, mechanistic interpretability of superposition in SAE features
2Climate Modeling & Carbon CaptureCMIP7 AR7 scenario SSP5-8.5 projection, direct air capture with solid amine sorbents, enhanced weathering of olivine for ocean alkalinity enhancement
3Quantum Computing & Error CorrectionSurface code logical error rates at 10⁻⁶ physical error threshold, cat qubit bias-preserving gates with autonomous stabilization, LDPC code benchmarks on IBM ibm_sherbrooke vs Google Willow
4CRISPR & Gene TherapyCRISPR-Cas12a multiplexed genome editing with AI-designed gRNA libraries, PCSK9 base editing for durable LDL cholesterol reduction, AAV9 capsid engineering for blood-brain barrier crossing
5Cybersecurity & Zero TrustNIST SP 800-207 Zero Trust Architecture deployment, post-quantum TLS 1.3 hybrid key exchange with Kyber-1024 + X25519, AI-driven SOC automation with graph neural network anomaly detection
6Neuroscience & Brain-Computer InterfacesHigh-density 1024-channel ECoG grid for speech decoding, latent diffusion models reconstructing perceived natural images from 7T fMRI BOLD signals
7Distributed Systems & Cloud NativeMulti-region Spanner-style TrueTime with bounded clock uncertainty, service mesh mTLS with SPIFFE identities, disaggregated memory pooling over CXL 3.0 fabrics
8NLP & Large Language ModelsLlama-4-Maverick MOE routing with 128 experts + top-8 gating, RLAIF vs RLHF head-to-head on MT-Bench and AlpacaEval 2.0, retrieval-augmented generation with late interaction ColBERTv2
9Robotics & Embodied AIDiffusion policy for dexterous manipulation with visuotactile feedback, sim-to-real transfer of quadruped locomotion via domain randomization
10Fusion Energy & Plasma PhysicsSPARC tokamak Q>1 breakeven experiments, stellarator coil optimization with adjoint methods
11Synthetic Biology & Metabolic EngineeringCell-free biosynthesis of taxol precursors, CRISPRi logic gates for genetic circuit design
12Protein Design & Drug DiscoveryRFdiffusion backbone generation + ProteinMPNN sequence design, PROTAC ternary complex prediction with AlphaFold3
13Battery Technology & Solid-State ElectrolytesLLZO garnet-type solid electrolyte ionic conductivity tuning, lithium metal anode dendrite suppression with ALD coatings
14Space Tech & Satellite ConstellationsStarlink V2 laser inter-satellite link mesh routing, lunar surface habitat construction with regolith 3D printing
15AI Safety (Frontier)Constitutional AI alignment workflows, mechanistic interpretability of attention head superposition, red-teaming procedures for CBRN knowledge boundary enforcement
16Science (Frontier)CRISPR-Cas12a multiplexed editing workflows, stem cell differentiation protocols, protein folding prediction pipelines with AlphaFold3
17Mathematics (Frontier)Category theory proof verification, homological algebra computation, topological data analysis with persistent homology
18Philosophy (Frontier)Analytic philosophy argument structure mapping, phenomenology consciousness studies, ethical framework deployment workflows
19Digital Humanities (Frontier)Text mining for corpus linguistics, digital archive curation workflows, computational narrative analysis
20CBT Psychology (Frontier)Cognitive restructuring session workflows, exposure therapy protocol management, mindfulness-based cognitive therapy deployment
21Aviation (Frontier)Flight deck procedure automation, air traffic control coordination protocols, aircraft maintenance scheduling workflows
22Space Exploration (Frontier)Mars mission planning workflows, orbital mechanics computation pipelines, satellite constellation deployment protocols

#Evaluation results — 25 tests, all passing (Modern + Extended + Frontier)

Layer 0: Workflow prediction (8 domains × 6 steps × 5 rounds = 240 observations)

DomainPredicted ProjectHit@3
AI SafetyProject: AI Safety Technical Report Q4 2025
Climate ModelingProject: Global Carbon Budget Analysis 2026
Quantum ComputingProject: Surface Code Error Correction Benchmark
CRISPR & Gene TherapyProject: CRISPR-Cas12a Off-Target Analysis Pipeline
CybersecurityProject: Zero Trust Architecture Security Audit
Neuroscience & BCIProject: High-Density ECoG Neural Decoding Pipeline
Distributed SystemsProject: Multi-Region Eventual Consistency Benchmark
NLP & LLMsProject: Multilingual LLM Evaluation Suite v3

Engine-wide Hit@3 = 0.7773. All 8/8 domains produce valid, domain-specific workflow predictions.

Layer 1: TM fuzzy match (multi-granularity white-box scoring)

Cross-domain TM recall consistently activates relevant memories:
  • fuzzy_match("RAG chunking vector store optimization") → hits AI/NLP entries with sim_token / sim_tfidf / sim_char / sim_ngram / sim_tokenset breakdowns
  • fuzzy_match("CRISPR knockout of PCSK9 gene") → hits CRISPR domain entries
  • fuzzy_match("neural decoding of brain signals") → hits neuroscience content via semantic overlap (threshold 0.15)

Layer 2: Term enforcement & TBX glossary

Loads 8 bilingual term entries from TBX format (en-USzh-CN), enforces term consistency on input text, and checks source-target alignment — covering retrieval-augmented generation, low-rank adaptation, surface code, enhanced weathering, guide RNA, zero trust, ECoG, and linearizability.

Layer 3: Cross-domain semantic recall

Multi-domain recall activates the correct domains for mixed queries:
  • "LLM safety benchmarking" → activates AI Safety + NLP domains
  • "gene editing + neural decoding" → activates CRISPR + Neuroscience domains
  • "quantum + distributed consensus" → activates Quantum Computing + Distributed Systems domains

Layer 4: Cold-start generalization

Unseen domains like Zero-Day Threat Intelligence Report and Perovskite Solar Cell Efficiency Roadmap correctly trigger D8 role abstraction to predict a sensible next step — confirming the engine generalizes beyond its training distribution.

Layer 5: Deep fuzzy match with white-box scoring

Cross-domain queries receive multi-granularity similarity breakdowns (token / TF-IDF / char / n-gram / token-set), while orthogonal queries (e.g., "Aristotle" against a technical corpus) correctly return zero results.

Layer 6: Deterministic serialization & JSON round-trip

Two independent engine instances with identical training produce byte-identical to_json() output; to_json → from_json → to_json round-trips are verified; prediction consistency across serialization boundaries is confirmed.

Layer 7: Consolidation, metrics & WAL event sourcing

Post-consolidation state is verified: nodes / edges pruning works correctly, WAL log and replay clone produce valid entries, and metrics report memories count and Hit@3 with expected values.

Layer 8: White-box explainability

explain_card returns rich JSON with activation_path (source→target node chain), prediction_path (step-to-step transitions), and value_breakdown (α·U + β·U_pred + γ·C_graph + δ·R − ε·Cost decomposition).

Layer 9: Attention-gated recall & domain bias modulation

Attention gating (α=0.3, β=0.2) + domain bias (+0.15 on quantum role) shifts recall ranking toward the preferred domain while preserving cross-domain awareness.

Layer 10: Active learning, federated export/import & distillation

Active learning candidates ranked by uncertainty + diversity; federated export produces increment diff; domain bias distilled at +0.25 for targeted roles; federated import merges external memory increments.

#Summary

MetricValue
Total modern corpus tests11/11 passing
Domains covered8 (AI safety, climate, quantum, CRISPR, cybersecurity, neuroscience, distributed systems, NLP)
Training observations240
Engine Hit@30.7773
Post-consolidation nodes/edges48/534
Cold-start generalization✅ Unseen domains produce valid predictions
Determinism✅ Byte-identical serialization, round-trip verified
White-box explainability✅ activation_path + prediction_path + value_breakdown
WAL event sourcing✅ Replay integrity confirmed


#Extended Corpus Evaluation — 6 New Frontier Domains

Beyond the 8 original domains, the engine is additionally validated on 6 emerging research domains sourced from 2025–2026 breakthroughs. All transcripts in yimai_prophecy_moonbit_extended_corpus_test.mbt.

#Additional domains

#DomainSample training content
1Robotics & Embodied AIDiffusion policy for dexterous manipulation with visuotactile feedback, sim-to-real transfer of quadruped locomotion via domain randomization
2Fusion Energy & Plasma PhysicsSPARC tokamak Q>1 breakeven experiments, stellarator coil optimization with adjoint methods
3Synthetic Biology & Metabolic EngineeringCell-free biosynthesis of taxol precursors, CRISPRi logic gates for genetic circuit design
4Protein Design & Drug DiscoveryRFdiffusion backbone generation + ProteinMPNN sequence design, PROTAC ternary complex prediction with AlphaFold3
5Battery Technology & Solid-State ElectrolytesLLZO garnet-type solid electrolyte ionic conductivity tuning, lithium metal anode dendrite suppression with ALD coatings
6Space Tech & Satellite ConstellationsStarlink V2 laser inter-satellite link mesh routing, lunar surface habitat construction with regolith 3D printing

#Evaluation results — 14 tests, all passing

Tests span the same Layer 0–10 framework, covering workflow prediction, TM fuzzy match across robotics and fusion pairs, extended TBX glossary enforcement (6 new terms), cross-domain recall, cold-start generalization, deterministic serialization, consolidation/WAL, explainability, attention-gated recall, and active learning/federated export/distillation.


#Frontier Corpus Evaluation — 10 Emerging Domains

2026-08 P4 增量新增:前 8 个(Modern + Extended)已覆盖 14 个前沿领域,Frontier Corpus 再增 10 个跨学科前沿领域,重点测试预测记忆引擎在超长上下文(步骤超过 10 步)复杂事实推理场景下的泛化能力。所有测试在 yimai_prophecy_moonbit_frontier_corpus_test.mbt

#Frontier domains

#DomainSample training content
1AI Safety (Frontier)Constitutional AI alignment workflows, mechanistic interpretability of attention head superposition, red-teaming procedures for CBRN knowledge boundary enforcement
2Science (Frontier)CRISPR-Cas12a multiplexed editing workflows, stem cell differentiation protocols, protein folding prediction pipelines with AlphaFold3
3Mathematics (Frontier)Category theory proof verification, homological algebra computation, topological data analysis with persistent homology
4Philosophy (Frontier)Analytic philosophy argument structure mapping, phenomenology consciousness studies, ethical framework deployment workflows
5Digital Humanities (Frontier)Text mining for corpus linguistics, digital archive curation workflows, computational narrative analysis
6CBT Psychology (Frontier)Cognitive restructuring session workflows, exposure therapy protocol management, mindfulness-based cognitive therapy deployment
7Aviation (Frontier)Flight deck procedure automation, air traffic control coordination protocols, aircraft maintenance scheduling workflows
8Space Exploration (Frontier)Mars mission planning workflows, orbital mechanics computation pipelines, satellite constellation deployment protocols

#Evaluation results — 14 tests, all passing

核心验证点
  • 超长上下文预测:每个 Frontier domain 训练 10+ 步超长工作流,验证 D3 Forward model 在深度上下文下的稳定性
  • 复杂事实推理:哲学/数学领域的抽象推理路径测试,验证 D1–D8 记忆网络在事实密集型场景的保持
  • 跨学科召回:测试 "AI Safety × Philosophy" 混合查询是否能正确激活两个领域的记忆节点
  • 指令/事实混合模式:Frontier Corpus 采用 instruction-fact 混合模式(add_tm 存储事实,observe 学习指令),更贴近真实工作流


#Service Layer — 纯 MoonBit HTTP API (cmd/service)

cmd/service纯 MoonBit 双层架构的 Layer 2:用 moonbitlang/async(http / fs / socket)起本地 HTTP server,把引擎能力以 REST API 暴露给前端工作台 / Agent / LLM 宿主。引擎到服务零桥接语言——同一门 MoonBit 完成全部。

#架构演进:旧架构 vs 新架构

本项目从「纯库」演进为「三层架构」。差异如下:

维度旧架构(v1)新架构(v2,当前)
总体形态单一纯库(零依赖内核)+ cmd/main demo三层:Layer0 零依赖内核 / Layer2 纯 MoonBit 服务层 / Layer1 知识层(文档 + 前端工作台已实现)
I/O 能力无 stdin / 无文件 I/O(wasm-gc 内存态)async/fs 原子写持久化tm_store.json,tmp+rename)+ 重启恢复闭环
对外接口仅 MoonBit 函数调用(moon add 后进程内调用)27 个 HTTP REST 端点,前端 / Agent / LLM 宿主可直接消费
集成路径2 条:库引用、算法移植4 条:A 构建运行 / B wasm-gc exports / C HTTP 服务(新增,已实测) / D 算法移植
运行形态wasm-gc 内存态(测试友好)native(Windows 需 MSVC)本地常驻服务,127.0.0.1:8787
语言栈单一 MoonBit(仅库)单一 MoonBit(库 + HTTP 服务 + 文件 I/O)——引擎到服务零桥接语言
状态持有调用方自管引擎实例Ref[ProphecyEngine] 服务内单例 + JSON 边界透出

演进动机:旧架构的引擎能力只能被「会 MoonBit 的程序」消费;新架构让任何会 HTTP 的宿主(浏览器前端、Agent 工具调用、LLM 函数调用)都能用上确定性记忆引擎——内核零依赖铁律不变,只是多了一层纯 MoonBit 的 I/O 壳。

#端点矩阵(27 个,全部 curl 实测通过)

端点方法请求体响应说明
/api/pingGET{"status":"ok"}健康检查
/api/add_tmPOST{"src","tgt"}{"id","status"}新增 TM,原子落盘
/api/fuzzy_matchPOST{"query","k","threshold"}Top-K(S1 四分量白盒)TM 模糊检索
/api/check_termsPOST{"source","target"}违规数组术语一致性校验
/api/concordancePOST{"term","k"}含术语 TM 段术语上下文检索
/api/qe_autoPOST{"source","target","match_rate"}{"qe_score","term_ok","mqm"}自动 QE 评分
/api/predictPOST{"k"}{"predictions","confidence","uncertainty"}下一步预测 + 白盒
/api/observePOST{"text","mtype"}{"mid","status"}记录真实步骤(学习/转移)→ 落盘
/api/recallPOST{"query","k"}Array[{id,text,score,via_edges}]激活扩散语义召回
/api/explainPOST{"mid"}白盒卡片value_breakdown / activation_path 证据链
/api/rewardPOST{"mid","score"}{"ok"}采纳/拒绝反馈 → predictive_value(闭环核心)
/api/consolidatePOST{"prune"}{pruned,nodes,edges,...}固化重放(价值重算 + 剪枝)
/api/retrieve_promptPOST{"query","k","threshold"}三段式TMPlm:suggestions/terms/glossary 供 LLM prompt 注入
/api/bleu / /api/chrfPOST{"ref","hyp"}{bleu} / {chrf}MT 质量评测(零依赖自实现)
/api/style_checkPOST{"text"}问题数组风格一致性(句长/标点/括号/术语命中)
/api/style_reportPOST{"text"?}{sentence_count,avg_src_len,avg_tgt_len,formal_score,distribution,term_variants,tips}风格一致报告(记忆库分布 + 术语变体族 + 新译文偏离建议)
/api/back_alignPOST{"source","target"}{align_score,misaligns,ops}回译 LCS 对齐(含字符级 ops 供热力图)
/api/term_conflictsPOST冲突数组一词多译 / 多词一译
/api/fed_export / /api/fed_importPOST{"added","updated"}{status}联邦增量导出/导入(FedAvg 端点层)
/api/distill_injectPOST{"table":{k:v}}{status,keys}蒸馏偏置表注入
/api/active_learningPOST{"k"}Array[{id,text,uncertainty,role}]主动学习推荐(uncertainty+diversity 待标注句)
/api/tm_countGET{"tm_count"}存量统计
/api/metricsGET引擎/服务指标可观测指标
/api/healthGET健康状态含 uptime / 上次落盘状态
/api/mqm_re_annotatePOST{"source","target","match_rate"}MQM 二次标注结果Critical 段强制重审

实测(中文 query,白盒分项全透出):

POST /api/fuzzy_match {"query":"电池包热管理方案","k":3,"threshold":0.5} → [{"id":"m3","source":"电池包温度管理方案","target":"Battery pack temperature management plan", "score":0.7377,"match_pct":73.7723,"sim_token":0.7125,"sim_tfidf":0.6796, "sim_char":0.7778,"sim_ngram":0.6667,"sim_tokenset":0.75}, ...]

#三层关联与记忆闭环

27 个端点不是孤立的——它们把三层连成记忆闭环(完整映射表见架构方案 §11):

前端操作 ──HTTP──▶ Layer2 服务端点 ──调用──▶ Layer0 引擎方法 ▲ │ │ ▼ └─── JSON 响应(白盒分数/证据链)◀── save_store() 原子落盘 ◀┘ │ ▼ 重启 load_store() → from_json() → 记忆不丢

两个闭环(实测):
  • 采纳闭环:前端「采纳译文」→ /api/reward{mid,+1}predictive_value 提升 → 下次 /api/predict 排序更优;
  • 学习闭环:前端「记录步骤」→ /api/observe{text} → 转移计数 → 落盘 → 重启恢复 → 预测更准(实测:observe 两步 → predict Top1 prob=1.0)。

#MCP Server(/mcp 端点)—— 供 Claude Desktop / 通用 MCP 客户端消费

cmd/service 同时暴露 MCP(Model Context Protocol)Server 变体(spec 2025-11-25,Streamable HTTP):挂 /mcp 端点,POST 单 JSON-RPC 消息、application/json 响应(无需 SSE)。**25 个引擎能力直接映射为 MCP tools(新增 mqm_re_annotate)`:

MCP tool参数说明
fuzzy_matchquery / k / thresholdTM 模糊检索(S1 四分量白盒)
add_tmsrc / tgt新增 TM 并落盘
check_termssource / target术语一致性校验
concordanceterm / k术语上下文检索
qe_autosource / target / match_rate自动 QE 评分
predictk下一步预测 + 白盒路径
observetext / mtype记录步骤(学习)并落盘
recallquery / k语义召回
explainmid白盒卡片
rewardmid / score采纳/拒绝反馈
consolidateprune固化重放
tm_count / ping存量 / 健康检查
retrieve_promptquery / k / thresholdTMPlm:为 LLM prompt 组装三段式检索上下文(suggestions/terms/glossary)
bleu / chrfref / hypMT 质量评测(零依赖自实现)
style_checktext风格一致性(句长/标点/括号/术语命中)
style_reporttext?风格一致报告:记忆库句长/正式度分布 + 术语变体族 + 新译文偏离建议
back_alignsource / target回译 LCS 对齐验证
term_conflicts术语冲突检测(一词多译/多词一译)
fed_export / fed_importadded / updated联邦增量导出/导入
distill_injecttable蒸馏偏置表注入
active_learningk主动学习推荐:uncertainty×0.6+novelty×0.4,角色去重,待标注句
mqm_re_annotatesource / target / match_rateMQM 二次标注:Critical 段强制重审

# MCP 握手(curl 模拟客户端) curl -X POST localhost:8787/mcp -d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{}}' curl -X POST localhost:8787/mcp -d '{"jsonrpc":"2.0","id":2,"method":"tools/list"}' curl -X POST localhost:8787/mcp -d '{"jsonrpc":"2.0","id":3,"method":"tools/call","params":{"name":"fuzzy_match","arguments":{"query":"电池","k":2}}}'

协议:initialize(protocolVersion 2025-11-25 + capabilities.tools)/ notifications/initialized(202)/ tools/list(25 tools,inputSchema JSON Schema 2020-12)/ tools/call(未知工具 -32602,引擎异常 isError:true);GET /mcp 回 405。实现为自建轻量 JSON-RPC 2.0 层(cmd/service/mcp.mbt,MoonBit 无现成 MCP 库),全程复用引擎 @lib.obj/str_json/num_json 构造(Json 为 FFI 类型)。

#持久化

  • 引擎状态经 to_json()@fs.write_file(tmp, create_mode=CreateOrTruncate)rename 原子落盘tm_store.json
  • 重启时 load_store()from_json() 完整恢复(实测 tm_count 持久化后重启一致);
  • 单例持有:engine_ref : @ref.Ref[@lib.ProphecyEngine](MoonBit 顶层无全局可变变量,Ref 是标准方案)。

#平台要求(Windows)

  • async 的 native 后端硬性要求 MSVCthread_pool.c: #error "Currently only MSVC is supported on Windows"),mingw gcc 不可用;wasm/js 后端暂不支持 socket server;
  • 构建需 MSVC 环境(INCLUDE/LIB)+ moon.pkg 配置 link.native.cc 指向 cl.exe
  • 工具链升级后需重建 core native bundle(cd ~/.moon/lib/core && moon clean --target-dir _build/native && moon bundle --target native --release)。


#API Reference

All public interfaces are methods of ProphecyEngine (encoding helpers in util.mbt are package-private).

#Core (D1–D8, persistence, feedback)

MethodSignatureDescription
make() -> ProphecyEngineCreate an empty engine.
remember(text, mtype, ctx : Array[String]) -> StringWrite/dedupe a memory node, build co-occurrence synapse, return node id.
observe(text, mtype) -> StringRecord a real next step: remember + update transition + hit accounting + advance context.
predict(k : Int) -> JsonPredict Top-K next steps from context window; returns prob/path/confidence/uncertainty.
recall(query : String, k : Int) -> Array[Json]Activation-spread associative recall; returns memory + activation + path.
consolidate(prune : Bool) -> JsonConsolidation replay: decay, recompute value, optional prune, meta-cognitive explore.
restore() -> JsonRoll back to the pre-consolidation snapshot (constraint contract).
reward(mid, score : Double) -> BoolSuccess/failure feedback into predictive value.
explain(mid : String) -> JsonWhite-box explanation of a node's edges/transitions/hit-rate.
end_episode() -> UnitEnd the current episode sequence.
hit_rate() -> DoubleCumulative prediction hit rate.
stats_view() -> JsonNode/edge/episode counts, type distribution, avg predictive value.
context_texts() -> Array[String]Texts in the current context window.
last_context_id() -> StringLast node id in the context window.
to_json() -> JsonExport full engine state (persistence).
from_json(data : Json) -> ProphecyEngineRestore engine state from JSON.

#Learning & explanation (准·美)

MethodSignatureDescription
set_domain_bias(role, delta : Double) -> UnitInject/accumulate per-domain bias ΔW (LoRA-style).
inject_distillation(table : Map[String, Double]) -> UnitInject read-only distilled bias table (neural-symbolic distillation, consume-side).
cl_step(anchor, positive, negative : String) -> UnitOnline contrastive learning: strengthen (anchor,pos), suppress (anchor,neg).
set_attention(alpha, beta : Double) -> UnitToggle attention-gated edge weights in recall (default off).
mark_term(mid : String) -> BoolMark a node as terminology (TermNode): boost recall + flag for explain card.
explain_card(mid : String) -> JsonWhite-box explainable card: activation/prediction paths + value breakdown.

#Incremental & collaborative (中/长周期)

MethodSignatureDescription
active_learning_candidates(k : Int) -> Array[Json]Top-K uncertain + role-diverse nodes for human labeling.
wal_replay() -> ProphecyEngineRebuild engine from the Write-Ahead Log (event sourcing).
wal_export / wal_compact / wal_clear / wal_len(…) -> Array[String] / Unit / IntWAL inspection & maintenance.
fed_export / fed_import() -> Json / (added, updated : Int) -> UnitFederated increment counter export/import (coordinator merges weights).

##22 TM/TB (检索 + 术语守门)

MethodSignatureDescription
add_tm(src, tgt : String) -> StringAdd a translation-memory entry (source→target), build the TF index, return the node id.
fuzzy_match(query : String, k : Int, threshold : Double) -> JsonTM fuzzy match Top-K. S1 升级评分 = 0.55·idf_dice + 0.20·char-2gram-dice + 0.15·token-set-dice + 0.10·position(IDF 加权让罕见术语优先、2-gram 捕捉形态变体、token-set Dice 容忍词序重排);returns match_pct / sim_token / sim_tfidf / sim_char / sim_ngram / sim_tokenset. (建议阈值 threshold = 0.70;MoonBit 无默认参数,调用方需显式传入。)
fuzzy_match_legacy(query : String, k : Int, threshold : Double) -> JsonA/B 对照基线,长期保留不删除(≥0.3.0 讨论移除):旧公式 0.7·token-cosine + 0.3·char-ratio,4 处引用(engine.mbt + 2 test + 本 README)。新代码请用 fuzzy_match(S1 公式 + IDF 倒排剪枝)。
concordance(term : String, k : Int) -> JsonConcordance search: returns all TM segments containing the query term, scored by term occurrence count (distinct from the fuzzy_match similarity score).
load_tbx(xml : String, src_lang~ : String = "en-US", tgt_lang~ : String = "zh-CN") -> IntParse a TBX (ISO 30042) termbase. Resolves source/target by each langSet's xml:lang (default en-US→zh-CN; falls back to document order when absent). Returns the number of concept entries loaded.
enforce_terms(text : String) -> JsonTerm enforcement: scan text for known terms (Latin terms require word boundaries, so log won't false-match logical), return hits with translation.
check_terms(source, target : String) -> JsonTerm-consistency check: for each source term whose translation is missing from the target, return a violation.


##22 TM/TB — Translation Memory & TermBase

This extension makes translation memory and terminology first-class citizens alongside the predictive core. It reuses the existing yimai_tokenize / tf_vector / cosine / align_diff / clamp01 primitives — no new dependencies, no re-invented wheel.

  • add_tm(src, tgt) builds a MemoryNode of mtype="tm" carrying text=source, translation=target (and maintains the TM document-frequency index for IDF).
  • fuzzy_match (S1 upgrade) scores 0.55·idf_dice + 0.20·char-2gram-dice + 0.15·token-set-dice + 0.10·position over mtype=="tm" nodes only. The IDF table (ln((N+1)/(df+1))+1) down-weights frequent words so rare domain terms dominate (R23); char 2-gram catches morphological variants; token-set Dice tolerates word reordering — a Chinese reordered query scores 0.91 with the new formula while the legacy one misses it entirely (R24). The old 0.7·token-cosine + 0.3·char-ratio formula remains as fuzzy_match_legacy for A/B comparison.
  • concordance counts query-term occurrences per TM segment — concordance % is term occurrence, distinct from the fuzzy similarity score (per the research baseline).
  • load_tbx parses TBX 2.0 (<ntig><termGrp><term> or simplified <tig><term>), language-aware via xml:lang, into mtype="term" nodes with is_term=true.
  • enforce_terms / check_terms provide terminology lock-in and missed-term detection, with word-boundary-aware matching for Latin terms.

All six methods are covered by regression tests R16–R22 (see Evaluation).


#Data Formats

#Engine persistence (to_json / from_json)

{ "memories": { "m1": { "id":"m1","text":"解析源文件结构","type":"step", "vec":{"解析":1,"源":1},"created":1,"last_used":3, "use_count":2,"feedback":0,"edges":{"m2":0.30}, "predictive_value":0.42,"hit_count":1,"predict_count":1, "is_term":false,"translation":"" } }, "transitions": { "m1": {"m2":1.0} }, "episodes": [["m1","m2","m3"]], "context": ["m1","m2","m3"], "stats": { "preds":1, "hits":1, "remembers":3, "evolutions":0 }, "seq": 3, "explore": 0.0, "clock": 3, "meta_hits": [1], "snapshot": {}, "role_trans": {}, "role_index": {} }

TM / TermBase nodes add two fields: a TM node carries "type":"tm","translation":"<target>"; a terminology node carries "type":"term","is_term":true,"translation":"<target term>". Both are round-trip preserved through to_json/from_json (covered by R20).

#predict(k) returns

{ "predictions": [ { "id":"m4", "text":"生成双语对照草稿", "prob":0.62, "path": [ { "from":"m3", "p":0.55 } ] } ], "confidence": 0.62, "uncertainty": 0.41 }

#recall(query, k) returns (array)

[ { "id":"m2", "text":"提取核心术语表并锁定", "type":"step", "score":0.71, "activation":1.0, "via_edges": [ { "to":"m1", "w":0.30 } ] } ]

#consolidate(prune) returns

{ "pruned": 0, "nodes": 3, "edges": 2, "explore": 0.0, "recent_hit_rate": 0.5 }

##22 — TM / TermBase outputs

fuzzy_match(query, k, threshold) (array, Top-K by score):

[ { "id":"m7", "source":"电池包热管理策略", "target":"Battery pack thermal management strategy", "score":0.91, "match_pct":91.0, "sim_token":0.90, "sim_tfidf":0.88, "sim_char":0.95, "sim_ngram":0.86, "sim_tokenset":0.93 } ]

concordance(term, k) (array):

[ { "id":"m11", "source":"打开设置菜单选择网络", "target":"Open Settings menu, choose Network", "hits":1 } ]

load_tbx input (TBX 2.0 fragment):

<martif><text><body> <termEntry id="1"> <langSet xml:lang="en-US"><ntig><termGrp><term>network logon</term></termGrp></ntig></langSet> <langSet xml:lang="zh-CN"><ntig><termGrp><term>网络登录</term></termGrp></ntig></langSet> </termEntry> </body></text></martif>

enforce_terms(text) (array — word-boundary matched):

[ { "term":"network logon", "translation":"网络登录", "mid":"m20" } ]

check_terms(source, target) (array — violations only):

[ { "term":"sensor", "expected":"传感器", "mid":"m21" } ]


#Evaluation & Test Results

All numbers below are produced by moon test --target wasm-gc and are reproducible.

Summary: Total tests: 175, passed: 175, failed: 0 (3 sub-packages: tests/core/ 69 + tests/corpus/ 54 + tests/feature/ 52; 4 quantitative acceptance + 16 API coverage + 6 benchmark + 7 golden + 12 long-text + 10 TM + 6 v2 + 3 whitebox + 1 main; 11 modern corpus + 11 extended corpus + 14 frontier corpus + 3 business corpus + 15 roadmap regression; 19 extension E1–E19 + 16 P4 quality T1–T37 + 12 routes_meta + 5 mqm_re_annotate).

LayerCheckResultEvidence
L1Batch Hit@3 > 0.8hit_rate = 0.8246 over 8 topics × 8 rounds
L1Determinism / reproducibilityTwo to_json calls are byte-identical
L1JSON round-tripto_json → from_json → to_json identical
L1Consolidation keeps core memory13 nodes → 13 nodes after consolidate
L2Known-project replay predicts correct nextobserve project → Top1 提取核心术语表并锁定
L2Cold-start generalizationunseen topic via D8 role-abstraction yields correct next step
L2White-box explainableexplain_card returns concrete activation_path / prediction_path / value_breakdown
L2Persistence after restartto_json → from_json Top1 unchanged
MCModern Corpus: 8 domains × 5 rounds11/11 tests passing; Hit@3=0.7773; see Modern Corpus Evaluation
MCCold-start on unseen domainsZero-Day Threat Intelligence / Perovskite Solar Cell → valid predictions
MCCross-domain semantic recallMulti-domain queries activate correct domain clusters
MCDeep fuzzy match (S1 upgrade)sim_token / sim_tfidf / sim_char / sim_ngram / sim_tokenset
MCAttention-gated recall + domain biasα=0.3, β=0.2 + ΔW=0.15 shifts ranking correctly
MCWAL event sourcing replay384 entries → replay clone produces 576 entries
MCFed export/import + distillationIncrement diff export → merge → distilled bias confirmed

#Credibility hardening

The engine's Hit@3 is verified across multiple independent evaluation surfaces:

  • Classic acceptance suite (4 tests): Layer1 batch training → Hit@3=0.8246, determinism, JSON round-trip, consolidation.
  • Modern corpus suite (11 tests): 8 cutting-edge English domains, 240 training observations, Hit@3=0.7773, cold-start generalization on unseen domains, cross-domain semantic recall, deep fuzzy match, attention-gated recall, WAL event sourcing, federated export/import, and distillation.
  • Regression suite (R1–R25): TM/TB fuzzy match % (S1 IDF + 2-gram + word-order), concordance, TBX load+enforce+check, word-boundary (loglogical), 3-language xml:lang, IDF discrimination, word-order tolerance, empty/short-query boundary.
  • Extension suite (E1–E18): QE+MQM, format-fidelity, multimodal-OCR-stub, batch-CI, TMS XLIFF/TMX, observability/drift.

Reproduce:

cd yimai_prophecy_moonbit moon test --target wasm-gc # all 175 tests (P6 hardened) moon test --target wasm-gc --filter Layer* # modern + extended + frontier corpus (25 tests) moon test --target wasm-gc --filter T* # P4 quality/security tests (T30–T51, 22 tests) moon build --target wasm-gc # library only cd cmd/main && moon build --target wasm-gc && moon run .


#Extension API (#2–#7)

All seven user-scoped extension capabilities are implemented in engine.mbt and regression-tested by yimai_prophecy_moonbit_extension_test.mbt (E1–E18). They are zero-dependency and deterministic — same input ⇒ same output.

#CapabilityKey methodsNotes
2Quality estimation + MQMqe_score, mqm_tags, qe_autoqe_score = 0.55·match_rate + 0.30·term_ok + 0.15·char_ratio (cross-language length penalty dropped — it dragged scores to ~0.7 and distorted QE). mqm_tags emits terminology / accuracy / fluency / omission with major / critical / minor severity.
3Format-fidelity round-tripcheck_format_fidelity, protect_tagsDetects missing (source tag absent in target) / extra (target-only) inline tags & placeholders; protect_tags masks them to __TAG__ so fuzzy token-cosine isn't polluted.
4Multimodal / screenshotocr_image (stub), align_regionsocr_image is an external boundary stub (real OCR = Tesseract / vision-LLM, injected by host). Regions flow as JSON {bbox, text}; the engine does region ↔ TM alignment purely.
5Localization CI / batchbatch_applyTop-1 TM match + term-gate per segment, threshold-driven → {total, passed, failed, items}. Drop-in for a CI localization gate.
6TMS interoperabilityparse_tmx, parse_xliff, export_tmxXLIFF 1.2 <trans-unit> and 2.0 <unit> both parsed; TMX 1.4 exported with XML escaping. Round-trip verified (E10).
7Observability & driftmetrics, drift_reportmetrics = tm/term counts + term-coverage; drift_report(before, after) diffs two to_json snapshots by (type\|is_term\|text\|translation) key to surface TM/term add/remove.

Boundary principle. Capabilities #4 (OCR) and any host persistence stay outside the zero-dependency engine. The engine speaks JSON at these boundaries, so the host (Node/Python/Agent) supplies OCR, files, and I/O — the MoonBit core stays 100% pure-stdlib and wasm-gc-testable.


#Roadmap & Extension Status

#Engine implementation status (内核「快·准·美」)

The "fast / accurate / beautiful" algorithm kernel is fully landed and contract-tested. Remaining items are front-end workbenches (ardot hand-off) or external services, not engine gaps.

Roadmap entryEngine implementationStatus
Role inverted-index + Top-K pruning + LRU pred cacherole_members / predict Top-K / pred_cache
2nd-order Markovtrans2
Multi-granularity rolesroles_of
Adaptive LR + elastic forgettinghebb_lr / consolidate
Domain bias ΔW (LoRA-style)domain_bias / set_domain_bias / inject_distillation
Online contrastive learningcl_step
Attention edge weightsattn_alpha/beta / set_attention
TermNode + explainable cardmark_term / explain_card
Incremental WALwal_*
Bilingual alignment (Myers/LCS)align_diff
Active-learning candidatesactive_learning_candidates
Federated incrementfed_export / fed_import
Neural-symbolic distillation (consume side)inject_distillation
#22 TM / TermBaseadd_tm / fuzzy_match / concordance / load_tbx / enforce_terms / check_terms✅ (new)
S1 fuzzy-match upgradefuzzy_match(IDF + 2-gram + word-order)/ fuzzy_match_legacy
Pure-MoonBit HTTP servicecmd/service:27 端点 + 记忆闭环 + 原子写持久化 + 重启恢复✅ (new)
MCP Server (/mcp)cmd/service/mcp.mbt:25 引擎能力 → MCP tools(spec 2025-11-25)✅ (new)
TMPlm 桥接 (M5)retrieve_for_prompt + /api/retrieve_prompt(三段式)✅ (new)
SKILL.md 编排壳docs/skill/SKILL.md(agent_created,安装见 For Agents)✅ (new)
Translator workbench (web front-end)cmd/service/web(四面板 + 记忆图谱)✅ (new,阶段 C 已交付)
Visual memory graphweb 第 4 面板(recall 节点 + via_edges 边 → SVG 确定性环布局)✅ (new)
S4 评测 / M1 风格 / M2 回译 / M3 冲突bleu_score/chrf_score/style_check/back_align/term_conflicts + 端点✅ (new)
风格一致报告 (风格一致)style_report(记忆库句长/正式度分布 + 术语变体族 + 偏离建议)+ /api/style_report + MCP + web ⑧ 面板✅ (new)
FedAvg / 蒸馏注入端点/api/fed_export /api/fed_import /api/distill_inject🔶 端点已落地,外部协调器/独立训练流仍缺

#Seven extension capabilities (user-scoped)

From the "translation-born skill" brainstorm — what's built vs. pending:

#CapabilityStatusNotes
1TM / TermBase first-class (fuzzy match %, concordance, TBX enforcement)✅ Donemoon test 175/175 (P6 hardening 后); reviewed + hardened (word-boundary, xml:lang); S1 fuzzy-match upgrade (IDF + 2-gram + word-order, R23–R25); open-code-review + MoA fixes for parse_tmx cross-language/</tu> split + mqm_tags cross-language false positives + empty-target/language-variant robustness; P0 长文 + 数字守门加固(MAX_TOKENS 截断 / numeric_consistency MQM 维度 / fuzzy_match 长 query 不崩 / L1–L12 长文回归); P4 fuzzy_match 抽公共 helper + drift_report.text_chrf_avg + MQM 严重度数值化; P6 last_body_oversize→Result enum 消 TOCTOU + NaN/Inf API 修正 + validate.mbt 同包。
2Quality estimation + MQM auto-eval✅ Doneqe_score (0.55·match + 0.30·term + 0.15·char) + mqm_tags (terminology/accuracy/fluency/omission w/ severity). Tested E1–E3.
3Format-fidelity round-trip✅ Donecheck_format_fidelity (missing/extra tag detection) + protect_tags (mask tags to __TAG__). Tested E4–E5.
4Multimodal / screenshot translation✅ Done (OCR external stub)ocr_image (external boundary) + align_regions (region ↔ TM align). Zero-dep engine speaks JSON at the OCR boundary; real OCR injected by host. Tested E6–E7.
5Localization CI / batch pipeline✅ Donebatch_apply (Top-1 TM + term-gate, threshold-driven) → {total, passed, failed, items}. Tested E8–E9.
6TMS interoperability✅ Doneparse_tmx / parse_xliff (XLIFF 1.2 <trans-unit> & 2.0 <unit>) + export_tmx (round-trip). Tested E10–E11; cross-language TMX correctness regression added as E14 (open-code-review fix).
7Observability & drift monitoring✅ Donemetrics (tm/term counts + coverage) + drift_report (before/after snapshot diff). Tested E12–E13.

This project is not packaged as a WorkBuddy skill yet. The dev loop for these is: research (Deep Research / WebSearch) → review (open-code-review) → verify (browser automation + moon test). MoA is intentionally not embedded inside the skill (kept as an external advisor).


#International Standards & Compliance (P4 增量)

2026-08 增补:随着项目演进,本节列出当前已声明对齐 / 仍属 roadmap 的国际/区域标准, 以及对应的本地化合规姿态。yimai 本身是技术构建块(library + local service)而非翻译服务 机构;本节为「adopter 集成指南」,非 ISO 认证声明。

#已对齐(algorithm / docs 层)

标准yimai 映射备注
ISO 17100:2015 (Translation services)observe / predict + reward 反馈闭环;retrieve_prompt 注入双语上下文译员能力、项目管理、技术资源、反馈机制由 yimai 闭环支撑;adopter 仍需认证译员/项目流程
ISO 18587:2017 (MT post-editing)qe_auto(QE+MQM 标签)+ bleu / chrf 度量MTPE 工作流核心;数字守门 P0 加固后覆盖本地化高危硬伤
ISO 30042:2019 / TBX3 (TermBase eXchange)load_tbx 解析 ISO 30042-compliant <martif>/<termEntry>TBX3 v3.0 dialect(非 TBX2 v2.0,namespace 不同)
ISO 11669:2024 (Translation projects — General guidance)predict 逐步推荐 + consolidate 项目收尾复盘完整标准(replaced ISO/TS 11669:2012)
ISO 5060:2024 (Translation services — Evaluation of translation output)qe_auto / mqm_tags / drift_report.text_chrf_avg 三层指标覆盖与 MQM Council 强对齐(MQM 官网声明 "Aligned with ISO 5060")
MQM Core (Lommel et al., 2014–present)mqm_tags 7 维度标签 + 严重度数值(severity_score: None=0 / Minor=1 / Major=5 / Critical=10)权威背书:https://www.themqm.org/
W3C ITS 2.0 (Internationalization Tag Set)由宿主 CMS/应用注入;protect_tags / mark_term 消费 in-text metadata不在引擎内;consume 边界由 adopter 决定
MCP 2025-11-25 (Model Context Protocol)/mcp 端点(Streamable HTTP + JSON-RPC 2.0)25 tools;2026-07-28 已发稳定版(无状态核心 / 移除 initialize 握手 / server/discover),但因纯本地无鉴权、现有 13 harness 均基于 2025-11-25 握手,yimai 停留在 2025-11-25(迁移 0.2.0 再议);已落实该版本的 Origin 头校验(DNS 重绑定防护)

#MQM 严重度尺度(与业界三方对齐,P5 增量)

yimai 采用 MQM Core 严重度数值化。业界主流 MQM 评分器使用三套不同的 penalty 数值 下表给出显式对照(便于跨工具数据交换):

Severityyimai severity_scorePhrase penaltyLokalise penalty (vs 100)用途
None000可接受变体,不扣分
Minor115局部小问题(拼写/标点)
Major5525影响理解(术语错/漏译)
Critical102575改变意义(negation flip / 数字错)

yimai vs Phrase 差异:Critical penalty yimai 用 10 而非 25。理由是 yimai 中 Critical 已自动触发 mqm_re_annotate 二次标注流程(Google 2025-10-28 论文对齐), 二次审后再被采纳的 Critical 段会被消费方拦截,因此 penalty 10 已足够震慑。 Phrase 走纯人工 review 路径,故用更重的 25 防止漏审。

⚠️ mqm_re_annotate 当前实现是 deterministic 自重审:因 qe_auto 算法本身确定性, re_severity == original_severity 恒成立,consistent: true 端点接口与 JSON schema(re_annotated / critical_count / re_annotations 三段式)已 按 Google 2025-10-28 论文语义设计;多标注员模型(multi-rater / Cohen's κ / Krippendorff's α) 只需替换 mqm_re_annotate 内部循环即可,外部契约不变。

yimai vs Lokalise 差异:Lokalise 走 100 - sum(penalties) 评分模式(满分 100), yimai 走 severity_score 原始累计 + qe_auto 综合公式(0.50·match_rate + 0.25·term_ok + 0.10·char + 0.15·bleu)。 两者数值不可直接比较,需要按公式反推。

参考:

#Roadmap(未在 0.1.0 落地,0.2.0 候选)

标准为什么重要状态
TMX 1.4bCAT-tool 互操作(Trados / memoQ / OmegaT)parse_tmx / export_tmxpub fn 但无 /api/* 入口
XLIFF 2.1 / ISO 21720:2024段级交换的事实标准(ISO 21720:2024 为第二版;OASIS XLIFF 2.2 2025-03 进入 CS)parse_xliffpub fn 但无入口
SRX 2.0跨工具段切规则可复现性未声明;需 /api/import_srx

#合规姿态:EU AI Act + GDPR(2026-08 声明)

  • GDPR Art. 28 要求翻译记忆若含 PII 必须本地化处理。yimai「纯本地、零云端依赖」天然合规—— TM/术语/学习闭环全在 127.0.0.1:8787 闭环,tm_store.json 原子写持久化在本地工作目录。
  • EU AI Act 2024-2026 高风险 AI 系统需可审计、可追溯。yimai 是 non-deployable unit adopter 集成进翻译产品时承担部署者责任。引擎提供白盒路径explain_card / activation_path / prediction_path / value_breakdown)+ MQM 严重度尺度作为 可审计凭证。
  • 不引入 OAuth / ID-JAG / Mcp-Session-Id 等云端鉴权方案——保持"纯本地"定位;如需 企业版鉴权,由 fork 独立分支承担。

#MCP server 集成 Claude Code / Cursor / Gemini CLI

yimai /mcp 是标准 MCP 2025-11-25 server,可被以下 harness 直接消费(同一份配置 schema 对所有 harness 透明;tools 列表一致):

Harness配置入口关键字段
Claude Code~/.claude/mcp.json 或项目级 .mcp.jsonmcpServers.yimai.{url,type:"http"}
Cursor~/.cursor/mcp.json同上
Gemini CLI~/.gemini/settings.jsonmcpServers.yimai.{url,type:"http"}

完整 13 个 harness 配置(Claude Desktop / Claude Code / Gemini CLI / Cursor / Cline / Continue.dev / Roo Code / Windsurf / OpenAI Codex CLI / Aider / Sourcegraph Cody / Zed / GitHub Copilot)见 docs/harness-configs/


#For Agents / Integration

This package delivers a zero-dependency, deterministic "Prophecy Memory Network" to other agents. Integration options (2026-08 更新,修正此前「无 I/O / 仅 IIFE」的过时声明):

AI agent 拆箱即用:克隆后先读 AGENTS.md(项目结构、构建/运行/消费指南、Windows 前置、MoonBit 坑、多 harness 接入),再跑 scripts/dev.ps1 一键起服务——无需人工配置。

  • Claude Code 用户:直接读 CLAUDE.md 即可
  • Gemini CLI 用户:直接读 GEMINI.md 即可
  • 其他 harness(Cursor / Cline / Continue.dev / Roo Code / Windsurf / Copilot / Codex / Devin):用 AGENTS.md 多 harness 接入表

安装为 WorkBuddy skill(可选):仓库内 docs/skill/SKILL.md 是编排手册(frontmatter 含 agent_created: true,触发词 + 26 端点 API 手册 + MCP 接入 + 数据契约)。复制到 ~/.workbuddy/skills/yimai-prophecy/ 并重启 WorkBuddy 后即成为可用技能。

Path A — build & run (agent has the MoonBit toolchain):

cd yimai_prophecy_moonbit moon test --target wasm-gc # green ⇒ engine is usable cd cmd/main && moon build --target wasm-gc && moon run .

Then moon add Across2005/yimai_prophecy_moonbit and call any of the ProphecyEngine methods.

Path B — wasm-gc exports (engine as a callable module): 新工具链(moonc v0.10.4+)支持在 moon.pkg.json 配置 link.wasm-gc.exports + use-js-builtin-string: true,让 JS 宿主直接调用 ProphecyEngine 方法(String 与 JS String 互通);JS 后端亦支持 format: esm/cjs(不再只有 IIFE)。详见 Package Configuration

Path C — service layer (纯 MoonBit HTTP server, cmd/service):已实测落地moonbitlang/async 提供 http / fs / socket,起本地服务并托管前端工作台;27 个 /api/* 端点(13 基础 + retrieve_prompt / bleu / chrf / style_check / style_report / back_align / term_conflicts / fed_export / fed_import / distill_inject / active_learning / metrics / health / mqm_re_annotate)全部 curl 通过,含记忆闭环(observe 学习 → predict 预测 → reward 反馈 → consolidate 固化)与原子写持久化 + 重启恢复(详见 Service Layer)。注意:Windows 上 async 的 native 后端仅支持 MSVC 编译thread_pool.c: #error "Currently only MSVC is supported on Windows");wasm/js 后端暂不支持 socket server(socket/unimplemented.mbt)。

Path D — algorithm port (agent has no MoonBit but needs the capability in-process):

engine.mbt is pure-stdlib, zero-I/O, with constants (HEBB_LR, EDGE_DECAY, BETA, …) that map 1:1 to D1–D8. It can be reimplemented in Python / TypeScript / Go by reading the source — the most portable route for cross-language agent loading.

Path E — MCP client (Claude Desktop / any MCP host): ✅ 已实测。cmd/service 暴露 /mcp 端点(MCP spec 2025-11-25 Streamable HTTP),25 个引擎能力映射为 MCP tools(fuzzy_match / add_tm / check_terms / concordance / qe_auto / predict / observe / recall / explain / reward / consolidate / tm_count / ping / retrieve_prompt / bleu / chrf / style_check / style_report / back_align / term_conflicts / fed_export / fed_import / distill_inject / active_learning / mqm_re_annotate)。Claude Desktop 配置:

{ "mcpServers": { "yimai": { "url": "http://127.0.0.1:8787/mcp" } } }

moon run cmd/service --target native 起服务,MCP 客户端即可经 initialize → tools/list → tools/call 消费全部引擎能力(详见 MCP Server)。


#Contributing

Issues and pull requests are welcome. The repo ships multiple test suites — please keep moon test --target wasm-gc green when you submit a change. For behavioural/evaluation changes, extend yimai_prophecy_moonbit_modern_corpus_test.mbt with real modern corpus so the evaluation stays honest.


#License

MIT — see LICENSE.


#Acknowledgements

  • Built on the MoonBit language and its core (json, math) packages.
  • README structure follows conventions of high-star open-source projects (e.g. sharkdp/bat, BurntSushi/ripgrep) and the idiomatic MoonBit library style of moonbit-community/moon_elk.
  • 完整架构方案(评审叙事 + 工程附录 13 章):译脉·先知2.0_完整架构方案.md(黑客松定位、三层架构、确定性/零依赖/白盒三大卖点、S1 实证、增强路线图 S/M/L → 端点映射)。

MemoryNode

pub struct MemoryNode {
id : String
text : String
mtype : String
vec : Map[String, Double]
created : Double
last_used : Double
use_count : Int
feedback : Double
edges : Map[String, Double]
predictive_value : Double
hit_count : Int
predict_count : Int
last_active : Double
is_term : Bool
translation : String
tm_toks : Array[String]
tm_ngrams : Map[String, Bool]
tm_toks_set : Map[String, Bool]
}

ProphecyEngine

pub struct ProphecyEngine {
memories : Map[String, MemoryNode]
transitions : Map[String, Map[String, Double]]
episodes : Array[Array[String]]
context : Array[String]
stats_preds : Int
stats_hits : Int
stats_remembers : Int
stats_evolutions : Int
seq : Int
last_pred : Array[String]
cur_episode : Array[String]
explore : Double
meta_hits : Array[Int]
snapshot : Map[String, MemoryNode]?
snap_trans : Map[String, Map[String, Double]]?
snap_role : Map[String, Map[String, Double]]?
role_trans : Map[String, Map[String, Double]]
role_index : Map[String, String]
role_members : Map[String, Array[String]]
trans2 : Map[String, Map[String, Double]]
domain_bias : Map[String, Double]
hebb_lr : Double
attn_alpha : Double
attn_beta : Double
pred_cache : Map[String, Json]
cache_epoch : Int
cl_buf : Array[Json]
wal_log : Array[String]
fed_add : Int
fed_upd : Int
clock : Double
tm_df : Map[String, Int]
tm_idf : Map[String, Double]
tm_idf_dirty : Bool
tm_count : Int
tm_postings : Map[String, Array[String]]
term_idx : Map[String, Array[String]]
term_idx_dirty : Bool
}

ProphecyEngine::active_learning_candidates

fn ProphecyEngine::active_learning_candidates(self : ProphecyEngine, k : Int) -> Array[Json]

ProphecyEngine::add_tm

fn ProphecyEngine::add_tm(self : ProphecyEngine, src : String, tgt : String) -> String

ProphecyEngine::align_regions

fn ProphecyEngine::align_regions(self : ProphecyEngine, regions : Array[Json], threshold : Double) -> Json

ProphecyEngine::back_align

fn ProphecyEngine::back_align(_self : ProphecyEngine, source : String, target : String) -> Json

ProphecyEngine::batch_apply

fn ProphecyEngine::batch_apply(self : ProphecyEngine, segments : Array[String], threshold : Double) -> Json

ProphecyEngine::check_format_fidelity

fn ProphecyEngine::check_format_fidelity(_self : ProphecyEngine, source : String, target : String) -> Json

ProphecyEngine::check_terms

fn ProphecyEngine::check_terms(self : ProphecyEngine, source : String, target : String) -> Json

ProphecyEngine::cl_step

fn ProphecyEngine::cl_step(self : ProphecyEngine, anchor : String, positive : String, negative : String) -> Unit

ProphecyEngine::concordance

fn ProphecyEngine::concordance(self : ProphecyEngine, term : String, k : Int) -> Json

ProphecyEngine::consolidate

fn ProphecyEngine::consolidate(self : ProphecyEngine, prune : Bool) -> Json

ProphecyEngine::context_texts

fn ProphecyEngine::context_texts(self : ProphecyEngine) -> Array[String]

ProphecyEngine::drift_report

fn ProphecyEngine::drift_report(_self : ProphecyEngine, before : Json, after : Json) -> Json

ProphecyEngine::end_episode

fn ProphecyEngine::end_episode(self : ProphecyEngine) -> Unit

ProphecyEngine::enforce_terms

fn ProphecyEngine::enforce_terms(self : ProphecyEngine, text : String) -> Json

ProphecyEngine::explain

fn ProphecyEngine::explain(self : ProphecyEngine, mid : String) -> Json

ProphecyEngine::explain_card

fn ProphecyEngine::explain_card(self : ProphecyEngine, mid : String) -> Json

ProphecyEngine::export_tmx

fn ProphecyEngine::export_tmx(self : ProphecyEngine, src_lang? : String, tgt_lang? : String) -> String

ProphecyEngine::fed_export

fn ProphecyEngine::fed_export(self : ProphecyEngine) -> Json

ProphecyEngine::fed_import

fn ProphecyEngine::fed_import(self : ProphecyEngine, added : Int, updated : Int) -> Unit

ProphecyEngine::from_json

fn ProphecyEngine::from_json(data : Json) -> ProphecyEngine

ProphecyEngine::fuzzy_match

fn ProphecyEngine::fuzzy_match(self : ProphecyEngine, query : String, k : Int, threshold : Double, weights? : (Double, Double, Double, Double)) -> Json

ProphecyEngine::fuzzy_match_full

fn ProphecyEngine::fuzzy_match_full(self : ProphecyEngine, query : String, k : Int, threshold : Double) -> Json

ProphecyEngine::fuzzy_match_legacy

fn ProphecyEngine::fuzzy_match_legacy(self : ProphecyEngine, query : String, k : Int, threshold : Double) -> Json

ProphecyEngine::hit_rate

fn ProphecyEngine::hit_rate(self : ProphecyEngine) -> Double

ProphecyEngine::inject_distillation

fn ProphecyEngine::inject_distillation(self : ProphecyEngine, table : Map[String, Double]) -> Unit

ProphecyEngine::last_context_id

fn ProphecyEngine::last_context_id(self : ProphecyEngine) -> String

ProphecyEngine::load_tbx

fn ProphecyEngine::load_tbx(self : ProphecyEngine, xml : String, src_lang? : String, tgt_lang? : String) -> Int

ProphecyEngine::make

ProphecyEngine::mark_term

fn ProphecyEngine::mark_term(self : ProphecyEngine, mid : String) -> Bool

ProphecyEngine::metrics

fn ProphecyEngine::metrics(self : ProphecyEngine) -> Json

ProphecyEngine::mqm_re_annotate

fn ProphecyEngine::mqm_re_annotate(self : ProphecyEngine, source : String, target : String, match_rate : Double) -> Json

ProphecyEngine::mqm_tags

fn ProphecyEngine::mqm_tags(_self : ProphecyEngine, source : String, target : String, _match_rate : Double, term_ok : Bool) -> Json

ProphecyEngine::observe

fn ProphecyEngine::observe(self : ProphecyEngine, text : String, mtype : String) -> String

ProphecyEngine::parse_tmx

fn ProphecyEngine::parse_tmx(self : ProphecyEngine, xml : String) -> Int

ProphecyEngine::parse_xliff

fn ProphecyEngine::parse_xliff(self : ProphecyEngine, xml : String) -> Int

ProphecyEngine::predict

fn ProphecyEngine::predict(self : ProphecyEngine, k : Int) -> Json

ProphecyEngine::protect_tags

fn ProphecyEngine::protect_tags(_self : ProphecyEngine, text : String) -> String

ProphecyEngine::qe_auto

fn ProphecyEngine::qe_auto(self : ProphecyEngine, source : String, target : String, match_rate : Double) -> Json

ProphecyEngine::qe_score

fn ProphecyEngine::qe_score(_self : ProphecyEngine, source : String, target : String, match_rate : Double, term_ok : Bool) -> Double

ProphecyEngine::recall

fn ProphecyEngine::recall(self : ProphecyEngine, query : String, k : Int) -> Array[Json]

ProphecyEngine::remember

fn ProphecyEngine::remember(self : ProphecyEngine, text : String, mtype : String, ctx : Array[String]) -> String

ProphecyEngine::restore

fn ProphecyEngine::restore(self : ProphecyEngine) -> Json

ProphecyEngine::retrieve_for_prompt

fn ProphecyEngine::retrieve_for_prompt(self : ProphecyEngine, query : String, k : Int, threshold : Double) -> Json

ProphecyEngine::reward

fn ProphecyEngine::reward(self : ProphecyEngine, mid : String, score_in : Double) -> Bool

ProphecyEngine::set_attention

fn ProphecyEngine::set_attention(self : ProphecyEngine, alpha : Double, beta : Double) -> Unit

ProphecyEngine::set_domain_bias

fn ProphecyEngine::set_domain_bias(self : ProphecyEngine, role : String, delta : Double) -> Unit

ProphecyEngine::stats_view

fn ProphecyEngine::stats_view(self : ProphecyEngine) -> Json

ProphecyEngine::style_check

fn ProphecyEngine::style_check(self : ProphecyEngine, text : String) -> Json

ProphecyEngine::style_report

fn ProphecyEngine::style_report(self : ProphecyEngine, text : String) -> Json

ProphecyEngine::term_conflicts

fn ProphecyEngine::term_conflicts(self : ProphecyEngine) -> Json

ProphecyEngine::to_json

fn ProphecyEngine::to_json(self : ProphecyEngine) -> Json

ProphecyEngine::wal_clear

fn ProphecyEngine::wal_clear(self : ProphecyEngine) -> Unit

ProphecyEngine::wal_compact

fn ProphecyEngine::wal_compact(self : ProphecyEngine, keep : Int) -> Unit

ProphecyEngine::wal_export

fn ProphecyEngine::wal_export(self : ProphecyEngine) -> Array[String]

ProphecyEngine::wal_len

fn ProphecyEngine::wal_len(self : ProphecyEngine) -> Int

ProphecyEngine::wal_replay

fn ProphecyEngine::wal_replay(self : ProphecyEngine) -> ProphecyEngine

align_diff

fn align_diff(a : String, b : String) -> Array[(Int, Int, Int)]

api_version

let api_version : String

arr_json

fn arr_json(items : Array[Json]) -> Json

baseline_paths

let baseline_paths : Array[String]

bleu_score

fn bleu_score(ref_text : String, hyp_text : String) -> Double

char_ratio

fn char_ratio(a : String, b : String) -> Double

chrf_score

fn chrf_score(ref_text : String, hyp_text : String) -> Double

decode_pct

fn decode_pct(s : String) -> String

error_codes

let error_codes : Array[String]

get_field

fn get_field(j : Json, key : String) -> Json

get_num

fn get_num(j : Json, key : String) -> Double

get_obj

fn get_obj(j : Json, key : String) -> Json

get_str

fn get_str(j : Json, key : String) -> String

levenshtein

fn levenshtein(a : String, b : String) -> Int

mqm_issue

fn mqm_issue(dimension : String, severity : String, detail : String) -> Json

mqm_severity_to_score

fn mqm_severity_to_score(sev : String) -> Double

num_json

fn num_json(d : Double) -> Json

numeric_tokens

fn numeric_tokens(text : String) -> Array[String]

obj

fn obj(pairs : Array[(String, Json)]) -> Json

ocr_image

fn ocr_image(_path : String) -> Array[Json]
外部边界:OCR 识别必须由外部视觉服务完成(Tesseract / 视觉大模型), 零依赖 MoonBit 引擎不内嵌 OCR。此桩返回空;调用方应替换为真实 OCR 调用, 返回的每个元素为 JSON 对象:{ "bbox": [x,y,w,h], "text": "..." }

role_of

fn role_of(text : String) -> String

roles_of

fn roles_of(text : String) -> Array[String]

routes_meta

let routes_meta : Array[(String, String)]

split_into_segments

fn split_into_segments(text : String, max_chars : Int) -> Array[String]

str_contains

fn str_contains(s : String, sub : String) -> Bool

str_json

fn str_json(s : String) -> Json

wal_escape_field

fn wal_escape_field(s : String) -> String

wal_unescape_field

fn wal_unescape_field(s : String) -> String

yimai_tokenize

fn yimai_tokenize(text : String) -> Array[String]