Sign in

    zxcvbn

    MoonBit port of Dropbox zxcvbn password strength estimator: attack-model scoring (guesses/entropy), dictionary/l33t/keyboard/date/sequence/repeat pattern matching, crack-time estimation.

    password
    strength
    entropy
    zxcvbn
    security
    crack-time
    Download zip
    Author
    Version
    0.1.0
    License
    MIT
    Last updated
    10 hours ago
    Downloads
    2

    #PaperY0/zxcvbn

    Dropbox zxcvbn 密码强度估计器的 MoonBit 移植。 输入密码字符串,输出 0–4 强度分 + 熵值(guesses/log10)+ 四场景破解时间 + 命中的模式明细 + 改进建议。

    #为什么是它

    MoonBit 生态里没有一个攻击模型驱动的密码强度估计器:现存 4 个名字带"密码强度"的库 (Q30399/moonvault、GMH233/moonpassword、ZijianFeng33/strongPasswordChecker、 kesmeey/tools)经逐个读源码确认,全部是 LUDS 字符计数式规则打分(长度档位 + 字符类别 存在性 + 重复扣分)——没有熵、没有词典、没有 l33t、没有键盘邻接、没有日期/序列/重复模式、 没有破解时间。zxcvbn 模拟真实破解器:对密码做全模式匹配后取最小猜测数,输出保守的强度 估计。三轮查重(mooncakes 2622 模块全量扫描 + GitHub 10 组检索 + awesome-moonbit + 官方 moon search zxcvbn → No modules found)确认这里是空白。

    #与 LUDS 规则打分的实证边界

    这不是"算法更高级"的口号,而是可复现的方向性错误。下表是本库的真实输出 (moon run cmd/main -- <password>),LUDS 列是典型 LUDS 实现(长度档位 + 字符类别计分) 会给出的分数:

    密码典型 LUDS 打分本库 zxcvbn真实情况
    Password1!4(4 类齐全、长度 10)1Password 是 rank 2 的高频词典词,大小写与符号几乎不增加成本
    correcthorsebatterystaple3(仅小写、无数字符号)44 个普通英文词、无任何模式,真正的离线破解下限是 2.7e14 次猜测
    qwerty1234563(长度 12、含数字)1键盘直行(qwerty)+ 数字序列(123456),1851 次即可命中
    Tr0ub4dor&344确实强(1e11 次暴力破解),但这是"碰巧",不是规则算出来的

    LUDS 的两个系统性失败:把"字符类别齐全"当作强度(Password1!), 以及把"只有小写字母"当作弱点(correcthorsebatterystaple)。 本库的做法是枚举真实破解器会先用的东西——词典、l33t 变体、键盘走位、日期、序列、重复—— 然后取最小猜测数,所以结论可以和实际抗破解能力对上。

    #能力范围

    • 8 个模式匹配器:6 个频率词典(30k 常见密码 / 30k 维基词 / 3712 女名 / 983 男名 / 10k 姓氏 / 19160 影视词,共 93,855 词)、 反向词典、l33t 替换(17 条替换表 + 子集枚举)、键盘空间(qwerty/dvorak/keypad/mac_keypad)、 重复(abcabcabc)、序列(abc/123/97531)、正则(年份)、日期(多格式 + 分隔符)
    • 最优匹配序列 DP:动态规划取最小 guesses 的非重叠序列
    • 输出:guesses / log10(guesses) / 0–4 分 / 四场景破解秒数(在线限流 100/h、在线 不限流 10/s、离线慢哈希 1e4/s、离线快哈希 1e10/s)/ warning + suggestions

    #API 设计

    pub fn zxcvbn(
    password : String,
    user_inputs : Array[String],
    reference_year? : Int = 2026,
    ) -> Entropy

    pub struct Entropy {
    password : String
    /// 保守估计的猜测次数(Double:长密码会达到 1e10^n 甚至上溢)
    guesses : Double
    /// log10(guesses)
    guesses_log10 : Double
    /// 0..4,阈值 1e3 / 1e6 / 1e8 / 1e10(与上游一致)
    score : Int
    /// 四场景破解秒数
    crack_times_seconds : CrackTimes
    /// 人类可读破解时间
    crack_times_display : CrackTimesDisplay
    /// 最优匹配序列
    sequence : Array[Match]
    feedback : Feedback
    }

    /// 一次模式匹配(按模式拆成 enum + 每模式 struct,仿 zxcvbn-rs 的强类型范式)
    pub enum Match {
    Bruteforce(BruteforceMatch)
    Dictionary(DictionaryMatch)
    Spatial(SpatialMatch)
    Repeat(RepeatMatch)
    Sequence(SequenceMatch)
    Regex(RegexMatch)
    Date(DateMatch)
    }
    pub fn Match::start_index(self) -> Int // 闭区间起点
    pub fn Match::end_index(self) -> Int // 闭区间终点
    pub fn Match::token(self) -> String // 命中的原文字符串
    pub fn Match::pattern(self) -> String // "dictionary" / "spatial" / ...
    pub fn Match::guesses(self) -> Double

    低层匹配器也可单独调用(上游签名的对应对齐,便于测试与二次封装):

    pub fn dictionary_match(Array<Char>, Array<(Dictionary, Map<String, Int])>, Int) -> Array<Match>
    pub fn reverse_dictionary_match(Array<Char>, Array<(Dictionary, Map<String, Int])>, Int) -> Array[Match]
    pub fn l33t_match(Array<Char], Array<(Dictionary, Map<String, Int])], Map<Char, Array<Char]], Int) -> Array<Match>
    pub fn spatial_match(Array<Char], Array<(String, Map<Char, Array<(Char, Char)?]])>) -> Array[Match]
    pub fn repeat_match(Array<Char], Array<(Dictionary, Map<String, Int])], Int, Map<Char, Array<Char]], Int) -> Array[Match]
    pub fn sequence_match(Array<Char]) -> Array<Match>
    pub fn regex_match(Array[Char]) -> Array<Match]
    pub fn date_match(Array<Char>, Int) -> Array<Match>

    #三个使用场景

    1. 注册/改密表单(服务端):zxcvbn(pw, [username, email]) —— 用户输入参与词典, 防止"用自己名字当密码"。
    2. CLI 工具:moon run cmd/main -- "Tr0ub4dor&3" 打印分数/熵/破解时间/模式明细/建议; --json 输出机器可读结果(供差分对拍消费)。
    3. 浏览器 Wasm 强度条:MoonBit 纯计算库直接编 wasm-gc,前端注册表单即时打分,密码不回传。

    #与上游的刻意差异(都是修正,逐条说明)

    1. 下标按 Unicode 标量值计,而非 UTF-16 码元。 上游在含非 BMP 字符的密码上有过经典的 UTF-16 边界 bug(HIV 密码一例);本实现用 MoonBit 的 Char(即标量值)遍历, 从根上避免。代价:token 长度与上游 String::length() 在含 emoji 时不同。
    2. guesses 用 Double 而非整数。 上游 JS 全程是浮点数;长密码的 bruteforce 猜测数会 达到 10^n 乃至上溢,用 Int 会静欲溢出。zxcvbn-rs 同样用 f64。
    3. reference_year 参数化(默认 2026)。 上游取运行时当年,MoonBit core 没有时钟 API, 故参数化——同时让测试完全确定。
    4. recent_year 正则升级为 19\d\d|20\d\d。 上游的正则是 2017 年的 19\d\d|200\d|201\d, 识别不出 2020 年之后的年份;这里沿用 zxcvbn-rs 已更新的范围。 (附带:MoonBit core 的 Regex 不支持 \d,须用 POSIX 类 [[:digit:]], 而正则也没有反向引用 \1——所以本项目的 repeat_match 用手写的周期检测, regex_match 用手写的四位年扫描,两者在语义上与上游等价。)
    5. 词典匹配的长度上限裁剪。 只枚举到"所有词典中最长词条"的长度:更长的子串不可能命中, 语义完全等价,但把词典匹配从 O(n²) 次哈希查询降到 O(n × 最长词长)。
    6. 无全局状态。 上游用模块级变量传递用户输入词典;本实现改为调用时传入,天然纯函数、 可重入、线程安全。

    #测试策略

    53 个测试全部通过,来源三层:

    1. 上游官方向量(test/test-matching.coffee 10 个块 + test/test-scoring.coffee 12 个块) 逐块转写:词典/反向/l33t(含子表与子集枚举)/键盘空间/序列/重复/正则/日期/omnimatch, 以及 nCk、log、DP 搜索、各模式 guesses、uppercase/l33t 变体。官方向量里的 genpws prefix/suffix 变体矩阵也一并转写。
    2. 性质与端到端:r0sebudmaelstrom11/20/91aaaa 精确复现上游 omnimatch 的四大匹配; correcthorsebatterystaple → score 4。
    3. 边界:空串、单字符、纯空白、1000 长度密码、非 BMP emoji(验证按标量值计长)、 user_inputs 命中自身、reference_year 参数化。

    moon check # 0 error,0 warning moon test # 53/53(默认 wasm 后端) # 三个后端都全绿(验证“纯计算、无 IO 依赖”的声明) moon test --target wasm # 53/53 moon test --target wasm-gc # 53/53 moon test --target js # 53/53 moon run cmd/main -- "correct horse battery staple"

    #非目标(明确不做)

    多语言词典、口令哈希(用 moonbitlang/x/bcrypt,互补)、分布式破解模拟、完整 UI 组件库、 Grapheme 级 Unicode 处理(按 Unicode 标量值遍历)。

    #开发

    moon check # 静态检查 moon info # 更新 .mbti 接口 moon fmt # 格式化 moon test # 测试 moon run cmd/main -- "password" python3 tools/gen_dictionaries.py # 从 data/upstream 重新生成词典数据模块 python3 tools/gen_adjacency_graphs.py # 从数据/upstream 重新生成键盘邻接图模块 # 两个生成器的产物需再跑一次 moon fmt(生成器不保证 fmt-clean,CI 的 fmt 门禁会校验)

    #数据与许可

    • 词典数据(6 个频率表,93,855 词条)与官方向量来自 dropbox/zxcvbn 的 src/frequency_lists.coffee(上游运行时实际加载的过滤后词典,与 zxcvbn-rs 逐词一致), MIT License, (c) Dropbox, Inc.,见 data/README.md。
    • 结构参考 shssoichiro/zxcvbn-rs(MIT)。
    • 本项目以 MIT 发布,LICENSE 保留上游版权声明。

    #当前状态(2026-09-30)

    ✅ 评分引擎完整实现(8 匹配器 + DP + 每模式 guesses + 破解时间 + 建议)· ✅ 上游 22 个官方向量 test 块全部转写通过 · ✅ 53/53 测试通过 · ✅ moon check 0 error / 0 warning · ✅ CLI 输出完整结果与 JSON · ✅ 键盘邻接图数据由脚本生成(可复现)· ✅ CI:wasm / wasm-gc / js 三后端测试矩阵 + fmt/check/接口校验门禁 · 🚧 Phase 3 剩余:DIFF-REPORT.md 系统性差分对拍、mooncakes 发布、demo/ wasm 演示页。


    #PaperY0/zxcvbn (EN)

    A MoonBit port of Dropbox's zxcvbn password strength estimator: pattern-matching based (dictionaries, l33t, keyboard adjacency, dates, sequences, repeats), minimum-guesses DP scoring, 0–4 score, crack-time estimation, and actionable feedback. Pure computation, no IO; the wasm / wasm-gc / js targets are verified in CI (the native target is also supported by the toolchain but not covered by CI).

    let result = @zxcvbn.zxcvbn("Tr0ub4dor&3", ["bob", "bob@example.com"])
    result.score // 0..4
    result.guesses // conservative attack-model lower bound (Double)
    result.sequence // optimal non-overlapping matches
    result.feedback // warning + suggestions

    Dictionary data and official test vectors are from dropbox/zxcvbn (MIT, (c) Dropbox, Inc.); structural reference: shssoichiro/zxcvbn-rs (MIT).

    AttackTimes

    pub struct AttackTimes {
    crack_times_seconds : CrackTimes
    crack_times_display : CrackTimesDisplay
    score : Int
    } derive(Eq,
    Debug
    )

    上游 time_estimates.estimate_attack_times 的完整返回。 注意 crack_times_display 是人类可读字符串,与秒数结构分开。

    AttackTimes::equal

    fn AttackTimes::equal(AttackTimes, AttackTimes) -> Bool

    AttackTimes::not_equal

    fn AttackTimes::not_equal(x : AttackTimes, y : AttackTimes) -> Bool

    BruteforceMatch

    pub struct BruteforceMatch {
    i : Int
    j : Int
    token : String
    guesses : Double
    guesses_log10 : Double
    } derive(Eq,
    Debug
    )

    暴力匹配的区间。

    BruteforceMatch::equal

    BruteforceMatch::not_equal

    fn BruteforceMatch::not_equal(x : BruteforceMatch, y : BruteforceMatch) -> Bool

    CrackTimes

    pub struct CrackTimes {
    online_throttling_100_per_hour : Double
    online_no_throttling_10_per_second : Double
    offline_slow_hashing_1e4_per_second : Double
    offline_fast_hashing_1e10_per_second : Double
    } derive(Eq,
    Debug
    )

    四类破解场景的耗时(秒)。

    对应上游的四档攻击模型:在线限流(100 次/小时)、在线不限流(10 次/秒)、 离线慢哈希(1e4 次/秒)、离线快哈希(1e10 次/秒)。

    CrackTimes::equal

    fn CrackTimes::equal(CrackTimes, CrackTimes) -> Bool

    CrackTimes::not_equal

    fn CrackTimes::not_equal(x : CrackTimes, y : CrackTimes) -> Bool

    CrackTimesDisplay

    pub struct CrackTimesDisplay {
    online_throttling_100_per_hour : String
    online_no_throttling_10_per_second : String
    offline_slow_hashing_1e4_per_second : String
    offline_fast_hashing_1e10_per_second : String
    } derive(Eq,
    Debug
    )

    四类破解场景的人类可读耗时。

    CrackTimesDisplay::equal

    CrackTimesDisplay::not_equal

    fn CrackTimesDisplay::not_equal(x : CrackTimesDisplay, y : CrackTimesDisplay) -> Bool

    DateMatch

    pub struct DateMatch {
    i : Int
    j : Int
    token : String
    separator : String
    year : Int
    month : Int
    day : Int
    guesses : Double
    guesses_log10 : Double
    } derive(Eq,
    Debug
    )

    日期匹配。上游不做真日期解析(feb 31st 合法、不查闰年)。

    DateMatch::equal

    fn DateMatch::equal(DateMatch, DateMatch) -> Bool

    DateMatch::not_equal

    fn DateMatch::not_equal(x : DateMatch, y : DateMatch) -> Bool

    Dictionary

    pub enum Dictionary {
    Passwords
    English
    FemaleNames
    MaleNames
    Surnames
    UsTvAndFilm
    UserInputs
    } derive(Eq,
    Debug
    )

    词典标识。命名与 zxcvbn-rs 的 DictionaryType 对齐,便于差分对拍时对照。

    Dictionary::equal

    fn Dictionary::equal(Dictionary, Dictionary) -> Bool

    Dictionary::not_equal

    fn Dictionary::not_equal(x : Dictionary, y : Dictionary) -> Bool

    DictionaryMatch

    pub struct DictionaryMatch {
    i : Int
    j : Int
    token : String
    matched_word : String
    rank : Int
    dictionary_name : Dictionary
    reversed : Bool
    l33t : Bool
    sub : Array[(Char, Char)]
    guesses : Double
    guesses_log10 : Double
    } derive(Eq,
    Debug
    )

    词典匹配。rank 为 1-based 行序;l33t/reversed 标记变体; sub 记录该匹配实际用到的 l33t 替换(subbed_char → 原字母)。

    DictionaryMatch::equal

    DictionaryMatch::not_equal

    fn DictionaryMatch::not_equal(x : DictionaryMatch, y : DictionaryMatch) -> Bool

    Entropy

    pub struct Entropy {
    password : String
    guesses : Double
    guesses_log10 : Double
    score : Int
    crack_times_seconds : CrackTimes
    crack_times_display : CrackTimesDisplay
    sequence : Array[Match]
    feedback : Feedback
    } derive(Eq,
    Debug
    )

    zxcvbn 评估结果。

    guesses / guesses_log10 / score 是保守的攻击模型下限估计: zxcvbn 不是"字符类别计数",而是做完全模式匹配后取最小猜测数。

    Entropy::equal

    fn Entropy::equal(Entropy, Entropy) -> Bool

    Entropy::not_equal

    fn Entropy::not_equal(x : Entropy, y : Entropy) -> Bool

    Entropy::to_repr

    Feedback

    pub struct Feedback {
    warning : String
    suggestions : Array[String]
    } derive(Eq,
    Debug
    )

    可执行的改进建议(与上游 feedback 对齐)。

    Feedback::equal

    fn Feedback::equal(Feedback, Feedback) -> Bool

    Feedback::not_equal

    fn Feedback::not_equal(x : Feedback, y : Feedback) -> Bool

    Feedback::to_repr

    Match

    pub enum Match {
    Bruteforce(BruteforceMatch)
    Dictionary(DictionaryMatch)
    Spatial(SpatialMatch)
    Repeat(RepeatMatch)
    Sequence(SequenceMatch)
    Regex(RegexMatch)
    Date(DateMatch)
    } derive(Eq,
    Debug
    )

    一次模式匹配。

    上游用动态对象承载所有字段;这里按 zxcvbn-rs 的强类型范式拆成 enum + 每模式 struct,让编译器捕获"某模式没有该字段"一类错误。

    Match::describe

    fn Match::describe(self : Match) -> String

    生成调试用的单行摘要(不含 guesses,避免噪声)。

    Match::end_index

    fn Match::end_index(self : Match) -> Int

    匹配终点(闭区间,Unicode 标量值下标)。

    Match::equal

    fn Match::equal(Match, Match) -> Bool

    Match::guesses

    fn Match::guesses(self : Match) -> Double

    已缓存的猜测数(0 表示尚未计算)。

    Match::guesses_log10

    fn Match::guesses_log10(self : Match) -> Double

    log10(guesses)(上游用于 feedback 文案)。

    Match::not_equal

    fn Match::not_equal(x : Match, y : Match) -> Bool

    Match::pattern

    fn Match::pattern(self : Match) -> String

    模式名(与上游 pattern 字符串一致)。

    Match::set_guesses

    fn Match::set_guesses(self : Match, guesses : Double) -> Unit

    写入猜测数缓存(DP 会多次调用 estimate_guesses,缓存避免重复计算)。

    Match::set_guesses_log10

    fn Match::set_guesses_log10(self : Match, log10 : Double) -> Unit

    写入 log10(guesses)。

    Match::span_length

    fn Match::span_length(self : Match) -> Int

    匹配长度(标量值个数)。

    Match::start_index

    fn Match::start_index(self : Match) -> Int

    匹配起点(闭区间,Unicode 标量值下标)。

    Match::to_repr

    Match::token

    fn Match::token(self : Match) -> String

    命中的原文字符串。

    MatchSequence

    pub struct MatchSequence {
    password : String
    guesses : Double
    guesses_log10 : Double
    sequence : Array[Match]
    } derive(Eq,
    Debug
    )

    上游 most_guessable_match_sequence 的结果。

    MatchSequence::equal

    MatchSequence::not_equal

    fn MatchSequence::not_equal(x : MatchSequence, y : MatchSequence) -> Bool

    Omnimatcher

    pub struct Omnimatcher {
    dictionaries : Array[(Dictionary, Map[String, Int])]
    max_word_length : Int
    reference_year : Int
    } derive(Eq,
    Debug
    )

    一次 omnimatch 需要的上下文(上游用模块级全局变量传递用户输入词典)。

    Omnimatcher::equal

    fn Omnimatcher::equal(Omnimatcher, Omnimatcher) -> Bool

    Omnimatcher::new

    fn Omnimatcher::new(user_inputs : Map[String, Int], reference_year : Int) -> Omnimatcher

    构造 omnimatch 上下文。user_inputs 已由调用方小写化并建表。

    Omnimatcher::not_equal

    fn Omnimatcher::not_equal(x : Omnimatcher, y : Omnimatcher) -> Bool

    Omnimatcher::run

    fn Omnimatcher::run(self : Omnimatcher, password : String) -> Array[Match]

    上游 omnimatch(password):跑完全部匹配器后按 sorted 排序。

    RegexMatch

    pub struct RegexMatch {
    i : Int
    j : Int
    token : String
    regex_name : String
    guesses : Double
    guesses_log10 : Double
    } derive(Eq,
    Debug
    )

    正则匹配。目前只有 recent_year。

    RegexMatch::equal

    fn RegexMatch::equal(RegexMatch, RegexMatch) -> Bool

    RegexMatch::not_equal

    fn RegexMatch::not_equal(x : RegexMatch, y : RegexMatch) -> Bool

    RepeatMatch

    pub struct RepeatMatch {
    i : Int
    j : Int
    token : String
    base_token : String
    base_guesses : Double
    base_matches : Array[Match]
    repeat_count : Int
    guesses : Double
    guesses_log10 : Double
    } derive(Eq,
    Debug
    )

    重复匹配。base_matches 是对 base_token 递归 omnimatch 的结果。

    RepeatMatch::equal

    fn RepeatMatch::equal(RepeatMatch, RepeatMatch) -> Bool

    RepeatMatch::not_equal

    fn RepeatMatch::not_equal(x : RepeatMatch, y : RepeatMatch) -> Bool

    SequenceMatch

    pub struct SequenceMatch {
    i : Int
    j : Int
    token : String
    sequence_name : String
    sequence_space : Int
    ascending : Bool
    guesses : Double
    guesses_log10 : Double
    } derive(Eq,
    Debug
    )

    序列匹配。sequence_space 为序列空间大小(lower/upper=26,digits=10)。

    SequenceMatch::equal

    SequenceMatch::not_equal

    fn SequenceMatch::not_equal(x : SequenceMatch, y : SequenceMatch) -> Bool

    SpatialMatch

    pub struct SpatialMatch {
    i : Int
    j : Int
    token : String
    graph : String
    turns : Int
    shifted_count : Int
    guesses : Double
    guesses_log10 : Double
    } derive(Eq,
    Debug
    )

    键盘空间匹配。turns 为转向次数,shifted_count 为按了 Shift 的键数。

    SpatialMatch::equal

    SpatialMatch::not_equal

    fn SpatialMatch::not_equal(x : SpatialMatch, y : SpatialMatch) -> Bool

    all_graphs

    fn all_graphs() -> Array[(String, Map[Char, Array[(Char, Char)?]])]

    全部四个键盘图(顺序与上游 GRAPHS 一致)。

    build_ranked_dict

    fn build_ranked_dict(words : Array[String]) -> Map[String, Int]

    由有序词表构建 rank 表。

    rank 为 1-based 行序;重复词后者覆盖前者——与上游 build_ranked_dict (JS 对象逐个赋值)及 zxcvbn-rs(HashMap collect)的语义保持一致。

    builtin_dictionaries

    let builtin_dictionaries : Array[(Dictionary,
    Lazy
    [Map[String, Int]])]

    全部内置频率词典,顺序与上游 RANKED_DICTIONARIES 的构造顺序一致。 匹配器(dictionary_match / reverse_dictionary_match / l33t_match)按这个顺序遍历。

    date_match

    fn date_match(password : Array[Char], reference_year : Int) -> Array[Match]

    上游 date_match。

    default_feedback

    let default_feedback : Feedback

    上游 feedback.default_feedback:没有任何可用建议时的兜底。

    dictionary_match

    fn dictionary_match(password : Array[Char], dictionaries : Array[(Dictionary, Map[String, Int])], max_word_length : Int) -> Array[Match]

    上游 dictionary_match。

    dictionary_match_inner

    fn dictionary_match_inner(password : Array[Char], password_lower : Array[Char], dictionaries : Array[(Dictionary, Map[String, Int])], max_word_length : Int) -> Array[Match]

    词典匹配主体。password_lower 由调用方传入(l33t 匹配要传"替换后"的小写形式)。

    english_wikipedia_list

    fn english_wikipedia_list() -> Array[String]

    英文维基常用词 的有序词表(rank = 下标 + 1)。

    estimate_attack_times

    fn estimate_attack_times(guesses : Double) -> AttackTimes

    上游 estimate_attack_times。

    factorial

    fn factorial(n : Int) -> Double

    上游 scoring.factorial。

    female_names_list

    fn female_names_list() -> Array[String]

    常见女性名 的有序词表(rank = 下标 + 1)。

    get_feedback

    fn get_feedback(score : Int, sequence : Array[Match]) -> Feedback

    上游 feedback.get_feedback。

    get_graph

    fn get_graph(name : String) -> Map[Char, Array[(Char, Char)?]]

    按图名取邻接图(spatial_match 的四个图)。图名与上游一致。

    graph_names

    let graph_names : Array[String]

    四个键盘图的名字(顺序与上游 GRAPHS 一致,决定 omnimatch 的匹配顺序)。

    guesses_to_score

    fn guesses_to_score(guesses : Double) -> Int

    上游 guesses_to_score:DELTA=5 的阈值分档。

    分档含义(原上游注释):
    • 0 too guessable:从在线限流攻击下也不安全
    • 1 very guessable:能防住在线限流攻击
    • 2 somewhat guessable:能防住不限流在线攻击
    • 3 safely unguessable:能防住离线慢哈希攻击(bcrypt/scrypt/PBKDF2/argon)
    • 4 very unguessable:同场景下更强的保护

    keyboard_average_degree

    let keyboard_average_degree : Double

    qwerty/dvorak 的平均度数。

    keyboard_starting_positions

    let keyboard_starting_positions : Int

    qwerty/dvorak 用的起始键位数。

    keypad_average_degree

    let keypad_average_degree : Double

    keypad/mac_keypad 的平均度数(上游注释:与 qwerty 略有不同,但取近似值)。

    keypad_starting_positions

    let keypad_starting_positions : Int

    keypad/mac_keypad 用的起始键位数。

    l33t_match

    fn l33t_match(password : Array[Char], dictionaries : Array[(Dictionary, Map[String, Int])], table : Map[Char, Array[Char]], max_word_length : Int) -> Array[Match]

    上游 l33t_match。

    log10

    fn log10(n : Double) -> Double

    上游 scoring.log10:Math.log(n) / Math.log(10)。

    log2

    fn log2(n : Double) -> Double

    上游 scoring.log2:Math.log(n) / Math.log(2)。

    male_names_list

    fn male_names_list() -> Array[String]

    常见男性名 的有序词表(rank = 下标 + 1)。

    max_dictionary_word_length

    let max_dictionary_word_length :
    Lazy
    [Int]

    内置词典中最长词条的字符数(按 Unicode 标量值计)。

    dictionary_match 只需枚举到这个长度:更长的子串不可能命中任何词条。 这是与上游语义完全等价的裁剪(上游枚举全部长度,结果必然为空), 但把词典匹配从 O(n²) 次哈希查询降到 O(n × 最长词长),长密码收益巨大。

    mod_impl

    fn mod_impl(n : Int, m : Int) -> Int

    上游 matching.mod:对负数也正确的取模。

    most_guessable_match_sequence

    fn most_guessable_match_sequence(password : String, matches : Array[Match], reference_year : Int, exclude_additive? : Bool) -> MatchSequence

    上游 scoring.most_guessable_match_sequence。

    把一组可能重叠的匹配,转换为非重叠且 guesses 最小的序列。 目标函数(与上游一致):

    g = l! * Prod(m.guesses for m in sequence) + D^(l - 1)

    其中 l 为序列长度,l! 是 l 个模式的排列数,D^(l-1) 是长度惩罚 (攻击者会先试更短的序列)。

    exclude_additive 仅供上游测试使用(测试里会把加性惩罚关掉)。

    nCk

    fn nCk(n : Int, k : Int) -> Double

    上游 scoring.nCk:组合数,官方向量 nCk(33,7) = 4272048。

    passwords_list

    fn passwords_list() -> Array[String]

    常见密码 的有序词表(rank = 下标 + 1)。

    ranked_lookup

    fn ranked_lookup(word : String) -> (Dictionary, Int)?

    在全部内置词典中查词,返回 (词典, rank)。

    查询顺序与上游 RANKED_DICTIONARIES 的构造顺序一致: passwords → english → female_names → male_names → surnames → us_tv_and_film。 惰性求值短路:在靠前的词典命中时不会强制构建后续词典。

    regex_match

    fn regex_match(password : Array[Char]) -> Array[Match]

    上游 regex_match。这里不用 Regex,直接手写等价扫描—— 一是 core 的正则不支持 \d,二是需要在 char 数组(标量值)下标上定位。

    relevant_l33t_subtable

    fn relevant_l33t_subtable(password : Array[Char], table : Map[Char, Array[Char]]) -> Map[Char, Array[Char]]

    上游 relevant_l33t_subtable:只保留密码里真的出现了的替换支。

    repeat_match

    fn repeat_match(password : Array[Char], dictionaries : Array[(Dictionary, Map[String, Int])], max_word_length : Int, table : Map[Char, Array[Char]], reference_year : Int) -> Array[Match]

    上游 repeat_match。

    reverse_dictionary_match

    fn reverse_dictionary_match(password : Array[Char], dictionaries : Array[(Dictionary, Map[String, Int])], max_word_length : Int) -> Array[Match]

    上游 reverse_dictionary_match。

    sequence_match

    fn sequence_match(password : Array[Char]) -> Array[Match]

    上游 sequence_match:靠相邻字符码点差不变来识别序列(支持跳跃与非拉丁字母)。

    sort_matches

    fn sort_matches(matches : Array[Match]) -> Array[Match]

    上游 matching.sorted:按 i 主序、j 次序排序。

    注意:这里不按 guesses 排序(上游同样如此)。DP 内部会按 j 分区再按 i 排序,所以全局顺序不影响结果,但保持确定性输出有助于测试与调试。

    spatial_match

    fn spatial_match(password : Array[Char], graphs : Array[(String, Map[Char, Array[(Char, Char)?]])]) -> Array[Match]

    上游 spatial_match。

    surnames_list

    fn surnames_list() -> Array[String]

    常见姓氏 的有序词表(rank = 下标 + 1)。

    to_ascii_lower

    fn to_ascii_lower(chars : Array[Char]) -> Array[Char]

    frequency_lists.coffee 在 ASCII 小写字母域内过滤,但词典里确实存在少量 非 ASCII 词条。这里把它们转成 ASCII 小写以对齐 Array[Char] 的匹配口径。

    us_tv_and_film_list

    fn us_tv_and_film_list() -> Array[String]

    美剧/电影词 的有序词表(rank = 下标 + 1)。

    zxcvbn

    fn zxcvbn(password : String, user_inputs : Array[String], reference_year? : Int) -> Entropy

    评估一个密码的强度。

    • user_inputs:与用户相关的词(用户名/邮箱等),命中时显著降低强度估计;
    • reference_year:日期匹配的参考年份。上游取运行时当年,MoonBit core 没有 时钟 API,故参数化并默认 2026,同时让测试完全确定。

    移植来源:dropbox/zxcvbn(MIT)。

    zxcvbn_password

    fn zxcvbn_password(password : String) -> Entropy

    便捷入口:不带用户输入、用默认参考年。