moon_scrub

    Offline sensitive-data detection, redaction, JSON and streaming pipelines for MoonBit applications.

    security
    redaction
    secrets
    pii
    logs
    streaming
    json
    detection
    Download zip
    Version
    0.3.0
    License
    Apache-2.0
    Last updated
    9 hours ago
    Downloads
    1

    #Moon Scrub

    Offline sensitive-data detection and redaction for MoonBit applications. Findings contain category, confidence, and UTF-16 offsets, but never retain the matched value.

    Version 0.1.0 is published on MoonCakes. Install it with moon add 123123213weqw/moon_scrub.

    #Public API

    Scanning and redaction:

    • scan(text, policy~) -> Array[Finding]
    • scan_with_config(text, config) -> Array[Finding]
    • scan_with_rules(text, rules, policy~) -> Array[Finding]
    • scan_full(text, rules, config) -> Array[Finding]
    • redact(text, style~, policy~) -> RedactionResult
    • redact_with_config(text, config, style~) -> RedactionResult
    • redact_full(text, rules, config, style~) -> RedactionResult
    • redact_with_rules(text, rules, style~, policy~) -> RedactionResult
    • verify_clean(text, policy~) -> Bool
    • verify_clean_with_config(text, config) -> Bool
    • redact_batch(lines, style~, policy~) -> BatchResult
    • redact_batch_with_config(lines, config, style~) -> BatchResult
    • redact_fields(fields, style~, policy~) -> StructuredResult
    • redact_fields_with_config(fields, config, style~) -> StructuredResult
    • redact_json(text, config~, style~) -> JsonRedactionResult
    • redact_json_with_rules(text, rules, config~, style~) -> JsonRedactionResult

    Streaming:

    • ChunkScanner::new(config~, overlap~)
    • ChunkScanner::push(chunk) -> Array[Finding]
    • ChunkScanner::push_lines(lines) -> Array[Finding]
    • ChunkScanner::finish() -> Array[Finding]

    Reporting and exceptions:

    • findings_summary(text, config~) -> Array[KindCount]
    • explain(text, config~) -> Array[String]
    • scan_except(text, keep, config~) -> Array[Finding]
    • redact_except(text, keep, style~, config~) -> RedactionResult

    Configuration:

    • ScanPolicy::secrets_only() and ScanPolicy::standard()
    • ScanConfig::secrets_only() and ScanConfig::standard(): per-family switches, false-positive suppressor toggles, and extra TokenPrefixRule prefixed-token rules
    • CustomRule exact-value rules

    Detectors cover structured access tokens across GitHub, GitLab, Slack, AWS, Google, Stripe, Anthropic, OpenAI-style, SendGrid, Doppler and npm prefixes, bearer and Basic credentials, JWTs, sensitive assignments (including JSON-style quoted keys), webhook URLs, email addresses, IPv4 and IPv6, Luhn-valid payment cards with major-network BIN prefixes (Amex spacing included), Chinese resident identity numbers with checksum validation, MAC addresses, UUIDs, SSH public keys and certificate blocks, wallet addresses, and opt-in E.164 phone numbers. This is a deterministic helper, not a complete DLP or compliance system. See the repository README and security model for limitations.

    BatchResult

    pub(all) struct BatchResult {
    lines : Array[String]
    summary : BatchSummary
    changed_indices : Array[Int]
    } derive(Eq,
    Debug
    )

    BatchResult::equal

    fn BatchResult::equal(BatchResult, BatchResult) -> Bool

    BatchResult::not_equal

    fn BatchResult::not_equal(x : BatchResult, y : BatchResult) -> Bool

    BatchSummary

    pub(all) struct BatchSummary {
    lines : Int
    changed_lines : Int
    findings : Int
    access_tokens : Int
    bearer_tokens : Int
    jwts : Int
    credentials : Int
    emails : Int
    ipv4s : Int
    payment_cards : Int
    private_keys : Int
    url_credentials : Int
    custom_secrets : Int
    resident_ids : Int
    mac_addresses : Int
    uuids : Int
    ipv6s : Int
    phones : Int
    public_keys : Int
    wallet_addresses : Int
    webhook_urls : Int
    } derive(Eq,
    Debug
    )

    BatchSummary::equal

    BatchSummary::not_equal

    fn BatchSummary::not_equal(x : BatchSummary, y : BatchSummary) -> Bool

    ChunkScanner

    type ChunkScanner

    Streaming scanner for input that arrives in chunks: log tailing, socket reads, or prompt fragments. Findings carry absolute offsets into the whole stream, never into the current chunk.

    The scanner keeps an overlap-sized tail of unseen data so secrets that straddle a chunk boundary are still detected. Correctness holds for any secret whose full span is no longer than overlap; longer spans can be split and missed, so pick overlap above the largest expected secret (PEM private-key blocks are the usual ceiling; the default fits 8 KB keys).

    ChunkScanner::finish

    fn ChunkScanner::finish(self : ChunkScanner) -> Array[Finding]

    Flush the remaining tail and close the stream. Chunks pushed afterwards are ignored.

    ChunkScanner::new

    fn ChunkScanner::new(config? : ScanConfig, overlap? : Int) -> ChunkScanner

    Build a chunked scanner. overlap is clamped to at least 64 code units.

    ChunkScanner::push

    fn ChunkScanner::push(self : ChunkScanner, chunk : String) -> Array[Finding]

    Consume one chunk and return every finding whose end position is settled: at least overlap code units of later data exist, so no future chunk can extend or start a match overlapping it.

    ChunkScanner::push_lines

    fn ChunkScanner::push_lines(self : ChunkScanner, lines : Array[String]) -> Array[Finding]

    Push complete log lines with their newline included and collect settled findings. Equivalent to one push per line, but a single buffer append.

    Confidence

    pub(all) enum Confidence {
    High
    Medium
    } derive(Eq,
    Debug
    )

    Confidence::equal

    fn Confidence::equal(Confidence, Confidence) -> Bool

    Confidence::not_equal

    fn Confidence::not_equal(x : Confidence, y : Confidence) -> Bool

    Confidence::to_string

    fn Confidence::to_string(self : Confidence) -> String

    CustomRule

    pub(all) struct CustomRule {
    label : String
    value : String
    } derive(Eq,
    Debug
    )

    A caller-supplied exact value that should never leave the process. Rules shorter than four UTF-16 code units are ignored to limit accidental broad matches. The label is metadata only and never appears in output markers.

    CustomRule::equal

    fn CustomRule::equal(CustomRule, CustomRule) -> Bool

    CustomRule::not_equal

    fn CustomRule::not_equal(x : CustomRule, y : CustomRule) -> Bool

    Finding

    pub(all) struct Finding {
    kind : SensitiveKind
    start : Int
    end : Int
    confidence : Confidence
    } derive(Eq,
    Debug
    )

    Half-open UTF-16 offsets. A finding never stores or returns the matched value, making reports safer to log than the original input.

    Finding::equal

    fn Finding::equal(Finding, Finding) -> Bool

    Finding::not_equal

    fn Finding::not_equal(x : Finding, y : Finding) -> Bool

    Finding::to_repr

    JsonFinding

    pub(all) struct JsonFinding {
    path : String
    finding : Finding
    } derive(Eq,
    Debug
    )

    A finding located inside a JSON document, addressed by a JSON-pointer-like path such as /user/0/email. Array positions appear as numeric segments. The field name is metadata; the matched value itself is never stored.

    JsonFinding::equal

    fn JsonFinding::equal(JsonFinding, JsonFinding) -> Bool

    JsonFinding::not_equal

    fn JsonFinding::not_equal(x : JsonFinding, y : JsonFinding) -> Bool

    JsonRedactionResult

    pub(all) struct JsonRedactionResult {
    text : String
    changed_fields : Int
    findings : Array[JsonFinding]
    parsed : Bool
    } derive(Eq,
    Debug
    )

    JsonRedactionResult::equal

    JsonRedactionResult::not_equal

    JsonRedactionResult::paths

    fn JsonRedactionResult::paths(self : JsonRedactionResult) -> Array[String]

    Distinct JSON-pointer-like paths that produced findings, sorted, so callers can compare document coverage across runs without values.

    KindCount

    pub(all) struct KindCount {
    kind : SensitiveKind
    count : Int
    } derive(Eq,
    Debug
    )

    One kind-to-count pair; summaries are ordered by kind name so identical inputs always produce identical reports.

    KindCount::equal

    fn KindCount::equal(KindCount, KindCount) -> Bool

    KindCount::not_equal

    fn KindCount::not_equal(x : KindCount, y : KindCount) -> Bool

    RedactionResult

    pub(all) struct RedactionResult {
    text : String
    findings : Array[Finding]
    changed : Bool
    } derive(Eq,
    Debug
    )

    RedactionResult::equal

    RedactionResult::not_equal

    fn RedactionResult::not_equal(x : RedactionResult, y : RedactionResult) -> Bool

    RedactionStyle

    pub(all) enum RedactionStyle {
    Marker
    Typed
    PreserveLast4
    } derive(Eq,
    Debug
    )

    RedactionStyle::equal

    RedactionStyle::not_equal

    fn RedactionStyle::not_equal(x : RedactionStyle, y : RedactionStyle) -> Bool

    ScanConfig

    pub(all) struct ScanConfig {
    detect_access_token : Bool
    detect_bearer : Bool
    detect_jwt : Bool
    detect_private_key : Bool
    detect_url_credential : Bool
    detect_assignment : Bool
    detect_email : Bool
    detect_ipv4 : Bool
    detect_payment_card : Bool
    detect_resident_id : Bool
    detect_webhook : Bool
    detect_wallet : Bool
    detect_public_key : Bool
    detect_phone : Bool
    detect_ipv6 : Bool
    detect_uuid : Bool
    detect_mac : Bool
    suppress_version_ipv4 : Bool
    require_known_card_prefix : Bool
    suppress_example_domains : Bool
    extra_prefixes : Array[TokenPrefixRule]
    } derive(Eq,
    Debug
    )

    Full detector configuration. Every detector family can be switched off independently, and the false-positive suppressors can be relaxed by callers who prefer recall over precision. Turning credential detectors off creates surprising leak paths, so the preset constructors keep them on.

    ScanConfig::equal

    fn ScanConfig::equal(ScanConfig, ScanConfig) -> Bool

    ScanConfig::not_equal

    fn ScanConfig::not_equal(x : ScanConfig, y : ScanConfig) -> Bool

    ScanConfig::pii_only

    fn ScanConfig::pii_only() -> ScanConfig

    Only the personal-identifier families run; every credential detector is off. Useful for telemetry scrubbing where secrets are handled elsewhere.

    ScanConfig::secrets_only

    fn ScanConfig::secrets_only() -> ScanConfig

    Only credential families run; emails, IPv4 addresses and payment cards are left untouched.

    ScanConfig::standard

    fn ScanConfig::standard() -> ScanConfig

    All detectors on with false-positive suppression enabled.

    ScanPolicy

    pub(all) struct ScanPolicy {
    detect_email : Bool
    detect_ipv4 : Bool
    detect_payment_card : Bool
    } derive(Eq,
    Debug
    )

    A policy controls which detector families run. Known credentials are always enabled because suppressing them creates surprising leak paths.

    ScanPolicy::equal

    fn ScanPolicy::equal(ScanPolicy, ScanPolicy) -> Bool

    ScanPolicy::not_equal

    fn ScanPolicy::not_equal(x : ScanPolicy, y : ScanPolicy) -> Bool

    ScanPolicy::secrets_only

    fn ScanPolicy::secrets_only() -> ScanPolicy

    ScanPolicy::standard

    fn ScanPolicy::standard() -> ScanPolicy

    SensitiveKind

    pub(all) enum SensitiveKind {
    AccessToken
    BearerToken
    Jwt
    CredentialAssignment
    Email
    Ipv4
    PaymentCard
    PrivateKey
    UrlCredential
    CustomSecret
    ResidentId
    MacAddress
    Uuid
    Ipv6
    Phone
    PublicKey
    WalletAddress
    WebhookUrl
    } derive(Eq,
    Debug
    )

    Detection categories intentionally stay small and auditable. This library detects structured secrets and common identifiers; it is not an NLP system.

    SensitiveKind::equal

    SensitiveKind::not_equal

    fn SensitiveKind::not_equal(x : SensitiveKind, y : SensitiveKind) -> Bool

    SensitiveKind::to_string

    fn SensitiveKind::to_string(self : SensitiveKind) -> String

    StructuredField

    pub(all) struct StructuredField {
    path : String
    value : String
    } derive(Eq,
    Debug
    )

    Domain-neutral structured field adapter. path may be a JSON pointer, dotted property path, form field name, or database column name.

    StructuredField::equal

    StructuredField::not_equal

    fn StructuredField::not_equal(x : StructuredField, y : StructuredField) -> Bool

    StructuredFieldResult

    pub(all) struct StructuredFieldResult {
    path : String
    value : String
    findings : Array[Finding]
    changed : Bool
    sensitive_path : Bool
    } derive(Eq,
    Debug
    )

    StructuredFieldResult::equal

    StructuredFieldResult::not_equal

    StructuredResult

    pub(all) struct StructuredResult {
    fields : Array[StructuredFieldResult]
    changed_fields : Int
    findings : Int
    } derive(Eq,
    Debug
    )

    StructuredResult::equal

    StructuredResult::not_equal

    fn StructuredResult::not_equal(x : StructuredResult, y : StructuredResult) -> Bool

    TokenPrefixRule

    pub(all) struct TokenPrefixRule {
    prefix : String
    min_length : Int
    } derive(Eq,
    Debug
    )

    A caller-defined prefixed-token rule. min_length is the minimum total match length including the prefix. Prefixes shorter than two code units or minimums below the prefix length are ignored to limit accidental broad matches.

    TokenPrefixRule::equal

    TokenPrefixRule::not_equal

    fn TokenPrefixRule::not_equal(x : TokenPrefixRule, y : TokenPrefixRule) -> Bool

    explain

    fn explain(text : String, config? : ScanConfig) -> Array[String]

    Human-readable, value-free finding report: one line per finding in the form KIND start-end CONFIDENCE. Safe to print because it never contains matched text.

    findings_summary

    fn findings_summary(text : String, config? : ScanConfig) -> Array[KindCount]

    Count findings by kind without producing redacted text. Ordered by kind name for deterministic output; counts only, never values.

    redact

    fn redact(text : String, style? : RedactionStyle, policy? : ScanPolicy) -> RedactionResult

    Redact all enabled findings. Already-produced markers contain no detector pattern, so applying the same policy again is idempotent.

    redact_batch

    fn redact_batch(lines : Array[String], style? : RedactionStyle, policy? : ScanPolicy) -> BatchResult

    Redact independent log/API/prompt lines and return aggregate counts. The summary contains categories and counts only, never matched values.

    redact_batch_with_config

    fn redact_batch_with_config(lines : Array[String], config : ScanConfig, style? : RedactionStyle) -> BatchResult

    Batch redaction under a full configuration, enabling post-0.1.0 families such as resident ids and MAC addresses in log pipelines.

    redact_except

    fn redact_except(text : String, keep : Array[String], style? : RedactionStyle, config? : ScanConfig) -> RedactionResult

    Redact while sparing caller-certified literals, such as a load balancer IP that is safe to publish. Everything else runs under the given config.

    redact_fields

    fn redact_fields(fields : Array[StructuredField], style? : RedactionStyle, policy? : ScanPolicy) -> StructuredResult

    Redact already-parsed structured fields. Sensitive paths are fail-closed: their complete value is replaced even when its shape matches no detector. Other fields use content scanning. Paths are retained for caller-side joins.

    redact_fields_with_config

    fn redact_fields_with_config(fields : Array[StructuredField], config : ScanConfig, style? : RedactionStyle) -> StructuredResult

    Redact structured fields under a full configuration, enabling post-0.1.0 families for callers that normalize records themselves.

    redact_full

    fn redact_full(text : String, rules : Array[CustomRule], config : ScanConfig, style? : RedactionStyle) -> RedactionResult

    Redact with both built-in detectors under full configuration and caller-supplied exact-value rules.

    redact_json

    fn redact_json(text : String, config? : ScanConfig, style? : RedactionStyle) -> JsonRedactionResult

    redact_json_with_rules

    fn redact_json_with_rules(text : String, rules : Array[CustomRule], config? : ScanConfig, style? : RedactionStyle) -> JsonRedactionResult

    redact_with_config

    fn redact_with_config(text : String, config : ScanConfig, style? : RedactionStyle) -> RedactionResult

    Redact with full detector configuration, suppressor toggles and extra prefixed-token rules.

    redact_with_rules

    fn redact_with_rules(text : String, rules : Array[CustomRule], style? : RedactionStyle, policy? : ScanPolicy) -> RedactionResult

    Redact text with built-in and caller-supplied exact-value rules.

    scan

    fn scan(text : String, policy? : ScanPolicy) -> Array[Finding]

    Scan text without retaining sensitive values in the returned findings.

    scan_except

    fn scan_except(text : String, keep : Array[String], config? : ScanConfig) -> Array[Finding]

    Scan while keeping caller-certified literals in place: any finding that overlaps a kept literal is dropped. Keep values shorter than four code units are ignored so a tiny string cannot neuter whole families.

    scan_full

    fn scan_full(text : String, rules : Array[CustomRule], config : ScanConfig) -> Array[Finding]

    scan_with_config

    fn scan_with_config(text : String, config : ScanConfig) -> Array[Finding]

    Scan with full control over detector families, suppressors and extra prefixed-token rules.

    scan_with_rules

    fn scan_with_rules(text : String, rules : Array[CustomRule], policy? : ScanPolicy) -> Array[Finding]

    Add caller-known exact secrets to the built-in detector set. The returned findings do not carry rule labels or values. Empty and very short values are deliberately ignored.

    verify_clean

    fn verify_clean(text : String, policy? : ScanPolicy) -> Bool

    Return true only when a fresh scan finds no enabled sensitive values.

    verify_clean_with_config

    fn verify_clean_with_config(text : String, config : ScanConfig) -> Bool

    verify_clean_with_rules

    fn verify_clean_with_rules(text : String, rules : Array[CustomRule], config? : ScanConfig) -> Bool

    Return true only when neither built-in detectors under the configuration nor caller-supplied exact-value rules find anything.

    Powered by MoonBit

    Site sourceReport issuePackagesBuild queueSkillsStatistics

    © 2026 mooncakes.io