segy

    MoonBit SEG-Y file exchange and trace inspection

    Download zip
    Version
    0.2.0
    License
    MIT
    Last updated
    6 hours ago
    Downloads
    1

    #MoonSEGY

    评审/首次使用请先看实际任务、替代方案与可运行证据:读入SEG-Y,扫描头与道索引,按道集分组检查振幅/非有限值/分析范围,输出带原始道号与字节位置的问题清单,再由调用方筛选或导出。

    MoonBit 原生 SEG-Y 文件交换、道筛选和原始振幅质量检查库。面向反射地震记录与道集数据,不是 miniSEED/SAC 工具、反演程序或地球物理结论生成器。头、样本编解码、索引、变换与分析全部在 MoonBit;Node 只负责文件和 CLI 参数。

    chenglong-coding/segy 是尚未发布的本地模块。MIT;AI 辅助开发并保留真实作者。本地完成不等于远程 CI、注册表发布或正式比赛验收。原范围见 SCOPE,实测边界见 TESTING。

    #本轮公开道集与兼容边界

    已消费USGS DS259的一条完整测线:19,158道、14,368,500样本,全局/分组统计与独立原始解码一致,识别2,667条零振幅道。完整文件及精确验证范围见 FULL-USGS。原300道节选的225000值逐值核对、36条零道与当时403记录保留在 PUBLIC-USGS,可离线复现。旧版本字仍须显式兼容选项;默认严格拒绝,报告保留假设,原字节不变。此方向作为旧IRC替代候选,尚无正式仓库或获准换题结论。

    #支持能力

    • 常用 rev0、rev1.0、受限 rev2.0;大/小端识别;ASCII 和 EBCDIC CP037;计数或 EndText 结束的扩展文本块。
    • IBM32、IEEE32/64、有符号8/16/24/32/64位样本。int64 保持整数和原字节,不先变成 Double。
    • 标准240字节道头,定长/变长扫描索引、原字节复制、未知普通头字节保留;受限 rev2 固定道支持扩展计数/采样间隔。
    • 坐标/高程/时间标量,明确角度单位下的 DMS 解码;道选择、头值筛选、道集分组、样本窗口和样本码转换。
    • 原始或显式权重振幅统计、RMS、总体标准差、死道/用户阈值裁剪检查;CSV、脚本无关 SVG 波形。
    • 全道集/指定原始道号的QC:按头字段分组,原始字节位置、非有限/整数范围/权重问题定位,完整汇总与限额问题明细;JSON/CSV。

    #快速运行

    需要 MoonBit 与 Node.js 24。当前固定工具链见.moonbit-version;0.2.0检查见docs/TESTING.md,旧0.1.0日志不当作当前版本结论。核心运行不需要 Python/segyio。

    moon check --target all moon test --target js moon test --target wasm-gc moon build --target js --release node tools/segy.mjs create examples/create.json demo.sgy node tools/segy.mjs validate demo.sgy node tools/segy.mjs stats demo.sgy examples/stats.json node tools/segy.mjs groups demo.sgy examples/groups.json node tools/segy.mjs select demo.sgy examples/select.json selected.sgy node tools/segy.mjs window demo.sgy examples/window.json window.sgy node tools/segy.mjs csv demo.sgy examples/select.json samples.csv node tools/segy.mjs svg demo.sgy examples/trace.json trace.svg node tools/segy.mjs qc demo.sgy examples/quality.json moon run examples/quality_check --target js

    示例生成三道合成数据,包含两个同道集波形和一条零幅度道。输出必须是新路径,不覆盖既有文件;错误写 stderr、退出码2。SVG 必须显式选道,超过20000样本请先窗口,不用未经说明的降采样隐藏细节。 qc 示例故意包含零幅度道和阈值问题,正常输出报告后退出3;这不是命令崩溃。qc/qc-csv 无发现退出0,参数/结构/IO错误退出2。纯MoonBit示例正常退出并打印发现结果,可用JS/Wasm-GC运行。 单样本道使用可见圆点表示(包括零值),不虚构时间跨度;多样本道保持逐点折线。

    普通命令形式为 COMMAND INPUT [OPTIONS.json] [OUTPUT],选项是 JSON 文件;create OPTIONS.json OUTPUT 例外。读取报告默认打印 JSON,CSV/SVG 打印文本;二进制操作必须给 OUTPUT。

    命令选项/行为
    inspect / validate结构信息 / 另外遍历样本拒绝非有限 IEEE 值
    tracetrace,可选 first、count(每次最多100000);含时间和头字段
    statstrace,可选 dead_threshold、clip_threshold、weighted
    groups / findfield / 再加 minimum、maximum,原始有符号字段闭区间
    coordinatetrace、field;可选 degrees:true 将显式角度单位转十进制度
    copy / select / filter精确复制 / traces / field、minimum、maximum
    window / convertfirst、count / sample_code
    csv / svgtraces 数组 / trace
    qc / qc-csv可选 traces、group_by、dead_threshold、clip_threshold、weighted、max_details、max_groups;见 QC指南

    可用原始字段名:sequence_line、sequence_file、field_record、trace_in_record、source_point、ensemble、trace_in_ensemble、identification、offset、receiver_elevation、source_elevation、source_depth、elevation_scalar、coordinate_scalar、source_x/y、group_x/y、coordinate_units、delay_ms、weighting;rev1/2另有 cdp_x/y、inline、crossline、shotpoint、shotpoint_scalar、time_scalar。未列字段仍保留原字节,但不解释其语义。

    所有读命令可用 endian:"big"/"little"、text_encoding:"ASCII"/"EBCDIC" 显式覆盖。覆盖端序不能与识别到的 rev2 标志冲突。自动文本检测是可报歧义的启发式,不把所有 EBCDIC 变种都当 CP037。

    #库 API 与数据语义

    根包为纯 MoonBit。公共签名见 pkg.generated.mbti。parse_header / decode 读取,make_header / make_trace / create 新建。Dataset 提供 trace_index、trace、select、find_traces、groups、window、convert_samples、csv 等;Trace 提供 sample、field、scaled_field、coordinate_degrees、time_seconds、statistics、svg。

    let dataset = @segy.decode(bytes)
    let selected = dataset.select(dataset.find_traces("ensemble", 10, 20))
    let windowed = selected.window(2, 100)
    let output = windowed.encode()

    调用方需处理可抛出的错误。道和样本索引从0开始,代码中的字节偏移从0开始,规范表从1开始。采样间隔单位微秒;延迟头字段是带时间标量的毫秒;CSV/time_seconds 为相对震源时间的秒,不猜测绝对日期/时区。

    encode() 原样返回不可变字节;select/window 重算结构计数,保留其他采集元数据。select 更新 sequence_file,原始 sequence_line、记录/道集标识保留。window 对每道用同一样本索引区间,不把短道悄悄补零。改变起点若不能由既有时间标量精确表达,就报错,不同时破坏其他 lag/mute 时间字段。

    原始振幅不是物理标定值。weighted:true 仅应用 2^(-weighting),不自动应用传感器换能系数。标量坐标保持声明的米/英尺/角单位;不从数值猜 CRS。clipped_samples 只表示超过用户阈值,dead 只表示幅度阈值意义上的平坦零道,不等价于设备损坏诊断。

    Dataset::quality_check 能报告并继续扫描不可分析样本。任何样本不可可靠分析时,该整道统计为显式null,不进入分组/全局振幅汇总,完整原因与样本计数保留;不偷偷算有效子集均值。合并统计按样本数量加权,不按每道平均值或时间长度加权。完整64位原值与文件保持不变。CLI道号/头字段/预算必须为精确整数,不接受小数截断。

    #精度与明确限制

    • Sample::Signed(Int64) 保存完整64位整数;转 Double 分析保守拒绝绝对值超过2^53。CLI 的整数样本使用十进制字符串输出;新建JSON大整数也必须使用字符串。
    • IBM32 写入规范化、最近舍入(中点远离零);浮点下溢至零/溢出报错,IEEE64 原字节复制不损精度。Float→Integer 只接受精确整数,不隐式截断。
    • 非有限 IEEE 样本可结构读取和原字节复制,validate、逐道统计、样本CSV、SVG拒绝;QC会报告并继续扫描,不修改原样本。不能因 inspect 成功声称数值有效。
    • 文件最大256 MiB,1000000道,单道1000000样本,合计16000000样本,1024扩展文本块,CSV1000000行;这些是数据限额,不承诺恒定内存。
    • 不支持额外240字节道头、头布局重映射、二进制用户 stanza、pair-swapped、tape label、data trailer、rev2.1、无符号/过时定点样本码。未知布局明确拒绝,普通未解释头字节可保留。
    • 不提供任意原始头的跨端序/跨版本重写,因为未知字段无法安全重排。可新建两种端序,现有文件的选择、窗口和样本格式转换保持原端序/版本。
    • rev2 宽计数/非整数间隔仅支持固定道;变长道须能由标准16位样本数/间隔表达。rev2 新建版本为独立 major/minor 字节,读取兼容旧 little-endian uint16 写法。

    #复现验证

    moon fmt --check moon info node tools/check-cli.mjs python -m pip install -r tools/requirements.txt python tools/verify-reference.py python tools/verify-quality.py

    增强前基线为13组测试、CLI16项、独立332场景/1171项,保留在 reference.json。当前QC增强的源码绑定与复查结果见 quality-20260922.json 和 TESTING,不以旧散列证明新代码。segyio 1.9.14不支持24位码7,且不能替代变长道/rev2完整验证:这些用独立 struct/CP037参考,绝不标成 segyio 通过。CI 已配置三系统,远程未执行。 SVG 的六类合成样本已在本机真实 Chromium 浏览器中检查,发现并修复单样本不可见问题;见 图形输出核查。

    依据:SEG-Y rev2.0 原始规范、segyio。CP037 数据表与 Python 独立编码器全256字节互核,非 Python 包装实现。复用本批自有二进制读取设计,未复制第三方库核心。

    验证工具许可证和合成样例来源见 SOURCES。

    初始查重见五项目选型记录;2026-09-22刷新 MoonBit+SEG-Y/注册表网页/GitLink 索引未命中同向项目,不是全球不存在的证明。已有 miniSEED/SAC 项目不是本库的 SEG-Y 交换工作流。

    #本地验收与公开交付(2026-09-28)

    核心实现使用 MoonBit;固定编译器为 moonc 0.10.14+7d59c7ec9。先按本文安装宿主依赖、运行 moon update,再从仓库根目录执行以下与 CI 对齐的检查;可运行任务和适用边界见本文前面的示例与说明。

    moon check --target all --deny-warn moon test --target js --deny-warn moon test --target wasm-gc --deny-warn moon build --target js --release --deny-warn moon package

    跨平台复核(2026-09-28,本地 Ubuntu-D 26.04 WSL2):从当时的源码归档全新解包,固定 moonc 0.10.14+7d59c7ec9 下通过 moon update、moon fmt --check、moon info、严格检查、JS/Wasm-GC 测试及 JS release 构建;Node 24.21.0 跑通本仓一条宿主入口。本次补记仅修改文档,代码与 CI 未变;复核日志在本地交接包中,公开提交后的 GitHub Actions 仍须单独核对。

    本地核验:JS/Wasm-GC 各 21 项测试、release 构建和 CLI 示例通过。 moon package 已完成离线打包预检,它不等于已发布到 Mooncakes。

    公开交付(2026-09-28 核对):尚无本项目正式公开仓库 URL 或 Mooncakes 版本;模块名 chenglong-coding/segy 是拟交付账号形式的本地名称,正式发布前须核实账号归属和发布权限;换题资格、仓库、公开 CI 和首次发布均待团队办理,不能沿用旧题仓库链接。相关远端 CI 与赛事结果仍需以实际记录核对。项目许可见 LICENSE;如使用第三方材料,其来源和许可见仓内相应说明。

    AmplitudeStatistics

    pub struct AmplitudeStatistics {
    samples : Int
    minimum : Double
    maximum : Double
    mean : Double
    rms : Double
    standard_deviation : Double
    } derive(ToJson)

    Sample-weighted amplitude moments over fully analyzable traces only. No duration weighting, resampling or physical-unit calibration is implied.

    AmplitudeStatistics::to_json

    Dataset

    pub struct Dataset {
    header : FileHeader
    // private fields
    }

    Dataset::convert_samples

    fn Dataset::convert_samples(self : Dataset, code : Int) -> Dataset raise

    Convert samples without endian/revision changes. No silent integer rounding or lossy int64-to-Double conversion. Identical format returns the same data.

    Dataset::csv

    fn Dataset::csv(self : Dataset, indices : Array[Int]) -> String raise

    Long CSV with exact integer spelling and seconds. Selected traces use original zero-based indices; maximum one million output samples.

    Dataset::encode

    fn Dataset::encode(self : Dataset) -> Bytes

    Exact byte copy (Bytes is immutable); does not canonicalize unknown fields.

    Dataset::extended_headers

    fn Dataset::extended_headers(self : Dataset) -> Array[Bytes]

    Dataset::find_traces

    fn Dataset::find_traces(self : Dataset, field : String, minimum : Int, maximum : Int) -> Array[Int] raise

    Raw signed header interval, inclusive. Returns [] when no trace matches.

    Dataset::groups

    fn Dataset::groups(self : Dataset, field : String) -> Array[TraceGroup] raise

    Group by an unscaled header value; groups retain first-occurrence order.

    Dataset::nonfinite_samples

    fn Dataset::nonfinite_samples(self : Dataset) -> Int raise

    Inspect all samples; returns the count of nonfinite IEEE values. Exact int64 samples are valid even when Double analysis would be refused.

    Dataset::quality_check

    fn Dataset::quality_check(self : Dataset, traces? : Array[Int], group_by? : String, dead_threshold? : Double, clip_threshold? : Double, weighted? : Bool, max_details? : Int, max_groups? : Int) -> QualityReport raise

    Scan the whole dataset or a unique nonempty original-index selection. Continue after un-analyzable amplitudes; structural parse errors still raise. Thresholds use raw or explicitly weighting-scaled amplitudes: dead if ALL |x|<=dead_threshold; clip if |x|>=clip_threshold (0 disables clipping). Statistics exclude whole un-analyzable traces, not just their bad samples. Limits: max_details 0..100000, max_groups 1..100000; data limits unchanged.

    Dataset::select

    fn Dataset::select(self : Dataset, indices : Array[Int]) -> Dataset raise

    Select/reorder unique zero-based trace indices. Preserve raw headers and samples; sequence_file is re-numbered from one, original line IDs remain.

    Dataset::trace

    fn Dataset::trace(self : Dataset, index : Int) -> Trace raise

    Dataset::trace_count

    fn Dataset::trace_count(self : Dataset) -> Int

    Dataset::trace_index

    fn Dataset::trace_index(self : Dataset) -> Array[TraceIndex]

    Dataset::window

    fn Dataset::window(self : Dataset, first : Int, count : Int) -> Dataset raise

    Same sample-index window on every trace; rejects any too-short trace. Raw sample bytes remain exact, including int64 and nonfinite IEEE bit patterns.

    Endian

    pub(all) enum Endian {
    Big
    Little
    } derive(Eq)

    Endian::equal

    fn Endian::equal(Endian, Endian) -> Bool

    Endian::not_equal

    fn Endian::not_equal(x : Endian, y : Endian) -> Bool

    FileHeader

    pub struct FileHeader {
    endian : Endian
    revision : Int
    raw_revision_word : Int
    assumed_legacy_revision_one : Bool
    sample_code : Int
    samples : Int
    interval_us : Double
    fixed_length : Bool
    extended_count : Int
    declared_traces : UInt64
    first_trace : Int
    text_encoding : String
    raw : Bytes
    }

    Immutable decoded file prefix. Raw is the original 3600-byte prefix; extended blocks and traces are parsed separately. Units are microseconds.

    QualityGroup

    pub struct QualityGroup {
    value : Int
    first_trace : Int
    summary : QualitySummary
    } derive(ToJson)

    Groups are unscaled header values, in first SELECTED occurrence order. No geometric grouping, coordinate-unit conversion or CRS inference.

    QualityGroup::to_json

    QualityReport

    pub struct QualityReport {
    status : String
    format_assumptions : Array[String]
    source_trace_count : Int
    weighted : Bool
    dead_threshold : Double
    clip_threshold : Double
    clipping_checked : Bool
    group_by : String?
    summary : QualitySummary
    groups : Array[QualityGroup]
    details : Array[TraceQuality]
    max_details : Int
    details_truncated : Bool
    }

    status clear/issues refers only to the explicitly enabled checks. Neither is instrument certification, a geological conclusion or acceptance proof. Details keep the first max_details problem traces in selection order.

    QualityReport::issues_csv

    fn QualityReport::issues_csv(self : QualityReport) -> String

    Retained problem-trace rows only. Use JSON for aggregate/group statistics, exclusions, format assumptions and truncation; an empty CSV is NOT evidence of a clear dataset. Keep the JSON report alongside an exported issue CSV.

    QualityReport::to_json

    fn QualityReport::to_json(self : QualityReport) -> Json

    QualitySummary

    pub struct QualitySummary {
    traces : Int
    samples : Int
    analyzed_traces : Int
    analyzed_samples : Int
    problem_traces : Int
    nonfinite_samples : Int
    integer_range_samples : Int
    negative_weighting_traces : Int
    weighting_range_samples : Int
    dead_traces : Int
    header_dead_traces : Int
    clipped_traces : Int
    clipped_samples : Int
    minimum_samples : Int
    maximum_samples : Int
    minimum_interval_us : Double
    maximum_interval_us : Double
    statistics : AmplitudeStatistics?
    }

    Counts cover ALL selected traces, even when details are truncated. analyzed_* and pooled statistics exclude a WHOLE trace if any amplitude is unavailable. Nonfinite/integer-range/weighting-range counts are sample counts; negative weighting counts traces. dead/clipped counts use fully analyzable traces.

    QualitySummary::to_json

    fn QualitySummary::to_json(self : QualitySummary) -> Json

    Sample

    pub(all) enum Sample {
    Signed(Int64)
    Real(Double)
    }

    Integers remain exact, including int64. Real includes raw IEEE nonfinite values for inspection/copy; numerical analysis must call as_double().

    Sample::as_double

    fn Sample::as_double(self : Sample) -> Double raise

    Trace

    pub struct Trace {
    // private fields
    }

    Owned raw trace. Sample bytes are retained exactly until explicitly changed.

    Trace::amplitude

    fn Trace::amplitude(self : Trace, index : Int, weighted? : Bool) -> Double raise

    Optional 2^(-weighting) header gain only; does not apply transduction/units.

    Trace::coordinate_degrees

    fn Trace::coordinate_degrees(self : Trace, name : String) -> Double raise

    Angular units 2 (arc-seconds), 3 (degrees), 4 (signed DDDMMSS.ss). Unknown/linear coordinates raise; do not assume a geographic CRS.

    Trace::field

    fn Trace::field(self : Trace, name : String) -> Int raise

    Raw signed header value, before coordinate/elevation/time scaling.

    Trace::interval_us

    fn Trace::interval_us(self : Trace) -> Double

    Trace::raw_header

    fn Trace::raw_header(self : Trace) -> Bytes

    Trace::raw_samples

    fn Trace::raw_samples(self : Trace) -> Bytes

    Trace::sample

    fn Trace::sample(self : Trace, index : Int) -> Sample raise

    Trace::samples

    fn Trace::samples(self : Trace) -> Int

    Trace::scaled_field

    fn Trace::scaled_field(self : Trace, name : String) -> Double raise

    Apply only the designated SEG header scalar. No CRS conversion or unit guessing: length coordinates stay in declared meters/feet, angular fields in their recorded units. Use coordinate_degrees for explicit angular units.

    Trace::statistics

    fn Trace::statistics(self : Trace, dead_threshold? : Double, clip_threshold? : Double, weighted? : Bool) -> TraceStatistics raise

    Raw (or optionally weighting-scaled) amplitudes, population deviation. Dead means every |value|<=dead_threshold; clipping is a caller threshold, not evidence of instrument saturation. clip_threshold=0 disables it.

    Trace::svg

    fn Trace::svg(self : Trace) -> String raise

    Standalone, script-free waveform SVG for <=20000 samples. Normalized height

    Trace::time_seconds

    fn Trace::time_seconds(self : Trace, index : Int) -> Double raise

    Time relative to trace source time; scaled delay (ms) plus index*dt (us).

    TraceGroup

    pub(all) struct TraceGroup {
    value : Int
    indices : Array[Int]
    } derive(ToJson)

    TraceGroup::to_json

    fn TraceGroup::to_json(TraceGroup) -> Json

    TraceIndex

    pub(all) struct TraceIndex {
    header_offset : Int
    sample_offset : Int
    samples : Int
    interval_us : Double
    } derive(ToJson)

    TraceIndex::to_json

    fn TraceIndex::to_json(TraceIndex) -> Json

    TraceQuality

    pub struct TraceQuality {
    trace : Int
    header_offset : Int
    sample_offset : Int
    sequence_line : Int
    sequence_file : Int
    field_record : Int
    ensemble : Int
    identification : Int
    samples : Int
    interval_us : Double
    weighting : Int
    flags : Array[String]
    nonfinite_samples : Int
    first_nonfinite_sample : Int?
    integer_range_samples : Int
    first_integer_range_sample : Int?
    weighting_range_samples : Int
    first_weighting_range_sample : Int?
    statistics : TraceStatistics?
    }

    Retained problem trace. Offsets refer to this Dataset's encoded bytes, not a rewritten selection. Header identifiers are raw and need not be unique. statistics is absent if ANY sample cannot be analyzed; no partial mean. integer_range_samples counts |integer|>2^53, a conservative analysis boundary, NOT a claim that every such integer is inexact or malformed.

    TraceQuality::to_json

    fn TraceQuality::to_json(self : TraceQuality) -> Json

    TraceStatistics

    pub(all) struct TraceStatistics {
    samples : Int
    minimum : Double
    maximum : Double
    mean : Double
    rms : Double
    standard_deviation : Double
    dead : Bool
    clipped_samples : Int
    clipping_checked : Bool
    } derive(ToJson)

    TraceStatistics::to_json

    create

    fn create(header : FileHeader, traces : Array[Trace], extended_headers? : Array[Bytes]) -> Dataset raise

    Assemble same-format traces, derive structural counts and validate by scan.

    decode

    fn decode(data : Bytes, endian? : Endian, text_encoding? : String, legacy_revision_one? : Bool) -> Dataset raise

    Scan fixed/variable traces with strict bounds; does not expand all samples. Undefined extended-text counts stop only at the SEG EndText marker.

    decode_sample

    fn decode_sample(data : Bytes, offset : Int, code : Int, endian : Endian) -> Sample raise

    Decode one unweighted sample at a byte offset; not a sample index.

    decode_text

    fn decode_text(data : Bytes, encoding : String) -> String raise

    Decode exactly one 3200-byte block. No trimming or guessed line endings.

    encode_samples

    fn encode_samples(samples : Array[Sample], code : Int, endian : Endian) -> Bytes raise

    Encode <= one million samples; no integer coercion/truncation is implicit.

    encode_text

    fn encode_text(text : String, encoding : String) -> Bytes raise

    Pad a flat <=3200-character header with spaces. Supply your own 80-column

    make_header

    fn make_header(samples : Int, interval_us : Double, sample_code : Int, endian? : Endian, revision? : Int, fixed_length? : Bool, text? : String, text_encoding? : String, extended_count? : Int) -> FileHeader raise

    Construct a standard prefix; revision 0 forbids extended headers and fixed flag. Full dataset writers later fill first-trace offset and trace count.

    make_trace

    fn make_trace(header : FileHeader, samples : Array[Sample], interval_us? : Double, fields? : Map[String, Int]) -> Trace raise

    Build a trace at the file's byte order/format. fields use named signed raw values, not scaled coordinates. Unknown fields and narrowed overflow raise.

    parse_header

    fn parse_header(data : Bytes, endian? : Endian, text_encoding? : String, legacy_revision_one? : Bool) -> FileHeader raise

    Detect byte order from rev2 sentinel or supported format code. Explicit endian still must agree with a recognized sentinel. Pair-swapped is rejected.

    sample_width

    fn sample_width(code : Int) -> Int raise

    Codes 1,2,3,5,6,7,8,9. Obsolete fixed-with-gain and unsigned codes rejected.