15 KiB
Calculet NPU calbin 产物字段字典与静态验收规范
日期:2026-08-02
样本:Qwen3-30B-A3B、batch 1、context 40960、双芯粒产物(67 个文件)
1. 目的
本规范回答四个问题:每个文件是什么、每个字段如何解释、哪些关系必须相互一致、拿到一个新模型包如何在不上卡的情况下拒绝坏产物。
结论分级:字段在当前样本出现为 C0;字段单位/语义如果只来自注释或推导则保持 C3 待厂商确认。静态验收通过只证明包结构自洽,不证明数值正确或设备可运行。
2. 当前目录清单与所有权
| 类别 | 当前数量 | 文件/模式 | 生产是否必需 | 所有者 |
|---|---|---|---|---|
| 根 manifest | 1 | <model>.yaml |
是 | compiler |
| Runtime 顶层 metadata | 1 | model_memory_reserved_info.txt |
是 | compiler/calbin packer |
| Runtime 测试参数 | 1 | hparam.yaml |
runner/golden 需要 | compiler/test |
| 参数块 | 2 | param_blk0.bin, param_blk1.bin |
是 | quant/codegen |
| 子模型目录 | 2 | *_prefill_*, *_decode_* |
是 | compiler |
| 子模型 metadata | 2 | */submodel_memory_reserved_info.txt |
是 | compiler |
| CCU ELF | 8 | 每子模型 2 core x 2 chip | 是 | codegen |
| CPU ping/pong 命令 | 8 bin | 每子模型 2 engine x 2 chip | 是 | command generator |
| PLD 命令 | 8 bin | 同上 | 当前包是 | command generator |
| 命令文本 | 16 txt | CPU/PLD x ping/pong x chip x submodel | 诊断需要 | command generator |
| CPU helper SO | 2 | prefill 的 riscv/x86 SO | 当前包是 | C oplib |
| 算子 I/O metadata | 2 | *_op_io_buf_ptr.yaml |
静态验收/调优需要 | compiler |
| per-chip op 顺序 | 4 | *.chip[01].profparts |
profiling/partition 需要 | compiler |
| tokenizer/config | 12 | __Tokenizer/* |
服务需要 | model export |
| 版本脚本 | 1 | version.py |
诊断 | packer |
当前参数文件精确大小:chip0 9,404,557,312 B,chip1 8,782,227,456 B,合计 18,186,784,768 B = 16.938 GiB。
3. 根 manifest <model>.yaml
3.1 one_model_cfg
每个 key 是编译前逻辑子图名,当前有 prefill 和 decode。字段:
| 路径 | 类型 | 当前值 | 含义/检查 |
|---|---|---|---|
model_path |
string | 厂商绝对 ONNX 路径 | 仅溯源;本地不存在不等于包坏,但交付必须另含 ONNX/hash |
input_shape.input_ids |
int[] | prefill [1,40960],decode [1,1] |
与 submodel ibuf、batch/context 一致 |
input_shape.position_ids |
int[] | 同 token | shape 和 dtype 都要对齐 Runtime tensor |
past_key_i |
int[4] | prefill [1,4,0,128] |
48 层,KV heads=4,head_dim=128 |
past_value_i |
int[4] | 同上 | prefill 初始 KV 长度为 0 |
decode past_*_i |
int[4] | [1,4,40959,128] |
预留一 token decode 到总长 40960 |
交叉规则:past_key_0..47 与 past_value_0..47 必须成对且编号连续;layer count 必须与 HF config、operator graph 和 KV reservation 相符。
3.2 model_info
| 字段 | 当前值 | 解释 | 拒绝条件 |
|---|---|---|---|
batch_num |
1 | 编译 batch | 与顶层 n_batch/submodel 不一致 |
q0_head_num |
32 | Q heads | 与模型 config 不一致 |
k0_head_num |
4 | K heads | 与 KV 公式/shape 不一致 |
v_head_num |
4 | V heads | 同上 |
*_head_dim |
128 | head dimension | 与 graph tensor layout 不一致 |
theta_base |
1000000.0 | RoPE theta | 与模型 config 不一致 |
max_seq_len |
40960 | 图上下文上限 | 与顶层 seq_len、memory metadata 不一致 |
is_encoder |
false | decoder-only | 模型架构冲突 |
is_prefill |
prefill true/decode false | 子图角色 | 子图角色冲突 |
alibi |
false | attention position mode | 与模型 config/graph 冲突 |
3.3 device_cfg 和 calcc_env_cfg
| 字段 | 当前值 | 能证明什么 | 不能证明什么 |
|---|---|---|---|
chip_nums |
2 | 该产物为双芯粒 | Runtime 支持多物理板 |
calcore_num |
2 | 每芯粒生成两个 CCU 程序 | 两个 host job 可并行 |
with_codegen |
true | 生成代码产物 | 可重建 compiler 环境 |
with_op_case |
true | 生成 op case 相关产物 | 每个 op 都有可独立 runner |
with_c_oplib |
true | 使用 C oplib | C oplib 已交付源码 |
use_auto_alloc |
true | compiler 自动地址规划 | 地址可在 Runtime 动态重定位 |
enable_flash_attention |
true | 图启用 FA | 所有 shape/dtype 都支持 FA |
enable_one_token |
true | decode one-token 路径 | dynamic batch 可用 |
dynamic_d2d |
prefill true/decode false | prefill 有动态 D2D 命令 | 实际短 prompt D2D bytes 等于静态 command sum |
has_loaded_onnx |
false | compiler 环境状态字段 | 不能断言最终图未从 ONNX 来 |
has_optimized_graph |
false | compiler 环境状态字段 | 不能断言无 graph fusion |
3.4 供应链字段
| 字段 | 当前值/形态 | 验收 |
|---|---|---|
model_name |
完整包名 | 与目录、model_arch 一致 |
seq_len |
40960 | 与两子图 max_seq 一致 |
onnx_hash[] |
两个相同 SHA-256 | 对交付 ONNX 重算;无 ONNX 时标记不可复现 |
git_commit_id.calcc |
575c2d... |
compiler container 应声明同 commit |
git_commit_id.lib/tvm |
7aebd1... |
同上 |
golden_hash.golden_prefill/decode |
hash 数组 | 必须说明每一项对应 tensor/layer/file;当前仅有 hash 无对象映射 |
4. model_memory_reserved_info.txt 字典
4.1 顶层字段
| 行格式 | 当前值 | 解释 | 校验 |
|---|---|---|---|
calcc_ver: VALUE |
v0.1 | metadata/compiler schema 版本线索 | 必须非空;需要正式 schema version |
device0: [ks1] |
ks1 | 兼容芯片类型 | 与硬件型号/Runtime compatibility 一致 |
device_memory_required: N MB |
16777216 MB | 单位明显可疑 | C3,禁止用于容量决策 |
model_arch NAME |
包名 | 架构标识 | 与根 manifest 一致 |
model_type llm |
llm | 顶层类别 | 当前必须为 llm |
n_batch 1 |
1 | 最大 batch | 与所有子图一致 |
max_seq_len 40960 |
40960 | 最大 context | 与根 manifest 一致 |
4.2 reservation 行
语法:
dram_rsvd vaddr <hex> size <decimal>
sram_rsvd vaddr <hex> size <decimal>
sync_rsvd vaddr <hex> size <decimal>
当前每芯粒逻辑地址计划:DRAM 15,113,064,448 B = 14.075 GiB,SRAM 17,662,336 B = 16.844 MiB,sync 20 B。这里的“每芯粒”来自相同 virtual address 在 chip mask 上分配的结构推断,应要求厂商明确物理/逻辑口径。
4.3 参数映射行
语法:
<file> offset <hex> vaddr <hex> size <hex> chip <mask95_64> <mask63_32> <mask31_0>
校验算法:
- file 必须存在且是普通文件,禁止
..、绝对路径和 symlink 逃逸。 offset <= file_size且size <= file_size-offset。- chip mask 非 0;当前 chip0 为低位
1,chip1 为低位2。 - 对同一个 file,区间允许不按 vaddr 排序,但文件 offset 区间不可越界。
- 对同一 chip/存储域,vaddr overlap 只有在 compiler 明确声明 alias/lifetime 复用时才允许;静态工具应报告而非自行判定非法。
- 所有映射 size 合计不一定等于文件大小,因为同一参数块含 DRAM/SRAM 段;需要统计 uncovered 和 duplicate ranges。
5. 子模型 metadata 字典
5.1 身份和执行类型
| 字段 | 当前值 | 解释/问题 |
|---|---|---|
model_arch |
顶层包名 | 必须一致 |
smodel_type |
prefill/decode 都写 llm_decode |
prefill metadata 明显错误,C3 修复 |
n_batch |
1 | 与顶层和 manifest 一致 |
calc_mchip |
mchip | 多芯粒编译产物,不等于多物理板 |
calc_utype |
[cpu,ccu] |
command 中包含 CPU/CCU 协作 |
workload |
17179869184 Byte | 两子图完全相同,含义/单位需厂商确认 |
ccu_stack_config location |
dram | CCU stack 放置域 |
5.2 命令和 ELF 行
cpu_chip0_ping.bin chip 00000000 00000000 00000001
ccu0_chip0.elf location dram chip 00000000 00000000 00000001
交叉检查:每个声明文件存在、非空;文件名 chip/core/engine 与 mask 对应;同一子图 Ping/Pong 文件都齐全;ELF 可由 readelf -h 解析,动态 SO 可由 readelf -d 检查依赖。不能只按文件名判断目标 ISA。
5.3 CSR 和 I/O 行
icsr <name> offset <hex>
ibuf <name> shape [dims] type <dtype> vaddr <ping> <pong> size <decimal>
obuf <name> shape [dims] type <dtype> vaddr <ping> <pong> size <decimal>
当前精确契约:
| 子图 | tensor | shape/type | bytes | ping/pong 地址 |
|---|---|---|---|---|
| prefill | inputs[0] token | [1,40960] int32 |
163840 | 不同 |
| prefill | inputs[1] position | [1,40960] int32 |
163840 | 不同 |
| prefill | outputs[0] logits | [1,1,151936] bfloat16 |
303872 | 不同 |
| decode | inputs[0] token | [1,1] int32 |
4 | 不同 |
| decode | inputs[1] position | [1,1] int32 |
4 | 不同 |
| decode | outputs[0] logits | [1,1,151936] bfloat16 |
303872 | 不同 |
必须验证 product(shape)*dtype_bytes == size。vaddr 是敏感内部地址,只进入离线报告,不进入普通服务 API。
6. 算子 I/O YAML 字典
每个顶层 key 是唯一 op instance,例如 nbu_fused_matmul_add_clib_4。每个 args_N/outs_N 是四元组:
- device_virtual_address
- shape_tuple
- PrimitiveType
- producer_or_role
解析规则:
const_N归为常量/权重,input_N归为模型输入,其他为 activation/中间结果。- 同地址不同 op 常表示 inplace、alias 或生命周期复用,不能把所有条目 size 简单相加。
- 地址覆盖统计要先把
[address,address+bytes)区间去重并合并。 - 地址并集也不是峰值活跃内存,因为缺少 tensor live interval。
PrimitiveType必须进入 dtype 字节表;未知 dtype 使容量统计失败,不能默认为 0。
当前图:prefill 1496 个顶层 op,decode 1398。地址并集 prefill 14.091 GiB、decode 10.634 GiB;这是静态地址覆盖,不是同时活跃内存。
7. profparts 与 command 文本
7.1 profparts
每行是某芯粒执行的 op instance。验证:
- 每行名称必须在对应 op I/O YAML 中存在。
- 同一芯粒不应出现无法解释的重名。
- chip0 prefill 1352 行、chip1 1351;decode 1254/1253。
- 两芯粒都覆盖 48 层 attention/MoE,embedding 只在 chip0,支持当前为层内 tensor parallel 的判断。
7.2 runtime commands
文本由注释行描述 command type,下一行含字段。重点字段:cmdId、chipMask*、srcAddr、dstAddr、dataSize、同步 register index。
静态汇总的边界:
- 当前最大 prefill 图的 PLD chip0 有 444,610 条带
dataSize的命令,合计 18.979 GiB;chip1 444,608 条,18.822 GiB。 - 这是 40960 上限图的命令描述,不是短 prompt 实测 PCIe/D2D 流量。
- ping/pong 文本相同或近似只证明双份 command 产物,不证明两任务可安全并行。
8. tokenizer 和服务一致性
必须 hash 并冻结:tokenizer.json、tokenizer_config.json、config.json、generation_config.json、merges.txt、vocab.json、vocab GGUF。
检查项:
- vocab size 151936 与 logits 最后一维一致。
- BOS/EOS/PAD、chat template、added tokens 与参考 runner 一致。
- config max position 40960 与 tokenizer 宣称 131072 冲突时,服务硬上限取 calbin 40960。
- 模型别名不能仅由目录名推断;发布 manifest 要绑定 tokenizer hash 和 calbin hash。
9. 静态验收流水线
Gate A:文件系统
目录可读 -> 无逃逸 symlink -> required files -> size>0 -> 文件数量/未知文件报告
大参数 full SHA-256 可在发布阶段做;日常开发 fast mode 可读取既有 hash manifest,但不能把 mtime/size 当完整性证明。
Gate B:schema
YAML 可安全解析 -> 必填 key/type -> submodel count/name -> metadata grammar
当前 txt metadata 并非标准 YAML,需用明确 grammar parser;不要用宽松 split 后猜字段。
Gate C:cross-reference
model name / batch / context / chip count
input shape / dtype / byte size
48 layers / Q-KV heads / head_dim
command + ELF declaration -> file exists
profpart op -> graph op exists
parameter map -> file range valid
Gate D:容量
parameter bytes by chip
DRAM/SRAM/sync reservation
KV formula
input/output/workspace address union
overlap/alias report
KV BF16 复算:
bytes = layers * kv_heads * 2(K,V) * head_dim * max_seq * dtype_bytes
= 48 * 4 * 2 * 128 * 40960 * 2
= 4,026,531,840 B = 3.750 GiB total
= 1.875 GiB/chip when 2 KV heads per chip
Gate E:供应链
要求 ONNX hash、compiler/TVM commit、compiler OCI digest、quant config、calibration manifest、oplib/kernel version、calbin schema version、Runtime/driver/firmware compatibility。
Gate F:Runtime 前检查
静态通过后由独立 runner CreateCalbin -> enumerate -> configure -> golden。任何 static warning 都要带到 Runtime 报告,不得因能 configure 就忽略 schema 错误。
10. 严格拒绝与警告策略
| 条件 | 等级 | 动作 |
|---|---|---|
| required file 缺失/空、YAML 解析失败 | ERROR | 拒绝 |
| 参数映射越界、I/O size 公式不符 | ERROR | 拒绝 |
| batch/context/model name 交叉不一致 | ERROR | 拒绝 |
| exact prefill/decode 唯一性不成立 | ERROR | 拒绝 |
| op YAML 引用未知 dtype | ERROR | 拒绝容量报告 |
| profparts 引用未知 op | ERROR | 拒绝 partition 报告 |
| ONNX/compiler/container 缺失 | BLOCKED_REPRODUCIBILITY | 可运行旧包,但禁止称可复现 |
prefill smodel_type=llm_decode |
WARNING_VENDOR | 当前 0.7.6 可继续隔离测试,必须厂商修复 |
| memory 单位不合理 | WARNING_VENDOR | 不使用该字段做容量判断 |
| tokenizer context > calbin context | WARNING_SERVICE | 服务强制使用 calbin 上限 |
11. 新产物报告格式
每次校验生成 validation.json:
{
"schema": "calculet-package-validation/v1",
"package": {"name": "...", "sha256_manifest": "..."},
"model": {"batch": 1, "max_seq": 40960, "chips": 2},
"submodels": [{"name": "...", "kind": "prefill", "inputs": [], "outputs": []}],
"memory": {"parameter_bytes": {}, "kv_bytes": 4026531840, "reservations": {}},
"graph": {"prefill_ops": 1496, "decode_ops": 1398, "op_histogram": {}},
"supply_chain": {"onnx_hash": [], "compiler_commits": {}},
"findings": [{"severity": "warning", "code": "META_SMODEL_TYPE", "evidence": "..."}],
"verdict": "pass_with_warnings"
}
地址只写到受控内部版;分享给业务/客户的净化版去掉 vaddr、源绝对路径和命令地址。
12. 新模型包完成定义
- 静态 Gate A-F 全部有机器可读结果,ERROR=0。
- 文件 hash manifest 覆盖全部 67 类似资产,不只参数文件。
- ONNX 到 calbin 能在固定 compiler container 重放。
- 生成的包在独立 runner 通过 prefill/decode golden、边界和失败恢复。
- package manifest 同时绑定 tokenizer、quant、compiler、Runtime、driver、firmware 和芯片版本。
- 所有 WARNING_VENDOR 有工单、厂商版本和关闭标准;不能永久作为“已知问题”漂移。