Files
calculet-npu-research-archive/reports/Calculet-NPU-Batch-KV与多模型调度实现规格-20260802.md
T

442 lines
16 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Calculet NPU Batch、KV 与多模型调度实现规格
日期:2026-08-02
基线:CalRT 0.7.6 单物理 VirtualDevice、当前双芯粒 Qwen3 calbin batch=1
## 1. 能力分层
| 层 | 当前/目标 | 状态 | 本规格如何处理 |
| --- | --- | --- | --- |
| 软件多 slot | llama.cpp 已有 | C0 | 保留,但不冒充 NPU batch |
| 单图 B1 串行 | 当前产品 | C0 | 稳定基线 |
| B1 ping/pong 两 job | SDK 有线索 | C1/C2 | runner 后灰度 |
| decode B4/B8/B16 | 需要新 calbin | C3 | 定义 scheduler/接口合同 |
| continuous batching | 需要 batch calbin+KV lane | C2/C3 | 实现目标 |
| 单芯粒小模型 | 需要新产物 | C3 | 只定义生命周期 |
| 同板多模型常驻 | 0.7.6 未证明 | C3 | 不当现有能力 |
| 多物理设备 | 0.7.6 明确不支持 | C4 | 只能多进程/新 Runtime |
## 2. 调度对象
```cpp
using RequestId = uint64_t;
using ModelGenerationId = uint64_t;
enum class RequestState : uint8_t {
New, WaitingModel, WaitingAdmission, PrefillReady, PrefillRunning,
DecodeReady, DecodeRunning, Finishing, Completed, Cancelled, Failed
};
struct SequenceState {
RequestId request_id;
calrt::KvSeqId seq_id;
ModelGenerationId generation;
RequestState state;
uint32_t prompt_tokens;
uint32_t generated_tokens;
uint32_t kv_length;
uint32_t max_new_tokens;
std::chrono::steady_clock::time_point deadline;
std::optional<uint16_t> hardware_lane;
uint64_t virtual_finish_time;
};
struct BatchProfile {
std::string exact_submodel;
uint16_t compiled_batch;
uint32_t max_seq;
PrimitiveType token_dtype;
PrimitiveType logit_dtype;
InactiveLanePolicy inactive_policy;
KvLayoutDescriptor kv_layout;
};
struct DecodeBatch {
ModelGenerationId generation;
const BatchProfile *profile;
std::vector<SequenceState *> active;
std::vector<int32_t> token_ids;
std::vector<int32_t> positions;
std::vector<uint8_t> lane_valid;
KvBatchReservation kv;
};
```
`compiled_batch` 来自经过校验的 calbin,不可由配置随意改。当前 profile 只能为 1。
## 3. 线程和组件
```text
API threads
-> RequestRegistry
-> AdmissionController
-> PrefillScheduler / DecodeScheduler
-> DeviceExecutor (one per physical device in 0.7.6 process)
-> CompletionDispatcher
-> SamplerPool
-> StreamWriter
ModelManager
-> immutable ModelGeneration
-> RuntimeSession (vdev/calbin/KV/pools)
-> HealthRecovery
```
0.7.6 下一个进程内只建一个 physical VirtualDevice。多板部署采用一进程绑定一板和上层负载均衡,设备选择 API 需厂商确认。
## 4. RequestRegistry
接口:
```cpp
StatusOr<RequestHandle> create(InferenceRequest request);
Status cancel(RequestId id);
Status on_prefill_complete(RequestId id, JobResult result);
Status on_decode_complete(RequestId id, JobResult result);
Status finish(RequestId id, FinishReason reason);
```
规则:
- state transition 使用 request-level mutex 或单线程 actor;不能由 HTTP、completion、sampler 同时直接改字段。
- cancel 幂等。New/Waiting/Ready 可立即完成取消;Running 只设 flag,底层 job 仍 drain。
- terminal state 只能进入一次;promise/stream close 只能触发一次。
- 所有 callback 带 generation/job id,丢弃过期 callback 但仍完成底层资源回收。
## 5. AdmissionController
### 5.1 预算
```cpp
struct CapacityBudget {
uint32_t max_sequences;
uint64_t max_kv_tokens;
uint32_t max_waiting_requests;
uint64_t max_waiting_prompt_tokens;
uint32_t max_inflight_jobs;
};
```
当前 Runtime `canAllocate(seq_num)` 只给 sequence resource 线索;服务仍需跟踪 token/KV 长度,避免接收必然超过 context 或内存的请求。
### 5.2 admission 结果
| 结果 | 对外 | 条件 |
| --- | --- | --- |
| Admit | 进入 prefill-ready | generation Ready 且预算充足 |
| Queue | 保持等待 | deadline 前有希望取得资源 |
| RejectInvalid | 400 | prompt/total context 越界 |
| RejectCapacity | 429 | 队列/token budget 已满 |
| RejectUnavailable | 503 | model draining/recovering/failed |
### 5.3 公平性
租户/模型使用 weighted fair queue
```text
cost = prompt_tokens + decode_weight * requested_output_tokens
virtual_finish = max(tenant_virtual_time, global_virtual_time) + cost/weight
```
长请求不能永久饿死,但也不能一次长 prefill 阻塞所有 decode。每租户限制 active sequences、waiting tokens 和 QPS;权重由控制面配置,不由请求参数指定。
## 6. PrefillScheduler
当前 B1 prefill:按 sequence 一次提交。队列分 `short <=512``medium <=4096``long >4096`,采用 weighted round robin,例如 decode:short:medium:long = `8:4:2:1` 的服务机会,具体值用负载实验调优。
如果图不可抢占,一个 40K prefill 仍会造成 head-of-line blocking。把 prompt 切块只有在以下条件全部成立才允许:
- compiler/Runtime 明确支持增量 prefill
- KV 写入与一次性 prefill 数值等价;
- position/CSR 语义清楚;
- chunk boundary 15/16/17 等通过;
- 端到端 token/golden 和性能通过。
否则 scheduler 只能通过 admission 和长任务配额控制,不能擅自 chunk。
## 7. DecodeScheduler
### 7.1 tick 算法
```cpp
StatusOr<DecodeBatch> build_decode_batch(const BatchProfile &p) {
auto candidates = ready_queue.take_eligible(p.compiled_batch);
DecodeBatch batch = allocate_host_tensors(p);
for (size_t lane = 0; lane < p.compiled_batch; ++lane) {
if (lane < candidates.size()) {
auto *seq = candidates[lane];
validate_generation_and_length(*seq, p);
batch.bind(lane, *seq, seq->last_token, seq->kv_length);
} else {
apply_inactive_lane_policy(batch, lane, p.inactive_policy);
}
}
CAL_ASSIGN_OR_RETURN(batch.kv, kv.reserve_batch(batch.active, p.kv_layout));
return batch;
}
```
### 7.2 选取策略
1. 排除 cancelled/deadline expired/generation mismatch。
2. 优先最早 deadline;同 deadline 采用 virtual finish/fair round robin。
3. 限制同一 tenant 占 lane 比例,除非无其他候选。
4. 可以等待短 `batching_window_us` 提高填充率,但等待不可让最早 deadline/SLA 超时。
5. 低流量直接发不满 batch,不能为追满 B16 无限等。
### 7.3 inactive lane
禁止默认填 token 0。必须从厂商 profile 得到:valid mask tensor/CSR、KV manager mode、logits 是否返回、inactive lane 是否写 KV。若没有 mask 契约,B>1 profile 不进入 Ready。
### 7.4 完成
按 lane 拆 logits,验证 request/generation/seq/lane 一致;每 lane 独立 sampling/finish。某一 lane 输入错误应在 submit 前拦截;设备级失败默认整个 batch 失败并标记所有相关 KV unknown,除非厂商提供 lane-level failure isolation。
## 8. Batch profile 选择
```cpp
const BatchProfile &select_profile(size_t ready, Duration oldest_wait) {
// 示例策略,不是固定生产阈值
if (ready >= 12 && profiles.has(16)) return profiles.b16();
if (ready >= 6 && profiles.has(8)) return profiles.b8();
if (ready >= 3 && profiles.has(4)) return profiles.b4();
return profiles.b1();
}
```
阈值由 B1/B4/B8/B16 active-lane A/B 得到。若 Runtime 不能同时常驻多个 calbin/profile,则不能动态选 profile,只能部署不同实例池,由上层路由低延迟/高吞吐流量。
## 9. KV lane 管理
### 9.1 状态
```cpp
enum class KvSequenceState : uint8_t { Free, Reserved, Active, Updating, Unknown, Releasing };
struct KvSequence {
KvSeqId seq_id;
uint16_t hardware_lane;
uint32_t committed_len;
uint32_t reserved_len;
ModelGenerationId generation;
KvSequenceState state;
};
```
### 9.2 不变量
- 同一 `seq_id` 在一个 generation 内只映射一个 active hardware lane。
- lane 复用前旧 sequence 已 Free 且 Runtime 确认释放。
- `committed_len <= reserved_len <= max_seq`
- batch 中 token position 等于该 lane `committed_len`(模型特殊语义另行 manifest)。
- submit 成功后 state=Updating;设备+输出成功才增加 committed_len。
- timeout/CCU fault 后 state=Unknown,不将 lane 分给其他请求。
### 9.3 context 操作支持
| 操作 | 当前策略 |
| --- | --- |
| append/decode | 当前主要路径,补事务 |
| remove tail | 只在现有 `RemoveTokensAtEnd` 边界通过后启用 |
| copy prefix/sequence | adapter 空实现,禁用 |
| keep one seq | 空实现,禁用 |
| position add/div | 空实现,禁用 |
| save/load session KV | 空实现,禁用 |
| shift | SDK 有 `DoShift`,但 llama 语义未闭环,C1/C2 |
Prompt cache、fork、beam-search KV copy 都依赖这些语义,不能在 API 上先声明支持。
## 10. 模型 generation
```cpp
enum class ModelState : uint8_t {
Unloaded, Validating, Configuring, Warming, Ready,
Draining, Unloading, Recovering, Failed
};
struct ModelGeneration {
ModelGenerationId id;
std::string alias;
std::string calbin_sha256;
ModelState state;
std::shared_ptr<RuntimeSession> runtime;
std::atomic<uint64_t> active_requests;
std::atomic<uint64_t> inflight_jobs;
};
```
请求创建时固定 generation,不在中途随 alias 切换。alias 指针只在新 generation golden/warmup 后原子替换。
## 11. load 实现
```text
load(alias, package)
1. resolve trusted package path
2. verify signature + all hashes
3. run static validator and compatibility
4. capacity plan (outside lifecycle lock)
5. acquire lifecycle write lock; stop new admission if replacement
6. drain old generation if 0.7.6 cannot co-reside
7. CreateCalbin -> configure -> KvManager -> pools
8. inspect exact tensors/CSR/profile
9. run golden + warmup
10. publish alias -> new generation
11. retain old signed artifact for rollback
```
步骤 7 后失败需清理新 generation;不能破坏仍可服务的旧 generation。0.7.6 若 configure 是全设备状态,则 replacement 前必须 drain 旧模型,无法做到真正零停机;用另一实例/板做蓝绿。
## 12. unload 实现
```text
Ready -> Draining
reject/route new requests
wait active_requests and inflight_jobs to zero until deadline
cancel waiting requests; submitted jobs still drain
free sequences/KV, destroy pools
destroy Runtime buffers/KvManager/Calbin
ResetConfiguration only per verified contract
verify device memory/health
-> Unloaded
```
若 drain deadline 到达但仍有 submitted job,不能释放其 buffer。进入 Recovering,按 timeout/reset 契约处理。强制 `delete` 不是取消。
## 13. 多模型策略
### 13.1 0.7.6 可落地方案
| 方案 | 同时常驻 | 隔离 | 适用 |
| --- | --- | --- | --- |
| 单实例单模型 | 否 | 清楚 | 当前推荐 |
| 同板时间切换 | 否 | 需 drain/configure | 低频模型 |
| 厂商组合 calbin | 可能 | compiler 统一规划 | 有正式产物时 |
| 一板一进程/模型 | 取决于板数 | 进程级 | 横向扩展 |
当前参考模型参数 16.938 GiBBF16 KV 3.750 GiB,总 DRAM reservation 推断 28.150 GiB/板,不能假设还有空间再常驻第二大模型。
### 13.2 厂商多模型 Runtime 契约
需要 model handle/namespace、独立 quota/address/KV/queue、job 显式 model handle、configure/unload 原子性、reset domain、参数共享生命周期。`current calbin` 全局状态不能支撑安全多模型。
### 13.3 两芯粒各一小模型
必须有两个单芯粒 calbin,以及 chip-scoped configure/queue/memory/reset 隔离。`ReadMem(...chipId)` 不证明这些能力。未取得 runner 和故障隔离前保持 C3。
## 14. 调度与 ping/pong 的组合
ping/pong 是 job engine/bank 选择,batch 是一个 job 内多 lane,二者不同:
```text
engine ping: DecodeBatch A [B lanes]
engine pong: DecodeBatch B [B lanes]
```
最大可能 in-flight 为 2 并不自动等于 2B 同时计算;需要 trace 证明 overlap。两个 batch 必须各有独立 InputBuf/OutputBuf/backing storage 和 KV transaction。
自动降级条件:任一输出串扰、status/counter 异常、P99 退化超门槛、连续 DeviceBusy/timeout、错误恢复不确定。降级只影响新 job;已提交 ping/pong 仍 drain。
## 15. 调度器配置
```yaml
schema: calculet-scheduler/v1
model_alias: qwen3
admission:
max_waiting_requests: 128
max_waiting_prompt_tokens: 262144
deadline_default_ms: 120000
queues:
decode_weight: 8
short_prefill_weight: 4
medium_prefill_weight: 2
long_prefill_weight: 1
short_prefill_max_tokens: 512
medium_prefill_max_tokens: 4096
batch:
profiles: [1]
batching_window_us: 0
executor:
max_inflight: 1
fixed_engine: false
recovery:
max_resets_per_10min: 2
restart_on_unknown_dma: true
```
当前产品只能 profiles `[1]`。配置 loader 必须将 requested profile 与 calbin manifest 交叉验证,配置中出现不存在的 B8 直接启动失败。
## 16. 指标
| 组件 | 指标 |
| --- | --- |
| admission | waiting requests/tokens、admit/reject/timeout、tenant quota |
| queues | depth/oldest wait by decode/short/medium/long |
| batch | compiled B、active lanes、fill ratio、batch wait、lane waste |
| KV | sequences by state、tokens used/reserved、lane wait、unknown、free latency |
| executor | in-flight、Ping/Pong、submit/Wait rc、status、queue counters |
| model | generation/state、load/configure/warmup/drain/unload duration |
| fairness | per-tenant service tokens、virtual lag、starvation events |
| recovery | fault reason、quarantined jobs/buffers/KV、reset/restart、recovery time |
所有 label 限制 cardinalityrequest_id/job_id 只进 trace/log,不作为 Prometheus label。
## 17. 单元和模型测试
### 17.1 纯调度单测
- Ready 队列 0..B+1 时 lane 选择。
- deadline、cancel、generation mismatch 被跳过。
- tenant 权重和最大 lane 比例。
- batching window 到期必发,不能饿死。
- long prefill 在持续 decode 下仍得到最低服务份额。
- profile 不存在时 fail-fast。
使用 fake clock,不用 sleep 驱动时间测试。
### 17.2 KV property test
随机生成 reserve/append/remove/free/cancel/fault 序列,持续断言 lane 唯一、length 范围、terminal 资源回收、Unknown 不复用。与一个纯内存 reference state machine 对照。
### 17.3 并发测试
TSAN build 下 1000 请求随机 cancel/deadline/shutdown;检查无 data race、double completion、promise double set、lease double free。Mock Wait 随机延迟打乱 completion 顺序,验证 generation/job correlation。
## 18. NPU 集成矩阵
| 测试 | 当前 B1 | B4/8/16 产物后 |
| --- | --- | --- |
| 单 seq prefill+decode | 必须 | 必须 |
| 两 seq 交替 | 必须 | 必须 |
| lane 1..B 填充 | 不适用 | 必须 |
| inactive lane 不写 KV | 不适用 | 必须 |
| mixed sequence length | 软件交替 | 同 batch 必须 |
| batch 某请求 cancel | queued/running | lane result 丢弃但整 batch drain |
| timeout/CCU fault | generation 恢复 | 全 lanes 一致恢复 |
| ping/pong two jobs | 实验 | 每 job 可含 B lanes |
| load/unload soak | 必须 | 每 profile/组合 |
## 19. 验收
M0 B1 稳定:24h、无 abort/忙等/串扰/泄漏,错误码全部可见。
M1 async in-flight=1:行为与同步 baseline 等价,P50 <=3%/P99 <=5% 退化。
M2 ping/pongaggregate +10% 以上、P99 退化 <=5%、24h 无错误,才默认启用。
M3 batch:所有 lane 正确,fill ratio/吞吐/P99 达项目目标,KV fault 可恢复;没有这些数据不宣传 continuous batching。
M4 multi-model lifecycle100 次 load/unload、8h 模型切换 soak、资源回落、generation 隔离、旧版本一键回滚。
## 20. 厂商必须提供
- B4/B8/B16 exact tensor/CSR/inactive mask/KV layout。
- 每 profile calbin/golden/capacity/performance 和 compatibility。
- PING/PONG/parallel 的时序与最大 in-flight。
- batch job 失败是整 job 还是 lane-levelKV 污染范围。
- 多 calbin 常驻、chip affinity、reset domain 的正式 API。
- 0.7.6 后多物理设备的版本/API/设备选择。
在这些交付前,设计可以实现和单测,但硬件能力仍是 C3;文档不得将设计完成写成设备已支持。