Files
calculet-npu-research-archive/reports/Calculet-NPU-Batch-KV与多模型调度实现规格-20260802.md
T

16 KiB
Raw Blame History

Calculet NPU Batch、KV 与多模型调度实现规格

日期:2026-08-02
基线:CalRT 0.7.6 单物理 VirtualDevice、当前双芯粒 Qwen3 calbin batch=1

1. 能力分层

当前/目标 状态 本规格如何处理
软件多 slot llama.cpp 已有 C0 保留,但不冒充 NPU batch
单图 B1 串行 当前产品 C0 稳定基线
B1 ping/pong 两 job SDK 有线索 C1/C2 runner 后灰度
decode B4/B8/B16 需要新 calbin C3 定义 scheduler/接口合同
continuous batching 需要 batch calbin+KV lane C2/C3 实现目标
单芯粒小模型 需要新产物 C3 只定义生命周期
同板多模型常驻 0.7.6 未证明 C3 不当现有能力
多物理设备 0.7.6 明确不支持 C4 只能多进程/新 Runtime

2. 调度对象

using RequestId = uint64_t;
using ModelGenerationId = uint64_t;

enum class RequestState : uint8_t {
    New, WaitingModel, WaitingAdmission, PrefillReady, PrefillRunning,
    DecodeReady, DecodeRunning, Finishing, Completed, Cancelled, Failed
};

struct SequenceState {
    RequestId request_id;
    calrt::KvSeqId seq_id;
    ModelGenerationId generation;
    RequestState state;
    uint32_t prompt_tokens;
    uint32_t generated_tokens;
    uint32_t kv_length;
    uint32_t max_new_tokens;
    std::chrono::steady_clock::time_point deadline;
    std::optional<uint16_t> hardware_lane;
    uint64_t virtual_finish_time;
};

struct BatchProfile {
    std::string exact_submodel;
    uint16_t compiled_batch;
    uint32_t max_seq;
    PrimitiveType token_dtype;
    PrimitiveType logit_dtype;
    InactiveLanePolicy inactive_policy;
    KvLayoutDescriptor kv_layout;
};

struct DecodeBatch {
    ModelGenerationId generation;
    const BatchProfile *profile;
    std::vector<SequenceState *> active;
    std::vector<int32_t> token_ids;
    std::vector<int32_t> positions;
    std::vector<uint8_t> lane_valid;
    KvBatchReservation kv;
};

compiled_batch 来自经过校验的 calbin,不可由配置随意改。当前 profile 只能为 1。

3. 线程和组件

API threads
  -> RequestRegistry
  -> AdmissionController
  -> PrefillScheduler / DecodeScheduler
  -> DeviceExecutor (one per physical device in 0.7.6 process)
  -> CompletionDispatcher
  -> SamplerPool
  -> StreamWriter

ModelManager
  -> immutable ModelGeneration
  -> RuntimeSession (vdev/calbin/KV/pools)
  -> HealthRecovery

0.7.6 下一个进程内只建一个 physical VirtualDevice。多板部署采用一进程绑定一板和上层负载均衡,设备选择 API 需厂商确认。

4. RequestRegistry

接口:

StatusOr<RequestHandle> create(InferenceRequest request);
Status cancel(RequestId id);
Status on_prefill_complete(RequestId id, JobResult result);
Status on_decode_complete(RequestId id, JobResult result);
Status finish(RequestId id, FinishReason reason);

规则:

  • state transition 使用 request-level mutex 或单线程 actor;不能由 HTTP、completion、sampler 同时直接改字段。
  • cancel 幂等。New/Waiting/Ready 可立即完成取消;Running 只设 flag,底层 job 仍 drain。
  • terminal state 只能进入一次;promise/stream close 只能触发一次。
  • 所有 callback 带 generation/job id,丢弃过期 callback 但仍完成底层资源回收。

5. AdmissionController

5.1 预算

struct CapacityBudget {
    uint32_t max_sequences;
    uint64_t max_kv_tokens;
    uint32_t max_waiting_requests;
    uint64_t max_waiting_prompt_tokens;
    uint32_t max_inflight_jobs;
};

当前 Runtime canAllocate(seq_num) 只给 sequence resource 线索;服务仍需跟踪 token/KV 长度,避免接收必然超过 context 或内存的请求。

5.2 admission 结果

结果 对外 条件
Admit 进入 prefill-ready generation Ready 且预算充足
Queue 保持等待 deadline 前有希望取得资源
RejectInvalid 400 prompt/total context 越界
RejectCapacity 429 队列/token budget 已满
RejectUnavailable 503 model draining/recovering/failed

5.3 公平性

租户/模型使用 weighted fair queue

cost = prompt_tokens + decode_weight * requested_output_tokens
virtual_finish = max(tenant_virtual_time, global_virtual_time) + cost/weight

长请求不能永久饿死,但也不能一次长 prefill 阻塞所有 decode。每租户限制 active sequences、waiting tokens 和 QPS;权重由控制面配置,不由请求参数指定。

6. PrefillScheduler

当前 B1 prefill:按 sequence 一次提交。队列分 short <=512medium <=4096long >4096,采用 weighted round robin,例如 decode:short:medium:long = 8:4:2:1 的服务机会,具体值用负载实验调优。

如果图不可抢占,一个 40K prefill 仍会造成 head-of-line blocking。把 prompt 切块只有在以下条件全部成立才允许:

  • compiler/Runtime 明确支持增量 prefill
  • KV 写入与一次性 prefill 数值等价;
  • position/CSR 语义清楚;
  • chunk boundary 15/16/17 等通过;
  • 端到端 token/golden 和性能通过。

否则 scheduler 只能通过 admission 和长任务配额控制,不能擅自 chunk。

7. DecodeScheduler

7.1 tick 算法

StatusOr<DecodeBatch> build_decode_batch(const BatchProfile &p) {
    auto candidates = ready_queue.take_eligible(p.compiled_batch);
    DecodeBatch batch = allocate_host_tensors(p);

    for (size_t lane = 0; lane < p.compiled_batch; ++lane) {
        if (lane < candidates.size()) {
            auto *seq = candidates[lane];
            validate_generation_and_length(*seq, p);
            batch.bind(lane, *seq, seq->last_token, seq->kv_length);
        } else {
            apply_inactive_lane_policy(batch, lane, p.inactive_policy);
        }
    }

    CAL_ASSIGN_OR_RETURN(batch.kv, kv.reserve_batch(batch.active, p.kv_layout));
    return batch;
}

7.2 选取策略

  1. 排除 cancelled/deadline expired/generation mismatch。
  2. 优先最早 deadline;同 deadline 采用 virtual finish/fair round robin。
  3. 限制同一 tenant 占 lane 比例,除非无其他候选。
  4. 可以等待短 batching_window_us 提高填充率,但等待不可让最早 deadline/SLA 超时。
  5. 低流量直接发不满 batch,不能为追满 B16 无限等。

7.3 inactive lane

禁止默认填 token 0。必须从厂商 profile 得到:valid mask tensor/CSR、KV manager mode、logits 是否返回、inactive lane 是否写 KV。若没有 mask 契约,B>1 profile 不进入 Ready。

7.4 完成

按 lane 拆 logits,验证 request/generation/seq/lane 一致;每 lane 独立 sampling/finish。某一 lane 输入错误应在 submit 前拦截;设备级失败默认整个 batch 失败并标记所有相关 KV unknown,除非厂商提供 lane-level failure isolation。

8. Batch profile 选择

const BatchProfile &select_profile(size_t ready, Duration oldest_wait) {
    // 示例策略,不是固定生产阈值
    if (ready >= 12 && profiles.has(16)) return profiles.b16();
    if (ready >= 6  && profiles.has(8))  return profiles.b8();
    if (ready >= 3  && profiles.has(4))  return profiles.b4();
    return profiles.b1();
}

阈值由 B1/B4/B8/B16 active-lane A/B 得到。若 Runtime 不能同时常驻多个 calbin/profile,则不能动态选 profile,只能部署不同实例池,由上层路由低延迟/高吞吐流量。

9. KV lane 管理

9.1 状态

enum class KvSequenceState : uint8_t { Free, Reserved, Active, Updating, Unknown, Releasing };

struct KvSequence {
    KvSeqId seq_id;
    uint16_t hardware_lane;
    uint32_t committed_len;
    uint32_t reserved_len;
    ModelGenerationId generation;
    KvSequenceState state;
};

9.2 不变量

  • 同一 seq_id 在一个 generation 内只映射一个 active hardware lane。
  • lane 复用前旧 sequence 已 Free 且 Runtime 确认释放。
  • committed_len <= reserved_len <= max_seq
  • batch 中 token position 等于该 lane committed_len(模型特殊语义另行 manifest)。
  • submit 成功后 state=Updating;设备+输出成功才增加 committed_len。
  • timeout/CCU fault 后 state=Unknown,不将 lane 分给其他请求。

9.3 context 操作支持

操作 当前策略
append/decode 当前主要路径,补事务
remove tail 只在现有 RemoveTokensAtEnd 边界通过后启用
copy prefix/sequence adapter 空实现,禁用
keep one seq 空实现,禁用
position add/div 空实现,禁用
save/load session KV 空实现,禁用
shift SDK 有 DoShift,但 llama 语义未闭环,C1/C2

Prompt cache、fork、beam-search KV copy 都依赖这些语义,不能在 API 上先声明支持。

10. 模型 generation

enum class ModelState : uint8_t {
    Unloaded, Validating, Configuring, Warming, Ready,
    Draining, Unloading, Recovering, Failed
};

struct ModelGeneration {
    ModelGenerationId id;
    std::string alias;
    std::string calbin_sha256;
    ModelState state;
    std::shared_ptr<RuntimeSession> runtime;
    std::atomic<uint64_t> active_requests;
    std::atomic<uint64_t> inflight_jobs;
};

请求创建时固定 generation,不在中途随 alias 切换。alias 指针只在新 generation golden/warmup 后原子替换。

11. load 实现

load(alias, package)
  1. resolve trusted package path
  2. verify signature + all hashes
  3. run static validator and compatibility
  4. capacity plan (outside lifecycle lock)
  5. acquire lifecycle write lock; stop new admission if replacement
  6. drain old generation if 0.7.6 cannot co-reside
  7. CreateCalbin -> configure -> KvManager -> pools
  8. inspect exact tensors/CSR/profile
  9. run golden + warmup
 10. publish alias -> new generation
 11. retain old signed artifact for rollback

步骤 7 后失败需清理新 generation;不能破坏仍可服务的旧 generation。0.7.6 若 configure 是全设备状态,则 replacement 前必须 drain 旧模型,无法做到真正零停机;用另一实例/板做蓝绿。

12. unload 实现

Ready -> Draining
  reject/route new requests
  wait active_requests and inflight_jobs to zero until deadline
  cancel waiting requests; submitted jobs still drain
  free sequences/KV, destroy pools
  destroy Runtime buffers/KvManager/Calbin
  ResetConfiguration only per verified contract
  verify device memory/health
  -> Unloaded

若 drain deadline 到达但仍有 submitted job,不能释放其 buffer。进入 Recovering,按 timeout/reset 契约处理。强制 delete 不是取消。

13. 多模型策略

13.1 0.7.6 可落地方案

方案 同时常驻 隔离 适用
单实例单模型 清楚 当前推荐
同板时间切换 需 drain/configure 低频模型
厂商组合 calbin 可能 compiler 统一规划 有正式产物时
一板一进程/模型 取决于板数 进程级 横向扩展

当前参考模型参数 16.938 GiBBF16 KV 3.750 GiB,总 DRAM reservation 推断 28.150 GiB/板,不能假设还有空间再常驻第二大模型。

13.2 厂商多模型 Runtime 契约

需要 model handle/namespace、独立 quota/address/KV/queue、job 显式 model handle、configure/unload 原子性、reset domain、参数共享生命周期。current calbin 全局状态不能支撑安全多模型。

13.3 两芯粒各一小模型

必须有两个单芯粒 calbin,以及 chip-scoped configure/queue/memory/reset 隔离。ReadMem(...chipId) 不证明这些能力。未取得 runner 和故障隔离前保持 C3。

14. 调度与 ping/pong 的组合

ping/pong 是 job engine/bank 选择,batch 是一个 job 内多 lane,二者不同:

engine ping: DecodeBatch A [B lanes]
engine pong: DecodeBatch B [B lanes]

最大可能 in-flight 为 2 并不自动等于 2B 同时计算;需要 trace 证明 overlap。两个 batch 必须各有独立 InputBuf/OutputBuf/backing storage 和 KV transaction。

自动降级条件:任一输出串扰、status/counter 异常、P99 退化超门槛、连续 DeviceBusy/timeout、错误恢复不确定。降级只影响新 job;已提交 ping/pong 仍 drain。

15. 调度器配置

schema: calculet-scheduler/v1
model_alias: qwen3
admission:
  max_waiting_requests: 128
  max_waiting_prompt_tokens: 262144
  deadline_default_ms: 120000
queues:
  decode_weight: 8
  short_prefill_weight: 4
  medium_prefill_weight: 2
  long_prefill_weight: 1
  short_prefill_max_tokens: 512
  medium_prefill_max_tokens: 4096
batch:
  profiles: [1]
  batching_window_us: 0
executor:
  max_inflight: 1
  fixed_engine: false
recovery:
  max_resets_per_10min: 2
  restart_on_unknown_dma: true

当前产品只能 profiles [1]。配置 loader 必须将 requested profile 与 calbin manifest 交叉验证,配置中出现不存在的 B8 直接启动失败。

16. 指标

组件 指标
admission waiting requests/tokens、admit/reject/timeout、tenant quota
queues depth/oldest wait by decode/short/medium/long
batch compiled B、active lanes、fill ratio、batch wait、lane waste
KV sequences by state、tokens used/reserved、lane wait、unknown、free latency
executor in-flight、Ping/Pong、submit/Wait rc、status、queue counters
model generation/state、load/configure/warmup/drain/unload duration
fairness per-tenant service tokens、virtual lag、starvation events
recovery fault reason、quarantined jobs/buffers/KV、reset/restart、recovery time

所有 label 限制 cardinalityrequest_id/job_id 只进 trace/log,不作为 Prometheus label。

17. 单元和模型测试

17.1 纯调度单测

  • Ready 队列 0..B+1 时 lane 选择。
  • deadline、cancel、generation mismatch 被跳过。
  • tenant 权重和最大 lane 比例。
  • batching window 到期必发,不能饿死。
  • long prefill 在持续 decode 下仍得到最低服务份额。
  • profile 不存在时 fail-fast。

使用 fake clock,不用 sleep 驱动时间测试。

17.2 KV property test

随机生成 reserve/append/remove/free/cancel/fault 序列,持续断言 lane 唯一、length 范围、terminal 资源回收、Unknown 不复用。与一个纯内存 reference state machine 对照。

17.3 并发测试

TSAN build 下 1000 请求随机 cancel/deadline/shutdown;检查无 data race、double completion、promise double set、lease double free。Mock Wait 随机延迟打乱 completion 顺序,验证 generation/job correlation。

18. NPU 集成矩阵

测试 当前 B1 B4/8/16 产物后
单 seq prefill+decode 必须 必须
两 seq 交替 必须 必须
lane 1..B 填充 不适用 必须
inactive lane 不写 KV 不适用 必须
mixed sequence length 软件交替 同 batch 必须
batch 某请求 cancel queued/running lane result 丢弃但整 batch drain
timeout/CCU fault generation 恢复 全 lanes 一致恢复
ping/pong two jobs 实验 每 job 可含 B lanes
load/unload soak 必须 每 profile/组合

19. 验收

M0 B1 稳定:24h、无 abort/忙等/串扰/泄漏,错误码全部可见。

M1 async in-flight=1:行为与同步 baseline 等价,P50 <=3%/P99 <=5% 退化。

M2 ping/pongaggregate +10% 以上、P99 退化 <=5%、24h 无错误,才默认启用。

M3 batch:所有 lane 正确,fill ratio/吞吐/P99 达项目目标,KV fault 可恢复;没有这些数据不宣传 continuous batching。

M4 multi-model lifecycle100 次 load/unload、8h 模型切换 soak、资源回落、generation 隔离、旧版本一键回滚。

20. 厂商必须提供

  • B4/B8/B16 exact tensor/CSR/inactive mask/KV layout。
  • 每 profile calbin/golden/capacity/performance 和 compatibility。
  • PING/PONG/parallel 的时序与最大 in-flight。
  • batch job 失败是整 job 还是 lane-levelKV 污染范围。
  • 多 calbin 常驻、chip affinity、reset domain 的正式 API。
  • 0.7.6 后多物理设备的版本/API/设备选择。

在这些交付前,设计可以实现和单测,但硬件能力仍是 C3;文档不得将设计完成写成设备已支持。