chore: Add Chinese README and refactor NPU backend to use Unix domain sockets

- Introduced a new Chinese version of the README (README_CN.md) to provide localized documentation for AgentOS.
- Refactored the NPU bridge to utilize Unix domain sockets instead of HTTP loopback, enhancing security and performance.
- Updated the NPU backend to include a server socket configuration, ensuring proper communication over Unix sockets.
- Modified the Ntex client to support both network and Unix socket transports, improving flexibility in backend communication.
- Adjusted validation logic to enforce the use of Unix sockets for local model endpoints, rejecting loopback HTTP addresses.
- Enhanced error messages and documentation throughout the codebase to clarify the new socket-based architecture.
This commit is contained in:
emmettlu
2026-08-02 15:48:09 +08:00
parent eee7fed161
commit f863f83960
8 changed files with 1157 additions and 65 deletions
+209
View File
@@ -0,0 +1,209 @@
# AgentOS 工程约定
本文件面向在本仓库内工作的代码代理。修改前先读本文件、根
`Cargo.toml`,再读目标 crate 的源码。不要根据旧 Python 版本、Octos 或
`npu_features/` 中的厂商样例臆测当前行为;`crates/` 才是实现事实来源。
## 项目目标
AgentOS 是运行在裸 Linux 上的 low-level OS agent,而不是 Linux 发行版、
容器沙箱或内核 fork。核心原则是尽量复用 Linux 原生能力:进程 UID/GID、
文件权限、Unix 凭据、pidfd、`no_new_privs`、Btrfs subvolume/snapshot 和设备
ioctl。
当前安全模型是稳定的 `AgentId -> (owner UID, agent UID, agent GID)` 绑定:
- agent 进程不能以 UID 0 或 GID 0 运行;
- 每次执行前必须核对进程的 effective UID/GID
- 一个状态数据库中,agent UID 只能绑定给一个 active `AgentId`
- 不额外构造 namespace、容器或通用 sandbox
- `no_new_privs` 是补充保护,不能替代 UID/GID 和工具策略;
- 模型不能获得任意 shell,更不能直接获得 root shell。
不要把 `owner_uid``agent_uid` 混为一谈。前者表示拥有者,后者才是内核
实际执行主体。新增 daemon 或 broker 时,Unix socket 对端必须用内核提供的
peer credentials 鉴权,不能信任请求体里的 UID。
## 当前成熟度边界
已经可用:
- `agentos` one-shot CLI、doctor 和确定性 fake backend
- OpenAI-compatible Chat Completions / Responses 非流式后端;
- `agentos.chat.v1` 内部协议、严格工具 schema、预算和循环检测;
- 只读 Linux 工具、SQLite 审计/记忆、Btrfs worldline 与 copy fallback
- CALCULET PCIe ABI、Rust ioctl/DMA 包装、Calbin 解析、张量/命令结构;
- Candle/Qwen3 host pipeline、tool-call 模板和 CALRT 张量适配;
- 旧 llama-server UDS bridgeHTTP 消息格式直接承载在 AF_UNIX 上,不经过 TCP。
尚未完成:
- 纯 Rust CALRT 的 CCU relocation、job launch/completion 和设备 KV reset
- 可在真实 NPU 上完成推理的 `npu_candle` backend
- typed mutation broker 和任何变更型系统工具;
- 长驻 AgentOS control daemon、control UDS RPC、服务安装和完整硬件端到端测试。
`ConfiguredRuntime::submit()` 当前必须 fail closed 并返回
`HardwareExecutionUnavailable`。在没有真实板卡证据前,不得把
`npu_candle` 标记为 ready,不得让 `auto` 选择它,也不得用 mock 测试宣称
硬件推理已验证。
## Workspace 与模块所有权
`agentos-cli` 外,`crates/` 下均为 library crate。
| Crate | 职责 |
| --- | --- |
| `agentos-cli` | `agentos``agentos-npu-bridge` 两个 binary;只做参数解析和装配入口 |
| `agentos-runtime` | composition root、环境配置、identity/backend/tool 装配、doctor |
| `agentos-agent` | 有预算的 agent loop、工具调度、循环检测、审计事件 |
| `agentos-protocol` | provider-neutral `agentos.chat.v1` 类型与 `ChatBackend` trait |
| `agentos-inference` | ntex HTTP、OpenAI-compatible、subprocess、fake/scripted backend |
| `agentos-tools` | fail-closed registry、schema 校验、审批契约和 Linux 只读工具 |
| `agentos-core` | ID、principal、execution identity、审计和 worldline 公共类型 |
| `agentos-kernel` | Linux/rustix 边界、凭据、pidfd、文件原子写、Btrfs ioctl |
| `agentos-memory` | SQLite WAL、FTS5、principal 绑定、memory 和 audit event |
| `agentos-worldline` | 类 Git 的 filesystem history、branch/commit/diff/rollback/recovery |
| `agentos-npu` | NPU probe、旧 bridge、Rust Candle NPU backend 装配 |
| `agentos-candle` | Qwen3 模板/tokenizer/sampling 与 Calbin prefill/decode host runner |
| `calculet-pcie-abi` | 驱动 0.9.0 / ABI 1.0.0 的 64-bit Linux ioctl 布局 |
| `calculet-pcie` | 安全的设备、BAR、DMA、MSI、reset 和 board/process API |
| `calculet-calrt` | Calbin、allocator、tensor、command stream、DeviceIo 和 runtime 骨架 |
保持依赖方向从高层向低层:
```text
agentos-cli -> agentos-runtime -> agentos-agent/tools/inference/npu/worldline/memory
agentos-npu -> agentos-candle -> calculet-calrt -> calculet-pcie -> calculet-pcie-abi
agentos-kernel/core/protocol 是底层公共边界
```
不要让底层 crate 反向依赖 CLI 或 runtime。跨 provider 的消息类型放在
`agentos-protocol`Linux syscall/ABI 放在 `agentos-kernel` 或对应的
`calculet-*` cratecomposition 逻辑只放在 `agentos-runtime`
## 不可破坏的设计约束
### Linux 与系统调用
- 项目只支持 Linux;当前 CALCULET ABI 只验证过 64-bit Linux。
- 项目自有的低层 syscall/ioctl 优先走 `rustix`,不要直接新增 `libc` 调用。
- 这不表示最终二进制完全不链接 libc;Rust `std`、OpenSSL 等依赖仍可使用
系统 ABI。
- `unsafe` 仅允许出现在无法避免的 ABI 边界,必须就结构布局、指针生命周期
和 opcode 写英文 `SAFETY` 注释,并在调用前完成长度、对齐和范围校验。
- 不要为了 worldline patch 内核。优先使用现有 Btrfs ioctl;非 Btrfs 环境保留
copy fallback。
### 工具与权限
- registry 必须显式 allowlist;未知工具一律拒绝。
- 所有 tool schema 默认 strictobject、列出全部 required、
`additionalProperties: false`
- 参数和输出必须有字节上限,外部命令必须有 deadline、`kill_on_drop`,并移除
API key 环境变量。
- 只有标记为 `ConcurrencyClass::Safe` 的只读工具可以并行;变更工具必须
exclusive。
- 现有 `ApprovalLedger` 只是进程内、单次、精确参数绑定的契约。实现 mutation
前还必须有 typed broker、内核凭据校验、持久审计和失败恢复,不能把审批 ID
当作任意命令授权。
- 不得增加通用 `shell``exec`、任意路径写入或任意 systemd unit 工具。
### 推理后端
- 内部统一使用 `agentos.chat.v1`provider 差异留在 backend adapter。
- 缺少 NPU 硬件时的 mock 是远程 OpenAI-compatible 模型,不是内置
`FakeBackend`。远程传输可以使用 HTTP 或 HTTPS;使用 HTTPS 时必须验证 peer。
- 本机常驻模型服务必须使用 filesystem Unix domain socket,禁止监听或连接
localhost、loopback 或其他 TCP 地址。纯 Rust Candle in-process 路径不需要 IPC。
- 本机 llama-server 可以保留 OpenAI-compatible HTTP 消息格式,但必须通过
ntex 自定义 connector 承载在 AF_UNIX 上,用户配置只能暴露 socket path,不能
接受本机 URL。
- OpenAI-compatible HTTP 和 UDS HTTP client 使用 `ntex`,不要引入 `reqwest`
- 禁止自动 redirect;响应大小、超时和 retry 必须有界。
- 支持 Chat Completions 与 Responses 两种方言,但不要假设所有兼容服务支持
完全相同的字段。新增兼容逻辑必须有请求构造和响应解析测试。
- subprocess backend 只通过 stdin/stdout 传 JSONstderr 仅用于有界错误信息;
不要把 secret 传给子进程。它只作为显式测试/bridge adapter,不能代表本机
常驻模型传输,也不能进入 `auto`
- `auto` 当前顺序是 UDS legacy NPU、远程 OpenAI-compatible。Fake 和
subprocess 必须显式选择,Rust Candle NPU 在硬件提交完成前不能进入 auto。
### NPU
- `npu_features/` 是忽略提交的厂商源码/部署抓取,仅作逆向参考,不能成为
发布包运行时依赖。
- 原始版本边界是 driver package 0.9.0、driver ABI 1.0.0、CALRT 0.7.6。
- ioctl struct 使用 `repr(C, packed)`,任何改动都必须同步 size/opcode 测试。
- DMA 单次上限 8 MiB,已知 H2C/C2H 各 8 个 channel;不要绕过现有验证。
- Calbin 参数部署会真实写设备。没有用户明确要求和真实硬件测试计划时,不要
默认开启 `deploy_parameters`
- captured Qwen3 fixture 的 logits 是 151936tokenizer vocab 是 151669;额外
267 个 padded logits 必须在采样前屏蔽。
- host-side `MockDevice` 测试只能验证解析、地址、buffer 和命令编码,不能证明
job submission、同步、KV cache 或输出数值正确。远程模型 mock 同样不能证明
NPU 硬件正确。
- legacy llama-server 的默认 socket 是 `/run/agentos/npu.sock`doctor 必须检查
它确实是可写的 Unix socket,不能仅检查路径存在。
### Worldline 与持久化
- Btrfs commit tree 使用 readonly snapshotbranch/current 使用 writable
snapshot。
- copy fallback 用于开发和无 Btrfs 环境,但当前不会提供内核强制的只读
commit tree。不要在文档里把它描述成与 Btrfs 等价的不可变性。
- commit 是内容/元数据哈希标识;rollback 必须创建新 commit,不能改写历史。
- checkout 要保留 transaction journal、目录 fsync 和可恢复的 replaced tree。
- 分支提交必须检查 base HEAD,禁止 stale branch 覆盖新 HEAD。
- SQLite principal 绑定不可静默重绑;schema 变更要考虑已有数据库升级。
## Rust 与依赖规则
- 使用 workspace 的 Rust edition,不添加 MSRV 或 `rust-toolchain.toml`
- 第三方依赖集中写在根 `Cargo.toml``[workspace.dependencies]`
- 版本只写主版本号,例如 `serde = "1"`,不要固定 `1.2.3`
- 内部 crate 统一用 `*.workspace = true`
- 读取依赖源码时,从 `$CARGO_HOME/registry/src` 找实际锁定版本,不靠记忆猜 API。
- 优先复用已有 crate 和抽象,不复制协议类型、HTTP client、schema validator、
runtime probe 或设备 ABI。
- 代码注释默认英文;用户可见文档分别维护英文与中文。
- 不要加入 Python、`reqwest` 或直接 `libc` 依赖。
- 保持开发/测试 profile 的快速编译取向,除非有基准数据,不要随意调高 dev
优化或减少 codegen units。
## 修改流程
1.`rg` 定位实现和调用方,先确认改动属于哪个 crate。
2. 先写清安全边界和失败模式;系统层能力默认 fail closed。
3. 修改公共协议时,同步所有 backend、tool adapter、CLI 和双语 README。
4. 新增环境变量时,同步 `RuntimeConfig::from_env`、doctor 和配置表。
5. 新增 NPU ABI 时,对照 `npu_features/cal-pcie-0.9.0` 或 CALRT 源码,并补布局、
opcode、边界和 mock 测试。
6. 不要改写或删除用户的 `npu_features/` 抓取、模型文件和未提交工作。
完成前至少执行:
```bash
cargo fmt --all --check
cargo test --workspace
cargo clippy --workspace --all-targets --all-features -- -D warnings
```
涉及 CLI 时还要实际运行相关命令;涉及 OpenAI-compatible 时至少覆盖两种 API
方言的 serialization/parsing;涉及硬件而本机无板卡时,明确报告未执行的验证。
部分 NPU 测试依赖本地、被 `.gitignore` 忽略的
`npu_features/snapshot_20260801`。fixture 不存在时,应将硬件抓取测试与普通
workspace 测试分层,而不是把模型数据提交进仓库或伪造通过结果。
## 完成标准
一次改动只有在以下条件都满足时才算完成:
- crate 边界和依赖方向没有被破坏;
- 非法输入、权限不足、后端缺失和硬件缺失均 fail closed
- 无 secret 出现在日志、tool output、子进程或 doctor 报告中;
- 单元/集成测试覆盖正常路径与关键拒绝路径;
- fmt、workspace test、严格 clippy 通过;
- README 与实际成熟度一致,未把 host-side `MockDevice` 或远程模型 mock 写成
NPU 硬件验证。
+351
View File
@@ -0,0 +1,351 @@
# AgentOS
[简体中文](README_CN.md)
AgentOS is a low-level, Linux-only operating-system agent written in Rust. It
combines a bounded tool-calling loop with native Linux identities, read-only
system inspection, durable audit state, Btrfs-backed filesystem history, and
multiple inference backends—including an in-progress pure-Rust path for a
CALCULET NPU.
> **Development status:** the remote OpenAI-compatible mock path and explicit
> subprocess/fake test adapters are implemented. Locally deployed model services
> use Unix domain sockets, never localhost or TCP. The Rust Candle/CALRT host
> pipeline can parse and validate the captured Qwen3 deployment, but CALRT
> hardware job submission is not implemented yet. `npu_candle` therefore fails
> closed and is not production-ready.
## Design
AgentOS is not a container sandbox or a kernel fork. Its isolation boundary is
the Linux process identity:
- every active `AgentId` is stably bound to an owner UID and a non-root agent
UID/GID;
- the runtime verifies the process effective UID/GID before every run;
- UID 0 or GID 0 agents are rejected;
- `no_new_privs` is enabled before model/tool execution;
- the model sees an explicit tool allowlist, never a raw root shell;
- system mutation is intentionally unavailable until a typed, authenticated
broker is implemented.
The current execution path is:
```text
agentos CLI
-> runtime + Linux principal
-> budgeted agent loop
-> ChatBackend (remote OpenAI-compatible / UDS NPU / in-process Candle)
-> strict ToolRegistry (read-only Linux tools)
-> SQLite audit log
filesystem state -> WorldlineStore -> Btrfs snapshots or copy fallback
NPU host path -> Candle/Qwen3 -> CALRT -> PCIe driver ABI
local model IPC -> Unix domain socket -> OpenAI-compatible HTTP payload
```
## What works today
- Provider-neutral `agentos.chat.v1` messages, tools, responses, usage, and
backend interface.
- Remote OpenAI-compatible non-streaming Chat Completions and Responses APIs
over `ntex`, with bounded responses, timeouts, retries, and verified TLS when
HTTPS is selected.
- OpenAI-compatible HTTP payloads over an AF_UNIX transport for the local legacy
llama-server, without a localhost/TCP hop.
- A JSON stdin/stdout subprocess adapter retained for explicit tests and custom
bridges; it is not the transport for a persistent local model service.
- Bounded agent iterations, tool calls, elapsed time, total tokens, parallel
read-only calls, duplicate call-ID checks, and no-progress/cycle detection.
- Strict JSON tool schemas, allowlists, argument/output limits, deadlines, and
hashed audit records.
- Read-only Linux, package, systemd, journal, and CALCULET NPU inspection.
- SQLite WAL state, FTS5 memory primitives, immutable principal bindings, and
append-only agent/tool audit events.
- Btrfs subvolume/snapshot worldlines with branch, diff, commit, log,
non-destructive rollback, stale-branch protection, and checkout recovery.
- Rust representations of CALCULET driver 0.9.0 / ABI 1.0.0, BAR and DMA access,
Calbin 0.7.6 parsing, memory allocation, tensors, and command encoding.
- A Candle-based Qwen3 host runner with chat templates, tool-call parsing,
sampling, prefill/decode contracts, BF16/F32 logits, and vendor tiled-logit
decoding.
## Workspace
`agentos-cli` is the only binary crate; it produces two binaries. Every other
workspace member is a library.
| Crate | Type | Responsibility |
| --- | --- | --- |
| `agentos-cli` | binary | `agentos` CLI and `agentos-npu-bridge` |
| `agentos-runtime` | library | configuration and composition root |
| `agentos-agent` | library | bounded agent loop and audit emission |
| `agentos-protocol` | library | provider-neutral chat/tool protocol |
| `agentos-inference` | library | ntex, remote, subprocess, and test backends |
| `agentos-tools` | library | tool policy, schema validation, approvals, Linux tools |
| `agentos-core` | library | identities, IDs, events, and shared state types |
| `agentos-kernel` | library | rustix-backed Linux and Btrfs boundary |
| `agentos-memory` | library | SQLite memory, audit, and principal state |
| `agentos-worldline` | library | filesystem branch/commit/rollback history |
| `agentos-npu` | library | NPU discovery, legacy bridge, Candle assembly |
| `agentos-candle` | library | Qwen3 host inference and CALRT tensor adapter |
| `calculet-pcie-abi` | library | captured 64-bit Linux ioctl ABI |
| `calculet-pcie` | library | safe PCIe device operations |
| `calculet-calrt` | library | Calbin/runtime/command/tensor implementation |
## Requirements
- 64-bit Linux. AgentOS itself is Linux-only, and the captured CALCULET ABI has
only been verified on 64-bit Linux.
- A current Rust toolchain with Rust 2024 edition support.
- A C toolchain, `pkg-config`, and OpenSSL development files. SQLite is built
from the bundled source.
- Btrfs is recommended for kernel-enforced read-only commit snapshots, but is
not required for development.
On Debian/Ubuntu, the native build dependencies are typically:
```bash
sudo apt install build-essential pkg-config libssl-dev
```
Build both binaries and all libraries:
```bash
cargo build --workspace
```
## Quick start without NPU hardware
First inspect the host and backend configuration:
```bash
cargo run -p agentos-cli --bin agentos -- \
--state-dir target/agentos-state --json doctor
```
Run a deterministic smoke test of the model/tool plumbing. This built-in fake
requests the real read-only `system_inspect` tool, but it is not the remote mock
used to substitute for unavailable NPU hardware:
```bash
cargo run -p agentos-cli --bin agentos -- \
--state-dir target/agentos-state --fake-model --json \
agent once "Inspect this Linux system"
```
Do not run the agent as root. By default, its principal uses the current real
UID as owner and the current effective UID/GID as the agent identity.
## Remote OpenAI-compatible mock inference
Hardware-independent mock testing uses a remote OpenAI-compatible model.
AgentOS supports either `/chat/completions` or `/responses` under a configured
API base URL. Both HTTP and HTTPS remote URLs are accepted; HTTPS should be used
whenever credentials or untrusted networks are involved. Localhost and loopback
model URLs are rejected because local model services must use UDS.
```bash
export AGENTOS_MODEL_BACKEND=openai_compatible
export AGENTOS_OPENAI_COMPATIBLE_API_KEY='...'
export AGENTOS_OPENAI_COMPATIBLE_BASE_URL='https://api.openai.com/v1'
export AGENTOS_OPENAI_COMPATIBLE_MODEL='your-model'
export AGENTOS_OPENAI_COMPATIBLE_API='responses' # or chat_completions
cargo run -p agentos-cli --bin agentos -- \
--state-dir target/agentos-state \
agent once "Summarize the current system state"
```
`OPENAI_API_KEY` is accepted as a fallback. A non-empty credential is currently
required by the remote adapter. Streaming is not implemented.
## Local model transport
A persistent model on the same Linux host must expose a filesystem Unix domain
socket. Do not bind it to localhost, a loopback address, or another TCP address.
The socket path and its owner/group/mode are part of the OS security boundary.
The legacy llama-server already supports this mode when `--host` ends in
`.sock`. AgentOS carries the familiar OpenAI-compatible HTTP request format over
AF_UNIX, so this is HTTP message framing over UDS—not HTTP over localhost.
The pure-Rust Candle backend runs in-process and therefore needs no IPC socket.
## Read-only tools
| Tool | Availability | Operation |
| --- | --- | --- |
| `system_inspect` | always | kernel, uptime, load, memory, root filesystem, OS release |
| `npu_inspect` | always | CALCULET device nodes, module/sysfs/proc monitor, CALRT library |
| `package_query` | always | exact package lookup via dpkg, rpm, or pacman |
| `service_inspect` | configured | bounded `systemctl show` for an allowlisted service |
| `journal_read` | configured | bounded journal read for an allowlisted service |
Enable service tools with an exact comma-separated allowlist:
```bash
export AGENTOS_SERVICE_INSPECT_ALLOWLIST='agentos.service,agentos-npud.service'
```
There are currently no mutation tools. The approval ledger in
`agentos-tools` is a contract primitive, not a complete privilege broker.
## Filesystem worldlines
Worldlines provide Git-like filesystem history without patching the kernel.
On Btrfs, current/branch trees are writable snapshots and commit trees are
read-only snapshots. Other filesystems use a copy fallback; that fallback does
not provide kernel-enforced immutability.
```bash
AGENTOS=(cargo run -q -p agentos-cli --bin agentos -- \
--state-dir target/agentos-state --json worldline)
"${AGENTOS[@]}" init
"${AGENTOS[@]}" branch
"${AGENTOS[@]}" status
"${AGENTOS[@]}" log --limit 20
```
`branch` returns a worldline ID and path. Modify that branch tree, then use:
```bash
"${AGENTOS[@]}" diff <worldline-id>
"${AGENTOS[@]}" commit <worldline-id> --message "describe the change"
"${AGENTOS[@]}" rollback <commit-id> --message "revert to known state"
"${AGENTOS[@]}" discard <worldline-id>
```
Rollback creates a new commit and retains history. Worldline operations are
currently explicit CLI operations; agent runs do not automatically create or
commit a branch.
## NPU paths
### Legacy llama-server bridge
The legacy `npu` backend executes `agentos-npu-bridge`, which connects only to a
filesystem Unix domain socket and sends `/v1/chat/completions` over that socket.
It additionally requires a ready CALCULET device, driver sysfs state, and vendor
CALRT shared library.
```bash
export AGENTOS_MODEL_BACKEND=npu
export AGENTOS_NPU_BRIDGE_COMMAND='agentos-npu-bridge'
export AGENTOS_NPU_SERVER_SOCKET='/run/agentos/npu.sock'
export AGENTOS_NPU_SERVER_MODEL='agentos-npu'
# Optional when CALRT is not in a probed system location:
export AGENTOS_CALRT_LIBRARY='/opt/calculet/lib/libcalrt-linux-x86_64.so'
```
Start the captured llama-server with the same socket path, for example
`--host /run/agentos/npu.sock`, and restrict access using the socket directory
and socket ownership/mode. No TCP listener is required.
### Pure-Rust Candle/CALRT path
The `npu_candle` path replaces the modified llama.cpp host code with Rust:
```bash
export AGENTOS_MODEL_BACKEND=npu_candle
export AGENTOS_NPU_CALBIN='/path/to/calbin-directory'
export AGENTOS_NPU_TOKENIZER='/path/to/tokenizer.json'
export AGENTOS_NPU_DEVICE_INDEX=0
```
`doctor` can parse these files and report the model plan without opening the
device. For the captured Qwen3 deployment, the host validates a 40,960-token
context, 151,936 logits, a 151,669-token vocabulary, and masks 267 padded logit
IDs.
This backend is deliberately not ready: `calculet-calrt` can deploy parameters
and transfer tensors, but `ConfiguredRuntime::submit()` still lacks CCU
relocation, job launch/completion, and device KV-cache reset. No real NPU result
has been validated on the current development machine.
## Configuration
| Variable | Default | Meaning |
| --- | --- | --- |
| `AGENTOS_STATE_DIR` | `$XDG_DATA_HOME/agentos` or `~/.local/share/agentos` | SQLite and worldline root |
| `AGENTOS_MODEL_BACKEND` | `auto` | `auto`, `openai_compatible`, `npu`, `npu_candle`, `subprocess`, `fake` |
| `AGENTOS_AGENT_ID` | `default-agent` | stable logical agent ID |
| `AGENTOS_OWNER_UID` | real UID | owning Linux UID |
| `AGENTOS_AGENT_UID` | effective UID | required execution UID |
| `AGENTOS_AGENT_GID` | effective GID | required execution GID |
| `AGENTOS_OPENAI_COMPATIBLE_API_KEY` | `OPENAI_API_KEY` | remote bearer credential |
| `AGENTOS_OPENAI_COMPATIBLE_BASE_URL` | `https://api.openai.com/v1` | remote compatible API root; no localhost/loopback |
| `AGENTOS_OPENAI_COMPATIBLE_MODEL` | empty | provider model name |
| `AGENTOS_OPENAI_COMPATIBLE_API` | `chat_completions` | `chat_completions` or `responses` |
| `AGENTOS_OPENAI_COMPATIBLE_MAX_RETRIES` | `2` | bounded HTTP retries |
| `AGENTOS_MODEL_COMMAND` | empty | explicit subprocess test/bridge command |
| `AGENTOS_MODEL_TIMEOUT_S` | `240` | model/bridge timeout |
| `AGENTOS_NPU_BRIDGE_COMMAND` | empty | legacy NPU bridge command |
| `AGENTOS_NPU_SERVER_SOCKET` | `/run/agentos/npu.sock` | legacy llama-server Unix socket |
| `AGENTOS_NPU_SERVER_MODEL` | `agentos-npu` | legacy llama-server model name |
| `AGENTOS_NPU_SERVER_TIMEOUT_S` | `240` | legacy llama-server UDS HTTP timeout |
| `AGENTOS_NPU_CALBIN` | empty | pure-Rust Calbin directory |
| `AGENTOS_NPU_TOKENIZER` | empty | pure-Rust tokenizer JSON |
| `AGENTOS_NPU_DEVICE_INDEX` | `0` | `/dev/calculetN` index |
| `AGENTOS_CALRT_LIBRARY` | probed | vendor CALRT shared library override |
| `AGENTOS_SERVICE_INSPECT_ALLOWLIST` | empty | enabled systemd service names |
| `AGENTOS_TOOL_OUTPUT_LIMIT_BYTES` | `65536` | per-tool serialized output limit |
| `AGENTOS_TOOL_ARGUMENT_LIMIT_BYTES` | `65536` | per-call argument limit |
| `AGENTOS_AGENT_MAX_ITERATIONS` | `8` | model/tool loop turns |
| `AGENTOS_AGENT_MAX_TOOL_CALLS` | `16` | total tool calls per run |
| `AGENTOS_AGENT_MAX_ELAPSED_S` | `120` | total run deadline |
| `AGENTOS_AGENT_MAX_OUTPUT_TOKENS` | `2048` | model output budget per turn |
| `AGENTOS_AGENT_MAX_TOTAL_TOKENS` | `65536` | cumulative token budget |
| `AGENTOS_AGENT_MAX_SAFE_WORKERS` | `4` | parallel safe-tool workers |
In `auto` mode, AgentOS selects the UDS-backed legacy NPU and then the remote
OpenAI-compatible backend. Subprocess and fake are explicit test adapters. The
incomplete Candle backend is not selected automatically.
## State layout
```text
AGENTOS_STATE_DIR/
agentos.sqlite3 # principals, audit events, memory, FTS index
worldline/
current/ # active filesystem tree
branches/ # writable candidate trees
commits/ # commit trees
metadata/ # commit and branch records
refs/ # HEAD references
replaced/ # recoverable previous checkouts
```
## Development
```bash
cargo fmt --all --check
cargo test --workspace
cargo clippy --workspace --all-targets --all-features -- -D warnings
```
Some NPU parser tests use the local, gitignored
`npu_features/snapshot_20260801` capture. Those tests validate host-side parsing
only. They are not hardware tests, and a clean clone may not contain the
fixture.
Project-owned low-level Linux calls use `rustix`; AgentOS does not directly add
`libc` syscall wrappers. This is not a claim that the final `std`/OpenSSL binary
has no libc ABI dependency.
## Known gaps
- No long-running AgentOS control daemon or privilege broker yet; the local
llama-server model transport already uses UDS.
- No mutation tools or complete privilege-approval path.
- No streaming remote inference.
- No completed pure-Rust NPU job submission or physical-board validation.
- The copy worldline backend lacks Btrfs's kernel-enforced read-only commits.
- Memory/recall primitives exist as a library but are not exposed as agent tools
or CLI commands yet.
## License notes
AgentOS crates declare Apache-2.0. The captured CALCULET PCIe ABI and device
wrapper crates declare GPL-2.0-only, while `calculet-calrt` retains its vendor
license file. Check the individual crate manifests before redistribution.
+330
View File
@@ -0,0 +1,330 @@
# AgentOS
[English](README.md)
AgentOS 是一个使用 Rust 编写、仅面向 Linux 的 low-level 操作系统 agent。它把
有界 tool-calling 循环与 Linux 原生身份、只读系统检查、持久审计、Btrfs 文件
系统历史和多种推理后端组合在一起,其中包括一条仍在开发中的 CALCULET NPU
纯 Rust 路径。
> **开发状态:**远程 OpenAI-compatible mock 路径和显式 subprocess/fake 测试
> adapter 已经实现。本机部署的模型服务只使用 Unix domain socket,绝不使用
> localhost 或 TCP。Rust Candle/CALRT host pipeline 可以解析并校验抓取到的
> Qwen3 部署,但 CALRT 硬件 job submission 尚未实现。因此 `npu_candle` 会
> fail closed,目前不能用于生产。
## 设计
AgentOS 不是容器 sandbox,也不是内核 fork。它以 Linux 进程身份作为隔离边界:
- 每个 active `AgentId` 稳定绑定 owner UID 和非 root agent UID/GID
- runtime 在每次运行前核对进程的 effective UID/GID
- 拒绝 UID 0 或 GID 0 的 agent
- 在模型和工具执行前启用 `no_new_privs`
- 模型只能看到显式工具 allowlist,永远拿不到原始 root shell
- 在 typed、可鉴权的 broker 完成前,系统变更能力保持不可用。
当前执行路径如下:
```text
agentos CLI
-> runtime + Linux principal
-> 有预算的 agent loop
-> ChatBackend(远程 OpenAI-compatible / UDS NPU / in-process Candle
-> strict ToolRegistryLinux 只读工具)
-> SQLite 审计日志
文件系统状态 -> WorldlineStore -> Btrfs snapshot 或 copy fallback
NPU host 路径 -> Candle/Qwen3 -> CALRT -> PCIe driver ABI
本机模型 IPC -> Unix domain socket -> OpenAI-compatible HTTP payload
```
## 当前已经实现
- Provider-neutral 的 `agentos.chat.v1` 消息、工具、响应、usage 和 backend 接口。
- 基于 `ntex` 的远程 OpenAI-compatible 非流式 Chat Completions 和 Responses
API,包含响应上限、超时、retry,以及选择 HTTPS 时的 TLS peer 校验。
- 通过 AF_UNIX 向本机 legacy llama-server 发送 OpenAI-compatible HTTP
payload,不经过 localhost/TCP。
- 保留 JSON stdin/stdout subprocess adapter,用于显式测试和自定义 bridge;
它不是常驻本机模型服务的传输方式。
- Agent iteration、tool call、耗时、总 token 预算,并行只读调用,重复 call ID
检查,以及无进展/周期检测。
- Strict JSON tool schema、allowlist、参数/输出上限、deadline 和哈希审计记录。
- Linux、包管理器、systemd、journal 和 CALCULET NPU 的只读检查。
- SQLite WAL 状态、FTS5 memory 基础能力、不可重绑的 principal 和追加式
agent/tool 审计事件。
- 基于 Btrfs subvolume/snapshot 的 worldline,包括 branch、diff、commit、log、
非破坏 rollback、陈旧分支保护和 checkout 恢复。
- CALCULET driver 0.9.0 / ABI 1.0.0 的 Rust 数据布局、BAR/DMA 访问、Calbin
0.7.6 解析、内存分配、tensor 和 command 编码。
- 基于 Candle 的 Qwen3 host runner,包括 chat template、tool-call 解析、采样、
prefill/decode 契约、BF16/F32 logits 和厂商 tiled-logit 解码。
## Workspace
`agentos-cli` 是唯一的 binary crate,并生成两个二进制。其他 workspace member
全部是 library。
| Crate | 类型 | 职责 |
| --- | --- | --- |
| `agentos-cli` | binary | `agentos` CLI 和 `agentos-npu-bridge` |
| `agentos-runtime` | library | 配置与 composition root |
| `agentos-agent` | library | 有界 agent loop 与审计事件 |
| `agentos-protocol` | library | provider-neutral chat/tool 协议 |
| `agentos-inference` | library | ntex、远程、subprocess 和测试 backend |
| `agentos-tools` | library | 工具策略、schema 校验、审批和 Linux 工具 |
| `agentos-core` | library | 身份、ID、事件和公共状态类型 |
| `agentos-kernel` | library | 基于 rustix 的 Linux/Btrfs 边界 |
| `agentos-memory` | library | SQLite memory、审计和 principal 状态 |
| `agentos-worldline` | library | 文件系统 branch/commit/rollback 历史 |
| `agentos-npu` | library | NPU 发现、legacy bridge 和 Candle 装配 |
| `agentos-candle` | library | Qwen3 host 推理与 CALRT tensor adapter |
| `calculet-pcie-abi` | library | 抓取到的 64-bit Linux ioctl ABI |
| `calculet-pcie` | library | 安全的 PCIe 设备操作 |
| `calculet-calrt` | library | Calbin/runtime/command/tensor 实现 |
## 环境要求
- 64-bit Linux。AgentOS 本身仅支持 LinuxCALCULET ABI 目前也只在 64-bit
Linux 上核对过。
- 支持 Rust 2024 edition 的当前 Rust toolchain。
- C toolchain、`pkg-config` 和 OpenSSL 开发文件。SQLite 使用 bundled 源码构建。
- 推荐使用 Btrfs,以获得内核强制只读的 commit snapshot;普通开发并不强制。
Debian/Ubuntu 通常需要:
```bash
sudo apt install build-essential pkg-config libssl-dev
```
构建两个 binary 和所有 library
```bash
cargo build --workspace
```
## 无 NPU 硬件快速开始
先检查主机和 backend 配置:
```bash
cargo run -p agentos-cli --bin agentos -- \
--state-dir target/agentos-state --json doctor
```
运行一次确定性的模型/工具 plumbing 冒烟测试。内置 fake 会请求真实的只读
`system_inspect` 工具,但它不是在 NPU 硬件缺失时使用的远程 mock:
```bash
cargo run -p agentos-cli --bin agentos -- \
--state-dir target/agentos-state --fake-model --json \
agent once "检查当前 Linux 系统"
```
不要以 root 运行 agent。默认 principal 使用当前 real UID 作为 owner,使用当前
effective UID/GID 作为 agent 身份。
## 远程 OpenAI-compatible Mock 推理
没有 NPU 硬件时,mock 测试使用远程 OpenAI-compatible 模型。AgentOS 支持
配置 API base URL 下的 `/chat/completions``/responses`。远程 URL 可以
使用 HTTP 或 HTTPS;涉及 credential 或不可信网络时应使用 HTTPS。模型 URL
不能是 localhost 或 loopback,因为本机模型服务必须使用 UDS。
```bash
export AGENTOS_MODEL_BACKEND=openai_compatible
export AGENTOS_OPENAI_COMPATIBLE_API_KEY='...'
export AGENTOS_OPENAI_COMPATIBLE_BASE_URL='https://api.openai.com/v1'
export AGENTOS_OPENAI_COMPATIBLE_MODEL='your-model'
export AGENTOS_OPENAI_COMPATIBLE_API='responses' # 或 chat_completions
cargo run -p agentos-cli --bin agentos -- \
--state-dir target/agentos-state \
agent once "总结当前系统状态"
```
也可以用 `OPENAI_API_KEY` 作为 fallback。远程 adapter 当前要求非空
credential。暂不支持 streaming。
## 本机模型传输
同一台 Linux 主机上的常驻模型必须暴露 filesystem Unix domain socket。不要
监听 localhost、loopback 或其他 TCP 地址。Socket path 及其 owner/group/mode
属于 OS 安全边界的一部分。
Legacy llama-server 的 `--host` 参数以 `.sock` 结尾时已经支持这种模式。
AgentOS 通过 AF_UNIX 承载熟悉的 OpenAI-compatible HTTP 请求格式,所以这是
UDS 上的 HTTP 消息 framing,而不是 localhost 上的 HTTP。纯 Rust Candle
backend 在进程内运行,不需要 IPC socket。
## 只读工具
| 工具 | 可用条件 | 操作 |
| --- | --- | --- |
| `system_inspect` | 始终 | kernel、uptime、load、内存、根文件系统、OS release |
| `npu_inspect` | 始终 | CALCULET 设备节点、module/sysfs/proc monitor、CALRT library |
| `package_query` | 始终 | 通过 dpkg、rpm 或 pacman 精确查询一个包 |
| `service_inspect` | 配置后 | 对 allowlist 内的 service 执行有界 `systemctl show` |
| `journal_read` | 配置后 | 对 allowlist 内的 service 执行有界 journal 读取 |
使用精确、逗号分隔的 allowlist 启用 service 工具:
```bash
export AGENTOS_SERVICE_INSPECT_ALLOWLIST='agentos.service,agentos-npud.service'
```
当前没有任何 mutation 工具。`agentos-tools` 中的 approval ledger 只是契约
基础设施,不是完整的权限 broker。
## 文件系统 Worldline
Worldline 在不 patch 内核的前提下提供类似 Git 的文件系统历史。在 Btrfs 上,
current/branch tree 是 writable snapshotcommit tree 是 readonly snapshot。
其他文件系统使用 copy fallback,但 fallback 不具备内核强制不可变性。
```bash
AGENTOS=(cargo run -q -p agentos-cli --bin agentos -- \
--state-dir target/agentos-state --json worldline)
"${AGENTOS[@]}" init
"${AGENTOS[@]}" branch
"${AGENTOS[@]}" status
"${AGENTOS[@]}" log --limit 20
```
`branch` 会返回 worldline ID 和路径。修改 branch tree 后执行:
```bash
"${AGENTOS[@]}" diff <worldline-id>
"${AGENTOS[@]}" commit <worldline-id> --message "描述变更"
"${AGENTOS[@]}" rollback <commit-id> --message "回退到已知状态"
"${AGENTOS[@]}" discard <worldline-id>
```
Rollback 会创建新 commit,不会改写历史。Worldline 当前仍由 CLI 显式管理;
agent run 不会自动创建或提交 branch。
## NPU 路径
### Legacy llama-server bridge
Legacy `npu` backend 会执行 `agentos-npu-bridge`bridge 只连接 filesystem
Unix domain socket,并通过该 socket 发送 `/v1/chat/completions`。此外还要求
CALCULET 设备、driver sysfs 状态和厂商 CALRT shared library 全部 ready。
```bash
export AGENTOS_MODEL_BACKEND=npu
export AGENTOS_NPU_BRIDGE_COMMAND='agentos-npu-bridge'
export AGENTOS_NPU_SERVER_SOCKET='/run/agentos/npu.sock'
export AGENTOS_NPU_SERVER_MODEL='agentos-npu'
# CALRT 不在自动检查路径时可覆盖:
export AGENTOS_CALRT_LIBRARY='/opt/calculet/lib/libcalrt-linux-x86_64.so'
```
抓取到的 llama-server 可使用 `--host /run/agentos/npu.sock` 启动,并通过 socket
目录及 socket 的 owner/group/mode 限制访问。不需要 TCP listener。
### 纯 Rust Candle/CALRT 路径
`npu_candle` 用 Rust 替代魔改 llama.cpp 的 host 代码:
```bash
export AGENTOS_MODEL_BACKEND=npu_candle
export AGENTOS_NPU_CALBIN='/path/to/calbin-directory'
export AGENTOS_NPU_TOKENIZER='/path/to/tokenizer.json'
export AGENTOS_NPU_DEVICE_INDEX=0
```
`doctor` 可以在不打开设备的情况下解析这些文件并报告 model plan。对抓取到的
Qwen3 部署,host 已校验 40,960 token context、151,936 logits 和 151,669
token vocabulary,并会屏蔽额外的 267 个 padded logit ID。
这个 backend 目前会刻意保持 not ready`calculet-calrt` 已能部署参数和传输
tensor,但 `ConfiguredRuntime::submit()` 仍缺少 CCU relocation、job
launch/completion 和设备 KV-cache reset。当前开发机没有验证过任何真实 NPU
推理结果。
## 配置
| 环境变量 | 默认值 | 含义 |
| --- | --- | --- |
| `AGENTOS_STATE_DIR` | `$XDG_DATA_HOME/agentos``~/.local/share/agentos` | SQLite 与 worldline 根目录 |
| `AGENTOS_MODEL_BACKEND` | `auto` | `auto``openai_compatible``npu``npu_candle``subprocess``fake` |
| `AGENTOS_AGENT_ID` | `default-agent` | 稳定的逻辑 agent ID |
| `AGENTOS_OWNER_UID` | real UID | owner Linux UID |
| `AGENTOS_AGENT_UID` | effective UID | 要求的执行 UID |
| `AGENTOS_AGENT_GID` | effective GID | 要求的执行 GID |
| `AGENTOS_OPENAI_COMPATIBLE_API_KEY` | `OPENAI_API_KEY` | 远程 bearer credential |
| `AGENTOS_OPENAI_COMPATIBLE_BASE_URL` | `https://api.openai.com/v1` | 远程 compatible API root;禁止 localhost/loopback |
| `AGENTOS_OPENAI_COMPATIBLE_MODEL` | 空 | provider model 名称 |
| `AGENTOS_OPENAI_COMPATIBLE_API` | `chat_completions` | `chat_completions``responses` |
| `AGENTOS_OPENAI_COMPATIBLE_MAX_RETRIES` | `2` | 有界 HTTP retry 次数 |
| `AGENTOS_MODEL_COMMAND` | 空 | 显式 subprocess 测试/bridge 命令 |
| `AGENTOS_MODEL_TIMEOUT_S` | `240` | model/bridge 超时 |
| `AGENTOS_NPU_BRIDGE_COMMAND` | 空 | legacy NPU bridge 命令 |
| `AGENTOS_NPU_SERVER_SOCKET` | `/run/agentos/npu.sock` | legacy llama-server Unix socket |
| `AGENTOS_NPU_SERVER_MODEL` | `agentos-npu` | legacy llama-server model 名称 |
| `AGENTOS_NPU_SERVER_TIMEOUT_S` | `240` | legacy llama-server UDS HTTP 超时 |
| `AGENTOS_NPU_CALBIN` | 空 | 纯 Rust Calbin 目录 |
| `AGENTOS_NPU_TOKENIZER` | 空 | 纯 Rust tokenizer JSON |
| `AGENTOS_NPU_DEVICE_INDEX` | `0` | `/dev/calculetN` 序号 |
| `AGENTOS_CALRT_LIBRARY` | 自动检查 | 厂商 CALRT shared library 覆盖路径 |
| `AGENTOS_SERVICE_INSPECT_ALLOWLIST` | 空 | 启用的 systemd service 名称 |
| `AGENTOS_TOOL_OUTPUT_LIMIT_BYTES` | `65536` | 单个工具序列化输出上限 |
| `AGENTOS_TOOL_ARGUMENT_LIMIT_BYTES` | `65536` | 单次调用参数上限 |
| `AGENTOS_AGENT_MAX_ITERATIONS` | `8` | 模型/工具循环轮数 |
| `AGENTOS_AGENT_MAX_TOOL_CALLS` | `16` | 单次 run 的 tool call 总数 |
| `AGENTOS_AGENT_MAX_ELAPSED_S` | `120` | 单次 run 总 deadline |
| `AGENTOS_AGENT_MAX_OUTPUT_TOKENS` | `2048` | 每轮模型输出预算 |
| `AGENTOS_AGENT_MAX_TOTAL_TOKENS` | `65536` | 累计 token 预算 |
| `AGENTOS_AGENT_MAX_SAFE_WORKERS` | `4` | 并行 safe-tool worker 数量 |
`auto` 模式依次选择 UDS legacy NPU 和远程 OpenAI-compatible。Subprocess 与
fake 都是显式测试 adapter,未完成的 Candle backend 不会被自动选中。
## 状态目录
```text
AGENTOS_STATE_DIR/
agentos.sqlite3 # principal、审计事件、memory、FTS index
worldline/
current/ # 当前文件系统 tree
branches/ # writable candidate tree
commits/ # commit tree
metadata/ # commit 和 branch 记录
refs/ # HEAD reference
replaced/ # 可恢复的旧 checkout
```
## 开发
```bash
cargo fmt --all --check
cargo test --workspace
cargo clippy --workspace --all-targets --all-features -- -D warnings
```
部分 NPU parser 测试使用本地、被 git 忽略的
`npu_features/snapshot_20260801` 抓取。这些测试只能验证 host 解析,不是硬件
测试;干净 clone 也可能不包含该 fixture。
项目自有的 low-level Linux 调用使用 `rustix`AgentOS 不直接添加 `libc`
syscall wrapper。这不代表使用 `std`/OpenSSL 的最终 binary 完全没有 libc ABI
依赖。
## 已知缺口
- 尚无常驻 AgentOS 控制 daemon 或权限 broker;本机 llama-server 模型传输已经
使用 UDS。
- 尚无 mutation 工具或完整的权限审批路径。
- 远程推理尚不支持 streaming。
- 纯 Rust NPU job submission 和真实板卡验证尚未完成。
- Copy worldline backend 不具备 Btrfs 的内核强制只读 commit。
- Memory/recall 基础能力已存在于 library,但尚未暴露为 agent 工具或 CLI。
## License 说明
AgentOS crates 声明为 Apache-2.0。抓取的 CALCULET PCIe ABI 与设备 wrapper
crates 声明为 GPL-2.0-only`calculet-calrt` 保留厂商 license 文件。重新分发前
请分别检查各 crate manifest。
@@ -1,4 +1,4 @@
use agentos_npu::{LoopbackServerConfig, bridge_chat};
use agentos_npu::{DEFAULT_NPU_SERVER_SOCKET, UnixSocketServerConfig, bridge_chat};
use agentos_protocol::ChatRequest;
use anyhow::{Context, Result, bail};
use clap::Parser;
@@ -30,9 +30,9 @@ async fn main() -> Result<()> {
serde_json::from_slice(&input).context("parse agentos.chat.v1 request")?;
request.validate()?;
request.parallel_tool_calls = false;
let config = LoopbackServerConfig {
endpoint: std::env::var("AGENTOS_NPU_SERVER_URL")
.unwrap_or_else(|_| "http://127.0.0.1:8031/v1/chat/completions".into()),
let config = UnixSocketServerConfig {
socket_path: std::env::var_os("AGENTOS_NPU_SERVER_SOCKET")
.map_or_else(|| DEFAULT_NPU_SERVER_SOCKET.into(), Into::into),
model: std::env::var("AGENTOS_NPU_SERVER_MODEL").unwrap_or_else(|_| "agentos-npu".into()),
timeout: Duration::from_secs(
std::env::var("AGENTOS_NPU_SERVER_TIMEOUT_S")
+207 -18
View File
@@ -1,4 +1,4 @@
//! OpenAI-compatible remote, subprocess/NPU bridge, and deterministic mock backends.
//! Remote OpenAI-compatible, Unix-socket, subprocess, and deterministic fake backends.
use agentos_protocol::{
BackendError, BackendStatus, ChatBackend, ChatMessage, ChatRequest, ChatResponse, MessageRole,
@@ -14,7 +14,8 @@ use openssl::ssl::{SslConnector, SslMethod, SslVerifyMode};
use serde::{Deserialize, Serialize};
use serde_json::{Map, Value, json};
use std::collections::VecDeque;
use std::path::PathBuf;
use std::os::unix::ffi::OsStrExt;
use std::path::{Path, PathBuf};
use std::sync::Mutex;
use std::time::Duration;
use tokio::io::AsyncWriteExt;
@@ -155,6 +156,17 @@ pub struct NtexJsonClient {
sender: tokio::sync::mpsc::UnboundedSender<NtexRequest>,
}
#[derive(Clone, Debug)]
pub struct NtexUdsJsonClient {
inner: NtexJsonClient,
}
#[derive(Clone, Debug)]
enum NtexTransport {
Network,
Unix(PathBuf),
}
struct NtexRequest {
endpoint: String,
payload: Value,
@@ -180,6 +192,22 @@ impl NtexJsonClient {
max_retries: usize,
user_agent: &'static str,
bearer_token: Option<String>,
) -> Result<Self, BackendError> {
Self::with_transport(
timeout,
max_retries,
user_agent,
bearer_token,
NtexTransport::Network,
)
}
fn with_transport(
timeout: Duration,
max_retries: usize,
user_agent: &'static str,
bearer_token: Option<String>,
transport: NtexTransport,
) -> Result<Self, BackendError> {
let (sender, receiver) = tokio::sync::mpsc::unbounded_channel();
let (startup_sender, startup_receiver) = std::sync::mpsc::sync_channel(1);
@@ -190,7 +218,7 @@ impl NtexJsonClient {
.name("agentos-ntex-client")
.build(rt::DefaultRuntime)
.block_on(async move {
match build_ntex_client(timeout, user_agent).await {
match build_ntex_client(timeout, user_agent, transport).await {
Ok(client) => {
let _ = startup_sender.send(Ok(()));
run_ntex_worker(
@@ -234,10 +262,47 @@ impl NtexJsonClient {
}
}
async fn build_ntex_client(timeout: Duration, user_agent: &str) -> Result<Client, BackendError> {
let tls = verified_tls_connector()?;
impl NtexUdsJsonClient {
pub fn new(
socket_path: impl Into<PathBuf>,
timeout: Duration,
max_retries: usize,
user_agent: &'static str,
) -> Result<Self, BackendError> {
let socket_path = socket_path.into();
validate_uds_path(&socket_path)?;
Ok(Self {
inner: NtexJsonClient::with_transport(
timeout,
max_retries,
user_agent,
None,
NtexTransport::Unix(socket_path),
)?,
})
}
pub async fn post(
&self,
endpoint_path: &str,
payload: &Value,
) -> Result<NtexJsonResponse, BackendError> {
let endpoint = uds_http_endpoint(endpoint_path)?;
self.inner.post(&endpoint, payload).await
}
}
async fn build_ntex_client(
timeout: Duration,
user_agent: &str,
transport: NtexTransport,
) -> Result<Client, BackendError> {
let connector = match transport {
NtexTransport::Network => Connector::default().openssl(verified_tls_connector()?),
NtexTransport::Unix(socket_path) => unix_http_connector(socket_path),
};
Client::builder()
.connector::<&str>(Connector::default().openssl(tls))
.connector::<&str>(connector)
.response_timeout(timeout)
.response_payload_limit(MAX_RESPONSE_BYTES)
.response_payload_timeout(Millis::from(timeout))
@@ -248,6 +313,50 @@ async fn build_ntex_client(timeout: Duration, user_agent: &str) -> Result<Client
.map_err(|error| BackendError::Transport(error.to_string()))
}
fn unix_http_connector(socket_path: PathBuf) -> Connector {
let factory = ntex::service::fn_factory_with_config(move |config: SharedCfg| {
let socket_path = socket_path.clone();
async move {
Ok::<_, std::io::Error>(ntex::service::fn_service(
move |_request: ntex::connect::Connect<ntex::http::Uri>| {
let config = config.clone();
let socket_path = socket_path.clone();
async move {
rt::unix_connect(socket_path, config)
.await
.map_err(ntex::connect::ConnectError::from)
}
},
))
}
});
Connector::new().connector(factory)
}
fn validate_uds_path(path: &Path) -> Result<(), BackendError> {
const LINUX_SUN_PATH_BYTES: usize = 108;
let bytes = path.as_os_str().as_bytes();
if !path.is_absolute()
|| bytes.is_empty()
|| bytes.len() >= LINUX_SUN_PATH_BYTES
|| bytes.contains(&0)
{
return Err(BackendError::NotReady(
"Unix model socket must be an absolute filesystem path shorter than 108 bytes".into(),
));
}
Ok(())
}
fn uds_http_endpoint(path: &str) -> Result<String, BackendError> {
if !path.starts_with('/') || path.contains(['\r', '\n', '?', '#']) {
return Err(BackendError::NotReady(
"Unix model endpoint must be an absolute HTTP path without query or fragment".into(),
));
}
Ok(format!("http://agentos.local{path}"))
}
async fn run_ntex_worker(
mut receiver: tokio::sync::mpsc::UnboundedReceiver<NtexRequest>,
client: Client,
@@ -482,7 +591,7 @@ impl ChatBackend for FakeBackend {
BackendStatus {
name: self.name().into(),
ready: true,
reason: "deterministic local mock".into(),
reason: "deterministic unit-test fake".into(),
details: json!({"hardware_required": false}),
}
}
@@ -495,7 +604,10 @@ impl ChatBackend for FakeBackend {
.find(|message| message.role == MessageRole::Tool)
{
return Ok(ChatResponse {
content: format!("Mock 已完成只读检查:{}", tool.content),
content: format!(
"Fake backend completed a read-only inspection: {}",
tool.content
),
usage: TokenUsage {
input_tokens: 12,
output_tokens: 8,
@@ -525,7 +637,7 @@ impl ChatBackend for FakeBackend {
});
}
Ok(ChatResponse {
content: "Mock backend is ready.".into(),
content: "Fake backend is ready.".into(),
..ChatResponse::default()
})
}
@@ -811,19 +923,32 @@ fn required_string(value: &Value, key: &str) -> Result<String, BackendError> {
fn validated_base_url(value: &str) -> Result<Url, BackendError> {
let url = Url::parse(value).map_err(|error| BackendError::NotReady(error.to_string()))?;
let loopback = url
.host_str()
.is_some_and(|host| matches!(host, "localhost" | "127.0.0.1" | "::1"));
let local = match url.host() {
Some(url::Host::Domain(host)) => {
host.eq_ignore_ascii_case("localhost")
|| host.to_ascii_lowercase().ends_with(".localhost")
}
Some(url::Host::Ipv4(address)) => address.is_loopback() || address.is_unspecified(),
Some(url::Host::Ipv6(address)) => {
address.is_loopback()
|| address.is_unspecified()
|| address
.to_ipv4_mapped()
.is_some_and(|mapped| mapped.is_loopback() || mapped.is_unspecified())
}
None => true,
};
if url.cannot_be_a_base()
|| url.host_str().is_none()
|| local
|| url.username() != ""
|| url.password().is_some()
|| url.query().is_some()
|| url.fragment().is_some()
|| !(url.scheme() == "https" || (url.scheme() == "http" && loopback))
|| !matches!(url.scheme(), "http" | "https")
{
return Err(BackendError::NotReady(
"invalid OpenAI-compatible base URL".into(),
"remote OpenAI-compatible base URL must use HTTP(S), not a local or loopback host"
.into(),
));
}
Ok(url)
@@ -950,10 +1075,12 @@ mod tests {
}
#[test]
fn only_https_or_loopback_http_is_accepted() {
fn remote_http_is_accepted_but_loopback_model_urls_are_rejected() {
assert!(validated_base_url("https://api.openai.com/v1").is_ok());
assert!(validated_base_url("http://127.0.0.1:8080/v1").is_ok());
assert!(validated_base_url("http://example.com/v1").is_err());
assert!(validated_base_url("http://model.example/v1").is_ok());
assert!(validated_base_url("http://127.0.0.1:8080/v1").is_err());
assert!(validated_base_url("http://localhost:8080/v1").is_err());
assert!(validated_base_url("http://[::ffff:127.0.0.1]:8080/v1").is_err());
}
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
@@ -1033,4 +1160,66 @@ mod tests {
}
server.await.unwrap();
}
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
async fn ntex_client_carries_local_http_over_a_unix_socket() {
use tokio::io::{AsyncReadExt, AsyncWriteExt};
async fn read_http_request(stream: &mut tokio::net::UnixStream) -> String {
let mut request = Vec::new();
loop {
let mut chunk = [0_u8; 4096];
let received = stream.read(&mut chunk).await.unwrap();
assert_ne!(received, 0);
request.extend_from_slice(&chunk[..received]);
let Some(header_end) = request.windows(4).position(|bytes| bytes == b"\r\n\r\n")
else {
continue;
};
let headers = String::from_utf8_lossy(&request[..header_end]);
let content_length = headers
.lines()
.find_map(|line| {
let (name, value) = line.split_once(':')?;
name.eq_ignore_ascii_case("content-length")
.then(|| value.trim().parse::<usize>().unwrap())
})
.unwrap_or(0);
let request_end = header_end + 4 + content_length;
if request.len() >= request_end {
return String::from_utf8(request[..request_end].to_vec()).unwrap();
}
}
}
let root =
std::env::temp_dir().join(format!("agentos-inference-uds-{}", std::process::id()));
let _ = std::fs::remove_dir_all(&root);
std::fs::create_dir_all(&root).unwrap();
let socket_path = root.join("model.sock");
let listener = tokio::net::UnixListener::bind(&socket_path).unwrap();
let server = tokio::spawn(async move {
let (mut stream, _) = listener.accept().await.unwrap();
let request = read_http_request(&mut stream).await;
assert!(request.starts_with("POST /v1/chat/completions HTTP/1.1"));
stream
.write_all(
b"HTTP/1.1 200 OK\r\nContent-Type: application/json\r\nContent-Length: 11\r\nConnection: close\r\n\r\n{\"ok\":true}",
)
.await
.unwrap();
});
let client =
NtexUdsJsonClient::new(&socket_path, Duration::from_secs(2), 0, "agentos-uds-test")
.unwrap();
let response = client
.post("/v1/chat/completions", &json!({"ping": true}))
.await
.unwrap();
assert!(response.is_success());
assert_eq!(response.body, json!({"ok": true}));
server.await.unwrap();
std::fs::remove_dir_all(root).unwrap();
}
}
-1
View File
@@ -16,7 +16,6 @@ rustix.workspace = true
serde.workspace = true
serde_json.workspace = true
tokio.workspace = true
url.workspace = true
[lints]
workspace = true
+46 -37
View File
@@ -5,16 +5,16 @@ use agentos_candle::{
HuggingFaceTokenizer, Result as CandleResult, TokenCodec,
};
use agentos_inference::{
NtexJsonClient, SubprocessBackend, chat_completion_messages, parse_chat_completions,
NtexUdsJsonClient, SubprocessBackend, chat_completion_messages, parse_chat_completions,
};
use agentos_protocol::{BackendError, BackendStatus, ChatBackend, ChatRequest, ChatResponse};
use async_trait::async_trait;
use rustix::fs::Access;
use serde::{Deserialize, Serialize};
use serde_json::{Value, json};
use std::os::unix::fs::FileTypeExt;
use std::path::{Path, PathBuf};
use std::time::Duration;
use url::Url;
use calculet_calrt::{Calbin, ConfiguredRuntime, PcieRuntimeDevice};
@@ -26,6 +26,7 @@ pub const CALRT_VERSION: &str = calculet_calrt::CALRT_COMPATIBLE_VERSION;
pub const H2C_DMA_CHANNELS: u8 = calculet_pcie_abi::WRITE_CHANNELS;
pub const C2H_DMA_CHANNELS: u8 = calculet_pcie_abi::READ_CHANNELS;
pub const MAX_DMA_TRANSFER_BYTES: usize = calculet_pcie_abi::MAX_DMA_TRANSFER_BYTES;
pub const DEFAULT_NPU_SERVER_SOCKET: &str = "/run/agentos/npu.sock";
pub type RustCandleNpuBackend = CandleChatBackend<CalrtCompiledModel<PcieRuntimeDevice>>;
@@ -229,10 +230,16 @@ impl NpuProbe {
pub struct NpuBackend {
bridge: SubprocessBackend,
probe: NpuProbe,
server_socket: PathBuf,
}
impl NpuBackend {
pub fn new(command: Vec<String>, timeout: Duration, probe: NpuProbe) -> Self {
pub fn new(
command: Vec<String>,
timeout: Duration,
probe: NpuProbe,
server_socket: PathBuf,
) -> Self {
Self {
bridge: SubprocessBackend {
name: "npu".into(),
@@ -241,6 +248,7 @@ impl NpuBackend {
serialized: true,
},
probe,
server_socket,
}
}
}
@@ -258,15 +266,24 @@ impl ChatBackend for NpuBackend {
fn status(&self) -> BackendStatus {
let bridge = self.bridge.status();
let npu = self.probe.probe();
let socket_ready = is_unix_socket(&self.server_socket);
BackendStatus {
name: self.name().into(),
ready: bridge.ready && npu.hardware_ready,
reason: if bridge.ready && npu.hardware_ready {
"NPU inference chain is ready".into()
ready: bridge.ready && npu.hardware_ready && socket_ready,
reason: if bridge.ready && npu.hardware_ready && socket_ready {
"NPU inference chain is ready over a Unix domain socket".into()
} else {
"NPU bridge, driver, device, or calrt is unavailable".into()
"NPU bridge, Unix model socket, driver, device, or calrt is unavailable".into()
},
details: json!({"bridge": bridge, "npu": npu}),
details: json!({
"bridge": bridge,
"model_socket": {
"path": self.server_socket,
"ready": socket_ready,
"transport": "unix",
},
"npu": npu,
}),
}
}
@@ -304,22 +321,21 @@ impl ChatBackend for NpuBackend {
}
#[derive(Clone, Debug)]
pub struct LoopbackServerConfig {
pub endpoint: String,
pub struct UnixSocketServerConfig {
pub socket_path: PathBuf,
pub model: String,
pub timeout: Duration,
}
pub async fn bridge_chat(
config: &LoopbackServerConfig,
config: &UnixSocketServerConfig,
request: &ChatRequest,
) -> Result<ChatResponse, BackendError> {
let endpoint = validate_loopback_endpoint(&config.endpoint)?;
let client = NtexJsonClient::new(
let client = NtexUdsJsonClient::new(
&config.socket_path,
config.timeout,
0,
concat!("agentos-npu-bridge/", env!("CARGO_PKG_VERSION")),
None,
)?;
let payload = json!({
"model": config.model,
@@ -339,10 +355,10 @@ pub async fn bridge_chat(
"temperature": 0,
"stream": false,
});
let response = client.post(endpoint.as_str(), &payload).await?;
let response = client.post("/v1/chat/completions", &payload).await?;
if !response.is_success() {
return Err(BackendError::Transport(format!(
"llama-server HTTP {}: {}",
"llama-server UDS HTTP {}: {}",
response.status,
response
.body
@@ -354,24 +370,10 @@ pub async fn bridge_chat(
parse_chat_completions(&response.body)
}
fn validate_loopback_endpoint(value: &str) -> Result<Url, BackendError> {
let url = Url::parse(value).map_err(|error| BackendError::NotReady(error.to_string()))?;
let loopback = url
.host_str()
.is_some_and(|host| matches!(host, "localhost" | "127.0.0.1" | "::1"));
if url.scheme() != "http"
|| !loopback
|| url.username() != ""
|| url.password().is_some()
|| url.query().is_some()
|| url.fragment().is_some()
|| !url.path().ends_with("/chat/completions")
{
return Err(BackendError::NotReady(
"NPU server URL must be a loopback HTTP chat endpoint".into(),
));
}
Ok(url)
fn is_unix_socket(path: &Path) -> bool {
path.is_absolute()
&& std::fs::symlink_metadata(path).is_ok_and(|metadata| metadata.file_type().is_socket())
&& rustix::fs::access(path, Access::WRITE_OK).is_ok()
}
fn read_small(path: PathBuf) -> Option<String> {
@@ -398,9 +400,16 @@ mod tests {
use super::*;
#[test]
fn bridge_never_accepts_a_remote_plaintext_endpoint() {
assert!(validate_loopback_endpoint("http://127.0.0.1:8031/v1/chat/completions").is_ok());
assert!(validate_loopback_endpoint("http://npu.example/v1/chat/completions").is_err());
fn local_model_endpoint_must_be_a_unix_socket() {
let root = std::env::temp_dir().join(format!("agentos-npu-uds-{}", std::process::id()));
let _ = std::fs::remove_dir_all(&root);
std::fs::create_dir_all(&root).unwrap();
let socket_path = root.join("model.sock");
let listener = std::os::unix::net::UnixListener::bind(&socket_path).unwrap();
assert!(is_unix_socket(&socket_path));
assert!(!is_unix_socket(Path::new("relative.sock")));
drop(listener);
std::fs::remove_dir_all(root).unwrap();
}
#[test]
+10 -5
View File
@@ -7,7 +7,10 @@ use agentos_inference::{
};
use agentos_kernel::{ProcessCredentials, enable_no_new_privileges, kernel_info};
use agentos_memory::MemoryStore;
use agentos_npu::{CandleNpuConfig, NpuBackend, NpuProbe, inspect_candle_npu, open_candle_npu};
use agentos_npu::{
CandleNpuConfig, DEFAULT_NPU_SERVER_SOCKET, NpuBackend, NpuProbe, inspect_candle_npu,
open_candle_npu,
};
use agentos_protocol::{BackendStatus, ChatBackend};
use agentos_tools::ToolRegistry;
use agentos_tools::linux::{LinuxToolConfig, build_linux_registry};
@@ -61,6 +64,7 @@ pub struct RuntimeConfig {
pub openai_compatible_api: OpenAiCompatibleApi,
pub openai_compatible_max_retries: usize,
pub npu_bridge_command: Vec<String>,
pub npu_server_socket: PathBuf,
pub npu_candle_calbin: Option<PathBuf>,
pub npu_candle_tokenizer: Option<PathBuf>,
pub npu_candle_device_index: u8,
@@ -112,6 +116,8 @@ impl RuntimeConfig {
openai_compatible_api,
openai_compatible_max_retries: env_usize("AGENTOS_OPENAI_COMPATIBLE_MAX_RETRIES", 2)?,
npu_bridge_command: command_env("AGENTOS_NPU_BRIDGE_COMMAND"),
npu_server_socket: std::env::var_os("AGENTOS_NPU_SERVER_SOCKET")
.map_or_else(|| PathBuf::from(DEFAULT_NPU_SERVER_SOCKET), PathBuf::from),
npu_candle_calbin: std::env::var_os("AGENTOS_NPU_CALBIN").map(PathBuf::from),
npu_candle_tokenizer: std::env::var_os("AGENTOS_NPU_TOKENIZER").map(PathBuf::from),
npu_candle_device_index: env_u8("AGENTOS_NPU_DEVICE_INDEX", 0)?,
@@ -266,6 +272,7 @@ pub fn doctor(config: &RuntimeConfig) -> Result<DoctorReport, RuntimeError> {
config.npu_bridge_command.clone(),
config.model_timeout,
probe.clone(),
config.npu_server_socket.clone(),
);
backends.push(npu_backend.status());
backends.push(candle_npu_status(config));
@@ -303,6 +310,7 @@ fn build_backend(config: &RuntimeConfig) -> Result<Arc<dyn ChatBackend>, Runtime
config.npu_bridge_command.clone(),
config.model_timeout,
npu_probe(config),
config.npu_server_socket.clone(),
));
let subprocess: Arc<dyn ChatBackend> = Arc::new(SubprocessBackend {
name: "subprocess".into(),
@@ -329,11 +337,8 @@ fn build_backend(config: &RuntimeConfig) -> Result<Arc<dyn ChatBackend>, Runtime
{
return Ok(Arc::new(openai_compatible));
}
if subprocess.status().ready {
return Ok(subprocess);
}
Err(RuntimeError::BackendUnavailable(
"no NPU, OpenAI-compatible, or subprocess backend is ready; use --fake-model only for tests".into(),
"no NPU or remote OpenAI-compatible backend is ready; local deployed models must use a Unix domain socket, and subprocess/fake backends are explicit test adapters".into(),
))
}
}