Compare commits

..

68 Commits

Author SHA1 Message Date
dadachann 8737b818b6 fix(cloud): render empty skills without sandbox grant 2026-07-30 03:21:26 +00:00
dadachann 4b2a628db6 fix(deps): pin green multi-tenant plugin SDK 2026-07-29 11:15:18 +00:00
Junyan Qin 610915b9c5 fix(runtime): bound tenant resource amplification 2026-07-29 18:27:44 +08:00
Junyan Qin e8d90c4259 fix(cloud): bound tenant maintenance and monitoring work 2026-07-29 16:02:44 +08:00
dadachann 2dfbe78271 fix(cloud): scope public login capability discovery 2026-07-29 05:58:30 +00:00
Junyan Qin c89e6f3bd2 fix(cloud): enforce instance capacity ceilings 2026-07-29 13:45:58 +08:00
Junyan Qin e52d6880f5 fix(cloud): eliminate periodic runtime CPU spikes 2026-07-29 12:47:53 +08:00
Junyan Qin aa342d9347 fix(cloud): bound runtime restart storms 2026-07-29 12:14:08 +08:00
Junyan Qin ae85ac2b16 feat(cloud): harden multi-tenant runtime resources 2026-07-29 11:32:26 +08:00
dadachann 32abbb636f fix(oss): resolve workspace owner in scoped session 2026-07-26 16:30:15 +08:00
dadachann f247a9d183 test(oss): cover invitation logout handoff 2026-07-26 16:12:03 +08:00
dadachann 624a197655 style: format OSS account service 2026-07-26 16:09:50 +08:00
dadachann 712f79ed77 feat(oss): enforce invitation account and owner billing flows 2026-07-26 16:04:33 +08:00
dadachann 602e10649b fix(cloud): recover box runtime without unscoped skill reload 2026-07-26 14:25:49 +08:00
dadachann d71bd571b1 style(web): format invitation flows 2026-07-26 14:01:20 +08:00
dadachann e90a1546de feat(cloud): complete secure invitation experience 2026-07-26 13:54:42 +08:00
dadachann ff068564ab test(cloud): require Space identity for invite registration 2026-07-26 12:32:02 +08:00
dadachann 94ff4fcd2d fix(cloud): preserve Core-owned collaboration state 2026-07-26 12:27:47 +08:00
dadachann 97b3e58884 fix(workspace): bind collaboration APIs to tenant UoW 2026-07-26 11:28:46 +08:00
dadachann 66a1ceac25 style: format collaboration changes 2026-07-26 10:54:21 +08:00
dadachann 8a445bfb22 feat(workspace): add in-product collaboration and direct Cloud launch 2026-07-26 10:40:36 +08:00
dadachann 276791e7af fix(ui): align workspace switcher with sidebar entries 2026-07-26 00:23:17 +08:00
dadachann baf7e86335 fix(ui): hide roles from workspace switcher 2026-07-26 00:07:44 +08:00
dadachann 741c20af07 fix(ui): widen and center workspace switcher 2026-07-25 23:44:49 +08:00
dadachann 5ac1ab3eac fix(plugin): keep runtime identity stable across restarts 2026-07-25 21:56:14 +08:00
dadachann f96116a050 fix(cloud): surface runtime and workspace plan status 2026-07-25 21:42:05 +08:00
dadachann 40abb03928 style(web): format workspace layout test 2026-07-25 21:04:23 +08:00
dadachann 59f68b8fb4 refactor(web): streamline workspace controls 2026-07-25 20:59:05 +08:00
dadachann 7c64cd9d51 feat(web): place workspace controls in sidebar 2026-07-25 16:26:13 +08:00
dadachann 9ea1a81048 test(web): cover Workspace dropdown menu 2026-07-25 14:11:28 +08:00
dadachann d3f08a90b1 feat(cloud): complete Workspace settings navigation 2026-07-25 13:51:37 +08:00
dadachann 64e772e32d fix(cloud): reuse authenticated account for user info 2026-07-25 02:41:31 +08:00
dadachann 84440df47f fix(cloud): preserve authenticated account context 2026-07-25 02:05:17 +08:00
dadachann c860159446 test(cloud): preserve minimal model manager fixtures 2026-07-25 00:45:42 +08:00
dadachann f977629a90 fix(cloud): skip legacy model sync during startup 2026-07-25 00:26:56 +08:00
dadachann ff13d52602 chore: update multi-tenant SDK pin 2026-07-24 23:31:08 +08:00
Junyan Qin 5beab49577 docs(cloud): update control plane verification 2026-07-24 22:58:56 +08:00
Junyan Qin e8a09b7537 fix(build): install git for pinned SDK 2026-07-24 19:29:14 +08:00
Junyan Qin 98f45aa88e feat(tenancy): connect cloud workspace control plane 2026-07-24 19:11:33 +08:00
Junyan Qin d7cdd206c2 docs(tenancy): record final isolation verification 2026-07-24 16:22:45 +08:00
Junyan Qin ac72563664 fix(tenancy): close isolation and permission gaps 2026-07-24 16:22:45 +08:00
Junyan Qin 64dc887b20 docs(tenancy): record final isolation verification 2026-07-24 16:22:45 +08:00
Junyan Qin 627eb6b8ef feat(tenancy): harden shared cloud runtime boundaries 2026-07-24 16:22:45 +08:00
Junyan Qin 3f01ffe63b feat(tenancy): establish cloud isolation foundations 2026-07-24 16:22:45 +08:00
Junyan Qin d7adbeec1e docs: finalize cloud v2 multi-tenant decisions 2026-07-24 16:22:44 +08:00
Junyan Qin abf77cecfa docs(tenancy): refine architecture options 2026-07-24 16:22:44 +08:00
Junyan Qin 270622ae9d docs(tenancy): revise single-instance SaaS topology 2026-07-24 16:22:44 +08:00
Junyan Qin 30f414a534 docs(tenancy): record verification evidence 2026-07-24 16:22:44 +08:00
Junyan Qin 8b7ce77cec feat(tenancy): implement workspace isolation 2026-07-24 16:22:44 +08:00
Junyan Qin 37099ddf7e docs: redesign multi-tenant workspace architecture 2026-07-24 16:22:44 +08:00
Junyan Qin ee59e2d3fd Add OSS and commercial workspace boundaries 2026-07-24 16:22:44 +08:00
Junyan Qin a4550350c0 Document multi-tenant workspace architecture 2026-07-24 16:22:44 +08:00
Hyu 38e35d328a feat(web): add marketplace likes (#2352)
Co-authored-by: dadachann <185672915+dadachann@users.noreply.github.com>
2026-07-24 15:16:37 +08:00
RockChinQ 4226f71f05 fix(runtime): make plugin and box connectors resilient (#2347)
Preserve the contributor-authored MCP timeout commit and the follow-up configurable timeout and runtime robustness fixes.
2026-07-23 20:25:21 +08:00
Junyan Qin 283c6949f4 fix(mcp): validate tool call timeout config 2026-07-23 18:26:41 +08:00
Junyan Qin 0dfae76e39 fix(mcp): make tool call timeout configurable 2026-07-23 18:26:37 +08:00
douxt 7677d1a288 fix(mcp): add 30s timeout to prevent MCP tool calls from hanging indefinitely
Previously, MCP tool calls via call_tool() had no timeout, so a hung MCP
server would block the entire session indefinitely (exacerbated by
concurrency.session=1). This wraps the call in asyncio.timeout(30) and
raises an Exception on expiry, letting the LLM recover gracefully.

Closes #2339
2026-07-23 18:26:18 +08:00
Hyu a53a41b1ca feat(web): trim dynamic form strings on save (#2349) 2026-07-23 18:14:54 +08:00
Junyan Qin 73e47ea2d4 fix(mcp): keep waiting for box recovery 2026-07-23 17:41:27 +08:00
Junyan Qin 06e03af994 fix(runtime): complete reconnect recovery paths 2026-07-23 17:28:31 +08:00
Junyan Qin 9cbbaf617b fix(runtime): consume sdk shutdown fix 2026-07-23 16:46:48 +08:00
Junyan Qin 210e5349d9 fix(runtime): preserve lockfile metadata 2026-07-23 16:34:03 +08:00
Junyan Qin a2f6814517 fix(runtime): pin robust sdk release 2026-07-23 16:28:33 +08:00
Junyan Qin 8511178666 fix(runtime): harden connector lifecycle 2026-07-23 16:03:26 +08:00
Junyan Qin 998b76d53a fix(runtime): make plugin and box connectors resilient 2026-07-21 18:43:24 +08:00
Junyan Qin 1765c43262 Revert "fix(runtime): make plugin and box connectors resilient"
This reverts commit 0b461e5830.
2026-07-21 18:43:01 +08:00
Junyan Qin 0b461e5830 fix(runtime): make plugin and box connectors resilient 2026-07-21 18:41:22 +08:00
DongXiaoming 76c5003c21 fix(localagent): stop re-seeding stream accumulator with previous round content (#2329)
When the LLM (e.g. MiniMax-M3) returns multiple rounds after tool calls, the _StreamAccumulator was initialized with initial_content=first_content, causing every subsequent round to repeat the entire first message. Remove the re-seeding so each round starts with a clean accumulator.

Add regression test verifying multi-round tool call content is not duplicated.
2026-07-20 16:38:11 +08:00
983 changed files with 95537 additions and 102822 deletions
+4
View File
@@ -37,6 +37,10 @@ jobs:
working-directory: web
run: pnpm install --frozen-lockfile
- name: Run frontend unit tests
working-directory: web
run: pnpm test:unit
- name: Install Playwright browsers
working-directory: web
run: pnpm exec playwright install --with-deps chromium
+10 -2
View File
@@ -44,7 +44,9 @@ jobs:
runs-on: ubuntu-latest
services:
postgres:
image: postgres:16
# Release migration 0013 installs the pgvector extension in the shared
# business database; CI must exercise the same extension availability.
image: pgvector/pgvector:pg16
env:
POSTGRES_USER: langbot
POSTGRES_PASSWORD: langbot
@@ -75,4 +77,10 @@ jobs:
- name: Run PostgreSQL migration tests
env:
TEST_POSTGRES_URL: postgresql+asyncpg://langbot:langbot@localhost:5432/langbot_test
run: uv run pytest tests/integration/persistence/test_migrations_postgres.py -q --tb=short
run: >-
uv run pytest
tests/integration/persistence/test_migrations_postgres.py
tests/integration/persistence/test_pgvector_postgres.py
tests/integration/persistence/test_release_migration_postgres.py
tests/integration/persistence/test_plugin_identity_migration.py
-q --tb=short
+1 -1
View File
@@ -83,7 +83,7 @@ Config keys to verify in `data/config.yaml` / `src/langbot/templates/config.yaml
## Change Rules
- HTTP API changes that should be agent-accessible must update the matching MCP tool in `src/langbot/pkg/api/mcp/server.py` and the relevant skill under `skills/` in the same pass.
- New schema changes use Alembic under `src/langbot/pkg/persistence/alembic/versions/`. LangBot 4.x does not support upgrading 3.x databases.
- New schema changes use Alembic under `src/langbot/pkg/persistence/alembic/versions/`; do not add legacy `dbmXXX` migrations.
- New platform behavior belongs in platform adapters only for platform translation; pipeline/business logic belongs in `pkg/pipeline/` or services.
- User-facing strings must support i18n (`en_US`, `zh_Hans`; include `ja_JP` where the repo already does).
- Code comments and docstrings must be English.
+16 -11
View File
@@ -52,13 +52,12 @@ LangBot/
│ │ ├── api/ # HTTP API + MCP server mount
│ │ ├── platform/ # IM adapters and runtime bot manager
│ │ ├── pipeline/ # Message routing and pipeline stages
│ │ ├── provider/ # Model providers and Host-owned tools
│ │ ├── agent/ # Agent/AgentRunner orchestration and run state
│ │ ├── provider/ # LLM runners, model manager, tools
│ │ ├── plugin/ # LangBot-side Plugin Runtime connector/handler
│ │ ├── box/ # LangBot-side Box service/connector
│ │ ├── skill/ # Skill metadata/activation integration
│ │ ├── rag/ , vector/ # Knowledge-base and vector DB integration
│ │ ├── persistence/ # SQLAlchemy/SQLModel and Alembic migrations
│ │ ├── persistence/ # SQLAlchemy/SQLModel, Alembic, legacy migrations
│ │ ├── storage/ # Local/S3 file storage abstraction
│ │ └── config/, entity/, utils/, telemetry/, survey/
│ ├── libs/ # Vendored third-party platform SDKs
@@ -81,7 +80,7 @@ Platform adapter
→ Controller
→ RuntimePipeline
→ PipelineStage chain
AgentRunner orchestrator / ToolManager / PluginRuntimeConnector / BoxService
RequestRunner / ToolManager / PluginRuntimeConnector / BoxService
→ response via adapter
```
@@ -108,7 +107,7 @@ Inbound platform messages enter through adapter-specific SDK callbacks. The comm
3. `MessageAggregator` batches/normalizes messages before adding a `Query` to `QueryPool`.
4. `Controller` in `pkg/pipeline/controller.py` selects queries subject to global pipeline concurrency and per-session concurrency.
5. `RuntimePipeline` in `pkg/pipeline/pipelinemgr.py` runs configured pipeline stages using a responsibility-chain style executor that supports generator stages.
6. The chat stage emits plugin events and projects the current query into the AgentRunner Host orchestrator. The selected plugin AgentRunner returns streaming or final results while the Host owns authorization, tools, telemetry, and conversation history.
6. The chat stage emits plugin events, calls a configured `RequestRunner`, handles streaming/non-streaming responses, records telemetry, and appends conversation history.
7. Output stages send text, cards, chunks, files, or error notices back through the original platform adapter.
Pipeline components are registered by decorators and package import side effects. When adding a new stage, loader, runner, or adapter, check the corresponding preregistration mechanism instead of inventing a second registry.
@@ -137,12 +136,12 @@ Important pieces:
Pipelines are configuration-driven. Prefer adding a stage or extending an existing stage family over hard-coding behavior in platform adapters.
## Agents, Providers, RAG, and Tools
## Provider, RAG, and Tools
Agent orchestration lives under `pkg/agent/`; model providers and tools live under `pkg/provider/`.
Provider code lives under `pkg/provider/`.
- `modelmgr/` manages configured model providers and requesters.
- `pkg/agent/runner/` discovers plugin AgentRunner components, resolves bindings, constructs run-scoped context/resources, and records execution state.
- `runners/` implements request runners such as the local agent runner and external workflow integrations.
- `tools/toolmgr.py` aggregates tools from native tools, plugin tools, external MCP servers, and skill-authoring tools.
- `tools/loaders/mcp.py` is the MCP client side: external MCP servers that LangBot connects to for agent tools.
- RAG lives across `pkg/rag/`, `pkg/vector/`, model services, and plugin KnowledgeEngine actions.
@@ -179,6 +178,12 @@ In this repo:
- `pkg/provider/tools/loaders/native.py`, `mcp_stdio.py`, and skill loaders depend on Box availability.
- `pkg/skill/manager.py` loads skills from the Box runtime, falling back to local `data/skills` when needed.
Durable Box Workspace storage is shared across placement generations, but
sandbox sessions and managed processes are generation-scoped. LangBot validates
the current execution binding before an MCP stdio relay attach and sends the
Workspace/generation binding in authenticated headers, so a placement cutover
retires stale processes and closes already-attached relays.
In `langbot-plugin-sdk`:
- `src/langbot_plugin/box/server.py` implements `lbp box` and the WebSocket endpoints on `:5410`.
@@ -207,8 +212,8 @@ Persistence is centered on `pkg/persistence/mgr.py`.
- SQLite is the default database; PostgreSQL is supported.
- Models live under `pkg/entity/persistence/`.
- Fresh schemas are created from current metadata, then Alembic migrations run to head. LangBot 4.x does not upgrade 3.x databases.
- New schema changes use Alembic under `pkg/persistence/alembic/versions/`; there is no legacy migration chain in 4.x.
- Fresh schemas are created from metadata, then legacy migrations run up to the frozen 3.x baseline, then Alembic migrations run to head.
- New schema changes should use Alembic under `pkg/persistence/alembic/versions/`; do not extend the frozen legacy migration chain.
Configuration starts from `src/langbot/templates/config.yaml` and is generated into `data/config.yaml` on first run. Most long-lived managers read from `ap.instance_config.data`.
@@ -240,7 +245,7 @@ When one of these changes, update the others if the behavior or contract changed
- New LLM tool source: extend `pkg/provider/tools/loaders/` and `ToolManager` intentionally.
- New plugin component/API/protocol: change `langbot-plugin-sdk` first or in lockstep, then update LangBot bridge code.
- New Box capability: change both `pkg/box/` and `langbot-plugin-sdk/src/langbot_plugin/box/`, plus config and tests.
- New database schema: add an Alembic migration.
- New database schema: add an Alembic migration, not a legacy `dbmXXX` migration.
## Design Biases
+3 -3
View File
@@ -38,7 +38,7 @@ COPY --from=node /app/web/dist ./web/dist
COPY --from=nsjail-build /usr/local/bin/nsjail /usr/local/bin/nsjail
RUN apt-get update \
&& apt-get install -y --no-install-recommends gcc ca-certificates curl gnupg \
&& apt-get install -y --no-install-recommends gcc ca-certificates curl git gnupg \
# nsjail runtime libraries (the build toolchain stays in the nsjail-build
# stage; only these shared libs are needed to execute the binary).
&& apt-get install -y --no-install-recommends libprotobuf32 libnl-route-3-200 \
@@ -63,8 +63,8 @@ RUN apt-get update \
&& rm -f /tmp/nodesource_setup.sh \
&& python -m pip install --no-cache-dir uv \
&& uv sync \
&& apt-get purge -y --auto-remove curl gnupg \
&& apt-get purge -y --auto-remove curl git gnupg \
&& rm -rf /var/lib/apt/lists/* \
&& touch /.dockerenv
CMD [ "uv", "run", "--no-sync", "main.py" ]
CMD [ "uv", "run", "--no-sync", "main.py" ]
+30 -2
View File
@@ -14,6 +14,13 @@ services:
restart: on-failure
environment:
- TZ=Asia/Shanghai
# Shared with the langbot service and sent only as a WebSocket handshake
# header. Generate with: openssl rand -hex 32
- LANGBOT_PLUGIN_RUNTIME_CONTROL_TOKEN=${LANGBOT_PLUGIN_RUNTIME_CONTROL_TOKEN:-}
# Process-wide admission for every asyncio.to_thread() call.
- LANGBOT_BLOCKING_EXECUTOR_MAX_WORKERS=${LANGBOT_BLOCKING_EXECUTOR_MAX_WORKERS:-8}
- LANGBOT_BLOCKING_EXECUTOR_MAX_PENDING=${LANGBOT_BLOCKING_EXECUTOR_MAX_PENDING:-128}
- LANGBOT_BLOCKING_EXECUTOR_MAX_INFLIGHT_PER_SCOPE=${LANGBOT_BLOCKING_EXECUTOR_MAX_INFLIGHT_PER_SCOPE:-4}
command: ["uv", "run", "--no-sync", "-m", "langbot_plugin.cli.__init__", "rt"]
networks:
- langbot_network
@@ -40,9 +47,19 @@ services:
restart: on-failure
environment:
- TZ=Asia/Shanghai
# Shared control-plane secret used to authenticate both the RPC socket
# and managed-process relay. Generate once (for example with
# ``openssl rand -hex 32``) and export it before enabling this profile.
# An empty value is accepted by Compose so Box can remain optional, but
# the Box runtime itself fails closed when the profile is started.
- LANGBOT_BOX_CONTROL_TOKEN=${LANGBOT_BOX_CONTROL_TOKEN:-}
# Box has its own process-wide blocking-work budget.
- LANGBOT_BLOCKING_EXECUTOR_MAX_WORKERS=${LANGBOT_BLOCKING_EXECUTOR_MAX_WORKERS:-8}
- LANGBOT_BLOCKING_EXECUTOR_MAX_PENDING=${LANGBOT_BLOCKING_EXECUTOR_MAX_PENDING:-128}
- LANGBOT_BLOCKING_EXECUTOR_MAX_INFLIGHT_PER_SCOPE=${LANGBOT_BLOCKING_EXECUTOR_MAX_INFLIGHT_PER_SCOPE:-4}
# The Box runtime does NOT read box.local.* from config.yaml or env; it
# receives its configuration from LangBot via the INIT RPC action.
# Do not add LANGBOT_BOX_* / BOX__* here — they would be silently ignored.
# receives its functional configuration from LangBot via the INIT RPC
# action. Do not add BOX__* here because those would be ignored.
# Launched through the same CLI entry point as the plugin runtime
# (`langbot_plugin.cli.__init__ <subcommand>`). WebSocket is the default
# control transport — mirrors `rt`, which also runs with no flag. Pass
@@ -60,6 +77,17 @@ services:
restart: on-failure
environment:
- TZ=Asia/Shanghai
# Must match langbot_plugin_runtime. Empty/missing values make the
# external control channel fail closed.
- LANGBOT_PLUGIN_RUNTIME_CONTROL_TOKEN=${LANGBOT_PLUGIN_RUNTIME_CONTROL_TOKEN:-}
# Must match the value supplied to langbot_box. The token is sent only
# in WebSocket handshake headers, never in URLs or action payloads.
- LANGBOT_BOX_CONTROL_TOKEN=${LANGBOT_BOX_CONTROL_TOKEN:-}
# Core process-wide blocking-work admission. These are native config
# overrides and are persisted with the effective data/config.yaml.
- SYSTEM__BLOCKING_EXECUTOR__MAX_WORKERS=${LANGBOT_BLOCKING_EXECUTOR_MAX_WORKERS:-8}
- SYSTEM__BLOCKING_EXECUTOR__MAX_PENDING=${LANGBOT_BLOCKING_EXECUTOR_MAX_PENDING:-128}
- SYSTEM__BLOCKING_EXECUTOR__MAX_INFLIGHT_PER_SCOPE=${LANGBOT_BLOCKING_EXECUTOR_MAX_INFLIGHT_PER_SCOPE:-4}
# Unified env-override convention: SECTION__SUBSECTION__KEY overrides the
# matching config.yaml field (see LoadConfigStage). These map onto
# box.* and are forwarded to the Box runtime via INIT RPC.
+94 -8
View File
@@ -4,6 +4,10 @@
# Full deployment guide (zh/en/ja): https://docs.langbot.app -> Installation -> Kubernetes
#
# Usage:
# kubectl -n langbot create secret generic langbot-plugin-runtime-control \
# --from-literal=token="$(openssl rand -hex 32)"
# kubectl -n langbot create secret generic langbot-box-control \
# --from-literal=token="$(openssl rand -hex 32)"
# kubectl apply -f kubernetes.yaml
#
# Prerequisites:
@@ -87,6 +91,12 @@ metadata:
data:
TZ: "Asia/Shanghai"
PLUGIN__RUNTIME_WS_URL: "ws://langbot-plugin-runtime:5400/control/ws"
SYSTEM__BLOCKING_EXECUTOR__MAX_WORKERS: "8"
SYSTEM__BLOCKING_EXECUTOR__MAX_PENDING: "128"
SYSTEM__BLOCKING_EXECUTOR__MAX_INFLIGHT_PER_SCOPE: "4"
LANGBOT_BLOCKING_EXECUTOR_MAX_WORKERS: "8"
LANGBOT_BLOCKING_EXECUTOR_MAX_PENDING: "128"
LANGBOT_BLOCKING_EXECUTOR_MAX_INFLIGHT_PER_SCOPE: "4"
# Box sandbox runtime endpoint. LangBot connects to the Box runtime over
# WebSocket. The hostname MUST match the langbot-box Service name. Note the
# in-container default ("langbot_box") uses an underscore, which is an
@@ -127,6 +137,26 @@ spec:
configMapKeyRef:
name: langbot-config
key: TZ
- name: LANGBOT_PLUGIN_RUNTIME_CONTROL_TOKEN
valueFrom:
secretKeyRef:
name: langbot-plugin-runtime-control
key: token
- name: LANGBOT_BLOCKING_EXECUTOR_MAX_WORKERS
valueFrom:
configMapKeyRef:
name: langbot-config
key: LANGBOT_BLOCKING_EXECUTOR_MAX_WORKERS
- name: LANGBOT_BLOCKING_EXECUTOR_MAX_PENDING
valueFrom:
configMapKeyRef:
name: langbot-config
key: LANGBOT_BLOCKING_EXECUTOR_MAX_PENDING
- name: LANGBOT_BLOCKING_EXECUTOR_MAX_INFLIGHT_PER_SCOPE
valueFrom:
configMapKeyRef:
name: langbot-config
key: LANGBOT_BLOCKING_EXECUTOR_MAX_INFLIGHT_PER_SCOPE
volumeMounts:
- name: plugin-data
mountPath: /app/data/plugins
@@ -139,7 +169,8 @@ spec:
cpu: "1000m"
# Liveness probe to restart container if it becomes unresponsive
livenessProbe:
tcpSocket:
httpGet:
path: /healthz
port: 5400
initialDelaySeconds: 30
periodSeconds: 10
@@ -147,7 +178,8 @@ spec:
failureThreshold: 3
# Readiness probe to know when container is ready to accept traffic
readinessProbe:
tcpSocket:
httpGet:
path: /healthz
port: 5400
initialDelaySeconds: 10
periodSeconds: 5
@@ -246,9 +278,28 @@ spec:
configMapKeyRef:
name: langbot-config
key: TZ
- name: LANGBOT_BOX_CONTROL_TOKEN
valueFrom:
secretKeyRef:
name: langbot-box-control
key: token
- name: LANGBOT_BLOCKING_EXECUTOR_MAX_WORKERS
valueFrom:
configMapKeyRef:
name: langbot-config
key: LANGBOT_BLOCKING_EXECUTOR_MAX_WORKERS
- name: LANGBOT_BLOCKING_EXECUTOR_MAX_PENDING
valueFrom:
configMapKeyRef:
name: langbot-config
key: LANGBOT_BLOCKING_EXECUTOR_MAX_PENDING
- name: LANGBOT_BLOCKING_EXECUTOR_MAX_INFLIGHT_PER_SCOPE
valueFrom:
configMapKeyRef:
name: langbot-config
key: LANGBOT_BLOCKING_EXECUTOR_MAX_INFLIGHT_PER_SCOPE
# The Box runtime does NOT read box.local.* / BOX__* from its own env;
# it receives its configuration from LangBot via the INIT RPC action.
# Do not add BOX__* here — they would be silently ignored.
# it receives its functional configuration from LangBot via INIT.
volumeMounts:
# Box workspace root — identical path on node, box, and sandbox
# containers (see the IMPORTANT note above).
@@ -265,14 +316,18 @@ spec:
memory: "1Gi"
cpu: "1000m"
livenessProbe:
tcpSocket:
httpGet:
path: /healthz
port: 5410
initialDelaySeconds: 20
periodSeconds: 10
timeoutSeconds: 5
failureThreshold: 3
readinessProbe:
tcpSocket:
httpGet:
# Unlike liveness, readiness validates the configured backend and
# all strict managed-mode isolation guarantees.
path: /readyz
port: 5410
initialDelaySeconds: 10
periodSeconds: 5
@@ -319,6 +374,10 @@ metadata:
app: langbot
spec:
replicas: 1
# Plugin Runtime has a single active LangBot control owner. Recreate avoids
# two LangBot pods fighting over that connection during a rolling update.
strategy:
type: Recreate
selector:
matchLabels:
app: langbot
@@ -352,6 +411,26 @@ spec:
configMapKeyRef:
name: langbot-config
key: PLUGIN__RUNTIME_WS_URL
- name: LANGBOT_PLUGIN_RUNTIME_CONTROL_TOKEN
valueFrom:
secretKeyRef:
name: langbot-plugin-runtime-control
key: token
- name: SYSTEM__BLOCKING_EXECUTOR__MAX_WORKERS
valueFrom:
configMapKeyRef:
name: langbot-config
key: SYSTEM__BLOCKING_EXECUTOR__MAX_WORKERS
- name: SYSTEM__BLOCKING_EXECUTOR__MAX_PENDING
valueFrom:
configMapKeyRef:
name: langbot-config
key: SYSTEM__BLOCKING_EXECUTOR__MAX_PENDING
- name: SYSTEM__BLOCKING_EXECUTOR__MAX_INFLIGHT_PER_SCOPE
valueFrom:
configMapKeyRef:
name: langbot-config
key: SYSTEM__BLOCKING_EXECUTOR__MAX_INFLIGHT_PER_SCOPE
# Box (sandbox) runtime endpoint. Connects LangBot to the langbot-box
# Service over WebSocket. Remove this (and the langbot-box Deployment)
# and set BOX__ENABLED=false if you do not want the sandbox.
@@ -360,6 +439,13 @@ spec:
configMapKeyRef:
name: langbot-config
key: BOX__RUNTIME__ENDPOINT
# Same Secret as langbot-box. It authenticates the RPC and managed-
# process relay handshakes and is never put in a URL or RPC payload.
- name: LANGBOT_BOX_CONTROL_TOKEN
valueFrom:
secretKeyRef:
name: langbot-box-control
key: token
# box.local.* config — forwarded to the Box runtime via INIT RPC. The
# host_root MUST match the box-root hostPath mountPath below AND the box
# Deployment's box-root mountPath, so that skill package paths resolve
@@ -392,7 +478,7 @@ spec:
# Liveness probe to restart container if it becomes unresponsive
livenessProbe:
httpGet:
path: /
path: /healthz
port: 5300
initialDelaySeconds: 60
periodSeconds: 10
@@ -401,7 +487,7 @@ spec:
# Readiness probe to know when container is ready to accept traffic
readinessProbe:
httpGet:
path: /
path: /healthz
port: 5300
initialDelaySeconds: 30
periodSeconds: 5
+31 -14
View File
@@ -8,13 +8,21 @@ API keys can be managed through the web interface:
1. Log in to the LangBot web interface
2. Click the "API Keys" button at the bottom of the sidebar
3. Create, view, copy, or delete API keys as needed
3. Create an API key and copy its secret immediately
4. Revoke keys that are no longer needed
Database-backed API-key secrets are returned exactly once. LangBot stores only
a SHA-256 lookup hash, so an existing secret cannot be displayed or recovered
later. Each key belongs to one Workspace, has explicit permission scopes, and
may have an expiry. The Workspace is derived from the authenticated key; an
`X-Workspace-Id` header cannot redirect it to another tenant.
## Global API Key (config.yaml)
In addition to web-UI-created keys (stored in the database, prefixed `lbk_`),
LangBot supports a **global API key** defined directly in `data/config.yaml`.
This is useful for automated deployments, infrastructure-as-code, and AI agents
This is a Community-edition bootstrap option for automated deployments,
infrastructure-as-code, and AI agents
that need API/MCP access **without a login session and without creating a
database record first**.
@@ -27,10 +35,12 @@ api:
Behavior:
- When `api.global_api_key` is a non-empty string, that exact value is accepted
anywhere a normal API key is accepted — the `X-API-Key` header or
`Authorization: Bearer <key>` — across the HTTP service API **and the MCP
server**.
- In Community edition's singleton Workspace, a non-empty
`api.global_api_key` is bound to that Workspace and accepted across the HTTP
service API and the MCP server.
- The global config key is rejected when multi-Workspace SaaS mode is enabled;
SaaS automation must use a database-backed Workspace key or a closed control
plane credential.
- The global key does **not** require the `lbk_` prefix; use any sufficiently
strong secret.
- Leave it empty (`''`, the default) to disable it entirely; only database-backed
@@ -38,9 +48,10 @@ Behavior:
- Existing installs are unaffected until you add the key — config completion only
backfills top-level keys, and the lookup is defensive when the field is absent.
> **Security:** the global key is stored in plaintext in `config.yaml`. Only
> enable it on trusted/internal deployments, keep the file permissions tight,
> always serve over HTTPS, and rotate the value if it may have leaked.
> **Security:** the global key is stored in plaintext in `config.yaml` and has
> the singleton Workspace's full fixed permission set. Only enable it on
> trusted/internal Community deployments, keep file permissions tight, always
> serve over HTTPS, and rotate it if it may have leaked.
## Using API Keys
@@ -60,7 +71,9 @@ Authorization: Bearer lbk_your_api_key_here
## Available APIs
All existing LangBot APIs now support **both user token and API key authentication**. This means you can use API keys to access:
Endpoints that declare API-key authentication accept either a user token or a
Workspace API key. The key must include the permission required by the route.
This includes:
- **Model Management** - `/api/v1/provider/models/llm` and `/api/v1/provider/models/embedding`
- **Bot Management** - `/api/v1/platform/bots`
@@ -227,6 +240,11 @@ or
}
```
### 403 Forbidden
The key is valid for its Workspace but does not include the fixed permission
required by the route.
### 500 Internal Server Error
```json
@@ -240,7 +258,7 @@ or
1. **Keep API keys secure**: Store them securely and never commit them to version control
2. **Use HTTPS**: Always use HTTPS in production to encrypt API key transmission
3. **Rotate keys regularly**: Create new API keys periodically and delete old ones
3. **Rotate keys regularly**: Create new API keys periodically and revoke old ones
4. **Use descriptive names**: Give your API keys meaningful names to track their usage
5. **Delete unused keys**: Remove API keys that are no longer needed
6. **Use X-API-Key header**: Prefer using the `X-API-Key` header for clarity
@@ -317,7 +335,6 @@ curl -X POST \
## Notes
- The same endpoints work for both the web UI (with user tokens) and external services (with API keys)
- API-key-enabled endpoints use the same resource shapes as the web UI
- No need to learn different API paths - use the existing API documentation with API key authentication
- All endpoints that previously required user authentication now also accept API keys
- API keys never select a Workspace from a request header; their persisted binding is authoritative
@@ -1,149 +0,0 @@
# Agent-owned Context 协议设计
本文档描述插件化 AgentRunner 场景下的上下文边界**设计理由**。结论先行:LangBot 不应成为最终 agentic context manager;它提供 context substrateAgentRunner 或其背后的 runtime 自己决定如何管理历史、压缩、召回和 KV cache。
> 涉及的数据结构(`AgentRunContext`、`ContextAccess`、`AgentRunAPIProxy` 等)唯一定义在 [PROTOCOL_V1.md](./PROTOCOL_V1.md)。本文只讲语义和约束,不重抄 schema。
## 1. 设计原则
### 1.1 Agent 拥有上下文策略
不同 runner 背后的 runtime 差异很大:
- 官方 local-agent 可能依赖 LangBot 的模型、工具、知识库和存储。
- Claude Code SDK / Codex 类 runtime 有自己的 session、transcript、tool loop 和上下文压缩。
- Pi Agent SDK 或外部 agent 平台可能只需要当前事件和一个外部 conversation key。
因此 LangBot 不应强行决定最终传给模型的历史窗口。Host 只提供:当前事件的完整结构化信息、稳定身份和会话引用、可授权读取的 history / event / state API、sandbox/workspace 文件能力、可投影给外部 harness 的 scoped context / SDK-owned MCP bridge / resource handles、payload hard cap 和权限 guardrail。
### 1.2 Host 不定义通用历史窗口
历史窗口策略不是 AgentRunner 协议或 Query entry adapter 的核心概念。Host 只提供 history pull API、cursor、hard cap 和权限边界;runner 自己决定是否读取、读取多少、如何截断和压缩。
正确的问题不是"LangBot 每轮裁几轮历史给 agent",而是:
- 这类 runner 是否自管 context
- 事件到来时 host 应 inline 哪些最小信息?
- agent 需要更多上下文时通过什么 API 拉取?
- host 如何保证安全、可审计和可分页?
### 1.3 Host 保存事实源,Agent 管理 working context
三类数据要分开:
- `EventLog`: Host 保存原始事件、工具调用、投递结果、错误和系统事件。
- `Transcript`: Host 从 EventLog 投影出的对话视图,用于 UI、审计和按需历史读取。
- `Working context`: Agent 本轮实际送进模型或 runtime 的上下文,由 AgentRunner 决定。
LangBot 不提供 host-side inline history window。简单 runner 如果需要历史窗口,应在 runner 内部通过 Host history API 拉取并裁剪。
## 2. Event 到来时传什么
默认 `AgentRunContext`PROTOCOL_V1 §5.2)应尽量小且稳定。默认规则:
- Host MUST NOT inline full history by default.
- Host SHOULD inline only current event / input and context handles.
- Runner owns working-context assembly.
- Runner MAY use Host history / event / state / storage API and sandbox/workspace file tools when authorized.
- Official runners MUST consume Host infrastructure through the same public API as third-party runners.
### 2.1 必须 inline 的内容
当前 event 的类型/id/时间/source;当前输入文本和结构化内容;附件/文件/图片的 metadata、path 或 URLactor / subject / conversation / thread / bot / workspacedelivery 能力;已授权资源列表;context cursors 和可用 API 能力;Agent/runner config。这些是 agent 决定下一步所需的最低信息。
### 2.2 默认不 inline 的内容
完整历史消息、大文件全文、大工具结果、全量知识库内容、平台原始 payload 大对象、每轮重新生成的大段 summary。这些会破坏跨进程序列化成本、泄露范围、KV cache 稳定性,也会迫使 host 替 agent 做 context 策略。
### 2.3 不提供 Host Inline History Window
`AgentRunContext` 不包含 `bootstrap` 字段。Host 不下发历史窗口,也不通过 Pipeline 配置决定窗口大小。runner 若需要类似 `recent_tail` 的策略,应在自己的 manifest/config schema 中声明参数,并在 runner 内部通过 history API 读取、裁剪和压缩。Host 只负责权限、分页、hard cap 和事实源。
## 3. ContextAccess 的作用
`ContextAccess`PROTOCOL_V1 §5.8)是 host 交给 agent 的上下文读取入口描述,告诉 agent:当前事件位于哪条 conversation / thread、若需要更多历史从哪个 cursor 开始拉、host inline 了什么没 inline 什么、当前 run 有哪些 context API 权限。
## 4. Agent 如何获取更多上下文
所有 API 都走 `AgentRunAPIProxy`PROTOCOL_V1 §8),由 host 用 `run_id` 校验。
外部 harness 不能直接访问 LangBot 资源。无论是 history、event、state、model、tool、knowledge base,还是 LangBot skills,都必须通过 SDK runtime 转发到 Host API,并由 Host 按 active `run_id`、runner identity、binding resource policy 和 caller plugin identity 校验。当前运行文件进入授权 sandbox/workspace 后,再由 runner 用 read/write/exec 类工具按需访问。harness 自己的 native tools 只属于 harness 执行环境,不能绕过 SDK runtime 访问 LangBot 内部资源。
### 4.1 History
```python
await api.history_page(conversation_id=ctx.context.conversation_id,
before_cursor=ctx.context.latest_cursor,
limit=50, direction="backward", include_attachments=False)
```
返回 `HistoryPage`schema 见 PROTOCOL_V1 §8)。
约束:`limit` 有 host hard cap;默认只能读当前 conversation / thread;跨会话读取需 binding policy / run authorization snapshot 授权;可返回 attachment ref,不默认返回大文件内容。
### 4.2 Search
```python
await api.history_search(query="用户之前提到的数据库连接信息",
filters={"conversation_id": ..., "event_types": ["message.received"]},
top_k=10)
```
Search 可先用数据库全文索引,后续接 embedding recall。它是 host 检索能力,不等于 agent 的长期记忆策略。
### 4.3 Event / State
- Event API`events.get` / `events.page`)用于读取非消息事件、工具事件、系统事件。Agent 不应把所有事件都当成 user/assistant message。
- State API`state.get` / `set`)是可选寄宿能力。自管 runtime 可以完全不用;依附 LangBot 的官方 runner 可以使用,例如 `external.session_id``summary.checkpoint`
### 4.4 大文件与工具协作
大文件、多模态输入和工具产物不要内联进 prompt 或 tool resultmessage/content 里只放小文本和必要摘要;当前事件附件由 Host staged 到授权 sandbox/workspace,并在 input attachment 中给出轻量 metadata/path。工具之间传递大结果时传 sandbox path 或 attachment ref,不传完整 blob。Host 只保证当前 run 授权范围,默认不允许插件直接读任意本地路径;临时文件由 sandbox 生命周期和清理机制管理。
### 4.5 External harness context projection
外部 harness 的总体边界以 [HOST_SDK_INFRASTRUCTURE.md](./HOST_SDK_INFRASTRUCTURE.md) §4.8 为准。本节只描述 context projection 的推荐形态。
Claude Code、Codex、Kimi Code 这类 runtime 通常已有自己的 session、工具 loop、MCP 加载、上下文压缩和工作目录。LangBot 不应把它们改造成"host prompt assembler",而应提供可审计的事件和资源投影。推荐 projection 形态:
- `agent-context.json`:结构化 JSON,包含 `run_id``event``actor``subject``input``delivery``resources``context``state``runtime`
- `LANGBOT_CONTEXT.md`:人类可读摘要。
- `resources`:只包含本次 run 授权后的资源句柄和能力摘要,不暴露 Host 内部私有对象、secret 或资源内容。
- `skills`LangBot skills 不是直接投影给 harness native tool loop 的文件能力,而是**一组被授权的 tool**。发现走 `list_skills`(或 `langbot_list_assets` 增加 skills 一类),激活/注册走 `activate` / `register_skill`,包内操作走 native exec/read/write,统一通过 `ctx.resources.tools``AgentRunAPIProxy` 或 SDK-owned MCP bridge 暴露。Host 不向 prompt 注入 skill 索引(无 progressive-disclosure 注入);harness 通过调用发现工具主动查询 skill 清单。`agent-context.json``skills` 字段仅作发现工具的数据来源与可选 `suggested_skill_prompt` 的输入。
- `MCP config`:只投影 per-run、scoped 的 SDK-owned bridge 或外部 MCP 连接配置;LangBot 资源访问必须回到 SDK runtime / Host API,不允许 harness 通过自带 MCP/native tool 直接读 Host 内部资源。
- `state pointers`:外部 session id、working directory、checkpoint 等小型 JSON 状态通过 Host state API 保存。
当前官方外部 harness 路径由 ACP / Claude Code / Codex 等 runner 插件承担(现状见 OFFICIAL_RUNNER_PLUGINS §7)。这类 projection 是"把 LangBot 事实源和授权资源句柄交给 harness",不是"把 LangBot 资源本体或内部权限交给 harness",也不是"由 LangBot 决定最终模型上下文"。
## 5. Runner 上下文边界
Host 只给当前事件、当前输入和 context handles。Runner 是否能拉取历史、事件、state 或 storage、是否能访问 sandbox/workspace 文件,以运行时 `ctx.context.available_apis` 和工具授权为准;runner 自己决定是否拉取历史、是否搜索、何时摘要、如何构造最终 prompt。
## 6. KV cache 友好的上下文管理
支持 Claude Code SDK、Codex、Pi Agent SDK 等 runtime 时,必须避免每轮由 LangBot 重组大块 prompt
- 稳定 session key`workspace/bot/binding/runner/conversation/thread`
- 每轮只传 delta:当前 event、attachment refs/path、少量 runtime metadata。
- 历史 append-only:不要每轮改写同一段 history 文本。
- Summary checkpoint 稳定:只有压缩发生时产生新 checkpoint。
- 大文件和工具结果写入 sandbox/workspace。
- Tool/context API schema 稳定,数据通过 API 拉取而非塞入 prompt。
- 对自管 runtime,优先让它复用自身 session/cache,而不是强制 LangBot 每轮重放 transcript。
- 模型窗口元信息应作为 resource/runtime metadata 暴露给 runner,由 runner 决定预算和压缩策略。
稳定 session key 的用途是隔离外部 runtime 的 resume/cache/state,不是改变 PROTOCOL_V1 §13 定义的 Agent 复用和 dispatch 边界。只有当某个外部 harness 的同一 native session 不支持并发 turn 时,runner 或 future runtime control plane 才应按 external session key 做 turn-level 串行化。
对长期运行的 external harness / daemon,推荐运行形态是 reader 与 writer 分离:一个 session reader 独占读取 stdout/SSE/native event stream,并把 native event 转成 `AgentRunResult` 或 task progress;用户输入只作为 turn write 进入该 session。当前一次性 CLI subprocess runner 可以继续在单次 `run(ctx)` 内同步收集 stdout,但后续改成长连接时不应让多个 request 同时读取同一 native stream。
## 7. Host guardrail
Agent 自管 context 不代表无限制访问。LangBot 仍必须控制:每次 run 的 active `run_id`、runner identity、当前 binding 的 resource policy、conversation / actor / subject scope、page size / sandbox file read size / API rate limit、跨会话读取权限、数据脱敏和敏感变量过滤、审计日志。Host 不负责"最佳上下文策略",但负责"不越权、不爆内存、不不可审计"。
外部 harness 的 native tools、shell、MCP 或 skill 机制不构成 LangBot 资源授权边界。只要访问的是 LangBot 持有的资源,就必须经 SDK runtime 转发并接受 Host 校验;完整边界见 HOST_SDK §4.8。
## 8. 官方 runner 与业务编排边界
官方 runner 插件可以把状态寄宿在 LangBot,但必须和第三方 runner 一样通过公开 Host API 消费。LangBot core 不内置官方 agent 的业务流程(prompt 组装、tool loop、RAG 编排、summary/compaction、"local-agent 专用"状态字段)。
官方 local-agent 应作为"依附 LangBot 基础设施的复杂 runner 参考实现"transcript/history 通过 `api.history_page()` / `api.history_search()` 读取,summary/checkpoint/外部 session id/用户偏好通过 `api.state_get()` / `api.state_set()` 或 storage 方法保存,图片/文件/工具大结果通过 sandbox/workspace read/write 工具访问,模型/工具/知识库通过 `api.invoke_llm()` / `api.call_tool()` / `api.retrieve_knowledge()` 调用。这样 LangBot 保持为通用 agent host,不变成内置 agent 框架。具体迁移要求见 [OFFICIAL_RUNNER_PLUGINS.md](./OFFICIAL_RUNNER_PLUGINS.md)。
@@ -1,227 +0,0 @@
# Agent Runner QA 指南
本文档是 agent-runner 插件化下一轮测试的唯一 QA 入口。它合并并取代旧的 Phase 1 验收矩阵与 2026-05-18 / 2026-05-29 两份本地 QA 报告。
目标不是保留完整历史流水账,而是指导测试 agent 用最小但高价值的路径判断当前分支是否仍然健康。
## 1. 测试边界
当前主线验证的是 AgentRunner Protocol v1
```text
event -> binding -> runner.run(ctx) -> result stream
```
本指南验证:
- Host 能通过当前 Query entry adapter 进入 event-first `run(event, binding)` 主链路。
- Runner 来自插件 registry,而不是旧内置 runner 分支。
- `local-agent` 能消费 Host 模型、工具、知识库、history、state、sandbox 文件等基础设施。
- 外部 harness runnerACP / Claude Code / Codex 等直接 runner 插件)能消费 event-first context,并把外部 session 指针写回 host-owned state。
- 错误、权限裁剪、无输出、timeout 等路径不会破坏主聊天流程。
本指南不验证:
- Runtime Control Plane v2。
- EventGateway / EventRouter 完整落地由外部 EBA 分支联调;本指南只验证本分支 Host 底座。
- 发布级 path isolation、secret filtering、MCP allowlist、资源配额和 workspace cleanup。
- 所有外部服务 runner 的真实凭据联调。
这些属于后续能力或发布门槛,分别见 [RUNTIME_CONTROL_PLANE_V2.md](./RUNTIME_CONTROL_PLANE_V2.md) 与 [SECURITY_HARDENING.md](./SECURITY_HARDENING.md)。
## 2. 状态定义
测试报告只使用以下状态:
| 状态 | 含义 |
| --- | --- |
| PASS | 按步骤执行,用户可见行为和日志证据都满足通过条件。 |
| FAIL | 环境可用,但行为不满足通过条件。 |
| BLOCKED | 凭据、CLI、外部服务、测试数据或本地配置缺失导致无法执行。必须写清阻塞原因。 |
| N/A | 当前 runner 或平台明确不支持该能力。必须引用 manifest、文档或配置说明。 |
不能使用“看起来正常”“大概通过”“基本没问题”等模糊状态。
## 3. 执行顺序
推荐按以下顺序执行,前一层失败时不要继续扩大测试面:
1. Host / SDK / runner 单测。
2. WebUI 登录与 Pipeline Debug Chat 基础 smoke。
3. `local-agent` 高价值场景。
4. 外部 code-agent harness smoke。
5. 权限和错误路径补充检查。
6. 汇总 PASS / FAIL / BLOCKED,并给出下一步建议。
用户可见流程必须通过 WebUI 或真实消息平台验证。API / curl 只能作为诊断证据,不能单独让 UI case PASS。
## 4. 必跑基线
### 4.1 单测基线
在 LangBot 仓库运行:
```bash
uv run --frozen pytest tests/unit_tests/agent
```
如果本次改动只触及默认配置或 API service,也至少补跑相关目标测试,例如:
```bash
uv run pytest tests/unit_tests/api/test_pipeline_service_defaults.py
```
通过条件:
- agent 单测全 PASS,或失败项已确认与本次 agent-runner 路径无关。
- 若失败来自 `context_builder``orchestrator``session_registry``resource_builder``plugin/handler.py` 的 run action 权限路径,不应进入 UI smoke。
### 4.2 环境基线
`langbot-skills` 做环境检查:
```bash
cd "$LANGBOT_SKILLS_REPO"
bin/lbs env doctor
bin/lbs case list
```
`LANGBOT_SKILLS_REPO` 指向当前工作区里的 `langbot-skills` 仓库。优先使用已有 case,而不是临时发明测试路径。
推荐首批 case
- `webui-login-state`
- `pipeline-debug-chat`
- `local-agent-basic-debug-chat`
- `local-agent-rag-debug-chat`(改动涉及 RAG / knowledge
- `local-agent-plugin-tool-call-debug-chat`(改动涉及 tool / resource policy
## 5. WebUI 主链路 Smoke
### 5.1 Runner registry
步骤:
1. 打开 WebUI Pipeline 配置页。
2. 查看 AI runner 下拉列表。
3. 选择 `plugin:langbot-team/LocalAgent/default`
4. 保存并刷新页面。
通过条件:
- runner 选项来自插件 registry。
- 保存后配置仍为 `ai.runner.id` + `ai.runner_config[id]`
- `runner_config` 表示 Agent/runner config,不表示插件实例状态。
- 不读取或回写旧 `ai.runner.runner` 字段。
- 不出现旧内置 runner stage 名(例如裸 `local-agent`)作为当前选中项或配置 surface。
- 插件没有循环重启或 metadata 加载失败。
### 5.2 主聊天路径
步骤:
1. 使用绑定 `plugin:langbot-team/LocalAgent/default` 的 Pipeline。
2. 在 Debug Chat 发送确定性普通文本。
3. 查看 WebUI 回复和后端日志。
通过条件:
- 用户可见回复正常。
- 后端日志显示走 `AgentRunOrchestrator` / `RUN_AGENT`
- 不走旧内置 local-agent 主执行分支。
- conversation transcript 写入用户消息和助手消息。
## 6. `local-agent` 高价值测试
只保留最能覆盖架构边界的场景。
| ID | 场景 | 操作 | 通过条件 |
| --- | --- | --- | --- |
| LA-01 | 绑定 prompt | 配置 system prompt 后发送文本。 | runner 使用 `ctx.config.prompt`,不读取 `ctx.adapter.extra["prompt"]`;回复体现绑定 prompt。 |
| LA-02 | history API | 连续两轮对话,第二轮引用第一轮 marker。 | runner 通过 Host history API 或自管上下文读取历史,不依赖 inline history window。 |
| LA-03 | 流式 / 非流式 | 分别用支持流式和关闭流式的路径发送文本。 | 流式 UI 不重复、不空白;非流式只输出最终消息。 |
| LA-04 | 工具调用 | 绑定测试工具,发送会触发工具的 prompt。 | `ctx.resources.tools` 只包含授权工具;工具调用 started/completed;最终回复包含工具结果。 |
| LA-05 | RAG | 绑定测试知识库,发送命中文档的 prompt。 | `ctx.resources.knowledge_bases` 包含所选知识库;runner 通过授权 API 检索;回复使用检索内容。 |
| LA-06 | 多模态 | 发送图片输入。 | `ctx.input.contents` 保留图片;支持视觉模型时正常处理,不支持时受控失败。 |
| LA-07 | fallback / 错误 | 模拟 primary 模型失败或 runner 抛错。 | fallback 或 `run.failed` 行为受控;后续请求不受影响。 |
| LA-08 | 无输出保护 | 测试 runner 完成但不产出消息。 | 不产生空白成功回复;按受控失败或明确缺陷处理。 |
| LA-09 | steering / 运行中追加消息 | 使用支持 steering 的 runner,第一条消息触发长 runrun 未结束时在同 conversation 追加第二条消息。 | 第二条消息被 active run claim,不启动并发 runrunner 通过 `steering_pull` 看到追加输入;EventLog 有 `queued` -> `steering.injected`,若未消费则有 `steering.dropped` 终态;后续普通消息仍可处理。 |
Rerank、remove-think、文件输入等场景只在本次改动直接涉及时补测,不作为每轮必跑项。
## 7. Code-agent Harness Smoke
这些测试用于验证 ACP、Claude Code、Codex 这类自管 runtime 能走同一条 Host 协议路径。若目标 harness 没有 CLI/daemon、登录态、代理配置或远端 workspace,标记 BLOCKED,不要伪造 PASS。
Smoke 前应优先保留一层轻量单测或 fixture 测试:session 创建/复用、消息发送、结果解析、`run_id` 注入和 LangBot MCP gateway 必须有稳定测试覆盖。WebUI smoke 证明真实链路可用,但不能替代转换层和错误映射测试。
### 7.1 外部 harness runner
步骤:
1. 确认目标 harness(例如 ACP daemon、Claude Code 或 Codex)在对应机器上可执行且已登录。
2. 绑定目标 runner,例如 `plugin:langbot-team/ACPAgentRunner/default``plugin:langbot-team/ClaudeCodeAgent/default``plugin:langbot-team/CodexAgent/default`
3. 配置 runner 必要字段,例如 remote target、workspace、provider、startup timeout、reuse session 等。
4. 在 Debug Chat 执行一次确定性真实 smoke。
5. 检查 LangBot MCP gateway、`run_id` 回填和 host-owned state。
通过条件:
- WebUI 可见回复包含预期 sentinel。
- 发送给 harness 的消息包含当前 LangBot `run_id` 和可访问资源摘要。
- Harness 通过 gateway 调用 `langbot_history_page``langbot_retrieve_knowledge``langbot_call_tool` 时必须携带正确 `run_id`;错误 run id 被拒绝。
- `external.session_id` 写入 host-owned state。
- 外部 harness 错误、timeout、empty output 都转成受控 `run.failed`
- resume 到同一 external session 时,全局锁边界符合 PROTOCOL_V1 §13。
### 7.2 API 型外部 runner
Dify、n8n、Coze、DashScope、Langflow、Tbox 等外部服务 runner 不作为每轮必跑项。只有在本次改动触及对应 runner 或凭据已经可用时执行 smoke。
通过条件:
- runner 可选,配置可保存。
- 请求成功,或外部服务错误被清晰返回。
- 外部服务凭据缺失时标记 BLOCKED,并记录缺失项。
## 8. 权限与隔离补充
以下优先用单测 / targeted fixture 覆盖,不要求每次通过 UI 人工构造恶意 runner。
| 场景 | 推荐证据 |
| --- | --- |
| 未授权模型调用被拒绝 | `plugin/handler.py` run action 权限测试或目标单测。 |
| 未授权工具调用被拒绝 | `ctx.resources.tools` 与 host action 拒绝日志。 |
| 未授权知识库检索被拒绝 | `ctx.resources.knowledge_bases` 与 host action 拒绝日志。 |
| run_id 结束后复用被拒绝 | session registry 注销测试。 |
| 插件身份不匹配被拒绝 | `caller_plugin_identity` mismatch 测试。 |
| 绑定插件身份的 run_id 省略 caller identity 被拒绝 | `_validate_run_authorization(..., caller_plugin_identity=None)` 返回错误。 |
| 未注册 Runtime 连接伪造插件身份被剥离 | SDK runtime forwarding 测试:请求自带 `caller_plugin_identity` 时,未注册连接转发前必须 `pop`,已注册连接必须覆盖为真实插件身份。 |
| storage/state scope 越权被拒绝 | state/storage proxy 单测。 |
| steering claim 异常不杀 consumer loop | controller 单测:无效 runner / registry 异常只让当前消息回到普通 session 槽位路径,消息消费循环继续。 |
| steering queue 未消费有终态 | session registry / orchestrator 单测:队列有上限;run unregister 时未 pull 项写 `steering.dropped` 审计。 |
如果这些单测失败,不能用 WebUI 正常回复替代。
## 9. 证据要求
每轮测试报告至少记录:
- LangBot commit、SDK commit、相关 runner 插件 commit。
- Pipeline UUID/name、runner id、关键 runner config 摘要。
- WebUI 截图或 Playwright 操作记录。
- 后端日志中对应 query id / run id 的关键行。
- `langbot-skills` case/report 路径。
- 外部 harness runner 的 context 文件、session id、working directory、CLI 错误摘要。
- FAIL/BLOCKED 的复现步骤和归属仓库建议。
报告结论必须回答:
- 是否建议继续进入下一阶段测试。
- 是否存在主聊天路径阻塞。
- 是否只是凭据 / 外部服务 / 本机 CLI 缺失导致 BLOCKED。
- 是否需要进入 [SECURITY_HARDENING.md](./SECURITY_HARDENING.md) 的发布级验收。
## 10. 历史高价值记录
历史高价值记录与当前 runner 验收状态见 [STATUS.md](./STATUS.md)。本指南只保留可重复执行的测试步骤和证据要求。
@@ -1,101 +0,0 @@
# Event Based Agent 接入设计
> 本文记录 EBA 如何接入当前 AgentRunner Protocol v1 / Host 底座。EventGateway、EventRouter、Event subscription/notification 由外部 EBA 分支实现并联调;本分支只保留 event-first 入口和 envelope/binding models。
>
> 数据结构唯一定义在 [PROTOCOL_V1.md](./PROTOCOL_V1.md)runner 可见)与 [HOST_SDK_INFRASTRUCTURE.md](./HOST_SDK_INFRASTRUCTURE.md)Host 内部模型);本文只讲 EBA 语义,不重抄 schema。
> 与当前 runner 外化分支、后续 Agent Platform / Runtime Control Plane 的边界见 [EXTENSION_SCOPE_MATRIX.md](./EXTENSION_SCOPE_MATRIX.md)。
本文描述 EBA 接入时,事件如何进入 LangBot、如何在平级的 Pipeline / Agent 处理器之间路由,以及 Agent 分支如何复用插件化 AgentRunner 基础设施。本分支不实现完整 EventBus / EventRouter / Platform API;这些能力正在外部 EBA 分支联调。这里的目标是把处理器路由与 runner 协议边界说清楚。
## 1. 设计目标
- 消息、撤回、入群、好友申请、定时任务、API 调用都能抽象为 host event。
- EventRouter 可以根据 event type、bot、workspace、conversation、actor、subject 选择一个 Pipeline 或 Agent 处理器。
- Pipeline 目标执行完整消息 Stage 链;Agent 目标通过统一 orchestrator 调用 AgentRunner。
- 非消息事件不伪造成用户文本消息。
- 平台动作执行通过显式 capability / permission / result type 预留,不混入普通文本回复。
## 2. 事件不是消息
`message.received` 只是事件的一种。协议不应假设:一定有用户文本、一定有 conversation history、一定要返回一条聊天消息、actor 一定等于 sender、subject 一定等于当前消息。
| event_type | actor | subject | input |
| --- | --- | --- | --- |
| `message.received` | 发消息的人 | 当前消息 | 文本、图片、文件等 |
| `message.recalled` | 撤回操作者,未知时为系统 | 被撤回消息 | 通常为空 |
| `group.member_joined` | 新成员或邀请人 | 群/成员关系 | 通常为空 |
| `friend.request_received` | 申请人 | 好友申请 | 验证消息或申请理由 |
| `schedule.triggered` | 系统 | 定时任务 | 任务 payload |
| `api.invoked` | API caller | API request | request payload |
## 3. 稳定事件名
先保留的稳定事件名(作为插件协议的一部分保持稳定):
- `message.received`
- `message.recalled`
- `group.member_joined`
- `friend.request_received`
平台原始事件名只能进入 `ctx.event.source_event_type` / `raw_ref`,不能成为 `ctx.event.event_type` 的公共契约。
## 4. Event Envelope 与 Binding
- 入口事件用 `AgentEventEnvelope`HOST_SDK §4.1)承载;顶层字段使用 LangBot 稳定协议名,平台原始事件名和原始 payload 放 `metadata` / `raw_ref`
- EBA 持久路由通过 `event_pattern``filters``target_type``target_uuid` 选择处理器。只有 `target_type=agent`,或 Pipeline AI Stage 需要调用 runner 时,才进一步解析 `AgentBinding`HOST_SDK §4.2)。
EBA 每个事件只选择一个有效处理器;AgentRunner 调用的基数、Agent 复用和 fan-out 边界以 PROTOCOL_V1 §13 为准。
路由 scope 示例:workspace 全局、bot 级、platform channel 级、conversation / group / thread 级、user / actor 级。Pipeline 是 `message.*` 场景的一等处理器,适合需要预处理、AI、后处理、扩展和输出控制的消息链路;Agent 是 runner 驱动的一等处理器,可处理其声明支持的消息与非消息事件。二者都不会被转换成对方。
Event Source 可包括:`platform_adapter`(飞书、QQ、微信、Telegram 等)、`webui``http_api``scheduler``system`。EventRouter 不应写死平台 adapter 的类名。
## 5. EventRouter 调用链
```text
Platform Adapter / WebUI / API
-> Event Gateway normalize payload
-> EventLog append raw event
-> EventRouter resolve one Processor target
-> target_type=pipeline: MessageAggregator -> QueryPool -> Pipeline stages
-> target_type=agent: resolve AgentBinding -> AgentRunOrchestrator
-> AgentRunContextBuilder -> PluginRuntimeConnector.run_agent()
-> AgentRunResult stream
-> DeliveryController render / platform action
```
约束:Pipeline 和 Agent 是 EventRouter 的平级目标;Pipeline 仅接受消息事件,Agent 受其事件能力声明约束。任何 AgentRunner 调用都必须复用现有 orchestrator,不能为 EBA 单独实现另一套 plugin runner 协议;非消息事件不能绕过 resource authorizationdelivery 和 platform action 走统一权限模型;外部 harness runner 也通过同一套 envelope/binding/context/result 协议接入。observer / fan-out / parallel arbitration 的额外语义仍按 PROTOCOL_V1 §13 处理。
## 6. 平台动作执行
EBA 后 `action.requested`PROTOCOL_V1 §7.3,当前仅 telemetry 不执行)将用于请求 host 执行平台动作:
```json
{ "type": "action.requested",
"data": { "action": "friend.request.accept",
"target": {"platform": "wechat", "request_id": "..."},
"payload": {"reason": "policy matched"} } }
```
Host 必须校验:binding / platform action policy 是否授权该 action、actor / bot / workspace 是否允许、是否需要人工审批,以及当前 run session / caller identity 是否匹配。EBA 还可能预留 `delivery.requested`(请求投递到某 surface)。
Delivery 方面,event 不一定回复到当前聊天窗口:消息事件通常带 reply target;系统事件可能没有默认 reply target,需要 runner 返回 `action.requested` 或由 binding 的 delivery policy 决定投递位置(`DeliveryContext` 见 PROTOCOL_V1 §5.7)。
当前 Host 会把 adapter 声明的通用 API 投影到
`DeliveryContext.platform_capabilities.supported_apis`,并据此设置
`supports_edit` / `supports_reaction`。该投影只供 runner 选择输出形态,不构成
平台动作授权;合成测试 adapter 会移除副作用能力并抑制实际出站调用。
## 7. 与 Context 协议的关系
EBA 事件进入 AgentRunner 时仍遵循 [AGENT_CONTEXT_PROTOCOL.md](./AGENT_CONTEXT_PROTOCOL.md)inline 当前事件、大 payload 用 raw/staged file ref、不默认 inline 完整 history、agent 按需通过 API 拉取、Host 保留 EventLog 和权限 guardrail。非消息事件可以被投影进 Transcript,但不能强制伪装为 user messageAgentRunner 根据 event type 自己决定是否纳入模型上下文。
## 8. 当前集成状态
当前分支已完成 EventRouter、Pipeline / Agent 平级处理器路由、Bot
`event_bindings` 持久化与 WebUI、AgentBinding 投影、路由 dry-run、合成测试事件、
运行状态和真实 OneBot 非消息事件到 Agent 的闭环。Pipeline 消息链和独立 Agent
均复用同一个 AgentRunner orchestrator / context / result 协议。
尚未落地的是 platform action permission model 和 `action.requested` 执行器;在显式
action allowlist、binding policy、adapter capability 和审批模型完成前,该 result 仍只
记录 telemetry,不执行平台副作用。
@@ -1,51 +0,0 @@
# AgentRunner 外化扩展边界矩阵
本文用于回答一个问题:本分支只做 AgentRunner 外化时,哪些能力已经作为扩展底座完成,哪些由外部 EBA / Agent Platform / Runtime Control Plane 分支接入,后续分支接入时应该走哪个扩展点。
结论:本分支不实现完整 Agent Platform,也不实现完整 EBA。EBA 完整事件网关与事件路由由外部 EBA 分支联调。本分支必须把 runner 外化的 Host / SDK 边界做干净,让外部分支只需要接入持久模型、事件路由或 runtime task,而不需要重写 `AgentRunner Protocol v1`
调度基数、Agent 复用、插件实例无状态、Pipeline adapter 和 fan-out 边界的单一事实源是 [PROTOCOL_V1.md](./PROTOCOL_V1.md) §13;本矩阵只说明后续能力应该接入哪个扩展点。
## 1. 分支边界
| 范围 | 本分支职责 | 不在本分支做 |
| --- | --- | --- |
| AgentRunner Protocol v1 | 定义 Host 调用 runner 的稳定合同:discovery、`AgentRunContext`、result stream、Host pull API、错误和权限边界。 | 不定义 Agent Platform 的产品数据库模型;不定义 runtime task queue。 |
| Host runner 外化底座 | 提供 `AgentEventEnvelope``AgentBinding` 运行投影、`run(event, binding)`、resource authorization、run-scoped session、EventLog / Transcript / State / sandbox 文件边界。 | 不实现 EventGateway、scheduler、integration provider、Agent 管控面 UI。 |
| Pipeline 的 AgentRunner 接入 | Pipeline 作为一等消息处理器执行完整 Stage 链;仅在 AI Stage 调用 runner 时,`QueryEntryAdapter` 把当前 Query/config 投影成 event + binding。 | 不把整个 Pipeline 当成临时 Agent;不复制 Pipeline 配置来自动创建 Agent。 |
| 官方 runner 插件 | 作为协议消费者验证 local-agent / 外部 harness runner 能接入 Host 基础设施。 | 不让官方 runner 的内部实现反向决定 Host / SDK 协议形态。 |
## 2. 扩展矩阵
| 能力 | 当前分支状态 | 后续归属 | 后续接入方式 | 禁止事项 |
| --- | --- | --- | --- | --- |
| Product `Agent` | 已有 `agents` 产品表 / API 和运行期 `AgentConfig` / `AgentBinding` 投影;完整 binding persistence / EventRouter / UI 闭环仍未完成。 | Agent Platform / binding persistence UI。 | 持久 Agent 保存 runner id、runner config、resource/state/delivery policy;运行前投影为 `AgentBinding`。 | 不把持久 Agent schema 加进 SDK 协议;插件实例边界见 PROTOCOL_V1 §13。 |
| Agent 处理器调用 runner | 已有单次运行前的 `AgentBinding` 解析投影;AgentRunner 调度语义见 PROTOCOL_V1 §13。 | EBA / Agent Platform。 | EventRouter 先选中 Agent 处理器,再根据 bot、channel、workspace、conversation、event type 解析有效 `AgentBinding`。Pipeline 目标走独立 Stage 链。 | 不用 `AgentBinding` 取代 EBA 的 Pipeline / Agent 处理器选择;不在本矩阵重定义 fan-out / observer 语义。 |
| Agent session / run | 已有持久 `AgentRun` / `AgentRunEvent` ledger 和 active `AgentRunSessionRegistry`;还没有独立 `AgentSession` / task 产品模型。 | Agent Platform / Runtime Control Plane。 | 如需要可新增 `AgentSession` / task 表,但执行仍回到 `run(event, binding)` 或 runtime-managed 等价入口。 | 不把持久 session 字段塞进 `AgentRunContext` 顶层;不要求所有 runner 长期持有 LangBot session。 |
| EventLog / Transcript / Sandbox files | 已完成 Host-owned store、history pull API 和 sandbox 文件边界;runner 不直接写 DB。 | 本分支持续维护底座;Agent Platform 可复用。 | 外部 EBA、scheduler、integration、runtime task 都写同一套 EventLog / Transcript;当前 run 文件通过 sandbox/workspace staging 共享。 | 不让 runner / sandbox 直接访问 Host DB;不把大 payload 内联进 prompt。 |
| Host-owned state / storage | 已有 state snapshot、`state.updated` 处理和 State APIstorage 作为授权能力保留。 | 本分支持续维护底座;Runtime / Platform 可复用。 | 外部 session id、working directory、checkpoint 等小 JSON 用 state;当前 run 大对象用 sandbox/workspace 文件。 | 不把跨轮次状态存在插件实例内;不绕过 run-scoped authorization。 |
| EventGateway / EventRouter | 本分支只提供 event-first envelope 和 `run(event, binding)` 入口。 | EBA 分支(联调中)。 | EventGateway 规范化平台/WebUI/API/scheduler 事件;EventRouter 解析一个 binding;调用现有 orchestrator。 | 不为 EBA 新增另一套 runner 调用协议;不把非消息事件伪装成 user message。 |
| Scheduler / Automation | 不实现。文档中只把 `scheduler` 作为 future event source。 | EBA / Agent Platform。 | 定时任务触发 `schedule.triggered` host event,复用 EventGateway -> EventRouter -> `run(event, binding)`。 | 不直接调用某个 runner 插件;不绕过 EventLog / authorization。 |
| Integration provider | 不实现。IM platform adapter 仍是当前平台接入系统。 | EBA / Agent Platform。 | OAuth/webhook/outbound provider 应先转成 canonical host event 或 platform action,再交给 AgentRunner。 | 不把 Linear/Slack/GitHub 等 provider 私有 payload 扩散到 runner 协议顶层。 |
| Platform action / delivery | `action.requested` 已预留但当前仅 telemetry,不执行。`DeliveryContext` 只作为上下文/策略投影。 | EBA / platform action executor。 | 后续 executor 校验 runner capability、binding policy、actor/bot/workspace 权限和审批后执行。 | 不让 runner 直接调用平台 adapter 私有 API;不把平台动作伪装成文本回复副作用。 |
| Runtime registry / worker / task queue | 已落地 Host-owned `AgentRun` / `AgentRunEvent`、run control primitives、最小 runtime registry / heartbeat / claim lease;当前官方外部 harness 仍通过 ACP、远端 daemon、本机 subprocess 或外部 HTTP API runner 调用目标运行环境,不在本分支维护完整通用 worker 队列。 | Runtime Control Plane v2。 | 后续可在现有 Host 事实源上补 queued run producer、daemon wakeup、claim execution loop、progress/audit 和运维诊断。 | 不把 heartbeat/task/warm pool 放进 Protocol v1;不让管理插件拥有 runtime/task 事实源。 |
| Warm pool / reconcile / diagnose | 不实现。 | Runtime Control Plane v2 / deployment layer。 | 作为 task/runtime 的运维能力,围绕 Host-owned runtime/task/audit 表实现。 | 不把 runtime 运维语义写进普通 runner 协议;不把 pod/task 细节泄漏给普通 runner。 |
| Agent memory | 不实现通用长期记忆产品层;提供 history/state/storage 和 sandbox 文件基础能力。 | Agent Platform 或具体 runner/plugin。 | 平台 memory 可通过 Host storage/state 或独立产品表实现,runner 通过授权 API 拉取。 | 不在 Host core 内置通用 agentic memory 策略;不默认把 memory 全量 inline 到 context。 |
| External harness native session | ACP / Claude Code / Codex 等 runner 支持 external session id state handoff 和 LangBot resource projection。 | 官方 runner 后续增强;Runtime Control Plane v2 可接管执行。 | 外部 harness 调用继续走 `runner.run(ctx)`;如后续引入长连接/daemon 模式,按 external session key 串行 turnreader 独占 native stream。 | 不把具体 provider native wire 变成 LangBot 协议;全局锁边界见 PROTOCOL_V1 §13。 |
## 3. 后续分支接入规则
外部 EBA、Agent Platform 或 Runtime Control Plane 分支接入时,默认遵守以下规则:
- 新入口只生产或解析 Host 内部模型:`AgentEventEnvelope`、持久 Agent 投影出的 `AgentBinding`、以及必要的 delivery/resource/state policy。
- runner 调用仍走 `AgentRunOrchestrator.run(event, binding)`,除非 Runtime Control Plane 明确引入 runtime-managed 执行模式;即便如此,runner 可见合同仍应保持 Protocol v1。
- Host-owned facts 继续写入 EventLog / Transcript / State,当前 run 文件继续走 sandbox/workspace;产品层可以新增更高阶视图,但不能替代这些事实源。
- 新能力如果需要持久化,优先加 Host-owned 表或 service;不要把事实源藏在插件 storage 或 runner subprocess 内。
- 新 result type 可以按 Protocol v1 的演进规则增加;不能用入口 adapter 私有字段绕过 schema。
- 任何 fan-out、observer agent、parallel arbitration、platform action execution 都必须单独定义 delivery、state conflict、approval 和 audit 语义。
## 4. 与 Agent Platform 产品层的关系
这里的 Agent Platform 指面向 agent 产品层的实体拆分:`Agent` 描述可配置 agent`Session` / `SessionMessage` 描述会话事实,`Automation` 描述自动触发,`IntegrationBinding` 描述外部集成连接,`Memory` 描述长期记忆,`WarmTask` 描述预热/后台任务。这些拆分对 LangBot 后续产品层有参考价值,但不能直接搬进本分支。
LangBot 当前分支的对应目标是更底层的:把 IM/WebUI/API 等入口统一投影到 Host event,把 Agent / binding 配置统一投影到 runner binding,把 runner 能力统一收束到 Protocol v1。完整 Agent Platform 可以在这个底座之上构建,而不应反过来污染本分支的 runner 外化边界。
@@ -1,266 +0,0 @@
# LangBot Host 与 SDK 基础设施设计
本文档描述 LangBot 作为 agent host 的内部能力与分层架构,以及 Host 内部模型。
- SDK ↔ Host 的协议数据结构(`AgentRunContext``AgentRunnerManifest``AgentRunResult``AgentRunAPIProxy` 等)的**唯一定义在** [PROTOCOL_V1.md](./PROTOCOL_V1.md);本文只引用,不重抄。
- 测试执行入口和 smoke 记录见 [AGENT_RUNNER_QA_GUIDE.md](./AGENT_RUNNER_QA_GUIDE.md);安全发布门槛见 [SECURITY_HARDENING.md](./SECURITY_HARDENING.md)。
- 本文定义的 Host 内部模型(`AgentEventEnvelope``AgentBinding``AgentRunnerDescriptor`)不属于 SDK 协议字段。
## 1. 目标
LangBot 要转为 agent host,而不是内置 runner 容器:
- 接收 IM、WebUI、API 和外部 EBA 分支 EventRouter 产生的事件。
- 接收 EBA 选中的 Agent 处理器,并根据事件、bot、workspace、scope 解析 AgentRunner binding。
- 发现、校验和调用插件提供的 AgentRunner。
- 为每次 run 提供受限资源、状态、存储、上下文引用和生命周期控制。
- 接收 AgentRunner 返回的事件流,投递到 IM、WebUI 或其他 output surface。
## 2. 非目标
- 不定义 Pipeline 的 Stage 编排语义;Pipeline 是 EBA 的同级处理器,其 AI Stage 只在需要 runner 时接入本 Host 边界。
- 不要求所有 AgentRunner 依赖 LangBot 的上下文管理。
- 不要求官方 local-agent 的旧行为反向塑造 host 协议。
- 不在 host 中实现通用 agentic prompt assembler。
- 不强制 runner 使用 LangBot state / storage;只提供可选、受控的寄宿能力。
- 不实现 EventGateway / EventRouter:它们由外部 EBA 分支提供并联调。本分支只定义 host-side envelope/binding models 和 `run(event, binding)` 入口。
## 3. 分层架构
```text
IM / WebUI / API / EventRouter (external EBA branch)
|
v
Event Gateway (external EBA branch)
|
v
EventRouter -> one Processor target
|-- target_type=pipeline -> Pipeline Stage chain
|
`-- target_type=agent -> AgentBindingResolver
|
v
AgentRunOrchestrator
|-- AgentRunnerRegistry
|-- AgentResourceBuilder
|-- AgentContextBuilder
|-- AgentRunSessionRegistry
|-- PersistentStateStore / EventLogStore / TranscriptStore
|-- Sandbox / workspace file tools
v
Plugin Runtime / AgentRunner
|
v
AgentRunResult stream
|
v
Delivery / Renderer / Platform API
```
Pipeline 与 Agent 是 EventRouter 的平级处理器目标。本文只定义 AgentRunner Host 边界:Agent 目标直接解析 `AgentBinding`;Pipeline 目标执行自己的完整 Stage 链,仅在 AI Stage 调用 runner 时通过 Query entry adapter 构造一次性 `AgentConfig` / `AgentBinding`。该 runner 调用投影不改变 Pipeline 的一等处理器地位,也不会把 Pipeline 持久化为 Agent。AgentRunner 的单绑定调度、Agent 复用、插件实例无状态和 fan-out 边界以 [PROTOCOL_V1.md](./PROTOCOL_V1.md) §13 为准。EventGateway / EventRouter 由外部 EBA 分支实现并联调。
## 4. LangBot 侧能力
### 4.1 Event Gateway / EventRouterExternal EBA Branch Integration Point
> EventGateway / EventRouter 由外部 EBA 分支实现并联调,不在本分支范围。本分支只保留 event-first 入口和 envelope/binding models。
Event Gateway 将把入口统一成 host eventIM 平台消息、WebUI debug chat、API 触发、后续非消息事件),输出稳定的 `AgentEventEnvelope`Host 内部模型):
```python
class AgentEventEnvelope(BaseModel):
event_id: str
event_type: str
event_time: int | None
source: str
bot_id: str | None
workspace_id: str | None
conversation_id: str | None
thread_id: str | None
actor: ActorRef | None
subject: SubjectRef | None
input: AgentInput # 见 PROTOCOL_V1 §5.6
delivery: DeliveryContext # 见 PROTOCOL_V1 §5.7
raw_ref: RawEventRef | None
metadata: dict[str, Any] = {}
```
`AgentEventEnvelope` 是 Host 内部入口模型;投影给 runner 的是 `ctx.event`PROTOCOL_V1 §5.4)。原始平台 payload 存为 raw event 或 staged file reference,不扩散到 runner 协议顶层。
**当前 adapter source**`QueryEntryAdapter.query_to_event(query)` 从 Query 生成 `AgentEventEnvelope`
### 4.2 AgentConfig 与 AgentBinding
`AgentConfig` 是 Host 内部的一次 AgentRunner 调用配置投影(不暴露给 SDK)。独立 Agent 从自己的持久配置生成它;Pipeline 只在 AI Stage 调用 runner 时,由 Query entry adapter 从该 Stage 的当前配置生成它。两种来源随后都由 BindingResolver 结合事件和 scope 解析为 `AgentBinding`。Pipeline 本身不是 `AgentConfig`,该调用投影也不会创建或更新持久 Agent。
```python
class AgentConfig(BaseModel):
agent_id: str | None = None
runner_id: str
runner_config: dict[str, Any] = {}
resource_policy: ResourcePolicy = ResourcePolicy()
state_policy: StatePolicy = StatePolicy()
delivery_policy: DeliveryPolicy = DeliveryPolicy()
event_types: list[str] = ["message.received"]
enabled: bool = True
metadata: dict[str, Any] = {}
```
`AgentBinding` 是"什么事件调用哪个 AgentRunner、带什么 Agent 配置"的 Host 内部运行投影(不暴露给 SDK)。它是 EventRouter / 当前 QueryEntryAdapter 在一次运行前解析出的有效绑定。
```python
class AgentBinding(BaseModel):
binding_id: str
enabled: bool
scope: BindingScope
event_types: list[str]
filters: list[EventFilter] = [] # EBA 阶段使用,见 EVENT_BASED_AGENT
runner_id: str
runner_config: dict[str, Any]
resource_policy: ResourcePolicy
state_policy: StatePolicy
delivery_policy: DeliveryPolicy
```
BindingResolver 的基数、fan-out 和冲突处理约束见 PROTOCOL_V1 §13;本节只定义 Host 内部投影形态。
**当前 adapter source**`QueryEntryAdapter.config_to_agent_config(query, runner_id)`
只在 Pipeline AI Stage 调用 runner 时,把 current config 临时投影为运行期 `AgentConfig`,再由
`AgentBindingResolver.resolve_one(event, [agent_config])` 解析出唯一
`AgentBinding`。该调用从 Pipeline 取得 AI runner config
→ runner_config、extension preference → resource_policy、output settings →
delivery_policy,但 Pipeline 仍执行并拥有完整 Stage/config 语义。该适配不会把 Pipeline 持久化为 Agent;独立 Agent 由用户自行新增和绑定。
### 4.3 AgentRunnerRegistry
Registry 收集 runner descriptor(来自插件 runtime、开发期本地插件):
```python
class AgentRunnerDescriptor(BaseModel):
id: str
source: Literal["plugin"]
label: I18nObject
description: I18nObject | None = None
plugin_author: str
plugin_name: str
runner_name: str
capabilities: AgentRunnerCapabilities # 见 PROTOCOL_V1 §4.3
permissions: AgentRunnerPermissions # 见 PROTOCOL_V1 §4.4
config_schema: list[DynamicFormItemSchema]
plugin_version: str | None = None
raw_manifest: dict[str, Any] = {}
```
职责:调用 `plugin_connector.list_agent_runners()` 拉取 runner、校验 typed `AgentRunnerManifest`、输出 descriptor、缓存 discovery 结果并提供 `refresh()`。单个插件 manifest 失败只记 warning,不影响其它 runner。`plugin:author/name/runner` 是稳定 id 格式;插件实例边界见 PROTOCOL_V1 §13。
Host 内置 runner / adapter 不能作为 `AgentRunnerDescriptor.source` 绕过插件
runtime、`run_id``ctx.resources``AgentRunAPIProxy` 权限链。若需要
开发期调试 adapter,应放在 Host 内部测试入口,不进入可选 runner 列表。
刷新触发点:插件安装/卸载/升级/重启后;Pipeline metadata 请求时发现缓存为空;可选 TTL(优先保证正确性)。
### 4.4 AgentRunOrchestrator
Orchestrator 是唯一运行入口:
```text
run(event, binding)
-> resolve runner descriptor
-> build resources
-> build context
-> register run session
-> call plugin runtime
-> normalize result stream
-> update state
-> unregister run session
```
它负责:`run_id` 生成和生命周期、timeout/deadline/cancellation、插件异常隔离、result schema 校验和大小限制、`state.updated` 处理、delivery backpressure 和 telemetry。
典型 run 时序:
```text
QueryEntryAdapter / EventRouter
-> AgentRunOrchestrator.run(event, binding)
-> AgentRunnerRegistry.resolve(runner_id)
-> AgentResourceBuilder.freeze_snapshot(binding, event)
-> AgentRunSessionRegistry.register(run_id, runner_id, snapshot)
-> AgentContextBuilder.build(event, binding, snapshot)
-> PluginRuntimeConnector.run_agent(ctx)
-> AgentRunAPIProxy action
-> validate active run session + caller identity + snapshot
-> Host API / Store
<- AgentRunResult stream
-> apply state.updated to PersistentStateStore
-> write message.completed to Transcript
-> keep current-run files and large tool outputs in sandbox/workspace
-> render delivery or raise RunnerExecutionError
-> AgentRunSessionRegistry.unregister(run_id)
```
`run_from_query()` 保留为 Query entry adapter 入口,但内部转换成 event + binding 后走统一 `run()`。约束:`ChatMessageHandler` 不解析 `plugin:*`、不实例化 wrapper、不知道 runner 组件细节;`PipelineService` 从 registry 读取 metadata,不直接访问插件 runtime;跨请求持久化状态必须走授权 storage / 外部服务。
### 4.5 Resource Authorization
LangBot 在每次 run 前生成 `ctx.resources`PROTOCOL_V1 §6),来自 manifest permissions 与 binding policy 的交集:
1. `descriptor.permissions` 声明 runner 需要的 LangBot 资源访问上限。
2. binding / resource policy 允许的资源范围。
3. Agent/runner config 中选择的模型、知识库、文件等资源。
4. 当前 event / actor / bot / workspace 的实际权限。
5. `ctx.context.available_apis` 暴露的 pull API 能力。
这次裁剪结果必须冻结为 run-scoped authorization snapshot,并由
`AgentRunSessionRegistry``run_id` 保存。`ctx.resources` 是投影给 runner
看的同一份授权结果;运行期每个 proxy action 只依据该 snapshot 校验 active
run session、caller plugin identity、resource id、scope、payload size、rate
limit 和 deadline。Handler 不应重新执行授权裁剪,否则 build-time 与 runtime
授权逻辑会漂移。
SDK 侧本地校验只用于开发体验,host 侧 run authorization snapshot 才是安全边界。`spec.capabilities` 只帮助 Host 判断 runner 是否需要 tool / knowledge 等资源投影,不能替代 permissions 或 binding policy。skill 不由独立 capability 决定是否投影——它通过统一 tool 授权(`resource_policy.allowed_tool_names`)消费,`skill_authoring` 仅作为「一键授权这组 skill tool + sandbox」的便捷开关。
资源裁剪应通用,不写死 local-agent。selector 与资源的映射示例:`model-fallback-selector` → primary/fallback LLM、`llm-model-selector` → LLM、`rerank-model-selector` → rerank 模型、`knowledge-base-multi-selector` → 知识库;新增 selector 时在 resource builder 中统一扩展。
构造 `ctx.resources.tools` 时,Host 一次塞齐每个工具的完整 schema(`ToolResource.parameters`),runner 不需再逐个 `get_tool_detail` 拉取,减少 N 次往返。
执行/文件/skill/MCP 等能力的接入方向:先由 Host / sandbox 封装成普通 scoped tool,再通过 `ctx.resources.tools` 和 SDK runtime 转发进入 runnerrunner 不应识别或硬编码执行环境 provider。外部 harness 的 native tools 不能直接访问 LangBot 资源。skill 的整个生命周期都走统一 tool:发现走 `list_skills` / `langbot_list_assets`,激活/注册走 `activate` / `register_skill`,包内操作走 native exec/read/write——runner 不需要独立的 skill 渲染或门控。
### 4.6 State / Storage
LangBot 可提供 host-owned state 让 runner 寄宿状态(conversation / actor / subject / runner / binding / workspace state),但**不是强制**。Host 只需提供:授权开关、scope key、get/set/list/delete API(见 PROTOCOL_V1 §8)、持久化 backend、审计和清理策略。外部 agent runtime 可维护自己的 session 和 memory。进程内 state store 只能作为过渡实现,不能作为正式生产语义。
部分 host-owned state 由 Host 自身直接写:例如 `activate` tool 在 Host 侧执行时,把已激活 skill 写入 conversation scope 的 `host.activated_skills`。host 直接写与 runner `state.updated` 写到同一 key 时按 **last-write-wins** 合并,runner 可覆盖。
### 4.7 EventLog / Transcript / Sandbox Files(事实源)
- `EventLog`: durable append-only,保存原始事件、系统事件、工具调用、投递结果、错误。
- `Transcript`: 从 EventLog 投影出的对话视图,用于 UI、审计和按需历史读取。
- `Sandbox / workspace files`: 当前 run 的上传文件、平台附件、工具大结果和临时产物。Host 负责 staging 与授权边界,runner 通过 read/write/exec 类工具按需访问。
三类数据与 working context 的边界、读取约束见 [AGENT_CONTEXT_PROTOCOL.md](./AGENT_CONTEXT_PROTOCOL.md)。AgentRunner 可读取这些能力,但不被迫使用 LangBot 作为唯一记忆系统。
### 4.8 External harness resource projection
Claude Code、Codex、Kimi Code 等外部 harness runner 可能不直接调用 LangBot 的 model/tool loop,而是把 LangBot 事件和授权资源句柄投影到自己的 harness 执行。Host 侧仍保持统一边界:Host 负责构造 event-first context、资源授权、state/storage、EventLog/Transcript、sandbox/workspace 文件边界和审计;Host 或 binding policy 决定哪些 MCP bridge、skill-backed tool、sandbox path、history/state 句柄可投影给 runnerrunner plugin 把 scoped projection 转成目标 harness 可消费形式;所有 LangBot 资源访问必须经 SDK runtime / `AgentRunAPIProxy` / SDK-owned MCP bridge 转发并接受 Host 校验;外部 harness 负责自己的 native session、tool loop、压缩、权限模式和 resume,但不能用 native tools 绕过 Host 授权。
投影的具体形态(context 文件、resource handles、LangBot MCP gateway、state pointers)见 AGENT_CONTEXT_PROTOCOL §4.5;当前 code-agent harness runner 形态见 OFFICIAL_RUNNER_PLUGINS §7。发布级隔离要求见 SECURITY_HARDENING。
## 5. SDK 侧协议
SDK 组件入口如下;所有数据结构定义见 PROTOCOL_V1。
```python
class AgentRunner(BaseComponent):
__kind__ = "AgentRunner"
@classmethod
def get_config_schema(cls) -> list[dict]: ...
async def run(self, ctx: AgentRunContext) -> AsyncGenerator[AgentRunResult, None]: ...
# ctx: PROTOCOL_V1 §5.2 ; AgentRunResult: PROTOCOL_V1 §7
```
- Manifest / capabilities / effective accessPROTOCOL_V1 §4。Capabilities 来自组件 manifest 的 `spec.capabilities`,不是 SDK 基类 classmethod。
- `AgentRunContext`PROTOCOL_V1 §5.2。`messages` / `bootstrap` 不是协议字段。
- `AgentRunResult`PROTOCOL_V1 §7。
- `AgentRunAPIProxy`PROTOCOL_V1 §8,是 runner 访问 host 能力的唯一入口,所有请求带 `run_id`
@@ -1,138 +0,0 @@
# 官方 AgentRunner 插件迁移计划
本文档描述内置 `RequestRunner` 迁出 LangBot 后,官方 runner 插件如何组织、迁移和验收。它是 [HOST_SDK_INFRASTRUCTURE.md](./HOST_SDK_INFRASTRUCTURE.md) 和 [AGENT_CONTEXT_PROTOCOL.md](./AGENT_CONTEXT_PROTOCOL.md) 的下游落地计划,不是 LangBot 宿主协议的设计前提。QA 入口和 smoke 记录见 [AGENT_RUNNER_QA_GUIDE.md](./AGENT_RUNNER_QA_GUIDE.md)。
官方 `local-agent` 可以外移,也可以重写。设计重点不是保留旧内置 runner 的内部结构,而是验证一个依附 LangBot host 基础设施的官方 agent 能否完整工作。同时,LangBot host 协议必须服务 Claude Code SDK、Codex、Pi Agent SDK、外部 Agent 平台等自管 context/runtime 的 runner,不能被官方插件的实现细节绑死。
## 1. 仓库组织
官方 runner 插件与 LangBot 主仓库、SDK 仓库以不同节奏迭代:LangBot 主仓库只维护宿主协议和调度,SDK 仓库维护 AgentRunner 组件和 runtime 协议,官方 runner 插件承载业务 runner 的具体实现和第三方平台适配。
当前推荐"官方插件可独立发布,必要时共享 SDK helper"。开发期采用本地多目录布局:
```text
langbot-app/
langbot-local-agent/ # plugin:langbot-team/LocalAgent/default
manifest.yaml
components/agent_runner/default.{yaml,py}
langbot-agent-runner/ # 外部服务 runner 仓库
acp-agent-runner/ claude-code-agent/ codex-agent/ dify-agent/ n8n-agent/ ...
```
后续可聚合进 monorepo,也可继续独立发布——这个选择不影响协议设计。重复逻辑优先沉淀到 SDK 或明确的共享 helper 包,不要把宿主私有结构泄漏给插件。旧 `src/langbot/pkg/provider/runners/*` 只作为历史行为对齐基准;当前未发布分支不提供旧内置 runner 的运行时 fallback。
## 2. 插件命名和 runner id
| 旧 runner | 官方插件 | runner id |
| --- | --- | --- |
| `local-agent` | `langbot-team/LocalAgent` | `plugin:langbot-team/LocalAgent/default` |
| `dify-service-api` | `langbot-team/DifyAgent` | `plugin:langbot-team/DifyAgent/default` |
| `n8n-service-api` | `langbot-team/N8nAgent` | `plugin:langbot-team/N8nAgent/default` |
| `coze-api` | `langbot-team/CozeAgent` | `plugin:langbot-team/CozeAgent/default` |
| - | `langbot-team/ACPAgentRunner` | `plugin:langbot-team/ACPAgentRunner/default` |
| - | `langbot-team/ClaudeCodeAgent` | `plugin:langbot-team/ClaudeCodeAgent/default` |
| - | `langbot-team/CodexAgent` | `plugin:langbot-team/CodexAgent/default` |
| `dashscope-app-api` | `langbot-team/DashScopeAgent` | `plugin:langbot-team/DashScopeAgent/default` |
| `deerflow-api` | `langbot-team/DeerFlowAgent` | `plugin:langbot-team/DeerFlowAgent/default` |
| `langflow-api` | `langbot-team/LangflowAgent` | `plugin:langbot-team/LangflowAgent/default` |
| `tbox-app-api` | `langbot-team/TboxAgent` | `plugin:langbot-team/TboxAgent/default` |
| `weknora-api` | `langbot-team/WeKnoraAgent` | `plugin:langbot-team/WeKnoraAgent/default` |
每个插件可后续提供多个 runner,但迁移目标的默认 runner 统一叫 `default`
## 3. 迁移批次
- **Batch 1(打通协议)**`local-agent`(能力最完整基准)、`acp-agent-runner` / `claude-code-agent` / `codex-agent`(外部 code-agent harness 路径)、`dify-agent`(传统 service API runner)。
- **Batch 2(外部 workflow**`n8n-agent``langflow-agent`webhook/workflow 输入输出、timeout、外部 conversation id)。
- **Batch 3(平台 Agent API**`coze-agent``dashscope-agent``tbox-agent``deerflow-agent``weknora-agent`(平台特有响应格式、引用资料、文件/图片输入、外部 thread/session 状态)。
## 4. 每个官方插件的组件要求
每个插件至少包含一个 `AgentRunner` 组件,manifest 示例:
```yaml
apiVersion: langbot/v1
kind: AgentRunner
metadata:
name: default
label: { en_US: Dify Agent, zh_Hans: Dify Agent }
description:
en_US: Run a Dify application as a LangBot AgentRunner.
zh_Hans: 将 Dify 应用作为 LangBot AgentRunner 运行。
spec:
config: []
capabilities: # 字段语义见 PROTOCOL_V1 §4.3
streaming: true
execution:
python: { path: ./main.py, attr: DefaultAgentRunner }
```
## 5. local-agent 插件方向
`local-agent` 是官方插件中能力最完整的消费者,但不是宿主协议的设计中心。它需要证明:一个主要依附 LangBot host 能力的 agent runner 可以通过公开协议完成模型、工具、知识库、状态、history、sandbox 文件访问、上下文压缩和消息投递。
迁移或重写需覆盖旧内置 runner 的用户可见能力:model primary/fallback 选择、prompt、knowledge-bases、rerank-model、rerank-top-k、function calling、streaming、multimodal input、conversation history、monitoring metadata。
责任边界与 Host API 消费方式见 AGENT_CONTEXT_PROTOCOL §8。关键约束:
-`ctx.config` 读取静态绑定 `prompt`**不**读取 `ctx.adapter.extra["prompt"]`;不消费 Query entry adapter 生成的历史窗口。
- 通过 `AgentRunAPIProxy.history` 拉取 transcript,而不是依赖 host 每轮强塞历史窗口。
- `ctx.input.contents` 保留图片/文件等多模态内容;RAG 只替换/插入文本部分,不丢图片/文件。
- 不能绕过 `ctx.resources` 调用未授权模型、工具或知识库。
- manifest 声明功能能力、LangBot 资源 permissions 和配置表单;实际授权来自 manifest permissions 与 binding resource policy、runner config、`ctx.context.available_apis` 和 Host run session snapshot 的交集。
### 5.1 Native Execution / Skills 后续接入
本阶段不把 sandbox/skills 做成 AgentRunner 协议字段。后续 sandbox/skills 分支合并后,命令执行、文件操作、skill、MCP managed process 应先由 Host / sandbox 封装成 scoped tools,再通过 `ctx.resources.tools` 和 SDK runtime 转发暴露给 runner。这让 local-agent 只消费授权后的 Host 基础设施,而不是直接持有宿主机执行能力。
## 6. 外部 runner 插件要求
外部平台 runner 迁移遵循:旧配置字段尽量保持同名便于 migration 复制;输出统一转换为 `AgentRunResult`;外部 API timeout 从 runner config 读取;平台 conversation id 存 plugin storage 或 context runtime state,不依赖 LangBot 内置 conversation uuid 私有结构;流式按平台能力声明,没有流式就只发 `message.completed`
### 6.1 Code-agent harness runner
Claude Code、Codex、Kimi Code 这类 runner 不一定通过 LangBot 的模型/工具 loop 执行,可以依赖自己的 harness,但仍必须遵守统一 Host 边界。总体边界见 [HOST_SDK_INFRASTRUCTURE.md](./HOST_SDK_INFRASTRUCTURE.md) §4.8context projection 形态见 [AGENT_CONTEXT_PROTOCOL.md](./AGENT_CONTEXT_PROTOCOL.md) §4.5;发布级要求见 [SECURITY_HARDENING.md](./SECURITY_HARDENING.md)。
本文件只补充官方 runner 的实现要求:输入来自 `ctx.event` / `ctx.input`,不依赖 Pipeline 私有 `Query`;外部 session id / workspace / checkpoint 写入 Host state 或 plugin storage;插件实例边界见 PROTOCOL_V1 §13CLI / subprocess runner 必须处理 timeout、取消、空输出、非零退出和 stderr 映射。
实现结构应把 provider-native output 解析与 LangBot result stream 组装分开:Claude stream-json、Codex JSONL、Kimi / OpenCode 事件等只在 runner adapter 内解析,输出统一归一为 `AgentRunResult``message.completed` / `message.delta``state.updated``run.completed` / `run.failed`)。文件和工具大结果留在当前 run 的 sandbox/workspace,通过消息 metadata、attachment ref 或 path 指向。未知 native event 不应导致 run 崩溃;应记录诊断 metadata 或 warning。新增 harness 时优先补 native fixture -> `AgentRunResult` 的转换测试,再接 WebUI smoke。
并发约束应按外部 session 粒度表达,而不是按 Agent / runner id / 插件实例表达;Agent 复用和全局锁边界见 PROTOCOL_V1 §13。若 runner 使用 `external.session_id` / `thread_id` resume 到同一 native session,且该 harness 不支持并发 turnrunner 应按稳定 external session key 串行写入;一次性 subprocess runner 可以只在单次 `run(ctx)` 内处理,长连接/daemon runner 则应采用 reader 独占 native stream、turn writer 串行写入的结构。
### 6.2 LangBot MCP gateway
外部 harness 不能直接持有进程内的 `plugin_runtime_handler`,也不能用自己的 native tools 直接访问 LangBot 资源。外部 harness runner 应通过稳定 HTTP MCP gateway 或 SDK-owned bridge 把 harness 的工具请求转回 SDK runtime / Host API
- Gateway 由 runner 插件启动,暴露稳定的 `langbot_history_page``langbot_retrieve_knowledge``langbot_call_tool` 等最小工具面。
- Harness 每次调用必须携带当前 LangBot `run_id`Host 仍按 run session、caller identity 和授权快照校验。
- Gateway 只转发 LangBot 资产访问,不承担外部 harness 的文件、进程或 native tool 权限边界。
第一批工具保持很小:history page、knowledge retrieve、authorized tool call。新增工具必须先有 Host action 权限与 run-scoped authorization,再由 gateway 投影。
## 7. Code-agent harness runner 当前形态
外部 code-agent harness 由直接 runner 插件承接,例如 `acp-agent-runner``claude-code-agent``codex-agent`,每个 runner 负责把目标 harness 的 native session、workspace、MCP bridge 和输出事件转换为统一 `AgentRunResult`。本地 smoke 验收入口与记录见 [AGENT_RUNNER_QA_GUIDE.md](./AGENT_RUNNER_QA_GUIDE.md)。
当前形态:
- Runner ID 示例:`plugin:langbot-team/ACPAgentRunner/default``plugin:langbot-team/ClaudeCodeAgent/default``plugin:langbot-team/CodexAgent/default`
- Runner 可通过 ACP、远端 daemon、本机 subprocess 或外部 HTTP API 调用 harnessharness 的安装、登录态、workspace 和 provider-native 权限由该运行环境负责。
- Runner 会把当前 LangBot `run_id`、可访问资源摘要和 gateway 使用规则注入本次消息;harness 通过 gateway 回填 `run_id` 后访问 LangBot 资产。
- 外部 session id / workspace / checkpoint 写回 Host state 或 plugin storage,后续轮次可复用目标 harness 会话。
### 7.1 当前限制
这不是发布级安全边界实现;LangBot 只约束 LangBot 持有资产的访问,外部 harness 的文件、进程、workspace、provider-native MCP 和模型凭据由对应 runner 的运行环境承担。当前 `run_id` 可由系统提示词、ACP metadata 或 runner 自有 session metadata 传递给 harness 并由 gateway 校验。runtime 管控面方向见 [RUNTIME_CONTROL_PLANE_V2.md](./RUNTIME_CONTROL_PLANE_V2.md)。
## 8. 发布和安装策略
最终 LangBot 安装/升级时需保证官方 runner 插件可用,可选方案:首次启动检测缺失并提示安装,或由用户从 marketplace 安装。当前分支未发布,因此不保留历史 Pipeline Agent 配置兼容、旧内置 runner fallback,也不把旧 Pipeline 内的 Agent 配置迁移成独立 Agent。4.x 只读取 `ai.runner.id``ai.runner_config[id]`;升级后由用户选择或安装需要的 AgentRunner。
## 9. 验收标准
- 每个目标 runner 都有对应官方 AgentRunner 插件和稳定 runner id;当前配置只使用 `ai.runner.id` + `ai.runner_config[id]`
- LangBot 主聊天路径不再通过 `RequestRunner` 执行业务 runner。
- 官方插件测试覆盖非流式、流式、错误、timeout、配置缺失。
- `local-agent` 能完成模型 fallback、tool calling、知识库检索、多模态输入、静态绑定 prompt 消费、history API 拉取、rerank。
- 外部 code-agent harness runner 能消费 event-first context、投影 scoped resources、保存 external session state,并通过 WebUI Debug Chat smoke。
- `local-agent` 覆盖旧内置 runner 的用户可见核心能力;代码结构和运行路径不需要相同。
@@ -1,820 +0,0 @@
# LangBot AgentRunner Protocol v1
本文档是 LangBot Host 与插件 SDK / Runtime / AgentRunner 之间协议合同的**唯一规范来源(single source of truth**。
- 本文件描述当前 Protocol v1 稳定合同,不混入验收流水。当前实现状态见 [STATUS.md](./STATUS.md),测试执行入口见 [AGENT_RUNNER_QA_GUIDE.md](./AGENT_RUNNER_QA_GUIDE.md),安全发布门槛见 [SECURITY_HARDENING.md](./SECURITY_HARDENING.md)。
- 本文件之外的任何文档**不得重新定义这里的数据结构**,只能引用,例如"见 PROTOCOL_V1 §4.2"。
- Host 内部模型(`AgentEventEnvelope``AgentBinding`、Descriptor、各 Store)不属于 SDK 协议,定义在 [HOST_SDK_INFRASTRUCTURE.md](./HOST_SDK_INFRASTRUCTURE.md)。
## 1. 协议目标
Protocol v1 只解决四件事:
- LangBot 如何发现插件提供的 AgentRunner。
- LangBot 如何把一次事件调用封装成 `AgentRunContext`
- AgentRunner 如何以事件流形式返回运行结果。
- AgentRunner 如何通过受限 API 访问 LangBot host 能力。
Protocol v1 **不定义**
- LangBot 内部如何持久化 `AgentBinding`(见 HOST_SDK)。
- AgentRunner 内部如何组装 prompt、压缩历史、管理 memory(见 [AGENT_CONTEXT_PROTOCOL.md](./AGENT_CONTEXT_PROTOCOL.md))。
- 官方 runner 的具体实现(见 [OFFICIAL_RUNNER_PLUGINS.md](./OFFICIAL_RUNNER_PLUGINS.md))。
- Pipeline 的长期配置模型。
- 发布级安全 hardening 的完整实现(见 [SECURITY_HARDENING.md](./SECURITY_HARDENING.md))。
## 2. 参与方
| 名称 | 职责 |
| --- | --- |
| LangBot Host | 事件入口、绑定解析、权限、资源、存储、生命周期、结果投递。 |
| Plugin Runtime | 加载插件,响应 Host 的 runner discovery 和 run 调用。 |
| AgentRunner | 插件提供的 agent 执行组件。 |
| AgentRunAPIProxy | AgentRunner 访问 Host 能力的受限 API。 |
| AgentBinding | Host 内部的事件到 runner 绑定配置,不直接暴露给 SDK(见 HOST_SDK §4.2)。 |
产品层同时保留 Pipeline 与独立 `Agent`:现有 Pipeline 不迁移为 Agent
用户可以新建 Agent 并绑定到 bot / IM channel,一个 Agent 可以被多个 bot / channel 复用。Host 内部的
`AgentBinding` 是一次事件运行前解析出的有效绑定,只影响 Host 构造出的
`ctx.config``ctx.resources``ctx.context``ctx.delivery`。SDK 不需要知道
Agent / binding 的持久化形态。
外部 harness runnerClaude Code、Codex、Kimi Code 等)也是 `AgentRunner`:它们消费 event-first `AgentRunContext`、返回 `AgentRunResult`,并通过 Host 授权的 state/storage API 保存跨轮次指针;当前运行文件和工具大结果进入 sandbox/workspace。它们内部可以继续使用自己的 session、tool loop、MCP、上下文压缩和权限模型。
## 3. 协议演进
当前 AgentRunner 合同不暴露显式 `protocol_version` 字段。协议演进先按字段级兼容规则处理:
- 新增可选字段保持向后兼容。
- 删除字段或改变既有字段语义,需要在 SDK 发布前完成;发布后应走新的显式兼容方案。
- 结果流演进:Host **必须忽略未知 result type 并记录 warning**(除非该 type 明确要求强校验)。SDK envelope 接收入站未知 `type` 字符串,runner 侧可按原字符串转发或忽略;新增 result type 不提升大版本。
- SDK 入站 context 类实体偏宽松,用于兼容 Host 附加的非核心字段;Host 返回 DTOhistory/event/steering/run/runtime/stats/error 等)忽略未知字段,保证 Host 增加可选返回字段时旧 SDK 不会解析失败;manifest、runner result payload 等由插件/runner 提交给 Host 的输入合同仍偏严格。安全边界仍在 Host,SDK 校验只提升开发体验。
## 4. Discovery 协议
### 4.1 LIST_AGENT_RUNNERS
Host 调用 Plugin Runtime 获取当前插件暴露的 runner 列表,请求无额外 payload。返回:
```python
class ListAgentRunnersResponse(BaseModel):
runners: list[AgentRunnerDiscovery]
class AgentRunnerDiscovery(BaseModel):
plugin_author: str
plugin_name: str
runner_name: str
manifest: AgentRunnerManifest
```
`manifest` 是 SDK typed `AgentRunnerManifest`,由 Runtime 从插件组件 manifest 解析并校验后返回。`plugin_author` / `plugin_name` / `runner_name` 保留为 transport 寻址字段;Host 以它们生成稳定 runner id,并把 `manifest.id` 校验为 `plugin:author/name/runner`。单个 runner manifest 解析失败时 Runtime/Host 记录 warning 并跳过该 runner,不影响同一插件或其它插件的 runner discovery。
### 4.2 AgentRunnerManifest
这里的 manifest 指 Runtime 返回给 Host 的 typed runner manifest
```python
class AgentRunnerManifest(BaseModel):
id: str
name: str
label: I18nObject
description: I18nObject | None = None
capabilities: AgentRunnerCapabilities = AgentRunnerCapabilities()
permissions: AgentRunnerPermissions = AgentRunnerPermissions()
config_schema: list[DynamicFormItemSchema] = []
metadata: dict[str, Any] = {}
```
- runner id 由 Host 生成,格式 `plugin:author/name/runner`
- `name` 是插件内 runner 名称,例如 `default`
- `config_schema` 只描述绑定配置表单,不代表插件实例状态。
- `capabilities` 是 Host 用于 UI 和资源投影的 typed bool model;它不是权限授予。
- `permissions` 是 runner 申请的 LangBot 资源访问上限;实际授权仍必须与 binding policy 求交。
- `metadata` 只放展示、诊断、非稳定扩展信息。
### 4.3 Capabilities
```python
class AgentRunnerCapabilities(BaseModel):
streaming: bool = False
tool_calling: bool = False
knowledge_retrieval: bool = False
multimodal_input: bool = False
skill_authoring: bool = False
interrupt: bool = False
steering: bool = False
interactions: bool = False
model_config = ConfigDict(extra="forbid")
```
- `streaming`: runner 可以返回 `message.delta`
- `tool_calling`: runner 可能调用 Host tool API。
- `knowledge_retrieval`: runner 可能调用 Host knowledge API。
- `multimodal_input`: runner 可以处理非纯文本 input / attachment。
- `skill_authoring`:(降级为便捷开关,非访问硬前提)声明该 runner 期望使用 LangBot skill 工具链。skill 本身通过**统一 tool 授权**获得——发现走 `list_skills` / `langbot_list_assets`,激活/注册走 `activate` / `register_skill`,操作走 native exec/read/write,全部计入 `resource_policy.allowed_tool_names`。该 capability 仅作为「一键授权这组 skill tool + sandbox」的便捷开关,不再单独决定 skill 是否可用。
- `interrupt`: runner 支持取消或中断。
- `steering`: runner 支持在 turn 边界通过 Host pull API 消费同 conversation 在途追加消息。
- `interactions`: runner 可以请求 Host 展示结构化交互,并处理后续 `interaction.submitted` 事件。
Capabilities 字段全部是 `bool`,未知 key 禁止进入 typed manifest。早期草案里的上下文/会话类 capability 已删除;对应语义由 event-first context 和 runner-owned context 原则表达。
### 4.4 Permissions 与 Effective Access
```python
class AgentRunnerPermissions(BaseModel):
models: list[Literal["invoke", "stream", "rerank"]] = []
tools: list[Literal["detail", "call"]] = []
knowledge_bases: list[Literal["list", "retrieve"]] = []
history: list[Literal["page", "search"]] = []
events: list[Literal["get", "page"]] = []
storage: list[Literal["plugin", "workspace"]] = []
files: list[Literal["config", "knowledge"]] = []
interactions: list[Literal["request"]] = []
model_config = ConfigDict(extra="forbid")
```
通用交互投递是当前唯一允许执行的 `action.requested` 白名单动作。Runner 必须同时声明
`capabilities.interactions=true``permissions.interactions=["request"]`Host 还必须将其与
当前 binding delivery policy、run authorization snapshot 和 adapter delivery capability 求交。
其它平台动作仍不属于当前 permissionsHost 收到后只记录 telemetry,不得执行。
Runner 实际可用 LangBot 资源来自 Host 在 run 前冻结的授权快照:
```text
effective_access = manifest.permissions ∩ binding.resource_policy ∩ current scope/config
```
具体落地:
1. `AgentResourceBuilder` 先用 manifest permissions 与 binding resource policy / runner config 求交,生成 `ctx.resources`
2. `AgentContextBuilder` 用 manifest permissions 与 binding state/storage policy 求交,生成 `ctx.context.available_apis`
3. `AgentRunSessionRegistry` 冻结 run-scoped resources 与 available APIs。
4. Runtime handler / `AgentRunAPIProxy` 按 active `run_id`、runner identity、caller plugin identity、resource id、scope、payload size、rate limit 和 deadline 校验每次调用。
反承诺:manifest permissions **只约束 LangBot 持有的资源访问**。它不承诺限制外部 harness 的 native shell、文件系统、CLI、MCP、网络或本机权限;这些能力由 operator/runtime/sandbox 另行约束,见 HOST_SDK §4.8 与 SECURITY_HARDENING。
默认原则:
- Host 不得默认 inline 全量历史。
- Host 只 inline 当前 event / input 和 context handles。
- Runner 拥有 working context assembly。
- Runner 可在授权后通过 Host history / event / state API 拉取更多上下文,并通过授权 sandbox/workspace 工具访问当前运行文件。
- 历史窗口策略不属于 Protocol v1 字段,也不属于 Host 通用语义。
context 边界的设计理由见 [AGENT_CONTEXT_PROTOCOL.md](./AGENT_CONTEXT_PROTOCOL.md)。
## 5. Run 协议
### 5.1 RUN_AGENT
Host 调用 Runtime
```python
class AgentRunRequest(BaseModel):
runner_id: str
runner_name: str
context: AgentRunContext
```
Runtime 返回 `AgentRunResult` 异步流。底层 transport 可继续用 `plugin_author` / `plugin_name` / `runner_name` 定位组件,但协议语义以 `runner_id``context` 为准。
### 5.2 AgentRunContext
这是 SDK 看到的**唯一权威 context 定义**。
```python
class AgentRunContext(BaseModel):
run_id: str
trigger: AgentTrigger
event: AgentEventContext
conversation: ConversationContext | None = None
actor: ActorContext | None = None
subject: SubjectContext | None = None
input: AgentInput
delivery: DeliveryContext
resources: AgentResources
context: ContextAccess
state: AgentRunState
runtime: AgentRuntimeContext
config: dict[str, Any] = {}
adapter: AdapterContext | None = None
metadata: dict[str, Any] = {}
```
核心约束:
- `event` 是必选字段,Protocol v1 是 event-first。
- `input` 表示当前事件的主输入,不等于历史消息。
- `bootstrap` / `messages` **不是协议字段**Host 不内联历史窗口。
- `adapter` 只放入口 adapter 的非核心元数据,runner 不应依赖它做长期能力。
- `config` 是 Agent/runner config,不是插件实例状态。
### 5.3 AgentTrigger
```python
class AgentTrigger(BaseModel):
type: str
source: Literal["platform", "webui", "api", "scheduler", "system", "host_adapter"]
timestamp: int | None = None
```
`trigger.type` 应与 `event.event_type` 一致或更粗粒度。例如入口适配器触发消息时:
```json
{ "type": "message.received", "source": "host_adapter" }
```
### 5.4 AgentEventContext
```python
class AgentEventContext(BaseModel):
event_id: str
event_type: str
event_time: int | None = None
source: str
source_event_type: str | None = None
raw_ref: RawEventRef | None = None
data: dict[str, Any] = {}
```
- `event_type` 使用 LangBot 稳定协议名,例如 `message.received`。稳定事件名清单见 [EVENT_BASED_AGENT.md](./EVENT_BASED_AGENT.md)。
- 平台原始事件名放入 `source_event_type`
- 大型原始 payload 必须放入 `raw_ref` 或 staged file,不应直接塞入 `data`
### 5.5 Conversation / Actor / Subject
```python
class ConversationContext(BaseModel):
conversation_id: str | None = None
thread_id: str | None = None
launcher_type: str | None = None
launcher_id: str | None = None
sender_id: str | None = None
bot_id: str | None = None
workspace_id: str | None = None
session_id: str | None = None
class ActorContext(BaseModel):
actor_type: str
actor_id: str | None = None
actor_name: str | None = None
metadata: dict[str, Any] = {}
class SubjectContext(BaseModel):
subject_type: str
subject_id: str | None = None
data: dict[str, Any] = {}
```
示例:
- 消息事件:actor 是发消息的人,subject 是当前消息。
- 入群事件:actor 是新成员或邀请人,subject 是群/成员关系。
- 定时事件:actor 可以是 systemsubject 是 schedule。
### 5.6 AgentInput
```python
class AgentInput(BaseModel):
text: str | None = None
contents: list[ContentElement] = []
attachments: list[InputAttachment] = []
interaction: InteractionSubmission | None = None
```
- 文本、多模态、附件都属于当前 event input。
- 大文件、图片、音频、工具大结果应进入授权 sandbox/workspaceinput attachment 只携带轻量 metadata/path/url/content。
- 平台原始消息链不属于 SDK `AgentInput`;需要诊断时放在 Host 内部 envelope 或 `ctx.adapter.extra` 的一次性兼容字段中,不作为长期 runner 合同。
- `interaction` 只在 `event.event_type == "interaction.submitted"` 时出现,承载经过 Host 校验和归一化的用户提交;平台原始 callback payload 不进入该字段。
### 5.7 DeliveryContext
```python
class DeliveryContext(BaseModel):
surface: str
reply_target: dict[str, Any] | None = None
supports_streaming: bool = False
supports_edit: bool = False
supports_reaction: bool = False
max_message_size: int | None = None
interactions: InteractionDeliveryCapabilities | None = None
platform_capabilities: dict[str, Any] = {}
```
Runner 可参考 delivery 能力决定返回 `message.delta``message.completed` 或交互请求。
`interactions is not None` 表示当前 surface 可以投递结构化交互;其中的 field/action 能力
决定 Host 可原生渲染的范围。Runner 仍必须提供 `fallback_text`,供不支持富交互或投递失败时降级。
平台事件进入独立 Agent 时,Host 会从当前 adapter 的 `get_supported_apis()` 投影
`supports_edit``supports_reaction`,并把去重后的 API 名称写入
`platform_capabilities.supported_apis`。这些字段只描述当前投递表面的能力,不授予
任意平台动作权限;交互请求按 §5.8 和 §7.3 的白名单约束执行。合成路由测试使用的
adapter 会过滤所有已知副作用 API,因此测试事件不会向 runner 宣告真实出站能力。
### 5.8 Structured Interaction
```python
InteractionFieldType = Literal[
"text", "textarea", "select", "multiselect", "number", "boolean", "file"
]
class InteractionOption(BaseModel):
value: str
label: str
description: str | None = None
class InteractionField(BaseModel):
id: str
label: str
type: InteractionFieldType
required: bool = False
options: list[InteractionOption] = []
placeholder: str | None = None
default: JSONValue | None = None
class InteractionAction(BaseModel):
id: str
label: str
style: Literal["default", "primary", "danger"] = "default"
class InteractionRequest(BaseModel):
interaction_id: str
kind: Literal["form", "confirmation", "choice"] = "form"
title: str
description: str | None = None
fields: list[InteractionField] = []
actions: list[InteractionAction] = []
expires_at: int | None = None
fallback_text: str
class InteractionSubmission(BaseModel):
interaction_id: str
action_id: str | None = None
values: dict[str, JSONValue] = {}
submitted_at: int | None = None
class InteractionDeliveryCapabilities(BaseModel):
field_types: list[InteractionFieldType] = []
action_styles: list[Literal["default", "primary", "danger"]] = []
supports_updates: bool = False
max_fields: int | None = None
```
Runner 使用 `AgentRunResult.interaction_requested()` 生成
`action.requested(action="interaction.requested")`。Host 只能把请求投递到当前 run 冻结的
delivery targetRunner 不得通过 `target` 改写 bot、conversation 或用户。Host 为请求保存
`interaction_id -> processor/binding/conversation/expiry` 关联;平台 callback 必须先经过签名、
目标、过期和幂等校验,再生成 `interaction.submitted` 事件。
Provider 私有 continuation token、workflow id、凭据或原始 callback payload 不属于该协议。
Runner 应把这些值保存到受授权的 Host state 或 plugin storage,并只用 `interaction_id` 关联提交。
Host 必须使用服务器接收时间判断 callback 是否过期,并覆盖 `InteractionSubmission.submitted_at`
平台事件时间只可作为原始事件诊断信息,不能控制 TTL 或进入 Runner 的可信提交字段。
### 5.9 ContextAccess
```python
class ContextAccess(BaseModel):
conversation_id: str | None = None
thread_id: str | None = None
latest_cursor: str | None = None
event_seq: int | None = None
transcript_seq: int | None = None
has_history_before: bool = False
inline_policy: InlineContextPolicy
available_apis: ContextAPICapabilities
class InlineContextPolicy(BaseModel):
mode: Literal["none", "current_event", "recent_tail", "summary_tail"]
delivered_count: int = 0
source_total_count: int | None = None
messages_complete: bool = False
reason: str | None = None
class ContextAPICapabilities(BaseModel):
prompt_get: bool = False
history_page: bool = False
history_search: bool = False
event_get: bool = False
event_page: bool = False
state: bool = False
storage: bool = False
steering_pull: bool = False
```
`ContextAccess` 告诉 runnerHost inline 了什么、没 inline 什么、需要更多上下文时走哪些 API。它是 runner 按需读取上下文的入口说明,不是 Host 的业务上下文编排策略。
### 5.9 AgentRuntimeContext
```python
class AgentRuntimeContext(BaseModel):
langbot_version: str | None = None
trace_id: str | None = None
deadline_at: float | None = None
metadata: dict[str, Any] = {}
```
### 5.10 AgentRunState
```python
class AgentRunState(BaseModel):
conversation: dict[str, Any] = {}
actor: dict[str, Any] = {}
subject: dict[str, Any] = {}
runner: dict[str, Any] = {}
```
State 是可选 host-owned snapshot。Runner 也可以完全自管状态。
## 6. Resources
```python
class ToolResource(BaseModel):
tool_name: str
tool_type: str | None = None
description: str | None = None
parameters: dict[str, Any] | None = None # 完整 JSON schema,由 Host 一次塞齐
operations: list[Literal["detail", "call"]] = []
class SkillResource(BaseModel):
skill_name: str
display_name: str | None = None
description: str | None = None
class AgentResources(BaseModel):
models: list[ModelResource] = []
tools: list[ToolResource] = []
knowledge_bases: list[KnowledgeBaseResource] = []
skills: list[SkillResource] = []
storage: StorageResource = StorageResource()
platform_capabilities: dict[str, Any] = {}
```
`tools` 携带每个授权工具的完整 schema(`parameters`),由 Host 在构造 `ctx.resources` 时一次塞齐,runner 不需再逐个调用 `get_tool_detail` 拉取,减少 N 次往返。
`skills` 是本次 run 中 pipeline-visible 的 skill facts`skill_name``display_name``description`)。**skill 通过统一 tool 形式消费,不是独立资源类别**:发现走 `list_skills` tool(或 `langbot_list_assets` 增加 skills 一类),激活走 `activate`,操作走 native exec/read/write。Host **不**把 skill 索引注入 system prompt,也不做 progressive-disclosure 注入;LLM 通过调用发现工具主动查询 skill 清单。Host **可选**在 ctx 提供预渲染的 `suggested_skill_prompt`(首轮延迟优化,runner 可忽略 / override),但它不是访问前提。`skills` 字段本身仅作为发现工具的数据来源与该可选预渲染的输入。
资源列表是本次 run 的授权结果。History / Event / State / Storage 访问通过 `ctx.context.available_apis` 和 Host 侧 run session 校验控制,不作为可枚举 resource list 暴露。Runner 只能通过 `AgentRunAPIProxy` 访问这些能力。当前事件的文件和工具大结果优先进入授权 sandbox/workspace,由 runner 通过 read/write/exec 类工具按需读取。
## 7. Result Stream
### 7.1 AgentRunResult envelope
```python
JSONValue = str | int | float | bool | None | list["JSONValue"] | dict[str, "JSONValue"]
ResultType = Literal[
"message.delta",
"message.completed",
"tool.call.started",
"tool.call.completed",
"state.updated",
"action.requested",
"run.completed",
"run.failed",
]
class AgentRunResult(BaseModel):
run_id: str
type: AgentRunResultType | str
data: dict[str, Any] = {}
usage: LLMTokenUsage | None = None
sequence: int | None = None
timestamp: int | None = None
```
SDK 当前实现是单一 envelope`type` 枚举 + `data` dict。Payload 由 SDK typed model 构造并 dump,但 wire 不改成 discriminated union;这样新旧版本偏斜时 Host 仍可按 §3 忽略未知 `type`
`usage` 是 runner 可选上报的 token 使用量,沿用 SDK `LLMTokenUsage`
```python
class LLMTokenUsage(BaseModel):
prompt_tokens: int | None = None
completion_tokens: int | None = None
total_tokens: int | None = None
# provider-specific detail/cached/reasoning counters are preserved as extra fields
```
约束:
- 运行时能观测到 provider/runtime usage 时,SHOULD 在 terminal `run.completed.usage` 上报本次 run 的最终聚合 token usage。
- `run.failed.usage` MAY 上报失败前已经产生的部分 usage。
- 不能观测 usage 的 runner 合法地省略该字段;缺失表示 unknown,Host 不得按 0 处理。
- ACP 等外部协议不保证统一 usageACP runner 只能在具体 provider/native event 提供 usage 时填充本字段。
- cost 不作为 runner result 的权威字段。Host 后续应基于 usage、model identity、时间和自身价格表计算账单成本;provider 原始 cost 如需保留,可放在 `usage` extra 字段中作为非权威 telemetry。
Host 边界分级校验:
- `message.delta``message.completed``state.updated``action.requested``run.completed``run.failed` 属于会影响投递或 Host 副作用的严格 payload;校验失败时丢弃该 result 并记录 warning。
- `tool.call.started``tool.call.completed` 当前只作为 telemetrypayload 宽松兼容。
- 未知 `type` 忽略并记录 warning。
### 7.2 稳定 result payloads
| type | `data` payload |
| --- | --- |
| `message.delta` | `{ "chunk": MessageChunk }` |
| `message.completed` | `{ "message": Message }` |
| `tool.call.started` | `{ "tool_call_id": str, "tool_name": str, "parameters": dict }` |
| `tool.call.completed` | `{ "tool_call_id": str, "tool_name": str, "result": dict \| None, "error": str \| None }` |
| `state.updated` | `{ "scope": "conversation" \| "actor" \| "subject" \| "runner", "key": str, "value": JSONValue }` |
| `action.requested` | `{ "action": str, "target": dict \| None, "payload": dict \| None }` |
| `run.completed` | `{ "finish_reason": str, "message"?: Message }` |
| `run.failed` | `{ "code": str, "error": str, "retryable": bool }` |
Runner 生成的大文件、工具输出和临时产物不通过 result event 回传;应写入当前 run 的授权 sandbox/workspace,再用消息文本、metadata 或 attachment reference 指向它们。
### 7.3 稳定 result types
| type | 说明 | 当前消费 |
| --- | --- | --- |
| `message.delta` | 流式消息片段。 | ✅ |
| `message.completed` | 完整消息。 | ✅ |
| `tool.call.started` | 工具调用开始的可观测事件。 | telemetry |
| `tool.call.completed` | 工具调用完成的可观测事件。 | telemetry |
| `state.updated` | runner 请求更新 host-owned state。 | ✅ |
| `action.requested` | runner 请求 Host 执行受控动作。 | `interaction.requested` 可执行;其它动作仅 telemetry |
| `run.completed` | run 正常结束。 | ✅ |
| `run.failed` | run 失败。 | ✅ |
`action.requested` 仍是严格白名单协议表面。当前 Host 只执行
`action="interaction.requested"`,并要求 payload 可校验为 `InteractionRequest`、Runner 声明
interaction capability/permission、当前 binding 允许交互且 delivery surface 支持交互。
Host 忽略 Runner 提供的外部 target,只使用 run authorization snapshot 中冻结的 delivery target。
其它 action 只记 telemetry,不执行;通用 platform action executor 仍不属于本协议。
Host 必须校验 `state.updated` 的 scope、key、value 大小和 JSON 可序列化性。Host 还必须限制
交互字段/动作数量、字符串长度、payload 总大小、过期时间和 interaction id 唯一性,并对 callback
应用默认 30 分钟、最长 24 小时的有效期;request 与 submission payload 均不得超过 256 KiB。
执行目标绑定、签名、幂等和过期校验。
除 runner 经 `state.updated` 写之外,Host 自身也可直接写部分 host-owned state。例如 `activate` tool 在 Host 侧执行时,直接把已激活 skill 写入 conversation scope 的 `host.activated_skills` 快照。当 host 直接写与 runner `state.updated` 写到同一 key 时,按 **last-write-wins** 合并——runner 可以覆盖 host 写的快照。
### 7.4 Stream delivery semantics
- Host 按 Runtime stream 顺序消费 result。当前 v1 不定义跨连接 replay,也不承诺 at-least-once;从 Host 视角,收到的 result 最多应用一次。
- `sequence` 是单个 `run_id` 内的结果序号。in-process / stdio 这类天然有序的在线 stream 可以省略;任何会缓冲、重放、跨进程队列或 runtime-managed task 的 transport 必须提供从 1 开始严格递增的 `sequence`
- Host 看到已提供 `sequence` 的 result 时,应按 `(run_id, sequence)` 做重复检测,并在缺号或乱序时记录 warning;除非 transport 明确声明 replay 语义,Host 不应自行等待缺失序号重排用户可见输出。
- `run.failed.data.retryable` 只表示整次 run 理论上可由上层重试;Protocol v1 不自动重试 run,也不自动重试 proxy action。
- History / Event / Transcript cursor 是 opaque token。runner 不得解析 cursor,也不得假设 cursor 在不同 API、conversation、thread 或 retention window 之间可比较;当前实现即使返回数字字符串,也只是实现细节。
### 7.5 示例
```json
{ "type": "message.delta", "data": { "chunk": { "role": "assistant", "content": "hel" } } }
{ "type": "message.completed", "data": { "message": { "role": "assistant", "content": "hello" } } }
{ "type": "state.updated", "data": { "scope": "conversation", "key": "external.session_id", "value": "abc" } }
{ "type": "action.requested", "data": { "action": "interaction.requested", "payload": { "interaction_id": "form_1", "kind": "choice", "title": "Approve?", "actions": [{"id": "approve", "label": "Approve", "style": "primary"}], "fallback_text": "Reply approve or reject." } } }
```
## 8. AgentRunAPIProxy
所有 proxy action 必须携带 `run_id`。Host 必须校验:active run session 存在、caller plugin identity 匹配、resource 在本次 `ctx.resources` 中授权、scope 不越界、payload size / rate limit / deadline 合法。
```python
# Model
await api.invoke_llm(llm_model_uuid, messages, funcs=None, extra_args=None)
await api.invoke_llm_with_usage(llm_model_uuid, messages, funcs=None, extra_args=None)
async for chunk in api.invoke_llm_stream(llm_model_uuid, messages, funcs=None, extra_args=None):
...
async for event in api.invoke_llm_stream_events(llm_model_uuid, messages, funcs=None, extra_args=None):
...
await api.invoke_rerank(rerank_model_id, query, documents, top_k=None)
# Tool
await api.get_tool_detail(tool_name)
await api.call_tool(tool_name, parameters)
# Knowledge
await api.retrieve_knowledge(kb_id, query_text, top_k=5, filters=None)
# History(返回 Transcript projection,不返回原始平台 payload
await api.get_prompt()
await api.history_page(conversation_id=None, before_cursor=None, after_cursor=None,
limit=50, direction="backward", include_attachments=False)
await api.history_search(query, filters=None, top_k=10)
# Event(返回稳定 event envelope 或受限 raw ref,不默认返回大 payload
await api.event_get(event_id)
await api.event_page(conversation_id=None, event_types=None, before_cursor=None, limit=50)
await api.steering_pull(mode="all", limit=None)
# State / Storage
await api.state_get(scope, key); await api.state_set(scope, key, value); await api.state_delete(scope, key)
await api.state_list(scope, prefix=None, limit=100)
await api.get_plugin_storage(key); await api.set_plugin_storage(key, value); await api.delete_plugin_storage(key)
await api.get_plugin_storage_keys()
await api.get_workspace_storage(key); await api.set_workspace_storage(key, value); await api.delete_workspace_storage(key)
await api.get_workspace_storage_keys()
# Host info
await api.get_langbot_version()
```
`invoke_llm()` / `invoke_llm_stream()` 的第一个参数在 SDK 中命名为
`llm_model_uuid`wire payload 字段也是 `llm_model_uuid`。该值对 runner
仍是 opaque identifier,不应解析其内部格式。
`invoke_llm()``invoke_llm_stream()` 保持兼容:前者返回 `Message`,后者只
yield `MessageChunk`。需要 provider 真实 token 计量的 runner 应使用
`invoke_llm_with_usage()``invoke_llm_stream_events()`。Host response 可在
原有 `{message: ...}` / `{chunk: ...}` 外额外携带可选 `usage` 字段;streaming
场景允许在所有 chunk 之后追加一个 usage-only event。`usage` 至少保留
OpenAI-compatible 的 `prompt_tokens``completion_tokens``total_tokens`
若 provider 返回 `prompt_tokens_details` / `completion_tokens_details`
cache token countersHost / SDK 不应丢弃这些字段。没有 usage 的 provider
必须继续返回成功响应,SDK 将 usage 置为 `None`
`get_prompt()` 返回当前 query-backed run 的 Host effective prompt messages
`list[Message]` 的 JSON 形式。该能力只在 `ctx.context.available_apis.prompt_get`
为 true 时可用;没有 query 缓存、prompt 已过期或非 query entry run 时 Host
可以返回错误或空列表。Runner 应在不可用时回退到自己的 config/prompt 策略。
`steering_pull(mode="all")` 是推荐默认:Host 按 claim 顺序返回全部 pending steering 输入并清空对应队列。`mode="one-at-a-time"` 仅用于 runner 主动节流,每次返回一条。Host 不合并多条用户消息;runner 负责在 turn 边界决定模型侧格式。
Steering 审计使用 EventLog 而不是 Transcript schema 扩展:被 active run 吸收的原始 `message.received` 事件保留原事件类型,并在 `metadata.steering` 标记 `status="queued"``trigger_behavior="absorbed_into_active_run"``claimed_by_run_id``claimed_runner_id``claimed_at`。Runner 成功 pull 后,Host 追加 `steering.injected` EventLog 记录,`metadata.steering.status="injected"` 并引用 `source_event_id`。若 run 结束时仍有已 claim 但未 pull 的 steering 输入,Host 追加 `steering.dropped` EventLog 记录,`metadata.steering.status="dropped"` 并引用 `source_event_id`;这不是用户消息事实的删除,只是 dispatch 终态。Transcript 继续只表示会话事实,不承担 dispatch 行为标记。
`state``storage` 的建议边界:`state` 放小型 JSONconversation / actor / subject / runner),`storage` 放 blob 或较大数据(插件私有数据、workspace 数据、checkpoint)。
Compaction checkpoint 的推荐 state 约定:
- scope: `conversation`
- key: `runner.compaction.checkpoint`
- value:
```json
{
"schema_version": "langbot.local_agent.compaction_checkpoint.v1",
"summary": "<conversation_summary>...</conversation_summary>",
"covers_until": "transcript-cursor-or-seq",
"tokens_before": 12345,
"created_at": 1710000000,
"conversation_id": "conv-..."
}
```
`covers_until` 是摘要覆盖到的 transcript 游标锚点。Runner 读取 checkpoint 后应只拉取该游标之后的 transcript;若 checkpoint 缺失、schema 不匹配、conversation 不匹配或游标不可用,应回退到无 checkpoint 的尾部历史拉取行为。
Proxy 返回数据结构也属于本协议:
```python
class TranscriptItem(BaseModel):
transcript_id: str
event_id: str
conversation_id: str | None = None
thread_id: str | None = None
role: str
item_type: str = "message"
content: str | None = None
content_json: dict[str, Any] | None = None
attachment_refs: list[dict[str, Any]] = []
seq: int | None = None
cursor: str | None = None
created_at: int | None = None
metadata: dict[str, Any] = {}
class HistoryPage(BaseModel):
items: list[TranscriptItem] = []
next_cursor: str | None = None
prev_cursor: str | None = None
has_more: bool = False
total_count: int | None = None
class HistorySearchResult(BaseModel):
items: list[TranscriptItem] = []
total_count: int | None = None
query: str
class AgentEventRecord(BaseModel):
event_id: str
event_type: str
event_time: int | None = None
source: str
bot_id: str | None = None
workspace_id: str | None = None
conversation_id: str | None = None
thread_id: str | None = None
actor_type: str | None = None
actor_id: str | None = None
actor_name: str | None = None
subject_type: str | None = None
subject_id: str | None = None
input_summary: str | None = None
input_ref: str | None = None
raw_ref: str | None = None
seq: int | None = None
cursor: str | None = None
created_at: int | None = None
metadata: dict[str, Any] = {}
class EventPage(BaseModel):
items: list[AgentEventRecord] = []
next_cursor: str | None = None
prev_cursor: str | None = None
has_more: bool = False
total_count: int | None = None
class SteeringInputItem(BaseModel):
claimed_run_id: str
runner_id: str
claimed_at: int | None = None
event: AgentEventContext
input: AgentInput
conversation: ConversationContext | None = None
actor: ActorContext | None = None
subject: SubjectContext | None = None
metadata: dict[str, Any] = {}
class SteeringPullResult(BaseModel):
items: list[SteeringInputItem] = []
```
## 9. 错误模型
```python
class AgentAPIError(BaseModel):
code: str
message: str
retryable: bool = False
details: dict[str, Any] = {}
```
| code | 说明 |
| --- | --- |
| `unauthorized` | 未授权访问资源或 scope。 |
| `not_found` | 资源不存在或对当前 runner 不可见。 |
| `deadline_exceeded` | 超过 run deadline。 |
| `payload_too_large` | 请求或响应过大。 |
| `rate_limited` | Host 限流。 |
| `invalid_argument` | 参数错误。 |
| `runtime_error` | Host 或下游能力错误。 |
SDK runner-facing proxy 在 Host 返回结构化错误或畸形响应时抛出 `AgentAPIException`,其中 `error` 字段为 `AgentAPIError`。Legacy transport 只返回字符串错误时,SDK 使用 `host.action_error` 包装,避免 runner 继续依赖裸 `KeyError` 或字符串匹配。
Runner 失败使用 `run.failed`
```json
{ "type": "run.failed", "data": { "code": "runner.error", "error": "failed to call external agent", "retryable": false } }
```
## 10. Timeout 与 Cancellation
- Host 在 `ctx.runtime.deadline_at` 下发总 deadlineSDK proxy 必须用该 deadline 限制单次 action timeout。
- Host 可以取消 active runRuntime 应尽力中断 runner。
- Protocol v1 的 run 绑定当前 Host 进程和当前 runtime channel,不保证跨 Host 重启恢复。Host 重启、runtime channel 断开或 run session 丢失时,Runtime / external harness connector 必须 fail-fast 并尽力取消仍在执行的 runner,不得继续使用旧 `run_id` 调用 Host API。
- Runner 支持中断时应返回或触发 `run.failed`code 为 `cancelled`
- Host 必须 unregister active run session。
## 11. Security 与 Guardrail(协议层)
Protocol v1 的安全边界在 Host
- Runner 不能直接访问未授权 model/tool/kb/history/storage/sandbox。
- SDK 本地校验只提升开发体验,不能替代 Host 校验。
- 所有 resource id 对 runner 来说都是 opaque。
- 默认只能访问当前 conversation / thread 的 history;跨会话、workspace 级访问必须额外授权。
- 大 payload 不应塞进 result event;当前 run 的文件和工具大结果应进入授权 sandbox/workspace,由 read/write/exec 类工具按需访问。
- Host 必须记录 run_id、runner_id、action、resource、scope、result。
Host 不负责业务编排:不拼接全量历史、不替 runner 做 prompt assembly、不内置 agent memory / tool loop / 上下文压缩策略。这些由官方或第三方 AgentRunner 插件实现。
外部 harness runner 的边界统一见 HOST_SDK §4.8。简言之:harness native permission mode、allowed/disallowed tools、shell/MCP 权限只是额外执行约束,不能替代 Host 对 LangBot 资源的授权。
> 发布级路径隔离、MCP allowlist、secret redaction、配额、workspace 清理等**不属于** v1 协议闭环,是生产默认启用前的 release gate,见 [SECURITY_HARDENING.md](./SECURITY_HARDENING.md)。
## 12. Pipeline AI Stage Adapter 边界
Pipeline 与 Agent 是 EBA 中平级的处理器:Pipeline 处理消息事件并执行完整
Stage 链,Agent 处理其声明支持的消息或非消息事件。本协议只约束 AgentRunner
调用,因此 Pipeline 仅在 AI Stage 调用 runner 时进入 Query entry adapter
该适配不会把 Pipeline 变成 Agent,也不会创建或更新持久 Agent。adapter 负责:
-`Query` 构造 `AgentEventContext` 和临时 `AgentBinding`(见 HOST_SDK §4.2)。
- 从当前 Agent/runner config 构造 `ctx.config`
- 将 Query-only 字段放入 `ctx.adapter`,例如 filtered params 放 `ctx.adapter.extra["params"]`
约束:
- adapter **不**定义历史窗口、prompt 组装或 agentic context 策略。
- `ctx.adapter.extra` 只允许承载一次性、JSON-safe、入口相关的非核心元数据,例如 `params`;不得承载 `prompt`、history window、RAG 结果、tool schema 或授权资源。
- 静态绑定 prompt 属于 `ctx.config.prompt`。preprocessing / hook 后的动态有效指令不通过 `ctx.adapter.extra` 主动推送;后续如需要保留这类能力,应通过 Host prompt/instruction pull API 暴露(占位见 HOST_SDK §4.8)。
- 新 runner 不应长期依赖 `adapter`,应只依赖 event-first context 和 Host API。
## 13. 已确认约束
- EBA 路由层是 `one event -> one Processor target (Pipeline | Agent)`;同一 bot / channel 可以让不同事件绑定不同类型的处理器。
- 进入 AgentRunner Protocol 后,调用基数是 `one AgentBinding -> one run_id -> one runner`。这既适用于独立 Agent,也适用于 Pipeline AI Stage 的单次 runner 调用。
- 一个 Agent 可以被多个 bot / channel 复用。如果 Agent 分支出现多个匹配 bindingBindingResolver 必须按明确规则选出一个或拒绝配置,不应默认 fan-out。
- observer agent、多 runner fan-out、并行裁决、result 合并等能力需要单独设计 delivery、state、platform action 和 audit 语义,不属于当前 v1 契约。
- `AgentRunnerDescriptor.source` 只允许 `plugin`Host 内置 adapter 不能作为 runner source 绕过插件/runtime/proxy 权限链。
- `ctx.resources` 与 proxy action 校验必须来自同一个 run authorization snapshotruntime handler 不应重新执行资源裁剪。
- v1 不要求 Agent、AgentRunner 插件实例或 runner id 全局串行。多个 bot / channel 可复用同一个 Agent;并发隔离依赖 `run_id`、binding、conversation / thread scope 和 Host authorization snapshot。
- 外部 harness runner 当前是 MVP / dev path,证明协议可接入,不代表发布级安全边界或 Docker 生产可用性完成。
## 14. 开放问题
- `AgentBinding` 是否需要进入 SDK 文档作为只读诊断信息,还是完全 Host 内部。
- State 与 Storage 的边界是否需要更强类型。
- platform action 的审批模型如何表达。
- Host 侧 scoped MCP / workspace projection 是否需要从 runner config 上移为一等 resource projection API。(skill 一项已收敛:skill 全 tool 化,作为被授权 tool 暴露,不再是独立 projection。)
-155
View File
@@ -1,155 +0,0 @@
# Agent Runner 插件化文档入口
本文档是 agent-runner 插件化工作的路由页。具体设计拆到独立文档中维护,避免把 LangBot 宿主架构、SDK 协议、上下文管理、EBA 接入边界和官方 runner 迁移混在同一份 README 里。
## 背景与问题
旧 runner 路径主要围绕 Pipeline / Query 和 `pkg/provider/runners` 内置实现展开,扩展外部 agent runtime 时容易把 runner 选择、上下文裁剪、资源授权和消息投递绑在同一条聊天链路里。这个分支要把 LangBot 收敛成 Agent Host:Host 负责事件、绑定、授权、事实源和结果投递;AgentRunner 作为插件或外部 harness 消费统一协议并自主管理 prompt / history / memory。
## 文档维护原则(单一事实源)
- **协议数据结构(schema)唯一定义在 [PROTOCOL_V1.md](./PROTOCOL_V1.md)。** 其他文档不得重抄 schema,只能引用,例如"见 PROTOCOL_V1 §4.2"。
- 当前实现状态、spec 差距与 runner 验收状态归 [STATUS.md](./STATUS.md);测试执行入口归 [AGENT_RUNNER_QA_GUIDE.md](./AGENT_RUNNER_QA_GUIDE.md),安全发布门槛归 [SECURITY_HARDENING.md](./SECURITY_HARDENING.md)。
- Host 内部模型(`AgentEventEnvelope``AgentBinding`、Descriptor、各 Store)定义在 [HOST_SDK_INFRASTRUCTURE.md](./HOST_SDK_INFRASTRUCTURE.md),不属于 SDK 协议。
- 其余专题文档只讲"为什么/边界/怎么用",避免重复叙述。
## 本分支目标
**本分支目标:AgentRunner 外化 / 插件化基础设施**
本分支只做 LangBot 作为 Agent Host 的基础能力建设,让现有 Pipeline 与用户新建的独立 `Agent` 都能调用插件化 AgentRunner;不负责把两者做持久化迁移:
- LangBot 与 SDK 的稳定协议合同(Protocol v1
- Host-side `AgentEventEnvelope` / `AgentBinding` 模型
- `run(event, binding)` event-first 入口
- `QueryEntryAdapter`Query → AgentEventEnvelope + AgentBinding
- EventLog / Transcript / PersistentStateStore
- History / Event / State pull APIs
- Sandbox/workspace read/write/exec 文件能力,用于当前 run 的上传文件、工具大结果和临时产物
- SDK runtime forwarding pull APIs + `caller_plugin_identity` 验证路径
## 本分支不实现
以下能力由其他分支负责,本分支只保留 integration point。EBA 完整事件网关与事件路由当前由外部 EBA 分支联调:
- **EventGateway / EventRouter**:完整事件网关实现、事件路由、事件持久化管理
- **Event subscription / Event notification**:事件订阅、推送通知
- **BindingResolver persistence UI**:绑定配置的持久化 UI 和 event router 集成(如由其他模块负责)
- **Scheduler / Background event source**:定时任务、后台事件源
- **完整 Agent Platform / daemon control plane**Host-owned `AgentRun` / `AgentRunEvent`、run control primitives、最小 runtime heartbeat/claim lease 已作为 v2 foundation 落地;业务队列、Platform UI、daemon supervisor、runtime wakeup channel 和分布式 runtime 管控仍不属于 Protocol v1 主线。
EventGateway / EventRouter 在本文档中描述为 **external EBA branch integration point**,由外部 EBA 分支提供并联调。本分支只定义 host-side envelope/binding models 和 `run(event, binding)` orchestrator 入口。
本分支与外部 EBA / Agent Platform / Runtime Control Plane 的扩展边界见 [EXTENSION_SCOPE_MATRIX.md](./EXTENSION_SCOPE_MATRIX.md)。
## 目标产品模型
产品层同时保留 Pipeline 与独立 `Agent`。现有 Pipeline 保持原实体和执行链,也可继续被 bot 绑定;新建 Agent 则携带 runner id、runner config、resource/state/delivery policy 等配置,并可被多个 bot / IM channel 复用。统一产品列表可以聚合展示两类处理器,但不会把 Pipeline 或其内嵌 runner 配置自动迁移成 Agent;用户需要 Agent 时自行新增并绑定。
调度基数、Agent 复用、插件实例无状态、Pipeline adapter 和 fan-out 边界的规范来源是 [PROTOCOL_V1.md](./PROTOCOL_V1.md) §13;README 不复写这些约束。
## Pipeline 与 AgentRunner 的关系
**Pipeline 与 Agent 是 EBA 中平级的处理器;`QueryEntryAdapter` 只适配 Pipeline 内部的 AgentRunner 调用。**
EBA 先根据 `target_type` 选择 Pipeline 或 Agent。Pipeline 目标执行完整 Stage 链;当 Pipeline 的 AI Stage 调用 runner 时,`run_from_query()``QueryEntryAdapter``Query` 转换为 `AgentEventEnvelope` + `AgentBinding`,再委托到统一的 `run(event, binding, ...)`。Agent 目标则直接从自己的持久配置构造 binding。两条路径可以复用同一套 AgentRunner Host capabilities,但 Pipeline 本身不会被投影或持久化为 Agent。
下一轮测试路径、状态定义和 smoke 记录见 [AGENT_RUNNER_QA_GUIDE.md](./AGENT_RUNNER_QA_GUIDE.md)。
## 术语表
| 术语 | 含义 |
| --- | --- |
| Protocol v1 | Host 调用 AgentRunner 的 runner 可见合同:discovery、`AgentRunContext`、result stream、Host pull API 和错误模型。 |
| Processor | EBA 的处理器上位概念;当前平级类型为 Pipeline 与 Agent。 |
| Agent | 目标产品层配置对象,保存 runner id、runner config 和资源/状态/投递策略;不等于插件实例。 |
| AgentConfig | Host 内部的单次 AgentRunner 调用配置投影,可由 Pipeline AI Stage 或持久 Agent 生成;投影本身不会创建 Agent。 |
| AgentBinding / binding | Host 在一次事件运行前解析出的有效绑定,决定调用哪个 runner 以及带什么策略。 |
| envelope | Host 内部事件封装,即 `AgentEventEnvelope`runner 看到的是由它投影出的 `ctx.event`。 |
| descriptor / manifest | runner discovery 的能力和配置描述;manifest 来自插件,descriptor 是 Host 校验后的注册表视图。 |
| EBA | Event Based Agent,把消息、撤回、入群、定时任务等都统一成 host event 的接入方向;完整网关和路由在外部 EBA 分支联调。 |
| harness runner | ACP、Claude Code、Codex 等已有自身 session / tool loop / MCP / 压缩机制的外部 runtime adapter。 |
| projection | Host 把内部事实源、授权资源或配置裁剪成 runner / harness 可消费视图的过程。 |
| Runtime Control Plane | v2 Host 能力层,当前已落地 Host-owned run/result ledger、run control primitives、最小 runtime heartbeat/claim lease;完整 daemon worker 管控、task wakeup 和 Agent Platform 产品形态不是 Protocol v1 主线。 |
## 设计文档
| 文档 | 关注点 |
| --- | --- |
| [PROTOCOL_V1.md](./PROTOCOL_V1.md) | **🔒 唯一 schema 事实源**。LangBot Host 与 SDK / Runtime / AgentRunner 的协议合同:版本协商、discovery、run context、result stream、proxy actions、错误和 adapter 边界。 |
| [HOST_SDK_INFRASTRUCTURE.md](./HOST_SDK_INFRASTRUCTURE.md) | LangBot 宿主能力与分层架构、Host 内部模型(`AgentEventEnvelope` / `AgentBinding` / Descriptor / 各 Store)、runner 发现、绑定、资源授权、状态、存储、生命周期和调用链。 |
| [AGENT_CONTEXT_PROTOCOL.md](./AGENT_CONTEXT_PROTOCOL.md) | Agent-owned context 方向:事件到来时 LangBot 传什么,agent 如何按需拉取更多历史 / state、如何访问 sandbox/workspace 文件,以及如何支持 KV cache 友好的上下文管理。 |
| [EXTENSION_SCOPE_MATRIX.md](./EXTENSION_SCOPE_MATRIX.md) | AgentRunner 外化与外部 EBA / Agent Platform / Runtime Control Plane 的扩展边界矩阵,说明哪些是本分支底座、哪些由外部分支接入。 |
| [EVENT_BASED_AGENT.md](./EVENT_BASED_AGENT.md) | EBA 接入边界:事件模型、事件来源、触发绑定、非消息事件如何复用 AgentRunner 调度;完整 EventGateway / EventRouter 由外部 EBA 分支联调。 |
| [eba-productization-release.md](./eba-productization-release.md) | EBA 适配器与 AgentRunner 插件化合并后的产品化 / 发布计划,说明非技术用户快速上手差距、Bot 与处理器边界、未来 Solution 分发标的,以及多 namespace SaaS 支持要求。 |
| [RUNTIME_CONTROL_PLANE_V2.md](./RUNTIME_CONTROL_PLANE_V2.md) | Agent Platform v2 / runtime 管控面决策:`AgentRun` / `AgentRunEvent` / run control 已作为 Host 事实源落地,最小 runtime heartbeat/claim lease 已落地;完整 runtime registry / daemon 管控仍是后续可选阶段。 |
| [OFFICIAL_RUNNER_PLUGINS.md](./OFFICIAL_RUNNER_PLUGINS.md) | 官方 runner 插件迁移,包括 local-agent 和外部 runner。它是下游落地计划,不是 LangBot 基础能力设计的前置约束。 |
| [RUN_STEERING_AND_CHECKPOINT.md](./RUN_STEERING_AND_CHECKPOINT.md) | 运行中消息注入(steering / follow-up)与压缩摘要持久化(compaction checkpoint)的设计与落地状态记录;schema 仍以 PROTOCOL_V1 为准。 |
| [STATUS.md](./STATUS.md) | 当前实现状态、spec 与实现已知差距、runner 验收状态和历史高价值记录。 |
| [AGENT_RUNNER_QA_GUIDE.md](./AGENT_RUNNER_QA_GUIDE.md) | Agent Runner QA 指南:保留最高价值测试路径,指导 agent 开展下一轮 WebUI / runner smoke 验证。 |
| [SECURITY_HARDENING.md](./SECURITY_HARDENING.md) | 安全发布级 hardening 的后续发布门槛:路径隔离、权限边界、secret、资源配额、MCP / skill 投影和审计。 |
## 工作拆分
### 1. LangBot + SDK 基础设施
目标是把 LangBot 从内置 runner 执行器变成 agent host
- LangBot 与 SDK 的稳定协议合同
- runner manifest / descriptor / registry
- Agent / binding 配置解析
- run orchestration 和生命周期管理
- resource authorization 与 `run_id` 级权限校验
- host-owned state / storage / event log / transcript 能力
- sandbox/workspace 文件 staging 与 read/write/exec 能力
- SDK `AgentRunner``AgentRunContext``AgentRunResult``AgentRunAPIProxy`
协议合同详见 [PROTOCOL_V1.md](./PROTOCOL_V1.md)。
详见 [HOST_SDK_INFRASTRUCTURE.md](./HOST_SDK_INFRASTRUCTURE.md)。
### 2. Agent-owned context
LangBot 不应成为最终 agentic context manager。它应提供事实源、默认上下文引用和按需读取 API;agent 或其背后的 runtime 负责历史剪裁、摘要、召回和 KV cache 策略。
Host 不定义通用历史窗口字段或策略;runner 通过 Host pull API 按需拉取历史并自行管理 working context。
详见 [AGENT_CONTEXT_PROTOCOL.md](./AGENT_CONTEXT_PROTOCOL.md)。
### 3. Event Based AgentExternal Branch
消息只是事件的一种。外部 EBA 分支中的 `message.received``message.recalled``group.member_joined``friend.request_received` 等事件都应能通过统一事件 envelope 触发 AgentRunner。
EBA dispatch 的基数和 fan-out 边界仍以 PROTOCOL_V1 §13 为准;本文档只列出本分支提供给外部 EBA 分支复用的入口点。
**本分支不实现 EBA 完整能力,只提供:**
- event-first envelope (`AgentEventEnvelope`)
- AgentBinding model
- `run(event, binding)` 入口
- QueryEntryAdapter(当前 AgentEventEnvelope / AgentBinding 的 Query entry adapter source
详见 [EVENT_BASED_AGENT.md](./EVENT_BASED_AGENT.md)。
### 4. 官方 runner 插件
官方 `local-agent` 和外部 runner 迁移是下游工作。它们需要依附 LangBot 提供的宿主能力,但不应反过来决定宿主协议。
`local-agent` 可以外移,也可以重写。验收重点是它能完整消费 LangBot 的模型、工具、知识库、存储、事件、history API 和 result stream,而不是保留旧内置 runner 的内部结构。
详见 [OFFICIAL_RUNNER_PLUGINS.md](./OFFICIAL_RUNNER_PLUGINS.md)。
### 5. Runtime Control Plane v2Foundation Partial
当前 AgentRunner v1 主线仍以 `event -> binding -> runner.run(ctx) -> result stream` 为 runner 可见合同。Host 侧已经新增持久 `AgentRun` / `AgentRunEvent`、result persistence、cancel/finalize/query 等通用 run control primitives,并提供受权限保护的最小 runtime register/heartbeat/list、claim/renew/release 和 reconcile 原语。
在这些 Host 能力之上,可以构建独立 agent 管控面插件;插件负责 UI、策略和编排体验,runtime/task 的事实源仍由 Host 持有。完整 daemon supervisor、任务唤醒/长轮询/WebSocket、跨 Host 分布式锁、provider 登录态诊断和产品化业务队列仍是后续工作。
详见 [RUNTIME_CONTROL_PLANE_V2.md](./RUNTIME_CONTROL_PLANE_V2.md)。
## 约束事实源
本分支已确认约束不在 README 重写:
- Runner 可见协议、result stream 和调度边界见 [PROTOCOL_V1.md](./PROTOCOL_V1.md)。
- Host 内部 `AgentConfig` / `AgentBinding` 投影见 [HOST_SDK_INFRASTRUCTURE.md](./HOST_SDK_INFRASTRUCTURE.md)。
- 外部 EBA / Agent Platform / Runtime Control Plane 接入边界见 [EXTENSION_SCOPE_MATRIX.md](./EXTENSION_SCOPE_MATRIX.md)。
@@ -1,541 +0,0 @@
# Agent Platform / Runtime Control Plane Decision Note
本文档记录 AgentRunner 插件化之后,LangBot 如何继续演进成 Agent Platform 基础设施层。这里讨论的是 Host capability layer,不是 `AgentRunner Protocol v2`,也不是把某个具体 Agent Platform 产品写进 LangBot core。
> 本文是当前决策版。协议数据结构仍以 [PROTOCOL_V1.md](./PROTOCOL_V1.md) 为准;测试执行入口见 [AGENT_RUNNER_QA_GUIDE.md](./AGENT_RUNNER_QA_GUIDE.md);扩展边界见 [EXTENSION_SCOPE_MATRIX.md](./EXTENSION_SCOPE_MATRIX.md)。
>
> 实现状态说明:本文描述的是 Runtime Control Plane v2 的目标能力和分阶段落地建议。当前 AgentRunner 插件化主线已经具备 event-first context、run-scoped authorization、EventLog / Transcript / State / sandbox 文件等 Host capability,并已落地持久 `AgentRun` / `AgentRunEvent` ledger、run control actions、最小 runtime heartbeat/claim lease 和 admin reconcile 原语。完整 Agent Platform 产品形态、daemon supervisor、runtime wakeup channel 和分布式 runtime 管控仍未完成。当前实现状态以 [STATUS.md](./STATUS.md) 为准。
## 1. 当前决策
LangBot 后续定位应更像 **Agent Host / infrastructure provider / transfer layer**,而不是把某个完整 Agent Platform 产品固化进 core。
结论:
- **Agent Platform 产品形态做成插件**。插件负责 agent 管理、策略、业务队列、UI、编排、多 agent 协作和产品体验。
- **Agent Platform 所需的基础事实源做进 Host**。当前 Host 已保存 event、state、transcript、sandbox 文件边界、active run 权限快照、持久 run/result ledger、审计关联和通用控制状态。
- **最小 runtime registry / heartbeat / claim lease 已作为 Host 原语落地,但不等于完整 daemon worker 管控**。远程 harness / daemon 的进程托管、wakeup channel、provider 登录态诊断和分布式调度仍可以先由 AgentRunner 插件和 SDK remote layer 自己维护。
- **不把业务调度写进 Host**。Host 提供通用 run/result/control primitivesPlatform 插件决定哪些事件触发哪些 agent、如何排队、如何分配、是否 fan-out。
推荐分层:
```text
LangBot Host
Current base: EventLog / runtime AgentBinding / State / Transcript / sandbox files / active run authorization
Current v2 foundation: Run / RunEvent / audit / result persistence / control primitives / minimal runtime heartbeat and claim lease
Planned: Agent / Binding persistence / daemon supervisor / wakeup channel / distributed runtime operations
Agent Platform plugin
Agent management UI / project-task model / event routing policy
Business queue / multi-agent orchestration / runtime selection policy
AgentRunner plugin / external harness runtime
Connects ACP / remote daemon / local subprocess / HTTP API
Executes and converts provider-native events to AgentRunResult
```
## 2. Platform 与非 Platform 的区别
当前 LangBot 已经具备 Agent Host 的核心特征:
- 抹平不同 AgentRunner。
- 从 IM / Pipeline 入口触发 runner。
- 有 event-first context 方向。
- 有 Host-owned EventLog / Transcript / State 和 sandbox/workspace 文件边界。
- 有 runner config 下发和 active run-scoped authorization。
-`run_id` 串联 event、transcript、state、sandbox 文件和内存授权上下文。
这还不是完整 Agent Platform。完整 Platform 至少还需要:
- 可管理的 agent 资产:agent profile、binding、resource policy、runner config、可用状态。
- 可观察的执行生命周期:run status、result stream、失败原因、文件引用、审计、回放。
- 可运营的控制面:取消、重试、排队、并发、超时、恢复、诊断。
- 可产品化的调度体验:事件订阅、路由策略、任务板、多 agent 协作、项目/工作区视图。
因此,区别不只是“有没有调度”,而是是否具备:
```text
managed agent assets + observable run lifecycle + operational run control
```
Host 负责这些能力的通用事实源和安全边界;Platform 插件负责把它们组装成具体产品。
### 2.1 当前实现边界
当前代码中的 `run_id` 已经连接 active run 授权、持久 run ledger 和多个 Host 事实源:
- `EventLog` 保存输入事件和审计入口,并记录 `run_id` / `runner_id`
- `Transcript` 保存对话历史投影,并用 `run_id` 关联 assistant 输出。
- Sandbox/workspace 保存当前运行输入文件和 runner 产物,并用 `run_id` 做访问边界的一部分。
- `PersistentStateStore` 保存 runner state,但不等同于 run lifecycle。
- `AgentRunSessionRegistry` 保存 active run 的内存态授权快照,用于 proxy action 校验;进程结束或 run 结束后不作为可回放事实源。
- `AgentRun` 保存 run lifecycle、scope、authorization snapshot、queue/claim 状态、cancel intent、usage/cost 和 metadata。
- `AgentRunEvent` 保存 runner/result/admin event stream,按 `run_id + sequence` 做可回放分页。
- `AgentRuntime` 保存最小 runtime registry / heartbeat 事实,用于 runtime list、stale mark 和 claim lease reconcile。
因此本文后续提到的 `AgentRun` / `AgentRunEvent``run_append_result``run_finalize``run_cancel``runtime_register``runtime_heartbeat``run_claim` 等基础原语已经存在。仍未完成的是独立 platform `run_create` action、Host-owned Agent / Binding 持久模型、业务队列产品形态、daemon supervisor、runtime wakeup channel、跨 Host 分布式锁和 provider/runtime 诊断面。
## 3. 基础概念
### 3.1 Event
Event 表示“发生了什么”:
```text
message.received
github.issue.opened
scheduler.tick
user.approved
system.webhook.received
```
EBA 负责把外部输入标准化成 event。Event 本身不是 queue,也不等同于一次 agent 执行。当前 `EventLog` 记录的是输入事件和审计事实;未来 `AgentRunEvent` 记录的是某次 run 的输出事件流,二者不能混用。
### 3.2 Run
Run 表示“某个 agent / binding / runner 针对某个 event 的一次执行”。
Run 应由 Host 持久化,成为执行状态、结果、权限和审计的事实源:
```text
run_id
event_id
agent_id / binding_id
runner_id
status
created_at / started_at / finished_at
error / failure_reason
delivery target
metadata
```
当前 `AgentRunSessionRegistry` 只保存 active run 的内存态授权信息,不足以支撑 Platform 的回放、审计、取消、重试和异步执行。
### 3.3 RunEvent / RunResult
RunEvent 是一次 run 过程中产生的结果事件流,对应 runner 返回的 `AgentRunResult`。它不同于 EBA/EventLog 的输入事件:
```text
message.delta
message.completed
tool.call.started
tool.call.completed
state.updated
action.requested
run.completed
run.failed
```
Host 应保存这些输出事件,按 `run_id + sequence` 可回放。Transcript、State 可以由这些 result event 触发写入现有 store,并保留能回溯到 `AgentRunEvent` 的关联。文件和工具大结果留在当前 run 的 sandbox/workspace 中,不作为 result event blob 回传。
### 3.4 Queue
Queue 不是 EBA 的替代品。
EBA 负责产生 eventqueue 负责处理“这个 event 对应的执行 work item 何时执行、谁来执行、如何取消/重试/恢复”。
队列可以分两层:
- **业务队列**:由 Platform 插件管理,例如项目任务、优先级、agent team、workflow、人工审批。
- **执行队列 / run queue**:可选 Host 原语,例如 queued / running / completed / failed / cancelled、claim lease、dispatch timeout、orphan recovery。
第一阶段不要求 Host 内置完整执行队列。Platform 插件可以先管理业务队列;在 Phase 1 / Phase 2 能力落地前,插件仍只能通过现有 `AgentRunOrchestrator.run(...)` 同步执行路径和现有 Host stores 获得有限的 run 关联能力。
### 3.5 Runtime / Daemon
Runtime / daemon 表示执行位置或执行能力,例如某台机器上的 Claude Code / Codex CLI。
当前决策:
- Host 不在第一阶段维护完整 runtime registry。
- AgentRunner 插件可以通过 SDK remote layer 与 daemon 保持连接、心跳和执行通道。
- 外部 harness / agent 不应直接访问 LangBot Host 或数据库。访问 LangBot 资源必须通过 daemon / AgentRunner plugin / SDK runtime / `AgentRunAPIProxy` / scoped MCP bridge,并接受 run-scoped authorization 校验。
- 如果后续多个插件都需要共享 runtime 状态,再把薄的 `RuntimeLease` / registry 下沉为 Host 通用能力。
## 4. Host 应新增的最小能力
第一阶段最重要的不是 daemon registry,而是让 Host 成为 run/result 的事实源。
### 4.1 AgentRun Store
新增持久 `AgentRun`
```text
id / run_id
event_id
agent_id
binding_id
runner_id
conversation_id / thread_id
workspace_id / bot_id
status
status_reason
created_at / started_at / finished_at / updated_at
deadline_at
cancel_requested_at
usage_json
cost_json
metadata_json
```
建议 status 至少包含:
```text
created
running
completed
failed
cancelled
timeout
```
如果后续加执行队列,再引入:
```text
queued
claimed
dispatching
```
### 4.2 AgentRunEvent Store
新增持久 `AgentRunEvent`
```text
id
run_id
sequence
type
data_json
usage_json
created_at
source
metadata_json
```
约束:
- 同一 `run_id``sequence` 单调递增。
- append 必须幂等,支持远程 daemon / plugin 重试。
- 未知 result type 可保存但 Host 只对已知类型执行副作用。
- 大 payload 仍应进入 sandbox/workspace,不直接塞入 result event。
- `usage_json` 保存 `AgentRunResult.usage` 原样结构;缺失表示 unknown,不等于 0。
### 4.3 Run Control API
Host 提供通用控制原语:
```text
run.create
run.get
run.list
run.events.page
run.cancel
run.append_result
run.finalize
```
语义:
- `run.create` 创建 Host-owned run 和授权快照。
- `run.append_result` 只允许受信 SDK/runtime 路径调用,必须绑定 run 创建时固化的授权快照,写入 `AgentRunEvent` 并触发 transcript/state/delivery 副作用。
- `run.finalize` 关闭 run,更新 terminal status。
- `run.cancel` 设置取消意图;同步 runner 通过 context/deadline 感知,远程 runner 通过插件/daemon 通道感知。
第一阶段可以只暴露给插件 runtime action,不一定先做公开 HTTP API。
### 4.4 Result Persistence In Orchestrator
当前 `AgentRunOrchestrator.run()` 已经处理:
```text
event -> binding -> context -> runner invocation -> result normalization
```
需要补齐:
- run 开始时创建 `AgentRun`
- 每个 `AgentRunResult` 进入 `AgentRunEvent`
- `run.completed` / 正常 generator 结束时标记 completed。
- `run.failed` / exception / timeout 标记 failed 或 timeout。
- terminal result 携带 usage 时,写入 `AgentRunEvent.usage_json` 并汇总到 `AgentRun.usage_json`
- `state.updated`、transcript 写入继续走现有 journal,但应与 `AgentRunEvent` 有可追踪关系。
### 4.5 Usage / Cost Accounting
SDK 侧 `AgentRunResult` 已提供可选 `usage` 字段,用于把不同 runner / external harness / provider-native event 的 token usage 归一到同一个 run result envelope。
语义:
- `run.completed.usage` SHOULD 表示本次 run 的最终聚合 token usage。
- `run.failed.usage` MAY 表示失败前已知的部分 token usage。
- 没有 usage 表示 upstream runtime 没有报告或 adapter 暂未接入;Host 不得按 0 计费或按 0 判断上下文消耗。
- Host 应把 event-level usage 原样写入 `AgentRunEvent.usage_json`,并在 terminal event 或 finalize 阶段汇总到 `AgentRun.usage_json`
- cost 应由 Host 根据 usage、runner/model identity、发生时间和价格表计算,写入 `AgentRun.cost_json`runner/provider 上报的 cost 只能作为非权威 telemetry 保留在 metadata 或 usage extra 中。
这层约束先解决协议位置和持久化位置;具体 ACP、remote daemon、local subprocess runner 如何从 native event 中抽取 usage,可在各插件后续适配。
### 4.6 Authorization Snapshot
异步或远程执行时,run 创建时必须固化授权快照:
- runner identity
- binding identity
- caller plugin identity
- resource policy
- allowed tools/models/files/knowledge bases/storage scopes
- state scopes
- conversation/thread/workspace scope
后续 append result、state API、history API 和 sandbox/workspace 文件访问都以这个 snapshot 校验,不重新扩大权限。
## 5. SDK 侧应新增的最小能力
SDK 不需要马上定义完整 daemon registry,但需要让插件和 runner 使用 Host run/result 能力。
### 5.1 Entities
新增或补齐:
```text
AgentRun
AgentRunStatus
AgentRunEvent
RunEventPage
RunCreateRequest / RunCreateResult
RunAppendResultRequest
```
这些是 Host control primitives,不替代 `AgentRunContext` / `AgentRunResult`
### 5.2 Proxy Methods
在 SDK proxy 中提供:
```python
create_run(...)
get_run(run_id)
list_runs(...)
page_run_events(run_id, cursor=None, limit=...)
cancel_run(run_id)
append_run_result(run_id, result, sequence=None)
finalize_run(run_id, status, error=None)
```
访问边界:
- 普通 AgentRunner 在同步 `run(ctx)` 内不一定需要直接调用这些 APIHost orchestrator 可自动记录。
- Platform 插件可以创建/查询/取消 run。
- AgentRunner 插件或 daemon bridge 可以 append/finalize 自己负责的 run。
- 外部 harness 仍不能直接调用 Host;必须经 SDK runtime / proxy / bridge。
### 5.3 Plugin-Daemon Heartbeat
远程 daemon 的初始心跳可以是 SDK / AgentRunner plugin 私有能力:
```text
daemon <-> AgentRunner plugin / SDK remote layer <-> LangBot plugin runtime <-> Host
```
Host 第一阶段只需要知道:
- 相关插件是否在线。
- run 是否有 progress/result。
- run 是否超时或取消。
如果后续需要跨插件共享 daemon 可用性,再把 heartbeat/registry 下沉为 Host 能力。
## 6. Platform 插件应负责什么
Agent Platform 插件可以负责:
- 管理哪些 agent 可用。
- 维护产品层 agent profile、项目、任务板、workflow、team。
- 订阅 EBA event,决定哪些 event 触发哪些 agent。
- 维护业务 queue:优先级、重试策略、人工审批、分配规则。
- 选择 runner / runtime / daemon。
- 在 Run Control API 落地后,调用 Host run API 创建、取消、查询执行。
- 展示 run status、result stream、文件引用、失败原因和审计。
Platform 插件不应负责:
- 在 Host Run Ledger 落地后,私有保存通用 run/result 事实源。
- 绕过 Host 直接写 transcript/state 或越权访问 sandbox/workspace 文件。
- 让外部 harness 直接访问 LangBot DB 或 Host 内部资源。
- 把某个业务队列语义强塞进 AgentRunner Protocol v1。
## 7. 与 EBA 的关系
EBA 做好后,事件流可以进入两种路径。
直接执行路径:
```text
EventGateway
-> EventRouter resolves AgentBinding
-> AgentRunOrchestrator.run(event, binding)
-> Host records AgentRun / AgentRunEvent (after Run Ledger lands)
-> delivery
```
Platform 插件编排路径:
```text
EventGateway
-> Platform plugin receives/subscribes event
-> plugin applies policy / business queue
-> plugin creates Host run (after Run Control API lands)
-> runner/plugin/daemon executes
-> Host records result and state
-> plugin displays / Host delivers
```
这两条路径最终应共享 Host run/result/state 事实源和 sandbox/workspace 文件边界。当前阶段可共享的是 event/transcript/state、sandbox 文件和同步执行链路;持久 run/result ledger 需要 Runtime Control Plane v2 Phase 1 补齐。区别在于是否有 Platform 插件参与产品化调度和业务队列。
## 8. 与 AgentRunner Protocol v1 的关系
本设计不改变 v1 的 runner 可见合同:
```text
AgentRunContext -> AgentRunner.run(ctx) -> AgentRunResult stream
```
必须保持:
- `AgentRunContext` 不塞入 daemon/worker/pod 细节。
- `AgentRunResult` 仍是 runner 输出的统一事件流。
- 普通 runner 不需要知道 task queue / runtime registry。
- 远程 harness 可以自管 session、tool loop、MCP、上下文压缩,但访问 LangBot 资源必须通过 SDK proxy / bridge。
- Runtime-managed execution 是 placement / transport 选择,不是普通 runner 协议的强制概念。
## 9. 分阶段实施建议
### Phase 1: Run LedgerFoundation Implemented
目标:Host 成为执行状态和结果事实源。
范围:
- `AgentRun` 表。
- `AgentRunEvent` 表。
- Orchestrator 自动创建/更新 run。
- Journal 持久化每个 `AgentRunResult`
- Run 查询和事件分页 API。
- SDK entities + proxy 方法。
复杂度:中等。
预计改动:
```text
Host: 12-20 个文件
SDK: 4-8 个文件
Tests: 8-15 个文件
```
### Phase 2: Platform Plugin Queue On Host Run PrimitivesControl Primitives Partially Implemented; Product Queue Pending
目标:Platform 插件管理业务 queueHost 提供 run/result/cancel 原语。
范围:
- `run.create`
- `run.cancel`
- `run.append_result`
- `run.finalize`
- result append 的 sequence/idempotency。
- 受权限保护的远程 append/finalize。
- Platform 插件可基于 Host run 构建任务板和调度体验。
复杂度:中等偏高。
预计改动:
```text
Host: 20-35 个文件
SDK: 8-14 个文件
Tests: 15-25 个文件
```
### Phase 3: Optional Host Execution Queue / Claim LeaseClaim Lease Primitive Implemented; Full Queue Pending
目标:当多个插件重复实现 claim/cancel/retry/recovery 时,再下沉执行队列到 Host。
范围:
- `queued/running/completed/failed/cancelled` 状态机扩展。
- `claim_run` / `lease_until`
- dispatch timeout。
- retry / orphan recovery。
- cancel propagation。
- 并发 claim 防重。
复杂度:高。
预计改动:
```text
Host: 35-55 个文件
SDK: 12-20 个文件
Tests: 25-40 个文件
```
### Phase 4: Optional Runtime RegistryMinimal Registry Implemented; Full Daemon Control Pending
目标:当 Host 需要统一管理多个 daemon / worker 时,再引入 runtime registry。
范围:
- runtime register / heartbeat / deregister。
- capability reportprovider、version、login status、workspace access、slot。
- runtime online/offline。
- runtime scoped auth。
- runtime audit。
- runtime gone recovery。
- task wakeup / long polling / websocket。
- 多 Host 实例下的 relay / distributed lock。
复杂度:很高。
预计改动:
```text
Host: 55-80+ 个文件
SDK: 18-30 个文件
Tests: 40+ 个文件
```
不建议现在直接进入此阶段。
## 10. 设计原则
- 先把 run/result 事实源做进 Host,再谈完整 runtime control plane。
- Agent Platform 产品做插件;Host 做基础设施。
- Host 不写业务调度策略,但要保存通用状态、结果、权限和审计。
- EBA event 不是 queuequeue 是执行生命周期问题。
- 业务 queue 可以先在 Platform 插件里;执行 queue 只有在复用需求明确后再下沉 Host。
- Daemon registry 不应污染 AgentRunner Protocol v1。
- 外部 harness 不直接访问 LangBot Host 或 DB。
- 所有 LangBot 资源访问必须走 SDK runtime / `AgentRunAPIProxy` / scoped MCP bridge。
- Docker / remote / local subprocess 只是 runtime placement,不是 runner 协议差异。
## 11. 非目标
当前阶段不做:
- 完整 Multica 式 runtime registry。
- Host 内置项目管理、任务板、agent team、workflow 产品逻辑。
- 把 daemon heartbeat / worker liveness 放进 `AgentRunContext`
- 把业务 queue 定义为 AgentRunner Protocol 字段。
- 让 Platform 插件私有保存 run/result 事实源。
- 让外部 agent/harness 直连 Host 内部资源。
## 12. 待定问题
- 已确认:Agent 与 Pipeline 都持久存在,并由 EBA 处理器 binding 指向其中之一;Pipeline 仅在 AI Stage 调用 runner 时投影一次性 `AgentBinding`,不会充当持久 Agent 的替代物。
- Platform 插件创建 run 时,是否传完整 `AgentBinding` snapshot,还是引用 Host-owned binding id。
- `AgentRunEvent` 与现有 `EventLog` / `Transcript` 的查询关系:直接 join,还是通过专门 view 聚合。
- `run.append_result` 的认证粒度:runner plugin identity、run token、scoped capability token,或 SDK runtime 内部 channel。
- 取消语义:同步 runner、external harness runtime/session 如何统一感知 cancel。
- 何时把插件私有 daemon heartbeat 提升为 Host `RuntimeLease`
- 若未来 Host 做 claim leasePlatform 插件业务 queue 与 Host execution queue 如何避免双队列混乱。
@@ -1,154 +0,0 @@
# Run Steering 与 Compaction CheckpointDesign Note
本文档记录两项 Host/runner 协作能力:**运行中消息注入(steering / follow-up**和
**压缩摘要持久化(compaction checkpoint**。两者来自官方 local-agent 对照
Pi agent harness`pi-mono/packages/agent`,下称 pi-agent-core)的差距分析:
local-agent 已移植 Pi 的事件生命周期、并行工具语义、hook 扩展点和压缩预算模型,
这两项需要 Host 协议、授权与 runner turn 边界协同才能闭环。
> 本文是设计备忘,不是 schema 事实源。涉及的数据结构最终落到
> [PROTOCOL_V1.md](./PROTOCOL_V1.md);上下文边界语义以
> [AGENT_CONTEXT_PROTOCOL.md](./AGENT_CONTEXT_PROTOCOL.md) 为准;
> run 持久化与控制原语以 [RUNTIME_CONTROL_PLANE_V2.md](./RUNTIME_CONTROL_PLANE_V2.md) 为准。
## 1. Run Steering / Follow-up(运行中消息注入)
### 1.1 问题
IM 场景下用户在 agent 运行中追加消息非常常见(补充信息、纠正方向、"算了别查了")。
EBA 先按事件选择一个 Pipeline 或 Agent 处理器;进入 AgentRunner 后,当前调用链是 `one AgentBinding -> one run_id -> one runner`
PROTOCOL_V1 §13):同会话的新消息要么等待当前 run 结束后触发新 run,
要么并发触发独立 run。两种行为都无法把新消息送进**正在执行的 tool loop**
用户体验是"agent 自顾自跑完过期任务,然后才看到新消息"。
cancelPROTOCOL_V1 §10)不解决这个问题:cancel 丢弃已完成的工作;
steering 是在保留当前进度的前提下改变后续方向。
### 1.2 Pi 的参考语义
pi-agent-core 区分两个队列,注入时机都在 turn 边界,不打断进行中的模型流或工具执行:
- **steering**:运行中插入。当前 assistant 消息的全部 tool call 完成后、
下一次模型调用前,注入排队的用户消息;模型在下一 turn 看到它们。
- **follow-up**:排队后续工作。仅当没有 pending tool call 且没有 steering 消息、
run 即将自然结束时检查;若有排队消息则注入并继续下一 turn,而不是结束 run。
两个队列各自支持 `one-at-a-time`(每次注入一条)和 `all`(一次注入全部)模式。
### 1.3 设计方向
职责划分遵循既有原则:Host 拥有事件路由和会话事实源,runner 拥有 turn 边界。
- **Host 侧**BindingResolver / dispatch 层识别"同 conversation 存在 active run
且 runner 声明支持 steering"的新消息事件,将其写入 run-scoped steering queue
并标记该事件已被在途 run 认领(不再触发新 run,避免破坏 §13 的基数约束)。
事件仍照常进 EventLog / Transcript(事实源不变,改变的只是触发行为)。
- **Runner 侧**:在 turn 边界(tool batch 完成后、下一次模型调用前,以及 run
即将自然结束前)通过 run-scoped pull API 拉取 pending steering 输入,
注入 working context。local-agent 的 `AgentLoopHooks.prepare_next_turn` /
`should_stop_after_turn` 已预留了对应的注入点。
- **能力协商**runner manifest 声明 `steering` capability(参照 PROTOCOL_V1 §4.3);
未声明的 runner 保持现状(新消息按现有规则另起 run)。
- **回执**:被 steering 消费的事件通过 EventLog 审计。原始 `message.received`
记录在 `metadata.steering` 标记 queued/absorbed 与 `claimed_by_run_id`
runner 成功 pull 后,Host 追加 `steering.injected` 记录并引用源事件。
run 结束时仍未被 pull 的已 claim 输入,Host 追加 `steering.dropped` 记录作为
dispatch 终态;原始 Transcript 事实不删除。
Transcript 继续只表示会话事实,不扩展 dispatch 行为字段。
已落地的协议面(最终定义归 PROTOCOL_V1):
1. `ContextAccess.available_apis` 增加 steering pull 能力位。
2. `AgentRunAPIProxy` 增加 steering 拉取 action:默认 `mode=all`Host 保序返回全部
pending 输入;`one-at-a-time` 仅作为 runner 主动节流选项。
3. dispatch 层的"认领"规则:`message.received` 可被同 conversation 的 active run
吸收,原事件写 EventLog / Transcriptdispatch 行为写入 EventLog metadata。
4. Host 对单 run steering queue 设置内存上限,队列满时不再 claim 新消息,消息回到
正常 dispatch 路径,避免 active run 无限吞入同会话输入。
### 1.4 边界
- 不引入 Host 替 runner 做 prompt 拼接:Host 只递队列,注入位置和格式由 runner 决定。
- 不与 observer / fan-out 混淆:steering 仍是单 run 内的输入补充,不产生第二个 runner。
- 远程 / 外部 harness runnerclaude-code、codex 等)若其底层 session 自带
steering 能力,adapter 可以直接转发;协议面保持一致。
## 2. Compaction Checkpoint 持久化
### 2.1 问题
local-agent 当前是无状态 runner:每次 run 重新拉取 transcript 尾部
(默认 50 条)、重新估算 token、重新生成压缩摘要。后果:
- 长会话中每 run 重复压缩计算,摘要每次重新生成,不同 run 之间措辞漂移,
对 provider KV cache 不友好(AGENT_CONTEXT_PROTOCOL §"Summary checkpoint 稳定"
已写明期望:只有压缩发生时才产生新 checkpoint)。
- 历史一旦超过 fetch limit,更早的内容永久不可见——没有 checkpoint 记录
"已压缩到哪里、压缩出了什么"。
pi-agent-core 把 compaction 条目持久化进 session tree:摘要带
`tokensBefore` 和覆盖范围,后续 turn 直接复用,只在再次越过阈值时增量压缩。
### 2.2 现状盘点
协议面和主消费路径已具备:
- State / Storage API 已定义(PROTOCOL_V1 §8 "State / Storage"),
且 AGENT_CONTEXT_PROTOCOL 已点名 `summary.checkpoint` 是 state 的预期用法。
- Host 会根据 binding state policy 暴露 `ContextAccess.available_apis.state`
- local-agent 会在 state API 可用时读取/写入 `runner.compaction.checkpoint`
缺失、schema 不匹配、conversation 不匹配或游标失败时回退尾部历史拉取。
- LLM 生成摘要**不依赖**本项 Host 能力——runner 用已授权的 `invoke_llm`
即可生成;checkpoint 只解决"存下来、下次复用"。
### 2.3 设计方向
- **存放位置**statescope=`conversation`(小 JSON,符合 PROTOCOL_V1 §8
对 state/storage 的边界建议)。若未来摘要膨胀,超出部分放 storage 并在
state 中留引用。
- **key 约定**`runner.compaction.checkpoint`runner 命名空间内)。
- **内容约定**schema 落 PROTOCOL_V1 或 runner 文档,此处只列语义):
- `schema_version`
- `summary`:压缩摘要文本(LLM 生成或确定性生成)
- `covers_until`:已被摘要覆盖的 transcript 游标(seq / message id),
是增量压缩和"从哪继续拉历史"的锚点
- `tokens_before` / `created_at`:诊断与失效判断
- **消费流程**run 开始时读 checkpoint → 只拉取 `covers_until` 之后的
transcript → 压缩触发时基于旧摘要增量生成新摘要、写回新 checkpoint。
checkpoint 缺失或解析失败时回退到现行为(全量拉尾部),保证向后兼容。
- **失效规则**`covers_until` 在 Host transcript 中不存在(会话被清理 / 重置)
即作废;runner 不得信任跨 conversation 的 checkpoint。
- **授权**Host 对声明需要 state 的 runner binding 开启
`available_apis.state`;校验沿用现有 run-scoped state 校验
scope、key、value 大小、JSON 可序列化,见 PROTOCOL_V1 §7.2 对
`state.updated` 的要求)。
### 2.4 相关但独立的工作
- **tokenizer / usage metadata 透传**runner 目前用 chars/4 启发式估 token
对 CJK 偏低 3-4 倍,压缩触发系统性偏晚。Host 应在模型响应或
`ctx.runtime.metadata` 透传 provider usageprompt/completion tokens)与
model context windowLiteLLM model-info 工作)。该项不阻塞 checkpoint
落地,但决定压缩触发的准确性。
## 3. 实施拆分
| 项 | 归属 | 依赖 |
| --- | --- | --- |
| steering queue、事件认领、基础审计 | LangBot Hostdispatch / binding 层) | 已落地,含队列上限与未消费 dropped 终态 |
| steering pull API + capability 位 | PROTOCOL_V1 + SDK proxy | 已落地 |
| turn 边界拉取与注入 | langbot-local-agent | 已落地 |
| local-agent 对 state API 的 checkpoint 读写 | langbot-local-agent | 已落地 |
| checkpoint key / 内容 / 失效约定 | PROTOCOL_V1 + local-agent README | 已落地 |
| LLM 压缩摘要生成 | langbot-local-agent | 已落地(`invoke_llm`,失败回退确定性摘要) |
| usage / context-window metadata 透传 | LangBot Hostmodel 层) | LiteLLM model-info |
剩余工作应优先补 usage / context-window metadata。streaming delivery 衔接依赖
`ctx.delivery` 编辑/追加语义,不建议在协议能力缺失时硬编码。
## 4. 开放问题
- streaming delivery 下 steering 注入后,前序 turn 已流出的内容与新 turn
输出在 IM 消息编辑面的衔接(涉及 `ctx.delivery` 能力,待 delivery 演进定)。
- checkpoint 是否需要 Host 侧主动失效通知(如会话清空时删除对应 state key)。
当前实现靠 runner 读取时校验并回退,功能不阻塞。
@@ -1,211 +0,0 @@
# Agent Runner Security Boundary
本文档记录 agent-runner 插件化后的安全边界和最小护栏。
## 状态
**当前结论:不采用高强度监管模型。**
LangBot 的目标不是托管一个强隔离、不可信 code runner 平台。AgentRunner 插件,尤其是 ACP / Claude Code / Codex / OpenCode / Kimi Code 这类外部 harness,默认视为 **operator-owned execution**:用户或部署者显式配置并承担其文件系统、进程、网络、workspace、provider 登录态和 native tool 风险。
LangBot 需要负责的是保护 **LangBot 自己持有的资源**,包括模型、知识库、LangBot tools、history、event、state、plugin/workspace storage、sandbox/workspace 文件访问等。只要这些资源访问是 run-scoped、permission-scoped、可校验、可诊断的,当前阶段即可接受。
这意味着:
- 不要求 LangBot 在应用层实现完整 OS sandbox、VM、cgroup、seccomp、CPU / memory / network quota。
- 不要求为 ACP runner 做复杂审批流;用户选择 ACP runner 即表示显式 opt-in。
- 不要求在非 Docker 进程部署里做强监管;只要文档明确风险归属即可。
- Docker / K8s 可以提供部署级隔离,但不是 LangBot agent-runner 协议发布的前置条件。
- 不能宣传 LangBot 已经提供 managed sandbox;除非未来真的提供受管执行环境。
## 责任边界
### LangBot Host 负责
- **资源授权**:根据 runner manifest permissions、binding resource policy、run scope 生成本次 run 可访问的资源快照。
- **运行期校验**:所有带 `run_id` 的 SDK / Host action 必须校验 active run session、caller plugin identity、resource id 和 operation。
- **Scoped projection**:只把授权后的资源摘要、MCP server config、context、attachment/path ref、state snapshot 投影给 runner。
- **LangBot 文件路径约束**LangBot 自己 staged 和读取的文件必须限制在声明 root 内,防止 path escape。
- **基础 secret 策略**:不要主动把 LangBot 持有的 API key / token / secret 投影给 runner;日志和错误里做常见 secret 字段脱敏。
- **基础运行约束**:提供 timeout、取消传播、输出大小限制或错误映射的基础能力。
- **audit-lite**:记录 event、run id、runner id、binding、资源授权摘要、关键失败、state/file/transcript 事实。
### Runner Plugin 负责
- 遵守 Host 下发的 `ctx.resources``ctx.context.available_apis`、runner config 和 state policy。
- 把 LangBot 资源投影成目标平台可消费的形式,例如 MCP config、context prompt、HTTP header、run token。
- 不绕过 SDK / Host action 直接访问 LangBot 内部资源。
- 对自己启动的外部进程做合理封装,包括参数构造、timeout、取消、输出解析和错误映射。
- 清楚记录自身 README 中的 provider 风险、部署假设和限制。
### 部署者 / 用户负责
- ACP / external harness 的 workspace 内容、文件系统访问、进程权限、网络访问、provider-native tool 权限。
- Docker / K8s 的 image、volume、secret、network policy、resource limit、namespace、service account 配置。
- 本机进程部署时的 OS 用户权限、PATH、HOME、CLI 登录态、全局配置和外部 MCP 配置。
- 是否允许 runner 对某个目录执行真实写操作。
### 外部 Harness 负责
Claude Code、Codex、OpenCode、Kimi Code、Gemini CLI 等外部工具继续使用自己的权限模型、MCP 加载策略、session/resume、sandbox 或 approval 能力。LangBot 不承诺约束这些工具对其所在容器或宿主 OS 用户本来可访问资源的能力。
## 部署场景策略
| 场景 | LangBot 策略 | 不由 LangBot 承担 |
| --- | --- | --- |
| 普通进程部署 | 文档提示 operator-owned executionHost 只保护 LangBot 资源。 | 阻止外部 CLI 读取同一 OS 用户可访问的文件、进程、HOME、全局 CLI 配置。 |
| Docker / K8s 部署 | 继续使用相同 Host 资源边界;容器隔离由部署环境提供。 | 应用层重复实现容器/VM/cgroup/seccomp/network quota。 |
| ACP runner | 用户显式选择 runner 和 workspaceLangBot 注入 scoped MCP / run token。 | ACP CLI native tools、workspace 写入、provider 登录态和外部 MCP 行为。 |
| 外部 SaaS runner,例如 Dify | LangBot 通过 run token / gateway 限制 LangBot 资产访问。 | SaaS 平台内部 agent 执行策略、模型工具消息格式、平台侧日志。 |
| 未来 managed runner | 只有当 LangBot 明确提供受管执行环境时,才需要单独定义强隔离 SLA。 | 当前协议闭环不承诺 managed sandbox。 |
## 最小护栏
以下是当前阶段需要维持的最小要求。它们是保护 LangBot 资源边界的要求,不是完整监管外部进程的要求。
### Resource Permission Boundary
每次 run 前必须冻结授权快照:
- runner manifest permissions 是资源访问上限。
- binding resource policy / runner config 决定本次实际授权。
- runtime action 按 `run_id` + `caller_plugin_identity` + resource id + operation 校验。
- manifest permissions 只约束 LangBot 持有资源,不约束 external harness native tools。
当前实现方向是正确的:`AgentRunSessionRegistry` 保存 run-scoped snapshot`plugin/handler.py` 对模型、工具、知识库、history、state、storage 等 action 做运行期校验,sandbox/workspace 文件访问由 scoped tool 边界控制。
**Skill 读写门控(不可弱化)**pipeline-visible 的 skill 一次性以 `rw` 挂进同一 sandbox,mount 层不区分「可见」与「已激活」;写类 native 操作(write/edit/exec)只放行 activated skill,读类放行 visible + activated——这层区分等同资产授权语义,必须保留。skill 全 tool 化后尤其注意:「都是 tool」不等于「只控资产授权即可」,native 层的 visible/activated 门控不能砍。可弱化的只是 realpath 越界字符串检查(有 chroot/namespace 兜底)。
### MCP / Asset Gateway Boundary
LangBot MCP / asset gateway 只暴露当前 run 授权的工具面:
- `langbot_list_assets`
- `langbot_get_current_event`
- `langbot_history_page`
- `langbot_retrieve_knowledge`
- `langbot_get_tool_detail`
- `langbot_call_tool`
外部平台需要使用短期 `run_token` 或 Authorization bearer token。token 缺失、错误或过期时必须拒绝访问。
不要求当前阶段实现 admin 级 MCP allowlist、dangerous tool approval 或复杂审批流。是否注册外部 MCP provider 是部署者/用户行为。
### Workspace / Path Boundary
LangBot 只需要约束自己管理的路径:
- Host staged 文件必须校验 `realpath` 和 root containment。
- Attachment/file metadata 不应暴露 Host-only storage key / host path。
- Context 文件、sandbox/workspace 文件如由 LangBot 创建,应放在可清理的位置。
用户配置给 ACP runner 的 workspace 不属于 LangBot 的强监管范围。Docker/K8s 下依赖 volume 挂载边界;普通进程部署下依赖 OS 用户权限和用户自担风险。
### Secret Handling
这里的 secret 指 API key、provider token、run token、MCP token、platform secret、数据库密码等。
当前阶段只要求基础策略:
- LangBot 不主动把自己持有的 secret 投影给 runner,除非这是 runner config 明确需要的外部服务凭据。
- run token 是短期、run-scoped 的,不应长期保存。
- 日志、错误、transcript、attachment/file metadata 尽量避免打印常见 secret 字段。
- 配置 UI / API 返回时继续沿用现有 secret masking 规则。
不要求当前阶段实现完整 DLP、全链路敏感数据追踪、secret lineage 或自动轮换体系。
### Process / Runtime Bounds
LangBot 需要提供基本可控性:
- Host run deadline / runner timeout。
- runner 侧请求 timeout。
- generator close / cancel 传播。
- 输出和 inline payload size 上限。
- 错误映射为受控 runner failure。
不要求 LangBot 为外部 harness 实现 CPU、内存、磁盘、网络、进程树强隔离。需要这些能力时由 Docker/K8s、systemd、容器平台或用户机器策略提供。
### UI / Admin Surface
前端可以展示 runner 权限摘要,但它是信息披露,不是审批系统。
权限摘要指 runner manifest 声明的 LangBot 资源权限,例如:
- `tools.detail`
- `tools.call`
- `knowledge_bases.retrieve`
- `history.page`
- `storage.plugin`
当前阶段不要求强制弹窗、管理员审批、dangerous tool approval 或生产禁用开关。可以在 runner 配置区展示简短提示:此 runner 能访问哪些 LangBot 资源,外部 harness 执行风险由用户/部署者承担。
### Audit Lite
需要记录足够排查问题的事实:
- run id、runner id、binding、event。
- 授权资源摘要。
- state update、file write/read event、transcript message。
- MCP / pull API 拒绝时的 warning。
- steering queued / injected / dropped。
不要求当前阶段建立独立安全审计产品、审批记录系统或 SIEM 级事件模型。
## 降级后的检查表
| 项目 | 当前要求 | 状态判断 |
| --- | --- | --- |
| Path isolation | 只约束 LangBot 管理的 context/sandbox 文件路径;runner workspace 归用户/部署环境。 | Minimal required |
| Permission boundary | 必须保护 LangBot 资源;不约束外部 CLI native 能力。 | Required |
| Secret handling | 基础不投影、基础 masking、run token 短期化。 | Basic required |
| MCP policy | run-scoped token + scoped tool surface;无复杂审批。 | Required |
| Skill access policy | skill 通过 Host 授权 tool 暴露(发现 / activate / register / native exec 走统一 tool 授权);**native 层 visible(只读)vs activated(可写)门控不可弱化**——所有 pipeline-visible skill 以 `rw` 挂进同一 sandbox,读写区分全靠 native 层;harness-native skill 文件不作为 LangBot 安全边界。 | Required |
| Process isolation | 由 Docker/K8s/用户机器负责。 | Out of scope |
| State lifecycle | scope 隔离、JSON size limit、基础 cleanup primitive。 | Basic required |
| Audit | 记录运行事实和拒绝原因。 | Audit-lite |
| UI / Admin control | 权限摘要可展示;不要求审批流。 | Optional |
| Test matrix | 覆盖 run auth、MCP token、permission deny、timeout、sandbox path、state size。 | Focused tests |
## 当前实现快照
截至 2026-06-15,已有实现覆盖:
- SDK typed AgentRunner manifest、capabilities、permissions。
- Host resource builder 按 manifest permissions 和 binding policy 生成 `ctx.resources`
- Active run session snapshot 和 `caller_plugin_identity` 校验。
- History / event / state / tool / knowledge runtime action 的 run-scoped 校验。
- Sandbox file path `realpath` + root containment。
- Persistent state scope 隔离和 JSON size limit。
- SDK-owned MCP bridge 和 long-lived asset gateway。
- Dify / ACP runner 对 LangBot asset gateway 的接入。
- Runner timeout、Dify HTTP timeout、ACP startup / initialize / request timeout。
仍可继续优化但不阻塞当前发布的事项:
- 前端展示 runner LangBot 资源权限摘要。
- 常见 secret 字段 redaction 收敛成统一 helper。
- Context/sandbox file TTL cleanup 调度。
- 更完整的 MCP 调用 audit。
- 更好的文档提示:ACP runner 是 operator-owned execution。
## 非目标
以下不属于当前 agent-runner pluginization 的安全目标:
- 防止 ACP / external harness 修改其 workspace。
- 防止外部 CLI 读取同一容器或 OS 用户本来可读的文件。
- 管控 external harness 的 provider-native tools、approval、MCP、browser、shell。
- 在 LangBot 应用层实现 VM / container / cgroup / seccomp / network policy。
- 为 Docker/K8s 部署替代平台自身的 secret、volume、network、resource limit 管理。
- 实现企业级审批系统、SIEM、DLP 或安全运营面板。
## 发布口径
可以对外说明:
> AgentRunner 插件通过 run-scoped authorization 和 scoped MCP gateway 保护 LangBot 持有资源。外部 code harness 的执行环境由用户或部署平台负责隔离;LangBot 当前不提供 managed sandbox。
不能对外说明:
> LangBot 已经安全沙箱化 Claude Code / Codex / OpenCode 等外部 runner。
-64
View File
@@ -1,64 +0,0 @@
# AgentRunner Pluginization Status
本文档是 `docs/agent-runner-pluginization/` 的状态事实源。协议 schema 仍以 [PROTOCOL_V1.md](./PROTOCOL_V1.md) 为准;测试步骤以 [AGENT_RUNNER_QA_GUIDE.md](./AGENT_RUNNER_QA_GUIDE.md) 为准;安全发布门槛以 [SECURITY_HARDENING.md](./SECURITY_HARDENING.md) 为准。
状态快照日期:2026-07-15。
## 实现状态
| 领域 | 状态 | 说明 |
| --- | --- | --- |
| SDK manifest schema | Done | `AgentRunnerManifest` 包含 typed `capabilities` / `permissions`;未知 capability / permission key 禁止进入 typed model。 |
| Runner discovery | Done | Runtime 返回 typed manifestHost registry 校验单个 runner,失败 warning + skip,不影响其它 runner。 |
| Host resource authorization | Done | `ctx.resources``ctx.context.available_apis` 由 manifest permissions 与 binding policy / run scope 求交后生成。 |
| Run authorization snapshot | Done | active run session 冻结 run-scoped resources 与 available APIsruntime handler 按 snapshot 校验 pull API。 |
| Result payload validation | Done | Wire 保持 `{type, data}`Host 对投递/副作用类 payload 严格校验,tool-call telemetry 宽松,未知 type 忽略并 warning。 |
| Old built-in runners | Done | 旧 `src/langbot/pkg/provider/runners/*``RequestRunner` 路径已从本分支删除。 |
| Official runner manifests | Done | `local-agent`、ACP / Claude Code / Codex 外部 harness runner、外部服务 runner 已重新声明真实生效的 LangBot resource permissions。 |
| Skill 链路 | Unit-pass; WebUI E2E pass | 已按 **skill 全 tool 化** 收敛:发现走 `list_skills` / `langbot_list_assets` 和 skill resources`activate` / `register_skill` 走统一 tool 授权;`skill_authoring` capability 降级为便捷开关。`activate` 会 best-effort 写入 conversation-scope `host.activated_skills`,后续 run 通过当前 pipeline-visible skill cache 恢复。新注册 Skill 在当前 Query 内立即获得临时可见性;Docker `exec` 产生的宿主侧不可写文件由 `write` / `edit` 回退到 Box 执行。2026-07-15 真实 LocalAgent Debug Chat 已完成创建、注册、同 Query 激活、编辑和执行闭环;非流式 runner turn 只向下游 Pipeline 产出一次,工具中间结果不再拆成额外 Bot 气泡。 |
| Runtime Control Plane v2 foundation | Partial | Host-owned `AgentRun` / `AgentRunEvent` ledger、orchestrator 自动建账、result event persistence、run get/list/event page/cancel/append/finalize actions 已落地;`agent_run:admin` / `runtime:admin` 控制权限、最小 runtime register/heartbeat/list/reconcile 和 run claim/renew/release 原语已落地。完整 Agent Platform 产品形态、daemon supervisor、任务唤醒/长轮询/WebSocket、分布式 runtime 管控仍未完成。 |
| Security boundary | Done | 当前口径降级为轻量边界:LangBot 保护自身持有资源;external harness 的 OS / process / network / workspace 风险由用户或部署环境承担;managed sandbox 不是当前承诺。 |
| Steering control path | Done | claim 异常不再逃逸 consumer loopqueue 有上限;未 pull 的 claimed 输入在 run 结束时写 `steering.dropped` 审计终态。 |
| SDK v1 contract closure | Done | SDK 提供 `AgentAPIError` / `AgentAPIException`、typed `SteeringPullResult`、未知 result type 宽容解析、result `sequence` 注入与取消传播。 |
| EBA processor routing | Done; release gate 5/5 pass | Bot `event_bindings`、Pipeline / Agent 平级路由、WebUI dry-run / 合成测试 / 状态、OneBot 非消息事件到 Agent 及平台回复已闭环;隔离空白实例已验证从 Space 安装并注册 LocalAgent。 |
| Structured interactions | Cross-repo unit-pass; provider E2E pending | Host 已完成 `interaction.requested` 白名单、持久化 callback correlation、TTL/作用域/幂等校验和 Pipeline/Agent 原处理器恢复;六个平台已接入按钮/单选投递,Lark 和 DingTalk 进一步支持原生单字段 `text` / `textarea` / `number` / `select` 控件。SDK typed contract、通用 Runner 脚手架和 DifyAgent `workflow_paused` plugin-storage continuation 已落入对应仓库。 |
## Spec 与实现已知差距
- `action.requested` 是严格白名单协议面:当前只执行 `interaction.requested`;其它 action 仍只记录 telemetry,不提供通用 platform action executor。
- 结构化交互 SDK typed contract 与 DifyAgent continuation 已实现;SDK 正式发布、真实 Dify 凭据 E2E,以及需要长驻双向进程的 Claude Code 权限确认仍是后续验收项。Host 不持有 provider 私有 token。
- State 与 storage 的长期类型边界仍可继续收窄;当前合同只要求 JSON-safe state 与受控 storage API。
- `ToolResource.parameters` 已作为 best-effort full schema 由 Host 在构造 `ctx.resources` 时一次塞齐;无 schema 时 runner 仍需兼容 `parameters=None` 或按需调用 detail API。
- EventLog / Transcript 已提供显式 cleanup primitive;长期 retention 默认值、TTL 调度接入和 sandbox/workspace 文件清理仍是运维收尾项,应在 Runtime Control Plane 产品化前补齐。
- External harness 的 native shell / filesystem / CLI / MCP 权限不受 manifest permissions 约束;manifest permissions 只约束 LangBot 持有的资源访问。
- LangBot 当前不承诺 managed sandboxexternal harness 的 OS/process/network quota、workspace GC、provider-native tool 权限由用户或部署环境承担。
- Runtime Control Plane v2 当前只落地 Host 事实源和控制原语;还没有内置 Agent Platform UI、业务队列、daemon 进程托管、runtime wakeup channel、跨 Host 分布式锁或 provider 登录态诊断。
## Runner 验收状态
| Runner | 状态 | 最近证据 |
| --- | --- | --- |
| `plugin:langbot-team/LocalAgent/default` | Unit-pass; Marketplace UI pass; Debug Chat E2E pass | 2026-07-12 隔离 first-run 实例从真实 AgentRunner catalog 安装 `langbot-team/LocalAgent` 0.1.0Host 注册 `plugin:langbot-team/LocalAgent/default`,Wizard 自动选中并解锁后续操作。2026-07-15 `2026-07-15-08-44-10-770-08-00-sandbox-skill-authoring-edit-existing-e2e` 使用真实 `gpt-5.5` 完成 Skill 创建、注册、同 Query 激活、已激活包编辑与脚本执行;三阶段 UI、浏览器诊断和结构化文件系统检查全部通过,每阶段恰好新增一个 Bot 气泡,p95 14.6 秒、错误率 0。 |
| `plugin:langbot-team/ACPAgentRunner/default` | Unit-pass; Debug Chat E2E pass | 2026-07-15 从本地 0.1.4 发布包安装并注册 PascalCase runnerremote-ssh Claude ACP 通过反向隧道调用 run-scoped `langbot_get_current_event`,97.8 秒返回可见结果;Host 将增量 delta 和 `message.completed` 聚合为一个完整 Bot 气泡。 |
| `plugin:langbot-team/ClaudeCodeAgent/default` / `plugin:langbot-team/CodexAgent/default` | Unit-pass; E2E pending | 通过 runner 仓库单测覆盖 session、run_id 注入和 LangBot MCP gateway;真实 harness E2E 取决于对应运行环境、CLI/daemon 可用性和 provider 登录态。 |
| Dify | Human-input unit-pass; credential E2E pending | `langbot-agent-runner/dify-agent` 已实现 `workflow_paused`、原子字段/确认交互、plugin-storage continuation、Dify submit/events 恢复与再次暂停;真实 Dify 凭据 E2E 待执行。 |
| n8n / Coze / DashScope / Langflow / Tbox / DeerFlow / WeKnora | Unit-pass; credential smoke optional | 2026-06-13 plugin layout / parser tests 通过;真实服务凭据 smoke 非每轮必跑。 |
## Host / SDK 验收状态
| 范围 | 状态 | 最近证据 |
| --- | --- | --- |
| LangBot Runtime Control Plane v2 foundation | Unit-pass; EBA release gate 5/5 pass; AgentRunner preflight pass | 2026-07-12 `eba-functional-20260712-release-gate-rerun` 通过 Quick Start 场景筛选、隔离实例 Runner Marketplace 安装、Runner 健康状态、事件路由 dry-run / 合成派发,以及真实 OneBot `group.member_joined` → Agent → `send_group_msg` 链路。2026-07-15 AgentRunner release preflight 16 项通过、0 warningfixture contract、5 类 behavior matrix、ledger schema / async DB readiness / 100-run stress / 120-run 8-worker contention / claim-lease-auth concurrency、SDK runtime chaos 探针全部通过。 |
| Host Skill / native tool integration | Unit-pass; WebUI E2E pass | 2026-07-15 provider / native / Skill / monitoring 定向测试 67 项通过,Pipeline / Chat / Wrapper 定向测试 61 项通过,Skills CLI 105 项通过;真实 Debug Chat 验证 `register_skill` 后同 Query `activate` 成功,监控工具调用不再把 SQL 行误取为字符串,结构化 JSON 文件检查不依赖格式空格,非流式多阶段 runner 结果只生成一个最终 Bot 气泡。 |
| SDK AgentRunner control entities / proxy | Unit-pass | 2026-06-23 SDK `tests/api/entities/builtin/agent_runner``tests/api/proxies``tests/api/test_agent_tools_mcp_bridge.py``tests/runtime/plugin/test_mgr_agent_runner.py``tests/runtime/test_pull_api_handlers.py``tests/runtime/io/handlers/test_plugin_handler.py`、EBA event entities 和 message tests 通过,覆盖 typed entities、AgentRunAPIProxy、MCP bridge、runtime manager 与 pull API handlers。 |
## 历史高价值记录
历史报告已合并为本状态页和 QA 指南,不再保留单独进度文档。后续若需要追溯,优先查看 `langbot-skills/reports/` 下的原始执行报告。
截至 2026-05-29,已有本地 smoke 证明:
- `local-agent` 可以通过 Pipeline Debug Chat 走插件化 `AgentRunOrchestrator` 主链路。
- 外部 harness runner 可以通过同一条 `run(event, binding)` 路径执行;当前官方实现已收敛到 ACP / Claude Code / Codex 等直接 runner 插件。
这些记录只证明本地协议闭环可用,不代表 LangBot 提供 managed sandbox 或 external harness OS 级隔离。
@@ -1,342 +0,0 @@
# EBA 产品化与发布计划
> 状态:规划草案,2026-07-01
>
> 范围:将已经合并的 AgentRunner 插件化和 Event Based Agent 适配器工作产品化,使非技术用户也能快速上手。本文聚焦产品缺口、发布门禁和 SaaS 多命名空间租户能力。本文不引入新的协议 schema;协议事实仍以 [PROTOCOL_V1.md](./PROTOCOL_V1.md) 为准,Host 模型事实仍以 [HOST_SDK_INFRASTRUCTURE.md](./HOST_SDK_INFRASTRUCTURE.md) 为准。
## 1. 产品方向
当前技术方向是正确的:LangBot 应该把平台输入视为事件,为每个事件解析出一个有效路由,并通过 AgentRunner Host 边界调用一个处理资产。但这还不是一个非技术用户无需理解内部架构就能采用的产品。
产品层模型应当是:
- **机器人(Bot)**:平台连接与事件路由入口。机器人拥有适配器凭据、平台权限、入站事件可见性,以及这些事件的路由表。
- **处理器(Processor)**:可复用的事件处理资产。当前处理器类型包括 **Agent****Pipeline**。未来可以增加 **Workflow**
- **Pipeline**:一等的无代码消息处理器,通过完整 Stage 链提供预处理、AI、后处理、扩展和输出控制。Pipeline 只处理消息,应当只能绑定到 `message.*` 事件。
- **Agent**:由 runner 驱动的事件优先处理器。Agent 可以根据自身声明的事件支持范围处理消息事件和非消息事件。
- **Solution(方案包)**:未来的分发/导出单元,包含处理器、路由模板、依赖清单、变量和文档。Solution 不应包含具体机器人凭据、租户密钥或已安装资产的 UUID 绑定。
`EBA` 是内部工程术语。它可以出现在内部设计文档中,但不应出现在主要产品流程里。面向用户的语言应优先使用“频道”“事件路由”“处理器”“消息流水线”“自动化”“路由模板”等表达。
## 2. 当前基础
已合并分支已经具备内部冒烟测试所需的技术基础:
- EBA 适配器可以把平台活动规范化为稳定的 Host 事件名。
- 机器人可以持久化 `event_bindings`,并将事件路由到 Agent、Pipeline 或丢弃目标。
- 旧消息输入可以投射到标准的 `message.received` 事件路径。
- Pipeline 仍可作为只处理消息的无代码处理器使用。
- AgentRunner 插件化提供了 Host 与 Runner 的契约、事件优先上下文、结果流和运行时集成边界。
- 官方 local runner 和 external runner 插件可以验证 runner 行为已经不再硬编码在 LangBot Core 中。
- WebUI 已具备 Agent/处理器管理页面,以及机器人侧事件路由页面。
- 机器人事件路由已具备 dry-run 诊断、运行时状态展示,以及安全的合成测试事件派发;测试事件会走已保存的 runtime 路由,但抑制真实平台出站动作。
- MCP 工具面已暴露机器人事件路由状态查询和合成测试事件派发,便于 QA agent 或外部调试工具复用。
这是一个技术收敛里程碑,还不是产品就绪版本。
## 3. 距离非技术产品的缺口
### 3.1 概念负担
当前用户仍需要理解过多内部概念:EBA、适配器事件名、runner 标识、插件运行时健康状态、事件模式、优先级和绑定目标。面向非技术用户的产品应当通过意图和结果来引导:
- “收到一条消息时,用这个处理器回复。”
- “有新群成员加入时,发送欢迎语。”
- “收到好友请求时,让 Agent 判断是否接受。”
`group.member_joined` 这类原始事件模式应继续保留在高级模式中,但默认 UI 应按友好名称和平台能力对事件进行分组。
### 3.2 上手路径
首次使用路径应当从用例开始,而不是从架构开始:
1. 选择一个频道。
2. 连接账号或 webhook。
3. 选择该频道支持的事件预设。
4. 选择或创建一个处理器。
5. 发送测试事件。
6. 如果失败,阅读简单的运行轨迹。
当前产品仍假设用户能诊断后端、插件运行时、Box 运行时、适配器和 runner 插件是否都已连接。对于开发者这是可接受的,但对非技术用户不可接受。
### 3.3 适配器就绪度
每个适配器都需要产品能力清单,而不仅是工程实现:
- 支持的事件列表及友好标签;
- 支持的出站动作;
- 所需凭据和配置步骤;
- 本地部署、自托管、SaaS 可用性;
- 测试信号可用性;
- 废弃/遗留状态;
- 已知限制。
废弃适配器应在产品中明确标记为“已废弃”或“遗留”。新的事件型适配器应按频道名称和能力描述,而不是使用 EBA 缩写。
### 3.4 处理器体验
处理器页面应管理可复用的 Agent 与 Pipeline,而不是让用户一次性理解所有事件路由决策。
- 创建 Agent 时应提供有倾向性的 runner 模板。
- 创建 Pipeline 时应继续保持无代码消息流水线路径。
- 未来 Workflow 的执行语义稳定后,可以作为另一种处理器类型引入。
- 支持的事件范围应作为能力信息和高级约束展示,而不是作为创建流程的主要概念。
- 依赖健康状态应可见:runner 插件是否安装、运行时是否连接、所需模型是否配置、所需资源是否可访问。
### 3.5 机器人事件路由体验
机器人页面应成为平台特定事件路由的主要配置位置,因为平台事件在已连接频道的上下文中最容易被用户理解。
最低产品要求:
- 基于适配器能力生成友好的事件选择器;
- 目标选择器按事件兼容性过滤;
- 对非消息事件隐藏 Pipeline;
- 对重叠路由给出冲突警告;
- 优先级先用视觉方式解释,而不是首先展示原始数字;
- 提供路由测试按钮,可以注入或重放样例事件;
- 每条路由展示状态,包括最近一次匹配的 run 和最近失败原因;
- 提供安全的兜底路由,包括显式丢弃。
### 3.6 可观测性
非技术用户需要的是简短运行轨迹,而不是原始日志:
```text
收到事件 -> 命中路由 -> 启动处理器 -> 动作已投递
```
当失败发生时,UI 应指出失败层级:
- 频道未连接;
- 适配器不支持该事件;
- 没有路由命中;
- 处理器已禁用;
- runner 插件不可用;
- 模型/资源缺失;
- 投递权限被拒绝。
### 3.7 文档和模板
发布需要面向产品场景的文档和模板:
- 客服机器人;
- 群欢迎和群管理;
- 好友请求审核;
- Dify 支持的外部 Agent
- 使用 LangBot 模型和知识库的本地 Agent;
- 用于多阶段消息处理的 Pipeline。
文档应先描述产品模型,只在高级架构章节中暴露内部术语。
## 4. 推荐 UX 边界
之前“把所有事件编排都放进 Agent”的方向应当收窄。更好的边界是:
- **机器人页面负责事件路由**,因为事件面是平台特定的,用户也自然会在机器人上配置频道行为。
- **处理器页面负责处理器资产**,因为 Agent 与 Pipeline 都应能跨机器人复用,并且未来可以被打包进 Solution。
- **Pipeline 保持为一种处理器类型**,而不是隐藏在 Agent 术语背后的历史对象。
这样可以降低心智负担:
- 用户在机器人页面问:“当这个机器人遇到某件事时,应该做什么?”
- 用户在处理器页面问:“我想复用什么处理逻辑?”
这也支持同一个机器人上的不同事件使用不同处理器类型:一个事件可以使用 Pipeline,另一个事件可以使用 Agent,未来另一个事件可以使用 Workflow。
## 5. 未来导出、分发和导入单元
导出/导入不在当前实现范围内,但产品边界不应阻塞它。
正确的未来分发单元是 **Solution**,不是机器人,也不是单独的 Agent。
Solution 应包含:
- 处理器:Agent、Pipeline、未来 Workflow 定义;
- 路由模板:事件模式、友好名称、目标逻辑引用、默认优先级和可选条件;
- 依赖清单:所需 runner 插件、适配器能力要求、模型、工具和资源;
- 变量:用户提供的值,例如 API key、频道选择、模型选择和 prompt 参数;
- 文档:配置意图和预期行为。
Solution 不应包含:
- 具体机器人凭据;
- 已安装运行时 token
- 租户或命名空间 UUID
- 密钥;
- 原始平台账号标识;
- 已解析的机器人事件绑定 UUID。
导入时,应在目标命名空间内解析路由模板。用户需要先选择机器人/频道,并授予所需权限。
## 6. SaaS 多命名空间架构
产品在公开 SaaS 发布前必须支持多命名空间 SaaS 架构。事件路由模型很敏感,因为适配器、凭据、处理器、运行时、状态和日志都会跨越信任边界。
### 6.1 命名空间模型
采用分层命名空间模型:
| 范围 | 用途 |
| --- | --- |
| 租户(Tenant) | 计费、法律归属、顶层隔离。 |
| 工作空间(Workspace) | 租户内的协作和产品工作区。 |
| 命名空间(Namespace) | 机器人、处理器、运行时 token、资源和日志的可部署隔离边界。自托管部署可以只有一个默认命名空间。 |
| 机器人范围 | 平台适配器实例和事件路由表。 |
| 处理器范围 | Agent、Pipeline、Workflow 及相关配置。 |
| 运行时范围 | 插件运行时、runner 注册、lease 和执行权限。 |
| 资源范围 | 知识库、模型凭据、文件、状态和密钥。 |
在 SaaS GA 前,核心持久化对象都应携带 `tenant_id``workspace_id``namespace_id`。自托管部署可以在迁移时种子化一个默认租户/工作空间/命名空间。
### 6.2 事件入口隔离
每个入站事件都必须先解析命名空间,再进行路由匹配:
```text
adapter ingress -> tenant/workspace/namespace resolution -> event normalization -> event log append -> route match -> processor run
```
Webhook 和回调端点应编码或查找命名空间范围内的适配器安装。来自一个命名空间的平台事件,绝不能匹配另一个命名空间的路由,即使适配器名称、机器人名称或原始平台 ID 发生碰撞。
### 6.3 路由目标规则
运行时路由绑定可以使用已安装 UUID,但导出的路由模板必须使用逻辑引用。在 SaaS 中:
- 默认情况下,机器人路由只能指向同一命名空间内的处理器;
- 跨命名空间目标默认禁止,除非由策略明确共享;
- Pipeline 目标仍然只允许处理消息;
- 路由冲突评估应限制在命名空间内;
- 路由审计事件必须包含租户、工作空间、命名空间、机器人、路由、目标和 run 标识。
### 6.4 运行时和插件隔离
插件运行时和 runner 注册表需要命名空间范围的授权:
- 运行时注册 token 只作用于一个命名空间,或显式允许的一组命名空间;
- runner 发现结果按命名空间权限过滤;
- lease 和 heartbeat 按命名空间隔离;
- run-scoped API token 不能访问 run 所在命名空间之外的对象;
- 插件存储、状态和临时文件的 storage key 应包含命名空间;
- 早期 SaaS 可以接受共享运行时,但每个 API 调用都必须被 scoped 并审计;
- 企业或高风险租户应支持专用运行时。
### 6.5 密钥、资源和状态
密钥和资源不能全局寻址:
- 适配器凭据存放在命名空间范围的密钥存储中;
- 模型提供商凭据可以根据策略按租户/工作空间/命名空间设定范围;
- 知识资源声明允许哪些命名空间使用;
- Agent 持久状态按租户/工作空间/命名空间/处理器分区,除非共享策略另有规定;
- 事件日志和会话记录按命名空间隔离,并受保留策略约束。
### 6.6 Marketplace 和已安装资产
Marketplace 包可见性不等于已安装资产可见性:
- Marketplace 包可以是公开、租户私有或工作空间私有;
- 安装会创建命名空间本地资产,或命名空间本地引用;
- 已安装处理器和路由模板默认复制,避免意外跨租户变更;
- 更新必须显式且可审计。
### 6.7 SaaS 测试要求
SaaS beta 前必须验证:
- 租户 A 的事件不能匹配租户 B 的路由;
- 命名空间 A 的运行时不能 claim 命名空间 B 的 run
- 命名空间 A 的处理器不能读取命名空间 B 的资源;
- Solution 导入不能保留源租户 UUID 或密钥;
- 路由重放不能暴露另一个命名空间的原始事件 payload;
- 管理员可以查看审计轨迹,但不能访问密钥值。
## 7. 发布计划
### Phase 0:技术收敛
目标:证明合并分支可以基于事件绑定和外置 runner 运行。
必要门禁:
- 机器人事件路由以 `event_bindings` 作为唯一路由来源。
- 旧机器人 pipeline 路由字段已移除或迁移。
- Pipeline 可作为只处理消息的处理器运行。
- Agent runner 插件冒烟测试通过,覆盖 local runner 和 Dify runner。
- 迁移命名和 downgrade 路径有效。
- 主 UI 不再在面向用户的适配器名称中暴露 “EBA”。
### Phase 1:面向技术用户的私有 Beta
目标:让贡献者和早期自托管用户可用。
必要门禁:
- 文档和 UI 中存在适配器能力矩阵;
- 路由编辑器会过滤不兼容目标;
- runner/plugin 健康检查可见;
- local Agent 和 Dify Agent 有引导式配置路径;
- 每条路由有运行轨迹;
- 废弃适配器被一致标记;
- 失败信息能指出失败层级。
### Phase 2:面向非技术用户的产品 Beta
目标:让用户无需阅读架构文档也能完成常见场景。
必要门禁:
- 首次使用机器人向导从用例和频道开始;
- 事件预设默认隐藏原始事件模式;
- 存在路由模拟或测试事件能力;
- 存在常见处理器模板;
- 冲突警告和兜底行为清晰;
- 文档先使用产品语言,再介绍高级术语;
- SaaS 命名空间 schema 已实现,或已经具备迁移准备。
### Phase 3SaaS Beta
目标:安全地为多个租户运行产品。
必要门禁:
- 所有路由、运行时、状态、日志和资源事实都具备租户/工作空间/命名空间字段;
- 命名空间范围的运行时注册和 run claim 已强制执行;
- 命名空间范围的密钥和适配器安装已强制执行;
- 路由匹配和审计限制在命名空间内;
- 配额、保留策略和管理员审计界面存在;
- 自托管默认命名空间迁移已有文档。
### Phase 4GA
目标:让产品具备广泛采用所需的可靠性。
必要门禁:
- Solution 导出/导入已实现,并支持依赖和变量解析;
- Marketplace 分发支持命名空间本地安装;
- 跨命名空间共享策略是显式的;
- 安全评审覆盖适配器、运行时 token、runner API、密钥、日志和路由重放;
- 升级和回滚流程已有文档;
- 产品遥测可以衡量上手流失和路由失败类别。
## 8. 验收清单
只有满足以下条件,才可以认为发布版本达到产品就绪:
- 非技术用户可以连接一个受支持频道,选择场景,绑定处理器,测试它,并在不编辑原始 JSON 的情况下理解结果;
- 产品 UI 在主要流程中避免使用 EBA 这类内部术语;
- 机器人页面负责平台事件路由,处理器页面负责可复用的 Agent 与 Pipeline
- Pipeline 仍作为无代码消息处理器可见,并且不会出现在非消息事件目标中;
- 同一个机器人的不同事件可以路由到不同处理器类型;
- 路由失败能按层级解释;
- 命名空间隔离通过 schema、service check、运行时 token、storage key 和测试强制执行;
- 导出/导入实现后,使用带路由模板的 Solution 包,而不是具体机器人绑定。
## 9. 待决策问题
- 组合处理器入口统一使用 “Processor / 处理器”;Agent 与 Pipeline 是其中平级的类型。
- Workflow 应作为独立持久化处理器类型,还是作为 Agent runner 类别。
- SaaS 命名空间初期是否与工作空间一一映射,还是高级租户在首个 SaaS beta 就需要一个工作空间下多个命名空间。
- 哪些适配器允许在 SaaS 共享运行时中运行,哪些需要专用运行时隔离。
- 未来 Solution 导出/导入的确切包格式。
-196
View File
@@ -1,196 +0,0 @@
# Event Based Agents 架构设计总览
## 1. 背景与动机
### 当前架构的局限性
LangBot 当前的平台适配器架构围绕**消息事件**单一场景设计:
- **事件层面**:只监听 `FriendMessage`(私聊消息)和 `GroupMessage`(群消息)两种事件
- **API 层面**:只暴露 `send_message``reply_message` 两个平台 API
- **处理层面**:所有消息统一进入 Pipeline 流水线处理,无法为不同事件类型配置不同处理逻辑
- **适配器结构**:每个适配器是单个 Python 文件(200-800 行),随着功能增加难以维护
这导致以下问题:
1. **无法处理非消息事件**:新成员入群、好友请求、消息撤回、消息编辑等大部分平台都支持的事件被完全忽略
2. **平台能力未充分利用**:编辑消息、撤回消息、获取群成员列表、管理群组等 API 无法使用
3. **插件能力受限**:插件只能监听消息事件、只能发送/回复消息,无法实现更丰富的交互
4. **处理逻辑不灵活**:所有消息走同一条 Pipeline,无法为入群欢迎、好友自动通过等场景配置独立的处理流程
### 设计目标
Event Based AgentsEBA)架构旨在将 LangBot 从"消息处理平台"升级为"事件驱动的智能代理平台":
- **丰富事件**:支持消息、群组、好友、Bot 状态等多种事件类型
- **丰富 API**:支持消息编辑/撤回、群组管理、用户信息查询等通用 API,以及适配器特有 API 的透传调用
- **灵活编排**:用户可在 WebUI 上为每个 Bot 的每种事件类型配置不同的处理器
- **可扩展**:适配器可声明自己支持的事件和 API,平台特有能力通过标准机制暴露
- **向后兼容**:现有插件无需修改即可在新架构下运行
## 2. 架构对比
### 现有架构
```
消息平台 (Telegram/Discord/...)
平台适配器 (单文件, 只处理消息)
│ FriendMessage / GroupMessage
RuntimeBot (注册 on_friend_message / on_group_message 回调)
MessageAggregator (消息聚合)
QueryPool → Controller → Pipeline (固定阶段链)
│ │
│ ▼
│ AgentRunner Host orchestrator
│ ▼
│ plugin AgentRunner
adapter.reply_message() / adapter.send_message()
```
关键代码路径:
- 适配器基类:`langbot-plugin-sdk/.../abstract/platform/adapter.py``AbstractMessagePlatformAdapter`
- 事件定义:`langbot-plugin-sdk/.../builtin/platform/events.py` — 仅 `FriendMessage` / `GroupMessage`
- Bot 管理:`LangBot/src/langbot/pkg/platform/botmgr.py``RuntimeBot` 只注册两个消息回调
- 流水线控制:`LangBot/src/langbot/pkg/pipeline/controller.py` — 从 QueryPool 消费并执行 Pipeline
### 新架构(Event Based Agents
```
消息平台 (Telegram/Discord/...)
平台适配器 (独立目录, 监听所有事件, 实现丰富 API)
│ MessageReceived / MemberJoined / FriendRequest / ...
EventBus (统一事件总线)
├─→ Plugin EventListener observers(始终广播,不参与响应仲裁)
EventRouter (读取 Bot 的 event_bindings)
├─→ Pipeline target — 完整 Stage 链,仅消息事件
├─→ Agent target — 独立 Agent,经插件 AgentRunner 执行
└─→ discard — 明确丢弃
统一平台 API
send / reply / edit / delete / getGroupInfo / getUserInfo / callPlatformApi / ...
```
## 3. 核心概念
### 3.1 统一事件体系
所有平台事件统一为命名空间式的事件类型:
| 命名空间 | 事件 | 说明 |
|----------|------|------|
| `message.*` | `message.received`, `message.edited`, `message.deleted`, `message.reaction` | 消息相关 |
| `feedback.*` | `feedback.received` | 用户对 Bot 回复的点赞、点踩、取消反馈等评价事件 |
| `group.*` | `group.member_joined`, `group.member_left`, `group.member_banned`, `group.info_updated` | 群组相关 |
| `friend.*` | `friend.request_received`, `friend.added`, `friend.removed` | 好友相关 |
| `bot.*` | `bot.invited_to_group`, `bot.removed_from_group`, `bot.muted`, `bot.unmuted` | Bot 状态 |
| `platform.*` | `platform.{adapter}.{action}` | 适配器特有事件 |
详见 [01-event-system.md](./01-event-system.md)。
### 3.2 统一平台 API
扩展适配器基类,提供通用 API + 透传机制:
| 类别 | API | 必需/可选 |
|------|-----|----------|
| 消息 | `send_message`, `reply_message`, `edit_message`, `delete_message`, `forward_message` | send/reply 必需,其余可选 |
| 群组 | `get_group_info`, `get_group_member_list`, `get_group_member_info`, `mute_member`, `kick_member` | 全部可选 |
| 用户 | `get_user_info`, `get_friend_list` | 全部可选 |
| 媒体 | `upload_file`, `get_file_url` | 全部可选 |
| 透传 | `call_platform_api(action, params)` | 可选 |
详见 [02-platform-api.md](./02-platform-api.md)。
### 3.3 适配器新结构
每个适配器从单文件迁移到独立目录:
```
pkg/platform/adapters/
├── _base/ # 基类和通用定义
│ ├── adapter.py
│ ├── events.py
│ ├── entities.py
│ └── api.py
├── telegram/
│ ├── __init__.py
│ ├── adapter.py # 主适配器类
│ ├── event_converter.py # 事件转换(多种事件类型)
│ ├── message_converter.py # 消息链转换
│ ├── api_impl.py # 通用 API 实现
│ ├── platform_api.py # 平台特有 API
│ ├── types.py # 平台特有类型
│ └── manifest.yaml
├── discord/
│ └── ...
```
详见 [03-adapter-structure.md](./03-adapter-structure.md)。
### 3.4 事件响应目标与观察者
Pipeline 与 Agent 是长期并存、场景不同的同级处理器。Pipeline 保留完整 Stage 链,面向消息处理;Agent 是独立配置对象,选择一个已安装的插件 AgentRunner,并可声明消息或非消息事件能力。Bot 的 `event_bindings` 只负责把事件绑定到既有 Pipeline、独立 Agent 或 `discard`
插件 EventListener 是观察者:事件先广播给有权限的监听器,随后路由器再选择一个响应目标。Webhook、Dify、n8n 等外部执行方式若需要作为响应者,应由对应 AgentRunner 插件表达,而不是增加另一套 Host Handler 主链。
现有 Pipeline 不会被转换为 AgentPipeline 内的 runner 配置也不会复制到独立 Agent。用户需要 Agent 时自行创建并绑定。
详见 [04-event-routing.md](./04-event-routing.md)。
### 3.5 插件 SDK 改造
- 新事件类型全部暴露给插件
- 新 API 全部通过 `LangBotAPIProxy` 暴露
- 兼容层保证现有插件零修改运行
详见 [05-plugin-sdk.md](./05-plugin-sdk.md)。
## 4. 关键设计决策
| # | 决策点 | 选择 | 理由 |
|---|--------|------|------|
| 1 | 事件处理器配置粒度 | 每个 Bot 独立配置 | Bot 是用户操作的核心单元,不同 Bot 可能对接不同业务场景 |
| 2 | 适配器特有 API | 统一抽象 + `call_platform_api` 透传 | 通用 API 覆盖大部分场景,透传机制保证灵活性,避免每个适配器导出独立的类型化 API 包 |
| 3 | 向后兼容策略 | 兼容层适配 | 保留旧事件类型和 API 作为新系统的 alias/wrapper,现有插件无需修改 |
| 4 | 处理器配置存储 | Bot 表使用 `event_bindings`,目标引用原始 Pipeline 或独立 Agent UUID | 路由关系不复制处理器配置,Pipeline/Agent 各自保持事实源 |
| 5 | Agent 处理器定位 | 独立 Agent + 插件 AgentRunner | Host 不再内置具体 runner;不同 AgentRunner 通过统一协议接入 |
| 6 | 事件命名方式 | 命名空间式(`message.received` | 清晰的分类层级,便于通配匹配(`message.*`),与 WebUI 配置天然对应 |
## 5. 文档索引
| 文档 | 内容 |
|------|------|
| [01-event-system.md](./01-event-system.md) | 统一事件体系:事件分类、定义、生命周期 |
| [02-platform-api.md](./02-platform-api.md) | 统一平台 API:通用 API、透传 API、实体定义 |
| [03-adapter-structure.md](./03-adapter-structure.md) | 适配器新结构:目录布局、基类、注册机制 |
| [04-event-routing.md](./04-event-routing.md) | 事件路由与编排:路由引擎、处理器类型、WebUI 数据模型 |
| [05-plugin-sdk.md](./05-plugin-sdk.md) | 插件 SDK 改造:新事件/API、兼容层 |
| [06-migration-plan.md](./06-migration-plan.md) | 分阶段迁移计划 |
| [07-agent-orchestration.md](./07-agent-orchestration.md) | **产品最终形态(2026-06 修订)**Pipeline / Agent 同级处理器编排、SDK Agent 组件契约、发布火车 |
## 6. 涉及的代码仓库
| 仓库 | 改动范围 |
|------|----------|
| **langbot-plugin-sdk** | 事件定义、实体模型、API 接口、适配器基类、通信协议扩展 |
| **LangBot**(后端) | 适配器实现、事件路由引擎、Bot/Agent 实体、AgentRunner Host 编排 |
| **LangBot**(前端) | Bot 事件处理器编排面板 |
| **langbot-wiki** | 新架构文档、插件开发指南更新、适配器开发指南 |
| **langbot-plugin-demo** | 示例更新(使用新事件和 API) |
-562
View File
@@ -1,562 +0,0 @@
# 统一事件体系
## 1. 设计原则
- **命名空间分类**:事件类型采用 `{namespace}.{action}` 格式,如 `message.received`
- **通用优先**:大部分平台都支持的事件抽象为通用事件,定义统一的字段格式
- **平台特有事件标准化**:各适配器的独有事件通过 `PlatformSpecificEvent` 承载,保留原始数据
- **向后兼容**:现有 `FriendMessage` / `GroupMessage` 通过兼容层映射到新的 `message.received` 事件
## 2. 事件基类层次
```
Event (事件基类)
├── MessageEvent (消息相关事件)
│ ├── MessageReceivedEvent # message.received
│ ├── MessageEditedEvent # message.edited
│ ├── MessageDeletedEvent # message.deleted
│ └── MessageReactionEvent # message.reaction
├── FeedbackEvent (用户反馈事件)
│ └── FeedbackReceivedEvent # feedback.received
├── GroupEvent (群组相关事件)
│ ├── MemberJoinedEvent # group.member_joined
│ ├── MemberLeftEvent # group.member_left
│ ├── MemberBannedEvent # group.member_banned
│ ├── MemberUnbannedEvent # group.member_unbanned
│ └── GroupInfoUpdatedEvent # group.info_updated
├── FriendEvent (好友相关事件)
│ ├── FriendRequestReceivedEvent # friend.request_received
│ ├── FriendAddedEvent # friend.added
│ └── FriendRemovedEvent # friend.removed
├── BotEvent (Bot 状态事件)
│ ├── BotInvitedToGroupEvent # bot.invited_to_group
│ ├── BotRemovedFromGroupEvent # bot.removed_from_group
│ ├── BotMutedEvent # bot.muted
│ └── BotUnmutedEvent # bot.unmuted
└── PlatformSpecificEvent # platform.{adapter}.{action}
```
## 3. 通用事件定义
### 3.1 事件基类
```python
class Event(pydantic.BaseModel):
"""事件基类"""
type: str
"""事件类型标识,如 'message.received'"""
timestamp: float
"""事件发生的时间戳"""
bot_uuid: str
"""接收到此事件的 Bot UUID"""
adapter_name: str
"""产生此事件的适配器名称"""
source_platform_object: typing.Optional[typing.Any] = None
"""原始平台事件对象,供适配器内部使用"""
```
### 3.2 消息事件
#### MessageReceivedEvent (`message.received`)
收到新消息。这是最核心的事件,替代现有的 `FriendMessage` / `GroupMessage`
```python
class MessageReceivedEvent(Event):
"""收到新消息"""
type: str = "message.received"
message_id: typing.Union[int, str]
"""消息 ID"""
message_chain: MessageChain
"""消息内容"""
sender: User
"""发送者"""
chat_type: ChatType # "private" | "group"
"""会话类型"""
chat_id: typing.Union[int, str]
"""会话 ID(私聊为对方用户 ID,群聊为群 ID)"""
group: typing.Optional[Group] = None
"""群信息(仅群聊时存在)"""
```
与现有类型的映射关系:
- `chat_type == "private"` → 等价于现有 `FriendMessage`
- `chat_type == "group"` → 等价于现有 `GroupMessage`
`ChatType` 枚举:
```python
class ChatType(str, Enum):
PRIVATE = "private"
GROUP = "group"
```
#### MessageEditedEvent (`message.edited`)
消息被编辑。
```python
class MessageEditedEvent(Event):
"""消息被编辑"""
type: str = "message.edited"
message_id: typing.Union[int, str]
"""被编辑的消息 ID"""
new_content: MessageChain
"""编辑后的新内容"""
editor: User
"""编辑者"""
chat_type: ChatType
chat_id: typing.Union[int, str]
group: typing.Optional[Group] = None
```
#### MessageDeletedEvent (`message.deleted`)
消息被删除/撤回。
```python
class MessageDeletedEvent(Event):
"""消息被删除/撤回"""
type: str = "message.deleted"
message_id: typing.Union[int, str]
"""被删除的消息 ID"""
operator: typing.Optional[User] = None
"""操作者(可能是发送者自己撤回,也可能是管理员删除)"""
chat_type: ChatType
chat_id: typing.Union[int, str]
group: typing.Optional[Group] = None
```
#### MessageReactionEvent (`message.reaction`)
消息收到表情回应。
```python
class MessageReactionEvent(Event):
"""消息收到表情回应"""
type: str = "message.reaction"
message_id: typing.Union[int, str]
"""被回应的消息 ID"""
user: User
"""回应者"""
reaction: str
"""回应的表情标识(emoji 或平台特定表情 ID)"""
is_add: bool
"""True 为添加回应,False 为移除回应"""
chat_type: ChatType
chat_id: typing.Union[int, str]
group: typing.Optional[Group] = None
```
### 3.3 用户反馈事件
#### FeedbackReceivedEvent (`feedback.received`)
用户对 Bot 回复提交反馈。该事件用于承载平台提供的点赞、点踩、取消反馈以及点踩原因等评价信息;典型来源包括企业微信 AI Bot 的 `feedback_event`、飞书卡片按钮回调、Web Embed 的反馈入口等。
```python
class FeedbackReceivedEvent(Event):
"""收到用户反馈"""
type: str = "feedback.received"
feedback_id: str
"""平台侧反馈 ID,用于幂等记录或取消反馈"""
feedback_type: int
"""1 = like, 2 = dislike, 3 = cancel/remove feedback"""
feedback_content: typing.Optional[str] = None
"""用户填写的自由文本反馈"""
inaccurate_reasons: typing.Optional[list[str]] = None
"""点踩时平台提供的预设不准确原因"""
user_id: typing.Optional[str] = None
"""提交反馈的用户 ID"""
session_id: typing.Optional[str] = None
"""会话 ID,例如 person_xxx 或 group_xxx"""
message_id: typing.Optional[str] = None
"""被评价的 Bot 回复消息 ID"""
stream_id: typing.Optional[str] = None
"""流式回复 ID,用于关联 streaming response"""
```
设计约定:
- `feedback_id` 是幂等键;同一个 `feedback_id` 的后续事件应更新已有记录。
- `feedback_type == 3` 表示用户取消/移除反馈,处理器可删除对应记录或标记为取消。
- 如果平台只能给出原始回调 payload,差异字段保留在 `source_platform_object``PlatformSpecificEvent.data` 中;通用字段仍优先映射到 `FeedbackReceivedEvent`
- 该事件保留向后兼容映射:EBA 事件可转换为旧的 `FeedbackEvent`,字段语义保持一致。
### 3.4 群组事件
#### MemberJoinedEvent (`group.member_joined`)
新成员加入群组。
```python
class MemberJoinedEvent(Event):
"""新成员加入群组"""
type: str = "group.member_joined"
group: Group
"""群组"""
member: User
"""加入的成员"""
inviter: typing.Optional[User] = None
"""邀请者(如有)"""
join_type: typing.Optional[str] = None
"""加入方式:'invite' / 'request' / 'direct' / None"""
```
#### MemberLeftEvent (`group.member_left`)
成员离开群组。
```python
class MemberLeftEvent(Event):
"""成员离开群组"""
type: str = "group.member_left"
group: Group
member: User
is_kicked: bool = False
"""是否被踢出"""
operator: typing.Optional[User] = None
"""操作者(踢出时为管理员)"""
```
#### MemberBannedEvent (`group.member_banned`)
成员被禁言。
```python
class MemberBannedEvent(Event):
"""成员被禁言"""
type: str = "group.member_banned"
group: Group
member: User
operator: typing.Optional[User] = None
duration: typing.Optional[int] = None
"""禁言时长(秒),None 表示永久"""
```
#### MemberUnbannedEvent (`group.member_unbanned`)
成员被解除禁言。
```python
class MemberUnbannedEvent(Event):
"""成员被解除禁言"""
type: str = "group.member_unbanned"
group: Group
member: User
operator: typing.Optional[User] = None
```
#### GroupInfoUpdatedEvent (`group.info_updated`)
群组信息被修改。
```python
class GroupInfoUpdatedEvent(Event):
"""群组信息被修改"""
type: str = "group.info_updated"
group: Group
"""更新后的群组信息"""
operator: typing.Optional[User] = None
"""操作者"""
changed_fields: list[str] = []
"""发生变更的字段名列表,如 ['name', 'description']"""
```
### 3.5 好友事件
#### FriendRequestReceivedEvent (`friend.request_received`)
收到好友请求。
```python
class FriendRequestReceivedEvent(Event):
"""收到好友请求"""
type: str = "friend.request_received"
request_id: typing.Union[int, str]
"""请求 ID,用于后续 approve/reject 操作"""
user: User
"""请求者"""
message: typing.Optional[str] = None
"""验证消息"""
```
#### FriendAddedEvent (`friend.added`)
成功添加好友。
```python
class FriendAddedEvent(Event):
"""成功添加好友"""
type: str = "friend.added"
user: User
"""新好友"""
```
#### FriendRemovedEvent (`friend.removed`)
好友被移除。
```python
class FriendRemovedEvent(Event):
"""好友被移除"""
type: str = "friend.removed"
user: User
"""被移除的好友"""
```
### 3.6 Bot 状态事件
#### BotInvitedToGroupEvent (`bot.invited_to_group`)
Bot 被邀请加入群组。
```python
class BotInvitedToGroupEvent(Event):
"""Bot 被邀请加入群组"""
type: str = "bot.invited_to_group"
group: Group
inviter: typing.Optional[User] = None
request_id: typing.Optional[typing.Union[int, str]] = None
"""邀请请求 ID,某些平台需要 Bot 确认才加入"""
```
#### BotRemovedFromGroupEvent (`bot.removed_from_group`)
Bot 被移出群组。
```python
class BotRemovedFromGroupEvent(Event):
"""Bot 被移出群组"""
type: str = "bot.removed_from_group"
group: Group
operator: typing.Optional[User] = None
```
#### BotMutedEvent / BotUnmutedEvent (`bot.muted` / `bot.unmuted`)
Bot 被禁言/解除禁言。
```python
class BotMutedEvent(Event):
"""Bot 被禁言"""
type: str = "bot.muted"
group: Group
operator: typing.Optional[User] = None
duration: typing.Optional[int] = None
class BotUnmutedEvent(Event):
"""Bot 被解除禁言"""
type: str = "bot.unmuted"
group: Group
operator: typing.Optional[User] = None
```
### 3.7 平台特有事件
对于无法抽象为通用事件的平台特有事件,使用统一的 `PlatformSpecificEvent` 承载:
```python
class PlatformSpecificEvent(Event):
"""平台特有事件
适配器无法映射到通用事件类型时,使用此类型承载。
插件可以通过 adapter_name + action 来识别和处理。
"""
type: str = "platform.specific"
action: str
"""平台特有的事件动作标识,如 'channel_created', 'pin_message'"""
data: dict = {}
"""事件数据,结构由具体适配器定义"""
```
事件类型字符串格式为 `platform.{adapter_name}.{action}`,例如:
- `platform.telegram.chat_member_updated` — Telegram 的群成员信息更新
- `platform.discord.channel_created` — Discord 的频道创建
- `platform.discord.voice_state_update` — Discord 的语音状态变更
- `platform.slack.app_home_opened` — Slack 的 App Home 打开
## 4. 各平台事件支持矩阵
下表标注各通用事件在主要平台上的支持情况:
| 事件 | Telegram | Discord | OneBot(QQ) | 飞书 | 钉钉 | Slack | 微信 | LINE | KOOK |
|------|----------|---------|-----------|------|------|-------|------|------|------|
| `message.received` | Y | Y | Y | Y | Y | Y | Y | Y | Y |
| `message.edited` | Y | Y | N | Y | N | Y | N | N | Y |
| `message.deleted` | Y | Y | Y | Y | N | Y | Y | N | Y |
| `message.reaction` | Y | Y | Y | Y | Y | Y | N | N | Y |
| `feedback.received` | N | N | N | Y | N | N | Y | N | N |
| `group.member_joined` | Y | Y | Y | Y | Y | Y | Y | Y | Y |
| `group.member_left` | Y | Y | Y | Y | Y | Y | Y | Y | Y |
| `group.member_banned` | Y | Y | Y | N | N | N | N | N | N |
| `group.info_updated` | Y | Y | Y | Y | Y | Y | N | N | Y |
| `friend.request_received` | N | Y | Y | N | N | N | Y | Y | Y |
| `friend.added` | N | Y | Y | N | N | N | Y | Y | N |
| `bot.invited_to_group` | Y | Y | Y | Y | Y | Y | Y | N | Y |
| `bot.removed_from_group` | Y | Y | Y | Y | N | N | Y | N | Y |
| `bot.muted` | Y | N | Y | N | N | N | N | N | N |
| `bot.unmuted` | Y | N | Y | N | N | N | N | N | N |
| `platform.specific` | Y | Y | Y | Y | Y | Y | Y | Y | Y |
> 注:此表为初步评估,具体以各平台 SDK/API 文档为准,实施时逐个确认。
## 5. 事件生命周期
```
1. 平台 SDK 回调触发
2. 适配器 EventConverter.target2yiri(raw_event)
│ 将平台原生事件转换为统一 Event 对象
│ 无法映射的事件 → PlatformSpecificEvent
3. 适配器回调注册的 listener(event, adapter)
4. RuntimeBot 接收事件
5. EventBus 分发
6. EventBus 向有权限的 Plugin EventListener 广播观察者事件
│ observer 不占用响应目标,也不参与 priority 仲裁
7. EventRouter 查询 Bot 的 event_bindings
│ 精确/通配匹配 event_pattern,应用 filters 与 priority
│ 选择一个 Pipeline、Agent 或 discard 目标
8. 目标处理事件
│ Pipeline → 进入完整 Pipeline 流水线(仅消息事件)
│ Agent → Host 编排已安装的插件 AgentRunner
│ discard → 不产生响应
9. 处理器执行完毕,可能通过 Host 授权 API 执行响应动作
(发消息、编辑消息、踢人、同意好友请求等)
```
## 6. 与现有事件类型的兼容映射
为保证现有插件不受影响,建立以下映射关系:
| 新事件 | 条件 | 旧事件 |
|--------|------|--------|
| `MessageReceivedEvent` (chat_type=private) | — | `FriendMessage` |
| `MessageReceivedEvent` (chat_type=group) | — | `GroupMessage` |
在插件 SDK 层面:
| 新事件 | 旧插件事件 |
|--------|-----------|
| `MessageReceivedEvent` (chat_type=private, 非命令) | `PersonNormalMessageReceived` |
| `MessageReceivedEvent` (chat_type=group, 非命令) | `GroupNormalMessageReceived` |
| `MessageReceivedEvent` (chat_type=private, 命令) | `PersonCommandSent` |
| `MessageReceivedEvent` (chat_type=group, 命令) | `GroupCommandSent` |
| `MessageReceivedEvent` (处理完毕后) | `NormalMessageResponded` |
兼容层在事件分发给插件 EventListener 时自动生成旧格式事件,确保监听旧事件类型的插件仍能正常工作。
## 7. 事件类型注册表
适配器在 manifest.yaml 中声明自己支持的事件类型:
```yaml
kind: MessagePlatformAdapter
metadata:
name: telegram
spec:
supported_events:
- message.received
- message.edited
- message.deleted
- message.reaction
- feedback.received
- group.member_joined
- group.member_left
- group.member_banned
- group.info_updated
- bot.invited_to_group
- bot.removed_from_group
- bot.muted
- bot.unmuted
- platform.specific
platform_specific_events:
- chat_member_updated
- chat_join_request
```
这份声明用于:
1. WebUI 在配置事件处理器时,只显示当前 Bot 的适配器支持的事件类型
2. EventRouter 在路由时校验事件类型有效性
3. 文档自动生成
-546
View File
@@ -1,546 +0,0 @@
# 统一平台 API 与实体定义
## 1. 设计原则
- **通用 API 抽象**:大部分平台都支持的操作(发消息、获取群信息等)定义为通用 API 方法
- **required / optional 标记**:每个 API 标记为必需或可选,适配器未实现可选 API 时抛出 `NotSupportedError`
- **透传机制**:适配器特有的操作通过 `call_platform_api(action, params)` 统一入口透传调用
- **能力声明**:适配器在 manifest 中声明自己支持的 API 列表,供 WebUI 和插件查询
- **实体统一**:通用实体(User、Group 等)在 SDK 层面统一定义,适配器负责转换
## 2. 通用实体定义
### 2.1 现有实体回顾
当前 SDK 已有以下实体(`langbot_plugin/api/entities/builtin/platform/entities.py`):
```python
Entity(id)
Friend(id, nickname, remark)
Group(id, name, permission)
GroupMember(id, member_name, permission, group, special_title)
```
### 2.2 新实体设计
扩展实体体系,保持向后兼容:
```python
class User(pydantic.BaseModel):
"""用户实体(统一表示)"""
id: typing.Union[int, str]
"""用户 ID"""
nickname: str = ""
"""昵称"""
avatar_url: typing.Optional[str] = None
"""头像 URL"""
is_bot: bool = False
"""是否为 Bot"""
# 以下为可选的扩展信息,不同平台可能部分为空
username: typing.Optional[str] = None
"""用户名(如 Telegram 的 @username"""
remark: typing.Optional[str] = None
"""备注名"""
class Group(pydantic.BaseModel):
"""群组实体"""
id: typing.Union[int, str]
"""群组 ID"""
name: str = ""
"""群组名称"""
description: typing.Optional[str] = None
"""群组描述"""
member_count: typing.Optional[int] = None
"""成员数量"""
avatar_url: typing.Optional[str] = None
"""群组头像 URL"""
owner_id: typing.Optional[typing.Union[int, str]] = None
"""群主 ID"""
class GroupMember(pydantic.BaseModel):
"""群成员实体"""
user: User
"""用户信息"""
group_id: typing.Union[int, str]
"""所属群组 ID"""
role: MemberRole
"""群内角色"""
display_name: typing.Optional[str] = None
"""群内显示名"""
joined_at: typing.Optional[float] = None
"""加入群组的时间戳"""
title: typing.Optional[str] = None
"""群头衔/特殊称号"""
class MemberRole(str, Enum):
"""群成员角色"""
OWNER = "owner"
ADMIN = "admin"
MEMBER = "member"
```
### 2.3 与现有实体的兼容映射
| 新实体 | 旧实体 | 映射方式 |
|--------|--------|----------|
| `User` | `Friend` | `User(id=friend.id, nickname=friend.nickname, remark=friend.remark)` |
| `Group` | `Group`(旧) | `Group(id=old.id, name=old.name)` + `permission` 字段弃用 |
| `GroupMember` | `GroupMember`(旧) | `GroupMember(user=User(...), role=..., display_name=old.member_name)` |
| `MemberRole` | `Permission` | `OWNER↔Owner`, `ADMIN↔Administrator`, `MEMBER↔Member` |
旧实体类保留,标记为 `@deprecated`,内部通过转换方法桥接到新实体。
## 3. 通用 API 定义
### 3.1 API 方法一览
#### 消息 API
| 方法 | 必需/可选 | 说明 |
|------|----------|------|
| `send_message(target_type, target_id, message)` | **必需** | 主动发送消息 |
| `reply_message(event, message, quote_origin)` | **必需** | 回复一个消息事件 |
| `edit_message(chat_type, chat_id, message_id, new_content)` | 可选 | 编辑已发送的消息 |
| `delete_message(chat_type, chat_id, message_id)` | 可选 | 删除/撤回消息 |
| `forward_message(from_chat, message_id, to_chat_type, to_chat_id)` | 可选 | 转发消息到另一个会话 |
| `get_message(chat_type, chat_id, message_id)` | 可选 | 获取指定消息的内容 |
#### 群组 API
| 方法 | 必需/可选 | 说明 |
|------|----------|------|
| `get_group_info(group_id)` | 可选 | 获取群组信息 |
| `get_group_list()` | 可选 | 获取 Bot 加入的群组列表 |
| `get_group_member_list(group_id)` | 可选 | 获取群成员列表 |
| `get_group_member_info(group_id, user_id)` | 可选 | 获取指定群成员信息 |
| `set_group_name(group_id, name)` | 可选 | 修改群名称 |
| `mute_member(group_id, user_id, duration)` | 可选 | 禁言群成员 |
| `unmute_member(group_id, user_id)` | 可选 | 解除禁言 |
| `kick_member(group_id, user_id)` | 可选 | 踢出群成员 |
| `leave_group(group_id)` | 可选 | Bot 退出群组 |
#### 用户 API
| 方法 | 必需/可选 | 说明 |
|------|----------|------|
| `get_user_info(user_id)` | 可选 | 获取用户信息 |
| `get_friend_list()` | 可选 | 获取好友列表 |
| `approve_friend_request(request_id, approve, remark)` | 可选 | 处理好友请求 |
| `approve_group_invite(request_id, approve)` | 可选 | 处理入群邀请 |
#### 媒体 API
| 方法 | 必需/可选 | 说明 |
|------|----------|------|
| `upload_file(file_data, filename)` | 可选 | 上传文件,返回可引用的文件 ID 或 URL |
| `get_file_url(file_id)` | 可选 | 获取文件下载 URL |
#### 透传 API
| 方法 | 必需/可选 | 说明 |
|------|----------|------|
| `call_platform_api(action, params)` | 可选 | 调用适配器特有 API |
### 3.2 API 方法签名详解
```python
class AbstractPlatformAdapter(pydantic.BaseModel, metaclass=abc.ABCMeta):
"""平台适配器基类(新版)"""
# ======== 必需方法 ========
@abc.abstractmethod
async def send_message(
self,
target_type: str, # "private" | "group"
target_id: typing.Union[int, str],
message: MessageChain,
) -> MessageResult:
"""主动发送消息
Returns:
MessageResult: 包含 message_id 等发送结果
"""
...
@abc.abstractmethod
async def reply_message(
self,
event: MessageReceivedEvent,
message: MessageChain,
quote_origin: bool = False,
) -> MessageResult:
"""回复一个消息事件"""
...
# ======== 可选消息方法 ========
async def edit_message(
self,
chat_type: str,
chat_id: typing.Union[int, str],
message_id: typing.Union[int, str],
new_content: MessageChain,
) -> None:
"""编辑已发送的消息"""
raise NotSupportedError("edit_message")
async def delete_message(
self,
chat_type: str,
chat_id: typing.Union[int, str],
message_id: typing.Union[int, str],
) -> None:
"""删除/撤回消息"""
raise NotSupportedError("delete_message")
async def forward_message(
self,
from_chat_type: str,
from_chat_id: typing.Union[int, str],
message_id: typing.Union[int, str],
to_chat_type: str,
to_chat_id: typing.Union[int, str],
) -> MessageResult:
"""转发消息"""
raise NotSupportedError("forward_message")
async def get_message(
self,
chat_type: str,
chat_id: typing.Union[int, str],
message_id: typing.Union[int, str],
) -> MessageReceivedEvent:
"""获取指定消息"""
raise NotSupportedError("get_message")
# ======== 可选群组方法 ========
async def get_group_info(
self,
group_id: typing.Union[int, str],
) -> Group:
"""获取群组信息"""
raise NotSupportedError("get_group_info")
async def get_group_list(self) -> list[Group]:
"""获取 Bot 加入的群组列表"""
raise NotSupportedError("get_group_list")
async def get_group_member_list(
self,
group_id: typing.Union[int, str],
) -> list[GroupMember]:
"""获取群成员列表"""
raise NotSupportedError("get_group_member_list")
async def get_group_member_info(
self,
group_id: typing.Union[int, str],
user_id: typing.Union[int, str],
) -> GroupMember:
"""获取指定群成员信息"""
raise NotSupportedError("get_group_member_info")
async def set_group_name(
self,
group_id: typing.Union[int, str],
name: str,
) -> None:
"""修改群名称"""
raise NotSupportedError("set_group_name")
async def mute_member(
self,
group_id: typing.Union[int, str],
user_id: typing.Union[int, str],
duration: int = 0,
) -> None:
"""禁言群成员,duration 为秒数,0 表示永久"""
raise NotSupportedError("mute_member")
async def unmute_member(
self,
group_id: typing.Union[int, str],
user_id: typing.Union[int, str],
) -> None:
"""解除禁言"""
raise NotSupportedError("unmute_member")
async def kick_member(
self,
group_id: typing.Union[int, str],
user_id: typing.Union[int, str],
) -> None:
"""踢出群成员"""
raise NotSupportedError("kick_member")
async def leave_group(
self,
group_id: typing.Union[int, str],
) -> None:
"""Bot 退出群组"""
raise NotSupportedError("leave_group")
# ======== 可选用户方法 ========
async def get_user_info(
self,
user_id: typing.Union[int, str],
) -> User:
"""获取用户信息"""
raise NotSupportedError("get_user_info")
async def get_friend_list(self) -> list[User]:
"""获取好友列表"""
raise NotSupportedError("get_friend_list")
async def approve_friend_request(
self,
request_id: typing.Union[int, str],
approve: bool = True,
remark: typing.Optional[str] = None,
) -> None:
"""处理好友请求"""
raise NotSupportedError("approve_friend_request")
async def approve_group_invite(
self,
request_id: typing.Union[int, str],
approve: bool = True,
) -> None:
"""处理入群邀请"""
raise NotSupportedError("approve_group_invite")
# ======== 可选媒体方法 ========
async def upload_file(
self,
file_data: bytes,
filename: str,
) -> str:
"""上传文件,返回文件 ID 或 URL"""
raise NotSupportedError("upload_file")
async def get_file_url(
self,
file_id: str,
) -> str:
"""获取文件下载 URL"""
raise NotSupportedError("get_file_url")
# ======== 透传 API ========
async def call_platform_api(
self,
action: str,
params: dict = {},
) -> dict:
"""调用适配器特有 API
Args:
action: 平台特有的 API 动作标识
params: 参数字典
Returns:
dict: 返回结果
Examples:
# Telegram: pin 消息
await adapter.call_platform_api("pin_message", {
"chat_id": 123456,
"message_id": 789
})
# Discord: 创建频道
await adapter.call_platform_api("create_channel", {
"guild_id": "...",
"name": "new-channel",
"type": "text"
})
"""
raise NotSupportedError("call_platform_api")
# ======== 流式输出(保留现有机制) ========
async def reply_message_chunk(
self,
event: MessageReceivedEvent,
bot_message: dict,
message: MessageChain,
quote_origin: bool = False,
is_final: bool = False,
):
"""流式回复消息"""
raise NotSupportedError("reply_message_chunk")
async def is_stream_output_supported(self) -> bool:
"""是否支持流式输出"""
return False
# ======== 生命周期方法(保留现有) ========
@abc.abstractmethod
async def run_async(self):
"""启动适配器"""
...
@abc.abstractmethod
async def kill(self) -> bool:
"""停止适配器"""
...
@abc.abstractmethod
def register_listener(self, event_type, callback):
"""注册事件监听器"""
...
@abc.abstractmethod
def unregister_listener(self, event_type, callback):
"""注销事件监听器"""
...
```
### 3.3 返回值类型
```python
class MessageResult(pydantic.BaseModel):
"""消息发送结果"""
message_id: typing.Optional[typing.Union[int, str]] = None
"""发送成功后的消息 ID"""
raw: typing.Optional[dict] = None
"""平台原始返回数据"""
class NotSupportedError(Exception):
"""适配器未实现此 API"""
def __init__(self, api_name: str):
self.api_name = api_name
super().__init__(f"API not supported by this adapter: {api_name}")
```
## 4. API 能力声明
适配器在 manifest.yaml 中声明支持的 API
```yaml
kind: MessagePlatformAdapter
metadata:
name: telegram
spec:
supported_apis:
required:
- send_message
- reply_message
optional:
- edit_message
- delete_message
- get_group_info
- get_group_member_list
- get_user_info
- upload_file
- get_file_url
- call_platform_api
platform_specific_apis:
- action: pin_message
description: "Pin a message in a chat"
params_schema:
chat_id: { type: "string", required: true }
message_id: { type: "string", required: true }
- action: unpin_message
description: "Unpin a message"
params_schema:
chat_id: { type: "string", required: true }
message_id: { type: "string", required: true }
```
用途:
1. **WebUI**:在配置界面展示当前 Bot 可用的 API 能力
2. **插件**:插件可查询某个 Bot 是否支持特定 API,据此决定行为
3. **文档**:自动生成各适配器的 API 支持矩阵
## 5. 各平台 API 支持矩阵
| API | Telegram | Discord | OneBot(QQ) | 飞书 | 钉钉 | Slack | 微信 | LINE | KOOK |
|-----|----------|---------|-----------|------|------|-------|------|------|------|
| `send_message` | Y | Y | Y | Y | Y | Y | Y | Y | Y |
| `reply_message` | Y | Y | Y | Y | Y | Y | Y | Y | Y |
| `edit_message` | Y | Y | N | Y | N | Y | N | N | Y |
| `delete_message` | Y | Y | Y | Y | N | Y | Y | N | Y |
| `forward_message` | Y | N | Y | Y | N | N | Y | N | N |
| `get_group_info` | Y | Y | Y | Y | Y | Y | N | Y | Y |
| `get_group_member_list` | Y | Y | Y | Y | Y | Y | N | Y | Y |
| `get_user_info` | Y | Y | Y | Y | Y | Y | N | Y | Y |
| `get_friend_list` | N | Y | Y | N | N | N | Y | N | N |
| `mute_member` | Y | Y | Y | N | N | N | N | N | N |
| `kick_member` | Y | Y | Y | N | N | N | N | N | Y |
| `upload_file` | Y | Y | Y | Y | Y | Y | Y | Y | Y |
| `call_platform_api` | Y | Y | Y | Y | Y | Y | Y | Y | Y |
> 注:此表为初步评估,具体以各平台 SDK/API 文档为准。
## 6. MessageChain 扩展
### 6.1 保留的通用组件
以下 MessageComponent 类型保持不变,继续作为通用消息元素:
- `Source` — 消息元信息
- `Plain` — 纯文本
- `Quote` — 引用回复
- `At` / `AtAll`@提及
- `Image` — 图片
- `Voice` — 语音
- `File` — 文件
- `Forward` — 合并转发
- `Face` — 表情
- `Unknown` — 未知类型
### 6.2 平台特有组件处理
当前 MessageChain 中存在大量微信特有的组件类型(`WeChatMiniPrograms`, `WeChatEmoji`, `WeChatLink` 等)。在新架构下:
- 这些类型**继续保留**在 SDK 中以保持兼容
- 新增的平台特有消息组件统一使用 `PlatformComponent` 基类:
```python
class PlatformComponent(MessageComponent):
"""平台特有的消息组件"""
type: str = "Platform"
platform: str
"""平台标识"""
component_type: str
"""组件类型"""
data: dict = {}
"""组件数据"""
```
适配器在转换消息链时,对于无法映射到通用组件的平台特有内容,使用 `PlatformComponent` 承载。
@@ -1,483 +0,0 @@
# 适配器新目录结构
## 1. 设计目标
- **模块化**:每个适配器从单文件拆分到独立目录,各模块职责清晰
- **可维护**:随着事件和 API 的增加,代码量会显著增长,目录结构有助于管理复杂度
- **一致性**:所有适配器遵循相同的目录布局和文件命名约定
- **兼容现有发现机制**:保持 YAML manifest + ComponentDiscoveryEngine 的注册体系
## 2. 新目录布局
### 2.1 整体结构
```
pkg/platform/
├── __init__.py
├── botmgr.py # PlatformManager + RuntimeBot(重构)
├── event_bus.py # EventBus(新增)
├── event_router.py # EventRouter(新增)
├── logger.py # EventLogger(保留)
├── webhook_pusher.py # WebhookPusher(重构为 WebhookHandler
├── adapters/ # 适配器(新目录)
│ ├── __init__.py
│ │
│ ├── telegram/
│ │ ├── __init__.py
│ │ ├── adapter.py # TelegramAdapter 主类
│ │ ├── event_converter.py # 平台事件 → 统一事件
│ │ ├── message_converter.py # MessageChain 互转
│ │ ├── api_impl.py # 通用 API 实现
│ │ ├── platform_api.py # call_platform_api 的动作映射
│ │ ├── types.py # 平台特有类型定义
│ │ └── manifest.yaml # 适配器清单
│ │
│ ├── discord/
│ │ ├── __init__.py
│ │ ├── adapter.py
│ │ ├── event_converter.py
│ │ ├── message_converter.py
│ │ ├── api_impl.py
│ │ ├── platform_api.py
│ │ ├── types.py
│ │ ├── voice.py # Discord 语音连接管理(特有)
│ │ └── manifest.yaml
│ │
│ ├── aiocqhttp/ # OneBot v11 (QQ)
│ │ └── ...
│ ├── qqofficial/
│ │ └── ...
│ ├── lark/ # 飞书
│ │ └── ...
│ ├── dingtalk/
│ │ └── ...
│ ├── slack/
│ │ └── ...
│ ├── wechatpad/
│ │ └── ...
│ ├── officialaccount/ # 微信公众号
│ │ └── ...
│ ├── wecom/ # 企业微信
│ │ └── ...
│ ├── wecombot/
│ │ └── ...
│ ├── wecomcs/
│ │ └── ...
│ ├── kook/
│ │ └── ...
│ ├── line/
│ │ └── ...
│ ├── satori/
│ │ └── ...
│ ├── websocket/ # 内置 WebSocket 适配器
│ │ ├── __init__.py
│ │ ├── adapter.py
│ │ ├── manager.py # WebSocket 连接管理
│ │ └── manifest.yaml
│ │
│ └── legacy/ # 旧版适配器(保留一段时间后移除)
│ ├── gewechat/
│ ├── nakuru/
│ └── qqbotpy/
└── handlers/ # 事件处理器实现(新增)
├── __init__.py
├── base.py # AbstractEventHandler 基类
├── pipeline_handler.py # PipelineHandler
├── agent_handler.py # AgentHandler
├── webhook_handler.py # WebhookHandler
└── plugin_handler.py # PluginHandler
```
### 2.2 适配器目录内各文件职责
以 Telegram 为例:
| 文件 | 职责 | 关键类/函数 |
|------|------|------------|
| `adapter.py` | 主入口,继承 `AbstractPlatformAdapter`,组装其他模块 | `TelegramAdapter` |
| `event_converter.py` | 将 Telegram 原生事件转换为统一事件类型 | `TelegramEventConverter` — 支持 Message/Edit/Delete/Reaction/MemberJoin 等所有事件 |
| `message_converter.py` | `MessageChain` 与 Telegram 消息格式互转 | `TelegramMessageConverter.yiri2target()` / `target2yiri()` |
| `api_impl.py` | 实现通用 API 方法(edit_message, delete_message, get_group_info 等) | 各 API 方法的 Telegram 实现 |
| `platform_api.py` | 实现 `call_platform_api` 的动作分发表 | `PLATFORM_API_MAP = {"pin_message": ..., "unpin_message": ...}` |
| `types.py` | 平台特有的类型定义 | Telegram 特有的枚举、配置结构等 |
| `manifest.yaml` | 适配器清单:名称、配置 schema、支持的事件和 API 列表 | — |
## 3. 新基类设计
### 3.1 AbstractPlatformAdapter
新基类继承自现有 `AbstractMessagePlatformAdapter` 并扩展,位于 `langbot-plugin-sdk` 中:
```python
# langbot_plugin/api/definition/abstract/platform/adapter.py
class AbstractPlatformAdapter(pydantic.BaseModel, metaclass=abc.ABCMeta):
"""平台适配器基类(EBA 版本)
相比旧版 AbstractMessagePlatformAdapter
- 新增通用 API 方法(edit_message, delete_message, get_group_info 等)
- 新增透传 APIcall_platform_api
- 新增能力声明(get_supported_events, get_supported_apis
- 事件监听器支持所有事件类型,不仅限于消息事件
"""
bot_account_id: str = ""
config: dict
logger: AbstractEventLogger = pydantic.Field(exclude=True)
class Config:
arbitrary_types_allowed = True
# ---- 能力声明 ----
def get_supported_events(self) -> list[str]:
"""返回此适配器支持的事件类型列表
默认实现从 manifest.yaml 读取。
适配器也可以 override 此方法动态声明。
"""
return ["message.received"]
def get_supported_apis(self) -> list[str]:
"""返回此适配器支持的 API 列表
默认实现从 manifest.yaml 读取。
"""
return ["send_message", "reply_message"]
# ---- 必需方法(抽象) ----
@abc.abstractmethod
async def send_message(self, target_type, target_id, message) -> MessageResult:
...
@abc.abstractmethod
async def reply_message(self, event, message, quote_origin=False) -> MessageResult:
...
@abc.abstractmethod
async def run_async(self):
...
@abc.abstractmethod
async def kill(self) -> bool:
...
@abc.abstractmethod
def register_listener(self, event_type, callback):
...
@abc.abstractmethod
def unregister_listener(self, event_type, callback):
...
# ---- 可选方法(默认抛 NotSupportedError ----
# edit_message, delete_message, forward_message,
# get_group_info, get_group_member_list, ...
# call_platform_api, ...
# (完整签名见 02-platform-api.md
# ---- 流式输出(保留) ----
async def reply_message_chunk(self, event, bot_message, message,
quote_origin=False, is_final=False):
raise NotSupportedError("reply_message_chunk")
async def is_stream_output_supported(self) -> bool:
return False
# ---- 消息卡片(保留) ----
async def create_message_card(self, message_id, event) -> bool:
return False
async def is_muted(self, group_id) -> bool:
return False
```
### 3.2 AbstractMessagePlatformAdapter 兼容
旧的 `AbstractMessagePlatformAdapter` 保留为 `AbstractPlatformAdapter` 的类型别名:
```python
# 向后兼容
AbstractMessagePlatformAdapter = AbstractPlatformAdapter
```
现有适配器代码中的 `AbstractMessagePlatformAdapter` 引用不需要立即修改。
### 3.3 EventConverter 新设计
现有 `AbstractEventConverter` 只有 `target2yiri``yiri2target` 两个静态方法,且只处理消息事件。
新设计支持多种事件类型:
```python
class AbstractEventConverter:
"""事件转换器基类(EBA 版本)
适配器需要实现此转换器,将平台原生事件转换为统一事件。
"""
@staticmethod
def target2yiri(raw_event: typing.Any) -> typing.Optional[Event]:
"""将平台原生事件转换为统一事件
Args:
raw_event: 平台 SDK 回调传入的原始事件对象
Returns:
统一 Event 对象,如果无法转换或不需要处理则返回 None
"""
raise NotImplementedError
@staticmethod
def yiri2target(event: Event) -> typing.Any:
"""将统一事件转换为平台原生事件(一般不需要)"""
raise NotImplementedError
```
具体适配器的 EventConverter 实现会是一个分发式的结构:
```python
class TelegramEventConverter(AbstractEventConverter):
"""Telegram 事件转换器"""
@staticmethod
def target2yiri(update: telegram.Update) -> typing.Optional[Event]:
# 消息事件
if update.message:
return TelegramEventConverter._convert_message(update)
# 消息编辑
if update.edited_message:
return TelegramEventConverter._convert_edited_message(update)
# 成员变动
if update.chat_member:
return TelegramEventConverter._convert_chat_member(update)
# 回调查询(按钮点击等)
if update.callback_query:
return TelegramEventConverter._convert_callback_query(update)
# 其他 → PlatformSpecificEvent
return TelegramEventConverter._convert_platform_specific(update)
@staticmethod
def _convert_message(update) -> MessageReceivedEvent:
...
@staticmethod
def _convert_edited_message(update) -> MessageEditedEvent:
...
@staticmethod
def _convert_chat_member(update) -> typing.Union[
MemberJoinedEvent, MemberLeftEvent, ...
]:
...
@staticmethod
def _convert_platform_specific(update) -> PlatformSpecificEvent:
...
```
## 4. Manifest 文件格式扩展
现有 manifest.yaml 只声明 `kind`, `metadata`, `spec.config`, `execution`
新增 `spec.supported_events``spec.supported_apis`
```yaml
apiVersion: v1
kind: MessagePlatformAdapter
metadata:
name: telegram
label:
en_US: Telegram
zh_Hans: Telegram
icon: telegram.svg
description:
en_US: Telegram Bot adapter
zh_Hans: Telegram Bot 适配器
spec:
config:
# 现有配置 schema(保持不变)
- key: token
label: { en_US: "Bot Token", zh_Hans: "Bot Token" }
type: string
required: true
sensitive: true
# ...
supported_events:
- message.received
- message.edited
- message.deleted
- message.reaction
- feedback.received
- group.member_joined
- group.member_left
- group.member_banned
- group.info_updated
- bot.invited_to_group
- bot.removed_from_group
- bot.muted
- bot.unmuted
- platform.specific
supported_apis:
required:
- send_message
- reply_message
optional:
- edit_message
- delete_message
- get_group_info
- get_group_member_list
- get_group_member_info
- get_user_info
- upload_file
- get_file_url
- call_platform_api
platform_specific_apis:
- action: pin_message
description: { en_US: "Pin a message", zh_Hans: "置顶消息" }
- action: unpin_message
description: { en_US: "Unpin a message", zh_Hans: "取消置顶" }
- action: get_chat_administrators
description: { en_US: "Get chat admins", zh_Hans: "获取群管理员列表" }
execution:
python:
path: pkg/platform/adapters/telegram/adapter.py
attr: TelegramAdapter
```
## 5. 适配器注册与发现
### 5.1 Blueprint 更新
`templates/components.yaml` 中更新扫描路径:
```yaml
kind: Blueprint
spec:
components:
MessagePlatformAdapter:
fromDirs:
- path: pkg/platform/adapters/ # 新路径
```
`ComponentDiscoveryEngine` 的递归扫描逻辑不变——它会扫描所有子目录中的 `.yaml` 文件。因此每个适配器目录下的 `manifest.yaml` 会被自动发现。
### 5.2 PlatformManager 适配
`PlatformManager.initialize()` 的核心逻辑基本不变:
```python
async def initialize(self):
# 1. 发现适配器组件(自动扫描新目录结构)
self.adapter_components = self.ap.discover.get_components_by_kind('MessagePlatformAdapter')
# 2. 动态导入适配器类
for component in self.adapter_components:
self.adapter_dict[component.metadata.name] = component.get_python_component_class()
# 3. 从数据库加载 Bot 并实例化适配器(不变)
await self.load_bots_from_db()
```
变更点:
- `execution.python.path``pkg/platform/sources/telegram.py` 变为 `pkg/platform/adapters/telegram/adapter.py`
- `get_python_component_class()` 正常工作,因为它按路径动态导入
## 6. RuntimeBot 重构
### 6.1 现有问题
当前 `RuntimeBot.initialize()` 硬编码注册了两个回调:
```python
# 现有代码
self.adapter.register_listener(platform_events.FriendMessage, on_friend_message)
self.adapter.register_listener(platform_events.GroupMessage, on_group_message)
```
### 6.2 新设计
`RuntimeBot` 改为注册一个通用的事件回调:
```python
class RuntimeBot:
async def initialize(self):
# 注册通用事件回调,接收所有事件类型
self.adapter.register_listener(Event, self._on_event)
async def _on_event(
self,
event: Event,
adapter: AbstractPlatformAdapter,
):
"""统一事件入口"""
# 1. 设置事件的 bot_uuid 和 adapter_name
event.bot_uuid = self.bot_entity.uuid
event.adapter_name = self.bot_entity.adapter
# 2. 日志记录
await self._log_event(event)
# 3. 提交给 EventBus
await self.ap.event_bus.emit(event, adapter)
```
适配器侧的 `register_listener` 实现也需调整:
-`event_type``Event`(基类)时,注册为"接收所有事件"的通配回调
- 适配器在收到平台原生事件时,通过 `EventConverter.target2yiri()` 转换后,调用所有匹配的回调
## 7. 从现有单文件适配器迁移
### 7.1 迁移模式
以 Telegram 为例,从 `sources/telegram.py`445 行)拆分:
| 原代码位置 | → 新文件 |
|-----------|----------|
| `TelegramMessageConverter` 类 | `telegram/message_converter.py` |
| `TelegramEventConverter` 类 | `telegram/event_converter.py`(扩展,支持更多事件) |
| `TelegramAdapter.__init__` / `run_async` / `kill` / `register_listener` | `telegram/adapter.py` |
| `TelegramAdapter.send_message` / `reply_message` / `reply_message_chunk` | `telegram/adapter.py`(消息方法保留在主类)+ `telegram/api_impl.py`(新增 API |
| 新增代码 | `telegram/api_impl.py`edit_message, delete_message, get_group_info 等) |
| 新增代码 | `telegram/platform_api.py`pin_message, unpin_message 等的映射) |
| `telegram.yaml` | `telegram/manifest.yaml`(扩展 supported_events/apis |
### 7.2 迁移顺序建议
1. **Telegram** — 功能最完整的适配器之一,适合作为模板
2. **Discord** — 第二个迁移,验证模式的通用性
3. **AioCQHTTP (OneBot)** — 国内最常用,确保兼容
4. **其他适配器** — 按使用频率排序
### 7.3 渐进式迁移
不需要一次性迁移所有适配器。可以采用渐进策略:
1. 先在 `adapters/` 下建立新适配器
2. `Blueprint` 同时扫描 `sources/``adapters/` 两个目录
3. 旧适配器在 `sources/` 中继续工作
4. 逐个迁移到新结构
5. 全部迁移完成后移除 `sources/` 目录
```yaml
# 过渡期的 Blueprint
kind: Blueprint
spec:
components:
MessagePlatformAdapter:
fromDirs:
- path: pkg/platform/sources/ # 旧路径(尚未迁移的适配器)
- path: pkg/platform/adapters/ # 新路径(已迁移的适配器)
```
-174
View File
@@ -1,174 +0,0 @@
# 事件路由与编排
> 状态:当前实施模型(2026-07-12)。本文以 Pipeline / Agent 平级并存为准,不再保留早期 `pipeline / agent / webhook / plugin` 四种 Handler 草案。
## 1. 路由边界
EBA 将事件处理拆成两个互不替代的阶段:
1. **观察者广播**:EventBus 把事件广播给有权限的插件 EventListener。观察者可记录、同步或执行受控副作用,但不占用响应目标。
2. **响应者仲裁**EventRouter 按 Bot 的 `event_bindings` 选择一个 Pipeline、独立 Agent 或 `discard`
Pipeline 与 Agent 是平级处理器:
| 处理器 | 配置事实源 | 执行路径 | 事件范围 |
| --- | --- | --- | --- |
| Pipeline | Pipeline 表与完整 Stage 配置 | MessageAggregator -> QueryPool -> RuntimePipeline | 消息事件,首版为 `message.received` |
| Agent | Agent 表中的 runner 与 runner config | AgentRunner Host orchestrator -> plugin AgentRunner | Agent/Runner 声明支持的消息或非消息事件 |
| discard | 无处理器配置 | 明确结束路由 | 任意事件 |
插件 EventListener 不是第三种响应目标。Webhook、Dify、n8n、Coze 等外部系统需要响应事件时,由对应 AgentRunner 插件承接。
## 2. 数据模型
### 2.1 独立 Agent
Agent 保存自己的 Runner 选择与配置,不嵌入 Pipeline,也不复制 Pipeline 的 AI stage
```python
class Agent(Base):
uuid: str
name: str
description: str
emoji: str
kind: str # 固定为 "agent"
component_ref: str # AgentRunner id
config: dict # runner + runner_config
enabled: bool
supported_event_patterns: list[str]
```
当前配置形状:
```json
{
"runner": {
"id": "plugin:langbot-team/LocalAgent/default"
},
"runner_config": {
"plugin:langbot-team/LocalAgent/default": {
"model": {
"primary": "model-uuid",
"fallbacks": []
}
}
}
}
```
Runner id 来自已安装插件的 AgentRunner manifest。Host 不维护 LocalAgent、Dify 或其他具体实现的内置分支。
### 2.2 EventBinding
Bot 维护事件到处理器的引用:
```python
class EventBinding(BaseModel):
id: str
event_pattern: str # 精确、namespace.* 或 *
target_type: str # agent | pipeline | discard
target_uuid: str | None # Agent/Pipeline 原始 UUID
filters: list[dict]
priority: int
enabled: bool
description: str
```
示例:
```json
[
{
"id": "binding-message",
"event_pattern": "message.received",
"target_type": "pipeline",
"target_uuid": "pipeline-uuid",
"filters": [],
"priority": 100,
"enabled": true
},
{
"id": "binding-member-joined",
"event_pattern": "group.member_joined",
"target_type": "agent",
"target_uuid": "agent-uuid",
"filters": [],
"priority": 50,
"enabled": true
}
]
```
Binding 只保存引用与路由条件。它不复制 Pipeline 或 Agent 配置。
## 3. 匹配与仲裁
事件模式支持:
- 精确匹配:`group.member_joined`
- 命名空间通配:`group.*`
- 全局通配:`*`
路由按以下顺序处理:
1. 忽略 `enabled = false` 的 binding。
2. 检查 `event_pattern` 与结构化 filters。
3. 校验目标存在、启用且声明支持该事件。
4.`priority` 从高到低选择;同优先级按稳定列表顺序。
5. 只执行一个响应目标。
Pipeline 目标只能匹配消息事件。非消息事件不得伪装成用户文本塞进 Pipeline。
## 4. 执行流程
```text
Platform adapter
-> normalized event
-> EventBus
-> authorized Plugin EventListener observers
-> EventRouter
-> Pipeline target -> full Pipeline stage chain
-> Agent target -> AgentRunner Host orchestrator
-> discard -> stop
-> Host delivery/platform API
```
### 4.1 Pipeline target
消息事件按原有方式构造 Query,经 MessageAggregator、QueryPool 和完整 Pipeline Stage 链执行。Pipeline 可以继续使用 AgentRunner 作为 AI stage 的实现,但 Pipeline 本身不会因此变成 Agent。
### 4.2 Agent target
Host 读取独立 Agent 的 Runner id/config,构造 event-first context、run-scoped resources 与 delivery policy,再调用插件 AgentRunner。Runner 输出由 Host 统一归一化、记录和投递。
AgentRunner 可通过 SDK/Python `AgentRunAPIProxy.call_tool` 或 SDK-owned scoped MCP bridge 回调 Host 能力。两条路径都映射到 `PluginToRuntimeAction.CALL_TOOL`,使用相同的 run authorization、Host execution Query、ToolManager 和 Box session 规则。Box session 是 Host canonical scope 的固定长度安全哈希;同一平台会话稳定、不同 scope 隔离、缺少 identity 时 fail closedRunner 不配置 sandbox scope。
### 4.3 Observer side effects
观察者与响应者可以同时工作。为避免编辑、reaction 等合成事件重复触发不可逆操作,Host 应按事件能力和授权过滤观察者可用 API,并记录副作用结果。Observer 广播不作为 fallback 响应。
## 5. Pipeline 与 Agent 的并存规则
1. Pipeline 与 Agent 保留各自的持久化、编辑和执行语义。
2. 处理器聚合页面可以统一展示二者,但不会创建第三份处理器记录。
3. 旧 Pipeline 仍是 Pipeline;其 runner config 不迁移、不复制为独立 Agent。
4. 需要 Agent 的用户新建 Agent、选择已安装 AgentRunner,再建立 event binding。
5. 一个 Bot 可按不同事件同时绑定 Pipeline 与 Agent。
## 6. WebUI 约束
处理器入口展示带类型标识的 Agent 与 Pipeline。Bot 事件编排器应:
- 按 adapter manifest 展示可用事件;
- 非消息事件不提供 Pipeline 目标;
- Agent 目标按 `supported_event_patterns` 过滤;
- 展示 priority、filters、enabled 与 discard
- 保存前校验目标 UUID 与 `target_type` 一致。
Pipeline 的配置、Debug Chat 和 Monitoring 继续使用 Pipeline 页面;Agent 使用独立表单配置 Runner 与事件能力。
## 7. 版本与迁移边界
此功能按当前 4.x schema 直接实现,不提供 LangBot 3.x 数据库或配置升级路径,也不读取旧 Runner 字段作为 fallback。旧 Pipeline 中的 runner 配置不会生成独立 Agent;用户按新产品模型添加 Agent 即可。
详见 [07-agent-orchestration.md](./07-agent-orchestration.md) 与 [08-agent-page-and-event-orchestration.md](./08-agent-page-and-event-orchestration.md)。
-738
View File
@@ -1,738 +0,0 @@
# 插件 SDK 改造
## 1. 概述
插件 SDK 需要配合 EBA 架构进行以下改造:
1. **新事件类型**:将所有通用事件暴露给插件
2. **新 API**:将新增的平台 API 通过 `LangBotAPIProxy` 暴露给插件
3. **兼容层**:保证现有插件零修改运行
4. **通信协议扩展**:新增 action 枚举支持新 API
## 2. 新事件类型暴露
### 2.1 插件事件模型扩展
当前插件 SDK 的事件模型(`api/entities/events.py`)只有消息相关事件。需要新增所有通用事件的插件级包装:
```python
# api/entities/events.py — 新增事件
# ---- 消息事件(扩展) ----
class MessageEditedReceived(BaseEventModel):
"""消息被编辑事件"""
launcher_type: str
launcher_id: typing.Union[int, str]
message_id: typing.Union[int, str]
editor_id: typing.Union[int, str]
new_content: MessageChain
chat_type: str # "private" | "group"
class MessageDeletedReceived(BaseEventModel):
"""消息被删除/撤回事件"""
launcher_type: str
launcher_id: typing.Union[int, str]
message_id: typing.Union[int, str]
operator_id: typing.Optional[typing.Union[int, str]] = None
chat_type: str
class MessageReactionReceived(BaseEventModel):
"""消息表情回应事件"""
launcher_type: str
launcher_id: typing.Union[int, str]
message_id: typing.Union[int, str]
user_id: typing.Union[int, str]
reaction: str
is_add: bool
# ---- 用户反馈事件 ----
class FeedbackReceived(BaseEventModel):
"""用户对 Bot 回复提交反馈"""
feedback_id: str
feedback_type: int # 1=like, 2=dislike, 3=cancel/remove feedback
feedback_content: typing.Optional[str] = None
inaccurate_reasons: typing.Optional[list[str]] = None
user_id: typing.Optional[str] = None
session_id: typing.Optional[str] = None
message_id: typing.Optional[str] = None
stream_id: typing.Optional[str] = None
# ---- 群组事件 ----
class GroupMemberJoined(BaseEventModel):
"""新成员加入群组"""
group_id: typing.Union[int, str]
group_name: str
member_id: typing.Union[int, str]
member_name: str
inviter_id: typing.Optional[typing.Union[int, str]] = None
join_type: typing.Optional[str] = None
class GroupMemberLeft(BaseEventModel):
"""成员离开群组"""
group_id: typing.Union[int, str]
group_name: str
member_id: typing.Union[int, str]
member_name: str
is_kicked: bool = False
operator_id: typing.Optional[typing.Union[int, str]] = None
class GroupMemberBanned(BaseEventModel):
"""成员被禁言"""
group_id: typing.Union[int, str]
member_id: typing.Union[int, str]
operator_id: typing.Optional[typing.Union[int, str]] = None
duration: typing.Optional[int] = None
class GroupMemberUnbanned(BaseEventModel):
"""成员被解除禁言"""
group_id: typing.Union[int, str]
member_id: typing.Union[int, str]
operator_id: typing.Optional[typing.Union[int, str]] = None
class GroupInfoUpdated(BaseEventModel):
"""群组信息被修改"""
group_id: typing.Union[int, str]
group_name: str
operator_id: typing.Optional[typing.Union[int, str]] = None
changed_fields: list[str] = []
# ---- 好友事件 ----
class FriendRequestReceived(BaseEventModel):
"""收到好友请求"""
request_id: typing.Union[int, str]
user_id: typing.Union[int, str]
user_name: str
message: typing.Optional[str] = None
class FriendAdded(BaseEventModel):
"""成功添加好友"""
user_id: typing.Union[int, str]
user_name: str
class FriendRemoved(BaseEventModel):
"""好友被移除"""
user_id: typing.Union[int, str]
user_name: str
# ---- Bot 状态事件 ----
class BotInvitedToGroup(BaseEventModel):
"""Bot 被邀请加入群组"""
group_id: typing.Union[int, str]
group_name: str
inviter_id: typing.Optional[typing.Union[int, str]] = None
request_id: typing.Optional[typing.Union[int, str]] = None
class BotRemovedFromGroup(BaseEventModel):
"""Bot 被移出群组"""
group_id: typing.Union[int, str]
group_name: str
operator_id: typing.Optional[typing.Union[int, str]] = None
class BotMuted(BaseEventModel):
"""Bot 被禁言"""
group_id: typing.Union[int, str]
operator_id: typing.Optional[typing.Union[int, str]] = None
duration: typing.Optional[int] = None
class BotUnmuted(BaseEventModel):
"""Bot 被解除禁言"""
group_id: typing.Union[int, str]
operator_id: typing.Optional[typing.Union[int, str]] = None
# ---- 平台特有事件 ----
class PlatformSpecificEventReceived(BaseEventModel):
"""平台特有事件"""
adapter_name: str
action: str
data: dict = {}
```
### 2.2 EventListener 注册方式
插件的 EventListener 继续使用 `@self.handler(EventType)` 装饰器注册,只是可以注册的事件类型大幅增加:
```python
class MyEventListener(EventListener):
def __init__(self, host):
super().__init__(host)
# 现有方式(继续工作)
@self.handler(PersonNormalMessageReceived)
async def on_person_message(ctx: EventContext):
...
# 新事件类型
@self.handler(GroupMemberJoined)
async def on_member_joined(ctx: EventContext):
group_name = ctx.event.group_name
member_name = ctx.event.member_name
await ctx.reply(MessageChain([
Plain(f"欢迎 {member_name} 加入 {group_name}")
]))
@self.handler(FriendRequestReceived)
async def on_friend_request(ctx: EventContext):
# 自动通过好友请求
await ctx.approve_friend_request(
ctx.event.request_id, approve=True
)
@self.handler(FeedbackReceived)
async def on_feedback(ctx: EventContext):
if ctx.event.feedback_type == 2:
await self.log_warning(
f"用户点踩了回复: {ctx.event.feedback_content or ''}"
)
@self.handler(PlatformSpecificEventReceived)
async def on_platform_event(ctx: EventContext):
if ctx.event.adapter_name == "telegram" and ctx.event.action == "chat_join_request":
...
```
## 3. 新 API 暴露
### 3.1 LangBotAPIProxy 扩展
`LangBotAPIProxy` 中新增以下方法,插件通过 `self.xxx()` 调用(在 BasePlugin 中继承):
```python
class LangBotAPIProxy:
# ---- 现有方法(保留) ----
# get_langbot_version, get_bots, get_bot_info,
# send_message, invoke_llm, get/set/delete_plugin_storage, ...
# ---- 新增消息 API ----
async def edit_message(
self,
bot_uuid: str,
chat_type: str,
chat_id: typing.Union[int, str],
message_id: typing.Union[int, str],
new_content: MessageChain,
) -> None:
"""编辑已发送的消息"""
...
async def delete_message(
self,
bot_uuid: str,
chat_type: str,
chat_id: typing.Union[int, str],
message_id: typing.Union[int, str],
) -> None:
"""删除/撤回消息"""
...
async def forward_message(
self,
bot_uuid: str,
from_chat_type: str,
from_chat_id: typing.Union[int, str],
message_id: typing.Union[int, str],
to_chat_type: str,
to_chat_id: typing.Union[int, str],
) -> dict:
"""转发消息"""
...
async def get_message(
self,
bot_uuid: str,
chat_type: str,
chat_id: typing.Union[int, str],
message_id: typing.Union[int, str],
) -> dict:
"""获取指定消息"""
...
# ---- 新增群组 API ----
async def get_group_info(
self,
bot_uuid: str,
group_id: typing.Union[int, str],
) -> dict:
"""获取群组信息"""
...
async def get_group_list(
self,
bot_uuid: str,
) -> list[dict]:
"""获取 Bot 加入的群组列表"""
...
async def get_group_member_list(
self,
bot_uuid: str,
group_id: typing.Union[int, str],
) -> list[dict]:
"""获取群成员列表"""
...
async def get_group_member_info(
self,
bot_uuid: str,
group_id: typing.Union[int, str],
user_id: typing.Union[int, str],
) -> dict:
"""获取指定群成员信息"""
...
async def mute_member(
self,
bot_uuid: str,
group_id: typing.Union[int, str],
user_id: typing.Union[int, str],
duration: int = 0,
) -> None:
"""禁言群成员"""
...
async def unmute_member(
self,
bot_uuid: str,
group_id: typing.Union[int, str],
user_id: typing.Union[int, str],
) -> None:
"""解除禁言"""
...
async def kick_member(
self,
bot_uuid: str,
group_id: typing.Union[int, str],
user_id: typing.Union[int, str],
) -> None:
"""踢出群成员"""
...
# ---- 新增用户 API ----
async def get_user_info(
self,
bot_uuid: str,
user_id: typing.Union[int, str],
) -> dict:
"""获取用户信息"""
...
async def get_friend_list(
self,
bot_uuid: str,
) -> list[dict]:
"""获取好友列表"""
...
async def approve_friend_request(
self,
bot_uuid: str,
request_id: typing.Union[int, str],
approve: bool = True,
remark: typing.Optional[str] = None,
) -> None:
"""处理好友请求"""
...
async def approve_group_invite(
self,
bot_uuid: str,
request_id: typing.Union[int, str],
approve: bool = True,
) -> None:
"""处理入群邀请"""
...
# ---- 新增透传 API ----
async def call_platform_api(
self,
bot_uuid: str,
action: str,
params: dict = {},
) -> dict:
"""调用适配器特有 API
Examples:
# Telegram: pin 消息
result = await self.call_platform_api(
bot_uuid, "pin_message",
{"chat_id": 123456, "message_id": 789}
)
# Discord: 创建频道
result = await self.call_platform_api(
bot_uuid, "create_channel",
{"guild_id": "...", "name": "new-channel"}
)
"""
...
# ---- 新增能力查询 API ----
async def get_supported_events(
self,
bot_uuid: str,
) -> list[str]:
"""获取指定 Bot 的适配器支持的事件类型"""
...
async def get_supported_apis(
self,
bot_uuid: str,
) -> list[str]:
"""获取指定 Bot 的适配器支持的 API"""
...
```
### 3.2 QueryBasedAPIProxy 扩展
在事件处理上下文中(EventContext),通过 `QueryBasedAPIProxy` 新增便捷方法:
```python
class QueryBasedAPIProxy:
# ---- 现有方法(保留) ----
# reply, get_bot_uuid, set_query_var, get_query_var,
# create_new_conversation, ...
# ---- 新增便捷方法 ----
async def edit_message(
self,
message_id: typing.Union[int, str],
new_content: MessageChain,
) -> None:
"""在当前会话中编辑消息(自动使用当前 bot_uuid 和 chat 信息)"""
...
async def delete_message(
self,
message_id: typing.Union[int, str],
) -> None:
"""在当前会话中删除消息"""
...
async def approve_friend_request(
self,
request_id: typing.Union[int, str],
approve: bool = True,
remark: typing.Optional[str] = None,
) -> None:
"""处理好友请求(上下文中自动获取 bot_uuid)"""
...
async def approve_group_invite(
self,
request_id: typing.Union[int, str],
approve: bool = True,
) -> None:
"""处理入群邀请"""
...
async def get_group_info(self) -> dict:
"""获取当前群组信息(仅群聊事件中可用)"""
...
async def get_group_member_list(self) -> list[dict]:
"""获取当前群组成员列表(仅群聊事件中可用)"""
...
async def call_platform_api(
self,
action: str,
params: dict = {},
) -> dict:
"""调用平台特有 API(自动使用当前 bot_uuid"""
...
```
## 4. 兼容层设计
### 4.1 事件兼容层
当 PluginHandler 将新的 `MessageReceivedEvent` 分发给插件时,需要同时生成旧格式事件:
```python
class PluginEventCompatLayer:
"""插件事件兼容层
将新的统一事件转换为旧的插件事件格式,
确保监听旧事件类型的插件仍能正常工作。
"""
@staticmethod
def convert_to_legacy_events(
event: Event,
) -> list[BaseEventModel]:
"""将统一事件转换为旧插件事件列表
一个统一事件可能生成多个旧插件事件。
例如 MessageReceivedEvent 会同时生成:
- PersonMessageReceived / GroupMessageReceived(总是生成)
- PersonNormalMessageReceived / GroupNormalMessageReceived(非命令时)
- PersonCommandSent / GroupCommandSent(命令时)
"""
legacy_events = []
if isinstance(event, MessageReceivedEvent):
if event.chat_type == ChatType.PRIVATE:
legacy_events.append(
PersonMessageReceived(
launcher_type="person",
launcher_id=event.chat_id,
sender_id=event.sender.id,
message_event=event.to_legacy_friend_message(),
message_chain=event.message_chain,
)
)
# 命令检测后还会生成 PersonNormalMessageReceived
# 或 PersonCommandSent,在 Pipeline 阶段处理
elif event.chat_type == ChatType.GROUP:
legacy_events.append(
GroupMessageReceived(
launcher_type="group",
launcher_id=event.chat_id,
sender_id=event.sender.id,
message_event=event.to_legacy_group_message(),
message_chain=event.message_chain,
)
)
# 新事件类型没有旧的对应物,不生成兼容事件
# 只有监听了新事件类型的插件才会收到
return legacy_events
```
### 4.2 分发流程
```
统一事件 (MessageReceivedEvent)
├─→ 转换为旧格式 (PersonMessageReceived / GroupMessageReceived)
│ └─→ 分发给监听旧事件类型的插件 EventListener
└─→ 直接分发为新格式 (MessageReceivedEvent → 对应的插件事件)
└─→ 分发给监听新事件类型的插件 EventListener
```
插件 Runtime 在分发事件时检查每个 EventListener 注册的事件类型:
- 如果注册的是旧类型(`PersonMessageReceived` 等),发送兼容层生成的旧格式事件
- 如果注册的是新类型(`GroupMemberJoined` 等),发送新格式事件
- 两者可以共存,同一个插件可以同时监听新旧类型
### 4.3 API 兼容层
现有插件使用的 API 不受影响:
| 现有 API | 新架构行为 |
|---------|----------|
| `self.send_message(bot_uuid, target_type, target_id, message_chain)` | 不变,直接调用适配器的 `send_message` |
| `ctx.reply(message_chain, quote_origin)` | 不变,在 MessageReceivedEvent 上下文中调用适配器的 `reply_message` |
| `self.get_bots()` | 不变 |
| `self.get_bot_info(bot_uuid)` | 不变 |
新 API 只是额外新增的方法,不影响现有方法。
## 5. 通信协议扩展
### 5.1 新增 Action 枚举
`entities/io/actions/enums.py` 中新增 action
```python
class PluginToRuntimeAction(str, Enum):
# ---- 现有 actions(保留) ----
REGISTER_PLUGIN = "register_plugin"
REPLY = "reply"
SEND_MESSAGE = "send_message"
# ...
# ---- 新增消息 API ----
EDIT_MESSAGE = "edit_message"
DELETE_MESSAGE = "delete_message"
FORWARD_MESSAGE = "forward_message"
GET_MESSAGE = "get_message"
# ---- 新增群组 API ----
GET_GROUP_INFO = "get_group_info"
GET_GROUP_LIST = "get_group_list"
GET_GROUP_MEMBER_LIST = "get_group_member_list"
GET_GROUP_MEMBER_INFO = "get_group_member_info"
MUTE_MEMBER = "mute_member"
UNMUTE_MEMBER = "unmute_member"
KICK_MEMBER = "kick_member"
# ---- 新增用户 API ----
GET_USER_INFO = "get_user_info"
GET_FRIEND_LIST = "get_friend_list"
APPROVE_FRIEND_REQUEST = "approve_friend_request"
APPROVE_GROUP_INVITE = "approve_group_invite"
# ---- 新增透传 API ----
CALL_PLATFORM_API = "call_platform_api"
# ---- 新增能力查询 ----
GET_SUPPORTED_EVENTS = "get_supported_events"
GET_SUPPORTED_APIS = "get_supported_apis"
class RuntimeToPluginAction(str, Enum):
# ---- 现有 actions(保留) ----
EMIT_EVENT = "emit_event"
# ...
# EMIT_EVENT 的 data 结构扩展以支持新事件类型
```
### 5.2 新增 Action 的请求/响应格式
`EDIT_MESSAGE` 为例:
```json
// 请求 (Plugin → Runtime)
{
"action": "edit_message",
"seq_id": 12345,
"data": {
"bot_uuid": "...",
"chat_type": "group",
"chat_id": "123456",
"message_id": "789",
"new_content": [
{ "type": "Plain", "text": "edited message" }
]
}
}
// 响应 (Runtime → Plugin)
{
"seq_id": 12345,
"code": 0,
"message": "ok",
"data": {}
}
```
`GET_GROUP_MEMBER_LIST` 为例:
```json
// 请求
{
"action": "get_group_member_list",
"seq_id": 12346,
"data": {
"bot_uuid": "...",
"group_id": "123456"
}
}
// 响应
{
"seq_id": 12346,
"code": 0,
"message": "ok",
"data": {
"members": [
{
"user": { "id": "111", "nickname": "Alice" },
"group_id": "123456",
"role": "admin",
"display_name": "管理员Alice"
},
...
]
}
}
```
`CALL_PLATFORM_API` 为例:
```json
// 请求
{
"action": "call_platform_api",
"seq_id": 12347,
"data": {
"bot_uuid": "...",
"action": "pin_message",
"params": {
"chat_id": "123456",
"message_id": "789"
}
}
}
// 响应
{
"seq_id": 12347,
"code": 0,
"message": "ok",
"data": {
"result": { ... }
}
}
```
### 5.3 LangBot 侧 Handler 实现
`ControlConnectionHandler`LangBot → Runtime 侧)和 `PluginConnectionHandler`Runtime → Plugin 侧)中新增对应的 action 处理逻辑:
```python
# PluginConnectionHandler 中新增
async def _handle_edit_message(self, data):
bot_uuid = data["bot_uuid"]
bot = await self.ap.platform_mgr.get_bot_by_uuid(bot_uuid)
await bot.adapter.edit_message(
chat_type=data["chat_type"],
chat_id=data["chat_id"],
message_id=data["message_id"],
new_content=MessageChain.model_validate(data["new_content"]),
)
return {}
async def _handle_call_platform_api(self, data):
bot_uuid = data["bot_uuid"]
bot = await self.ap.platform_mgr.get_bot_by_uuid(bot_uuid)
result = await bot.adapter.call_platform_api(
action=data["action"],
params=data.get("params", {}),
)
return {"result": result}
```
## 6. 插件开发者迁移指南
### 6.1 无需迁移(零修改运行)
以下场景的现有插件**不需要任何修改**:
- 使用 `PersonNormalMessageReceived` / `GroupNormalMessageReceived` 监听消息
- 使用 `PersonCommandSent` / `GroupCommandSent` 处理命令
- 使用 `ctx.reply()` 回复消息
- 使用 `self.send_message()` 主动发消息
- 使用 LLM / 存储 / RAG 等现有 API
### 6.2 推荐迁移(获得新能力)
如果插件希望利用新功能,可以:
1. **监听新事件类型**:在 EventListener 中注册新事件类型的 handler
2. **使用新 API**:调用 `self.edit_message()`, `self.get_group_info()`
3. **使用透传 API**:调用 `self.call_platform_api()` 使用平台特有功能
### 6.3 SDK 版本号
新功能通过提升 SDK minor 版本发布:
- 现有版本:`langbot-plugin-sdk >= x.y.z`
- 新版本:`langbot-plugin-sdk >= x.(y+1).0`
插件的 `manifest.yaml` 中的 `min_sdk_version` 决定是否能使用新 API。使用旧 SDK 版本的插件在新 LangBot 上正常运行(兼容层保证),只是无法调用新 API。
@@ -1,145 +0,0 @@
# EBA 分阶段实施计划
> 更新:2026-07-12。文件名沿用早期设计,但这里的“迁移”仅指代码架构逐步接入 EBA,不代表 LangBot 3.x 数据库或配置升级。
## 1. 发布边界
EBA 跨越 SDK、平台适配器、LangBot Host、WebUI 与插件生态,按可验证阶段落地。当前发布遵守以下硬边界:
- LangBot 4.x 不支持从 3.x 数据库或配置升级;不保留 legacy migration chain、旧 JSON 模板或旧 Runner 字段读取。
- Pipeline 与 Agent 平级且长期并存,分别保留持久化模型与执行链。
- 现有 Pipeline 不迁移为 AgentPipeline 内的 runner config 不复制到 Agent。
- 用户需要 Agent 时新建独立 Agent并选择已安装的 AgentRunner。
- Host 不按 LocalAgent id 做运行时、Box 或 WebUI 特判。
- AgentRunner 的 SDK/Python 与 scoped MCP bridge 回调共享 Host 授权与事件 session 规则。
## 2. 阶段总览
| 阶段 | 目标 | 主要仓库 | 完成条件 |
| --- | --- | --- | --- |
| P0 | SDK 事件、能力与 AgentRunner 协议 | `langbot-plugin-sdk` | typed entities、manifest、proxy、runtime action 通过测试 |
| P1 | 平台适配器 EBA 化 | LangBot + SDK | 事件转换、能力声明、通用/透传 API 通过 adapter checklist |
| P2 | Host 观察者与响应者路由 | LangBot backend | observer 广播 + Pipeline/Agent/discard 单目标仲裁可运行 |
| P3 | 独立 Agent 与 Runner 注册 | LangBot backend + plugins | Agent CRUD、registry、run authorization、delivery 可运行 |
| P4 | WebUI 处理器与事件编排 | LangBot web | Agent/Pipeline 聚合入口和 Bot binding 编辑器可用 |
| P5 | 发布门禁与文档 | LangBot skills/docs + runner plugins | 单测、UI E2E、真实 adapter smoke 和插件预检通过 |
## 3. P0SDK 契约
### 工作项
- 定义规范化平台事件、actor/subject/conversation/delivery context。
- 定义 adapter `supported_events``supported_apis` 与平台透传 API。
- 定义 AgentRunner manifest、run context/result、resource handles 和 pull/callback API。
- 提供 `AgentRunAPIProxy` 与 SDK-owned scoped MCP bridge。
- 保持协议传输与权限校验可测试,不把 Host 私有 Query 对象暴露给插件。
### 验收
- stdio/WebSocket runtime 对 typed action 的序列化一致。
-`run_id` 或越权资源调用被拒绝。
- Python proxy 与 MCP bridge 对同一 Host tool 呈现一致结果/错误形状。
## 4. P1:平台适配器
每个适配器按自己的能力增量接入,而不是要求所有平台一次性支持全部事件。
### 单适配器步骤
1. 将原生 SDK callback 转换为规范化事件。
2. 声明实际支持的事件与 API,不把缺失能力伪装为成功。
3. 实现 send/reply/edit/delete/group/user 等适用 API 与平台透传。
4. 对消息链、媒体、reply target 和错误语义写单元测试。
5.`adapters/acceptance-checklist.md` 记录真实平台 probe。
### 验收
- `message.received` 保持正常收发。
- 新事件不会重复转换或产生循环合成事件。
- `supported_apis` 与实际调用能力一致。
## 5. P2Host 路由
### 执行顺序
```text
adapter event
-> EventBus observer broadcast
-> EventRouter match event_bindings
-> one target: Pipeline | Agent | discard
-> Host delivery
```
### 约束
- Plugin EventListener 是 observer,不作为 priority fallback。
- Pipeline 只处理消息事件并复用完整 Stage 链。
- Agent 使用独立 Agent 配置和 AgentRunner Host orchestrator。
- edit/reaction 等事件的 observer 副作用能力按事件和 adapter 能力过滤。
- dry-run 与合成派发必须使用同一匹配器,避免 UI 预览与真实路由漂移。
### 验收
- 精确、namespace wildcard、全局 wildcard、filters 与 priority 有单元覆盖。
- 同一事件最多一个响应目标,但 observer 仍能收到事件。
- Pipeline 与 Agent 可以在同一个 Bot 的不同 binding 中同时生效。
## 6. P3:独立 Agent 与 AgentRunner
### 工作项
- `agents` 只保存 Agent;Pipeline 继续使用自己的表和 API。
- Agent config 使用 `runner.id``runner_config[runner_id]`
- registry 只展示已安装、有效的插件 AgentRunner。
- Host 构造 run-scoped resources、state、delivery 与 event log/transcript。
- SDK/Python `call_tool` 和 scoped MCP bridge 都回到同一个 Host ToolManager。
- Box session 由 Host 将 instance/workspace/bot/adapter/target/thread scope 规范化并哈希为固定长度 `lb-box-<sha256>`;同 scope 稳定、不同 scope 隔离、缺少 identity 时 fail closed。
### 验收
- 安装/卸载 Runner 后 metadata 与表单选项同步。
- Runner 无法调用未授权 model/tool/knowledge/state。
- 两种 Host callback transport 不能覆盖 sandbox session id。
- LocalAgent、ACP/ClaudeCode/Codex 与外部服务 Runner 不需要 Host id 特判。
## 7. P4WebUI
### 工作项
- `/home/agents` 聚合显示 Agent 与 Pipeline,并明确类型。
- 创建时选择 Agent 或 Pipeline,编辑时进入各自表单。
- Agent 表单读取动态 Runner metadata,保存当前 `runner` / `runner_config` 形状。
- Bot 事件编排器编辑 `event_pattern`、target、filters、priority 与 enabled。
- 非消息事件过滤 Pipeline;Agent 按声明事件能力过滤。
### 验收
- 页面不出现 LocalAgent 专属 banner、变量隐藏或 Box/Pipeline 注入逻辑。
- 空 Runner 市场状态给出可安装 AgentRunner 的正常路径。
- Pipeline Debug Chat/Monitoring 与 Agent 运行日志分别可用。
## 8. P5:发布门禁
### 自动化
- LangBot backend unit/integration tests 与 Ruff。
- SDK AgentRunner/proxy/MCP bridge tests。
- Web lint/build 与关键 Playwright cases。
- `skills/bin/lbs validate``skills/bin/lbs index --check`
- LocalAgent 与其他官方 Runner plugin package/test gate。
### 真实环境
- 至少一个消息事件走 Pipeline。
- 至少一个非消息事件走独立 Agent 并执行平台动作。
- SDK/Python 和 MCP bridge 各完成一次受权工具调用,并证明二者都进入同一个 `PluginToRuntimeAction.CALL_TOOL` Host handler。
- Box 可用/不可用的降级路径可观察且无 Runner 特判。
## 9. 非目标
- 不把 Pipeline 改名或包装成 Agent。
- 不自动把 Pipeline 内 runner 配置迁移成 Agent。
- 不为 3.x 保留数据库迁移、旧模板或配置 fallback。
- 不要求所有 Runner 经 MCP;本地 Python Runner 可以直接使用 SDK。
- 不允许 Runner 自定义 Box session scope。
- 不在首版实现多 Agent 串并联;多步骤编排留给后续 workflow。
@@ -1,185 +0,0 @@
# Agent 与 Pipeline 统一编排(产品最终形态)
> **状态**:方向修订稿(2026-06-12),供「适配器改造 / Agent 插件化 / 工作流引擎」三条工作线评审。
>
> 本文档修订 [00-overview.md](./00-overview.md) §3.4 与 [04-event-routing.md](./04-event-routing.md) 中"四种 Handler"的编排模型:**所有编排目标统一进入处理器选择与事件绑定界面,但独立 Agent 与现有 Pipeline 保持不同类型**。事件路由的匹配机制、数据迁移策略、WebUI 交互骨架等内容仍以 04 为准,仅 handler 分类法被本文档取代。
## 1. 产品最终形态
**适配器接收各种事件 → 用户编排处理逻辑 → 处理器统一选择表面**,实现从 0 代码到低代码再到全代码的全层面支持:
```
消息平台 (Telegram / Discord / 企微 / ...)
│ 各类平台事件
平台适配器(EBA 新结构,已迁移 12 个)
│ EBAEvent (message.* / group.* / friend.* / bot.* / feedback.* / platform.*)
EventRouter(事件 → 处理器绑定)
├─→ 选中的处理器(响应者,单一仲裁)
│ ├─ Pipeline:保留现有实体和执行链,仅处理消息事件
│ └─ Agent:用户新建并选择 AgentRunner 插件,可接本地、低代码或外部 runtime
└─→ 插件 EventListener(观察者,N 个广播,可 prevent_default
```
| 编写方式 | 处理器形态 | 代码化程度 |
|----------|-----------|-----------|
| WebUI 配置模型 + 提示词 + 工具 | 独立 Agent + LocalAgent 插件 | 0 代码 |
| 可视化消息编排 | Pipeline`kind=pipeline`,完整多 Stage 链) | 0 代码 |
| 市场安装 | Agent 插件(市场分发) | 0 代码(使用者视角) |
| 可视化工作流 | 工作流引擎定义的 Agent | 低代码 |
| 对接外部平台 | dify / n8n / coze / webhook 外部 Agent | 集成 |
| SDK 编写 | Agent 插件组件 | 全代码 |
### 1.1 三条并行工作线与汇合点
| 工作线 | 范围 | 在本架构中的位置 |
|--------|------|------------------|
| 适配器改造(refactor/eba,本分支) | 事件体系、适配器结构、平台 API、EventRouter | 事件的**生产侧** + 路由层 |
| Agent 插件化 | Agent 抽象、Agent 组件类型、市场分发 | 事件的**消费侧**统一抽象 |
| 工作流引擎 | 内部低代码工作流 | Agent 的一种**编写方式** |
**汇合点是 SDK 的 Agent 组件契约(§4)与 event→处理器绑定模型(§3)**。这两个接口冻结后,三条线可彼此 mock 独立推进。契约由本分支(EBA)牵头起草,三线评审后在 langbot-plugin-sdk 落地(发布通道:0.5.0aX pre-release 已打通)。
## 2. 从四种 Handler 到统一处理器表面
### 2.1 演进理由
04 文档中的 pipeline / agent / webhook / plugin 四种 handler_type,本质上都是"对事件作出响应的逻辑",差别只在编写和部署方式。产品层统一展示和绑定这些处理器,但不会把既有 Pipeline 持久化为 Agent
- **产品**:用户只需理解"给 Bot 的事件绑定处理器",处理器可以是 Pipeline 或 Agent
- **工程**:路由层按 `target_type` 分发到 Pipeline 或 AgentAgent 的扩展集中到 AgentRunner 抽象;
- **生态**:Agent 成为市场上可分发、可复用的一等公民。
### 2.2 收编映射
| 原 handler_type04 文档) | 收编后 |
|---------------------------|--------|
| `pipeline` | 保留 Pipeline 实体;binding 使用 `target_type=pipeline` 和原 `pipeline_uuid`,进程内直接复用 MessageAggregator → QueryPool → Pipeline 机制 |
| `agent`RequestRunner | 用户新建独立 Agent,并选择对应 AgentRunner 插件;不读取或复制旧 Pipeline 内嵌 runner 配置 |
| `webhook` | 外部 Agent 的一种:事件 POST 出去、响应解析为动作(保留 04 §5.4 的请求/响应格式) |
| `plugin`EventListener 分发) | **不收编**——角色不同,见 §2.3 |
### 2.3 响应者与观察者的角色切分
事件的消费方有两种角色,不应混为一谈:
- **响应者(Pipeline 或 Agent)**:路由选中**一个**,负责对事件作出回应(回复消息、执行动作)。多条绑定匹配同一事件时按 priority 仲裁,只取最高者。
- **观察者(插件 EventListener)**:**广播**给所有注册插件,做旁路逻辑(日志、审计、风控、统计)。沿用现有机制不变,包括 `prevent_default()`——观察者可拦截本次事件,使处理器不被调用(与现有"插件拦截流水线"行为完全兼容)。
执行顺序:事件到达 → 先广播观察者(按插件优先级)→ 若未被 prevent_default → 分发给选中的处理器。
## 3. 数据模型:event → 处理器绑定
### 3.1 独立 Agent 与现有 Pipeline
Agent 与 Pipeline 都是一等处理器。用户创建 Agent、选择已安装的 AgentRunner,再把适合的事件绑定到 AgentPipeline 继续保存在 Pipeline 表中,以完整 Stage 链处理消息事件。两者可在同一处理器列表中以不同 `kind` 展示和选择;这种聚合展示不会创建额外记录,也不会在两种模型之间复制配置。
```python
class Agent(Base):
"""Agent 实例:一个具体配置过的、可被事件绑定的响应者"""
uuid: str # 主键
name: str
kind: str # 固定为 "agent"Pipeline 使用自己的持久模型
component_ref: str # AgentRunner id,例如 plugin:<author>/<plugin>/<runner>
config: dict # JSON — runner id、runner config 与资源/状态/投递策略
# 多租户预留:归属主体字段(tenant/workspace),首版可空
```
Bot 上的绑定配置(替代 04 §2.2 的 EventHandlerConfig,沿用其匹配语义):
```python
class EventBinding(pydantic.BaseModel):
event_pattern: str # 精确 / "message.*" / "*",匹配规则同 04 §4
target_type: str # "agent" | "pipeline" | "discard"
target_uuid: str # Agent 或 Pipeline 的原始 UUIDdiscard 时为空
enabled: bool = True
priority: int = 0 # 多条匹配时取最高者(单一仲裁)
description: str = ''
```
`use_pipeline_uuid` 只迁移 Bot 的路由结构:写入 `{"event_pattern": "message.received", "target_type": "pipeline", "target_uuid": <原 pipeline uuid>}`,继续引用原 Pipeline。迁移不会创建独立 Agent,也不会复制 Pipeline 内嵌的 runner 配置;需要 Agent 的用户自行新增 Agent 并修改 binding。观察者广播不需要配置(始终发生),04 中"兜底 plugin 规则"不再需要。
## 4. SDK Agent 组件契约(草案)
Agent 成为插件系统的第七种组件(现有:Command / Tool / EventListener / KnowledgeEngine / Parser / Page)。
### 4.1 Manifest
```yaml
apiVersion: v1
kind: Agent
metadata:
name: group-assistant
label: { en_US: Group Assistant, zh_Hans: 群助理 }
spec:
handled_events: # 声明可处理的事件类型;绑定 UI 据此过滤
- message.received
- group.member_joined
config: # 实例化配置 schema,复用现有组件配置体系
- name: model
type: llm-model-selector
- name: persona
type: prompt-editor
execution:
python: { path: agent.py, attr: GroupAssistant }
```
### 4.2 运行时接口
```python
class Agent(BaseComponent):
async def handle(self, ctx: AgentContext) -> typing.AsyncGenerator[AgentChunk, None]:
"""处理一次事件,流式产出回复与动作。每次事件调用一次。"""
...
class AgentContext:
event: EBAEvent # 触发事件(统一事件体系)
bot: BotHandle # 来源 Bot 信息
session: SessionHandle # 会话句柄:历史消息、会话变量(LangBot 侧管理,Agent 保持无状态)
config: dict # 该 Agent 实例的配置
# 能力面(经 runtime RPC 回 LangBot 执行):
async def reply(self, chain: MessageChain, quote: bool = False): ...
async def send_message(self, target_type: str, target_id: str, chain: MessageChain): ...
async def call_platform_api(self, action: str, params: dict) -> dict: ...
async def invoke_llm(self, model_uuid: str, messages: list, funcs: list = None) -> dict: ...
# + 工具调用 / KB 检索 / 插件存储(沿用 LangBotAPIProxy 既有方法)
class AgentChunk:
delta_message: MessageChain | None = None # 增量回复(流式)
actions: list[dict] | None = None # 平台动作(同 webhook response_actions 格式)
final: bool = False
```
**流式**:复用 SDK 通信协议既有的 `chunk_status: continue/end` 机制,`handle()` 的每次 yield 对应一个 chunk。
**Pipeline 与 Agent 分流**Pipeline target 继续走 LangBot 进程内的 Pipeline 执行链;独立 Agent 经 AgentRunner 插件 runtime 分发。路由层通过 binding 的 `target_type` 明确区分二者。
### 4.3 执行语义与可靠性
| 关注点 | 约定 |
|--------|------|
| 仲裁 | 单响应者:priority 最高的匹配绑定生效,其余忽略 |
| 性能 | Pipeline 继续走进程内链路;插件 Agent 每事件过一次 RPC 边界,消息场景需设延迟预算(评审项:目标 P95 附加延迟) |
| 会话状态 | 归 LangBot 侧(SessionHandle),插件 Agent 原则上无状态,崩溃重启不丢会话 |
| 降级 | Agent 调用失败/超时:可配置 fallback(回错误提示,或指定备用处理器);不会隐式回退到旧 Pipeline 配置 |
| 多租户预留 | AgentContext / SessionHandle / 存储接口显式携带归属主体标识,禁止新增全局单例状态——为后续轻量 SaaS 多租户铺路 |
## 5. 发布火车
| 版本 | 内容 | 备注 |
|------|------|------|
| 4.11(可选) | 现状成果:12 个 EBA 适配器、插件全事件订阅、`call_platform_api` | 对用户不可见的管道工程 + 插件新能力,不动产品概念 |
| **5.0** | 产品形态首发:EventRouter + event→处理器绑定 + WebUI 编排 + 旧 Bot 路由迁移 + 独立 Agent / AgentRunner 插件 + SDK Agent 组件契约(可标 experimental | `use_pipeline_uuid` 仅改写为指向原 Pipeline 的 binding,不生成 Agent;配 SDK 0.5.0 正式版;走 beta 周期 |
| 5.x | 工作流 Agent(工作流引擎线挂入)、Agent 市场生态、剩余适配器(satori 等)、Agent 插件化收尾 | 验证开放注册机制 |
| 多租户 | 独立评估:仅数据隔离 → 5.x 部署选项;伴随权限/计费/产品定位变化 → 6.0 | 前置条件是 §4.3 的归属主体预留已落实 |
## 6. 开放问题(评审清单)
1. **webhook 的最终定位**:作为外部 Agent(响应者,现方案)之外,是否还需要"纯通知观察者"形态(现 WebhookPusher 的角色)?
2. **多 Agent 协作**:单一仲裁之外,是否需要"串联/并联多个 Agent"的场景?(建议 5.0 不做,留给工作流引擎表达)
3. **工作流引擎的宿主**:核心内置,还是自身也作为一个插件交付(解释工作流定义的 Agent 插件)?
4. **插件 Agent 的延迟预算**:消息主链路过 RPC 的 P95 目标值与压测方案。
5. **Workflow 的处理器边界**:未来 Workflow 应作为与 Pipeline、Agent 平级的处理器,还是作为其中一种处理器的内部编排实现。
6. **SDK 1.0 时机**:Agent 契约稳定后是否随 LangBot 5.x 给插件生态一个 API 稳定承诺。
@@ -1,181 +0,0 @@
# 处理器页面与事件编排产品设计
> 状态:实施稿(2026-06-23
>
> 本文档修订 [07-agent-orchestration.md](./07-agent-orchestration.md) 中“Agent 替代 Pipeline”的表述。当前产品形态保留两种长期并存的同级处理器:**Agent** 与 **Pipeline**。处理器页面只是共享入口,不改变二者各自的持久化模型和执行语义。
## 1. 产品边界
LangBot 的处理逻辑分成两种同级形态:
| 形态 | 定位 | 可处理事件 | 典型用户 |
| --- | --- | --- | --- |
| Agent | runner 驱动的事件优先处理器,承载 AgentRunner / 外部 runner | `message.*``group.*``friend.*``bot.*``feedback.*``platform.*` 等声明范围 | 需要直接处理多类平台事件或接入外部 agent runtime 的用户 |
| Pipeline | 可视化、可控、可组合的消息处理流水线,执行完整 Stage 链 | 仅 `message.*`,首版等价于 `message.received` | 需要预处理、AI、后处理、扩展和输出控制的消息场景 |
处理器页面负责统一管理这两种处理单元:
- 创建时选择 **Agent****Pipeline**
- 列表中清晰标注类型与事件能力;
- Pipeline 的编辑、调试、监控继续复用现有能力;
- Agent 保存 runner 配置与事件能力,并通过 Bot 事件绑定进入运行时执行。
## 2. 信息架构
### 2.1 处理器页面
路径:`/home/agents`
职责:
1. 展示所有可被事件绑定的处理单元,包括 Agent 与 Pipeline。
2. 创建时先选择类型:
- Agent:创建一条独立 Agent 配置对象,默认支持其 runner 声明的事件范围;
- Pipeline:创建一条独立 Pipeline,执行完整消息 Stage 链。
3. 编辑时按类型进入不同表单:
- Pipeline:沿用原 Pipeline 配置页,包括 AI、触发、安全、输出、扩展、Debug、Monitoring
- Agent:配置基础信息、runner、runner config 和事件能力。
`/home/pipelines` 继续提供 Pipeline 直接编辑路径;共享处理器入口当前使用 `/home/agents`。URL 是实现路径,不代表 Agent 包含 Pipeline。
### 2.2 Bot 的事件编排
Bot 上维护“事件 -> 处理单元”的绑定规则:
```text
Bot
└─ EventBinding[]
├─ event_pattern: message.received / group.member_joined / group.* / *
├─ target_type: agent / pipeline / discard
├─ target_uuid: Agent UUID 或 Pipeline UUID
├─ filters: 事件字段过滤条件
├─ priority: 数字越大越优先
└─ enabled
```
Pipeline 只能被绑定到 `message.*`。如果用户选择非消息事件,目标选择器不展示 Pipeline。
## 3. 持久化模型
### 3.1 Agent 实例与 Pipeline 实例
`agents` 表只保存 AgentPipeline 继续保存在 Pipeline 表中。当前物理表名 `legacy_pipelines` 是既有存储名称,不代表 Pipeline 在产品架构中是遗留或过渡形态。
```python
class Agent(Base):
uuid: str
name: str
description: str
emoji: str
kind: str # 首版固定为 "agent"
component_ref: str # runner id / workflow id / future external ref
config: dict # runner 与 runner_config
enabled: bool
supported_event_patterns: list[str]
created_at: datetime
updated_at: datetime
```
处理器聚合服务把 `agents` 与 Pipeline 表投影成同一个前端列表:
```json
{
"uuid": "...",
"name": "...",
"kind": "agent | pipeline",
"capability": {
"supported_event_patterns": ["*"],
"message_only": false
}
}
```
Pipeline 投影时固定:
```json
{
"kind": "pipeline",
"capability": {
"supported_event_patterns": ["message.*"],
"message_only": true
}
}
```
### 3.2 Bot 事件绑定
Bot 新增 `event_bindings` JSON 字段,首版作为轻量配置面。后续当 EventRouter 查询、审计和多作用域规则稳定后,再拆成独立表。
```json
[
{
"id": "uuid",
"event_pattern": "group.member_joined",
"target_type": "agent",
"target_uuid": "...",
"filters": [],
"priority": 100,
"enabled": true,
"description": "Welcome new group members"
}
]
```
## 4. 匹配规则
事件模式支持三层:
1. 精确匹配:`group.member_joined`
2. 命名空间通配:`group.*`
3. 全局通配:`*`
优先级:
1. `enabled = true`
2. event pattern 命中
3. filters 全部命中
4. `priority` 数值高者优先
5. 同优先级按列表顺序
## 5. 并存策略
1. Pipeline 与 Agent 长期并存,各自保存配置并执行自己的运行链路。
2. 现有 Bot 的 `use_pipeline_uuid` 转换为仍指向原 Pipeline 的消息事件绑定。
3. 现有 `pipeline_routing_rules` 仍只作用于消息事件。
4. `event_bindings` 允许 `target_type=pipeline|agent|discard`Pipeline 目标只限 `message.*`
5. Pipeline 与 Agent 保留各自的持久化和编辑语义;处理器聚合入口只负责统一展示和选择。
## 6. 分阶段落地
### P0:处理器入口统一
- 新增 `/home/agents`
- 侧边栏显示“处理器”,列表包含平级的 Agent 与 Pipeline。
- 创建时选择 Agent 或 Pipeline。
- `/home/pipelines` 继续作为 Pipeline 直接编辑路径。
### P1:配置模型落地
- 新增 `agents` 表与 `/api/v1/agents`
- Agent 可保存 runner 与 runner_config。
- Pipeline 继续使用原 Pipeline 表单与 API。
### P2:事件编排配置面
- Bot 表单新增事件编排编辑器。
- 读取 adapter manifest 的 `supported_events` 生成事件选项。
- 根据事件类型过滤可选目标:Pipeline 仅在 `message.*` 可选。
### P3EventRouter 执行接入
- EBA 事件先广播插件 observer。
- 然后按 `event_bindings` 的事件模式、filters、priority 和顺序选择一个处理器。
- Pipeline 目标通过 MessageAggregator 进入完整 Pipeline Stage 链;Agent 目标直接进入 AgentRunner 链路。
- 非消息事件只选择声明支持该事件的 Agent,不调用 PipelineAgentRunner 输出有平台 reply target 时会投递回平台。
## 7. 不做的事
- 不把 Pipeline 改名成 Agent,也不删除 Pipeline 的配置模型。
- 不把 Pipeline 降级为兼容层、Agent 子类型或临时运行入口。
- 不把非消息事件伪装成用户文本塞入 Pipeline。
- 不在首版做多 Agent 串并联;需要多步骤处理时留给后续 workflow。
@@ -1,41 +0,0 @@
# EBA Adapter Migration Records
This directory records adapter-level migration details for the Event-Based Agents architecture. Each adapter document should be kept close to the implementation and must answer four questions:
1. What changed in the adapter structure.
2. Which configuration fields are required.
3. Which events and APIs are supported.
4. What has been verified end to end.
## Adapter Documents
General acceptance checklist: [EBA Adapter Acceptance Checklist](./acceptance-checklist.md)
Current acceptance report: [EBA Adapter Acceptance Report](./acceptance-report.md)
| Adapter | Status | Document |
|---------|--------|----------|
| Telegram | Migrated; partial plugin E2E, real UI inbound image/file verified | [Telegram](./telegram.md) |
| Discord | Migrated; partial plugin E2E, media-inbound gaps remain | [Discord](./discord.md) |
| OneBot v11 / aiocqhttp | Migrated; Matcha UI plus protocol-level multi-component coverage | [OneBot v11 / aiocqhttp](./aiocqhttp.md) |
| DingTalk | Migrated; partial plugin E2E, real UI inbound image/file verified; group gap remains | [DingTalk](./dingtalk.md) |
| Lark / Feishu | Migrated; partial live text E2E, media-inbound gap remains | [Lark / Feishu](./lark.md) |
| WeCom | Migrated; private text plugin E2E verified, media/group gaps remain | [WeCom](./wecom.md) |
| WeComBot | Migrated; private text and outbound/API plugin E2E verified, feedback/group gaps remain | [WeComBot](./wecombot.md) |
| Official Account | Migrated; private text plugin E2E verified, proactive outbound not supported | [Official Account](./officialaccount.md) |
| QQ Official API | Migrated; WebSocket inbound reached LangBot, model config blocked reply | [QQ Official API](./qqofficial.md) |
| Slack | Migrated; private text and outbound/API plugin E2E verified | [Slack](./slack.md) |
| WeCom Customer Service | Migrated; customer-side UI text plugin E2E verified, inbound media and platform-API live coverage pending | [WeCom Customer Service](./wecomcs.md) |
| Kook | Migrated; unit/mocked converter and API coverage only, live acceptance pending | [Kook](./kook.md) |
## Documentation Checklist
When migrating a new adapter, add one document here with:
- Configuration table matching the adapter manifest.
- Supported event list.
- Supported common API list.
- Supported `call_platform_api` action list.
- Known unsupported APIs and the reason.
- Live test notes, including platform, channel type, destructive operations, and residual risks.
- A clear distinction between real UI inbound media, protocol-level injected inbound media, and bot outbound media.
@@ -1,208 +0,0 @@
# EBA Adapter Acceptance Checklist
This checklist is the architecture-level acceptance standard for every Event-Based Agents platform adapter. It is not platform-specific. Adapter migration is not complete until the adapter has a written result against this checklist.
## Evidence Levels
Use these evidence levels consistently in adapter records:
| Level | Meaning | Can Mark Complete |
|-------|---------|-------------------|
| `plugin-e2e-ui` | Real SDK plugin running through standalone runtime, LangBot core, the migrated adapter, and a real platform/simulator UI action. | Yes |
| `plugin-e2e-protocol` | Real SDK plugin running through standalone runtime, LangBot core, and the migrated adapter from a protocol-boundary event injection, such as a OneBot reverse WebSocket event. | Partial; must not be claimed as UI coverage |
| `plugin-e2e-outbound` | Real SDK plugin calls an API and the bot output is visible in the real platform/simulator UI. | Yes for send/API coverage only |
| `adapter-live` | Direct adapter probe connected to a real or simulator platform endpoint, bypassing plugin runtime. | No, auxiliary only |
| `unit` | Unit/API-shape tests with mocked platform SDK objects or mocked APIs. | No, auxiliary only |
| `not-supported` | Platform protocol or SDK has no equivalent capability. Must include reason and source. | Yes, as explicitly unsupported |
| `blocked` | Intended capability could not be verified because of credentials, permissions, endpoint gaps, or simulator gaps. | No |
The primary acceptance path must be `plugin-e2e-ui` for inbound UI-triggered behavior and `plugin-e2e-outbound` for bot send/API behavior. `adapter-live`, `plugin-e2e-protocol`, and `unit` tests are useful, but they must be labelled precisely.
## Required Architecture Path
Every adapter must prove this full path:
```text
Real platform / simulator UI
-> platform SDK native event
-> adapter event converter
-> unified EBA event/entity/message types
-> LangBot core event dispatch
-> standalone SDK runtime
-> real test plugin listener
-> plugin calls platform APIs through SDK
-> LangBot core API dispatch
-> adapter API implementation
-> real platform / simulator UI
```
The test plugin must record JSONL evidence containing:
- event class and `event.type`
- `bot_uuid` and `adapter_name` as received by the plugin
- adapter name
- chat type and chat ID
- sender/user/group IDs with secrets redacted
- message component list for received messages
- API action name, input summary, result or error
- raw unsupported/blocked reason when an item is skipped
## Required Message Receive Tests
For every adapter, inbound message conversion must be tested through `plugin-e2e-ui` for each component the platform can receive. If a protocol-level injection is used, label it `plugin-e2e-protocol`; it proves the adapter/core/plugin path, but it does not prove that the user-facing platform UI can send that component. If the platform UI/simulator cannot create a component, record it as `blocked` with the endpoint limitation.
| Component | Required Receive Assertion |
|-----------|----------------------------|
| `Source` | Message ID and timestamp are present and stable enough for reply/get/delete APIs. |
| `Plain` | Text is preserved exactly, including spaces and multi-line content. |
| `At` | Mentioned user ID is converted to common `At.target`. |
| `AtAll` | Broadcast mention is converted to common `AtAll`, if platform supports it. |
| `Image` | Image ID, URL, path, or base64 is represented without leaking platform-native segment shape. |
| `Voice` | Voice/audio component is represented as `Voice` when the platform exposes it. |
| `File` | File name, ID/URL, and size are represented as `File` when available. |
| `Quote` | Reply/quote source ID and origin content are represented when the platform exposes it. |
| `Face` | Native emoji/sticker/dice/rps-like components are represented as `Face` or documented as platform-specific. |
| `Forward` | Merged/forwarded messages are represented as `Forward` when the platform exposes structured content. |
| `Unknown` | Unsupported native segments become `Unknown` or `PlatformSpecificEvent` data, not crashes. |
| Mixed chain | A message containing multiple component types preserves order. |
The plugin must subscribe to `MessageReceivedEvent` and assert that `message_chain` contains common `langbot_plugin.api.entities.builtin.platform.message` components, not platform-native SDK objects.
## Required Message Send Tests
For every adapter, outbound message conversion must be tested through `plugin-e2e-outbound` by having the plugin call SDK platform APIs and verifying the platform UI/simulator receives the expected message.
| Component | Required Send Assertion |
|-----------|-------------------------|
| `Plain` | Text appears exactly on the platform. |
| `At` | User mention renders as a mention or platform equivalent. |
| `AtAll` | Broadcast mention renders or is explicitly unsupported. |
| `Image` | URL, path, or base64 image sends and renders/downloads correctly. |
| `Voice` | Voice/audio sends when supported. |
| `File` | File sends with name and content/link when supported. |
| `Quote` | Quoted reply points to the original message when supported. |
| `Face` | Native emoji/sticker/dice/rps sends or is explicitly unsupported. |
| `Forward` | Forward/merged-forward sends when supported; otherwise fallback behavior is documented. |
| Mixed chain | A mixed chain preserves component order as closely as the platform allows. |
If a platform supports a component only in one direction, the adapter record must say so explicitly.
## Required Event Tests
The plugin must subscribe to every event declared in `manifest.yaml -> spec.supported_events` and record one of `plugin-e2e-ui`, `plugin-e2e-protocol`, `not-supported`, or `blocked`.
| Event | Required Assertion |
|-------|--------------------|
| `message.received` | Real message reaches plugin as `MessageReceivedEvent`. |
| `message.edited` | Edited message reaches plugin with message ID and new content, if declared. |
| `message.deleted` | Deleted/recalled message reaches plugin with message ID and operator when available, if declared. |
| `message.reaction` | Reaction add/remove reaches plugin with message ID, user, reaction, and direction, if declared. |
| `feedback.received` | Feedback payload reaches plugin with feedback type and message/session IDs, if declared. |
| `group.member_joined` | Join event reaches plugin with group and member. |
| `group.member_left` | Leave/kick event reaches plugin with group, member, and kick flag. |
| `group.member_banned` | Mute/ban event reaches plugin with group, member, operator, and duration. |
| `group.info_updated` | Group metadata update reaches plugin with changed fields, if declared. |
| `friend.request_received` | Friend request reaches plugin with request ID and message. |
| `friend.added` | Friend-added event reaches plugin. |
| `friend.removed` | Friend-removed event reaches plugin, if declared. |
| `bot.invited_to_group` | Bot invite/join request reaches plugin with group and inviter/request ID. |
| `bot.removed_from_group` | Bot removal reaches plugin with group and operator when available. |
| `bot.muted` | Bot mute reaches plugin with duration. |
| `bot.unmuted` | Bot unmute reaches plugin. |
| `platform.specific` | At least one unmapped native event is delivered as structured platform-specific data, if declared. |
Do not declare an event in the manifest unless there is an implementation path and an acceptance entry.
## Required Common API Tests
The plugin must call every common API declared in `manifest.yaml -> spec.supported_apis.required` and `optional`. Each call must be recorded with input summary and result.
| API | Required Assertion |
|-----|--------------------|
| `send_message` | Plugin sends to private and group/channel targets where supported. |
| `reply_message` | Plugin replies to the triggering message, with quoted mode tested when supported. |
| `edit_message` | Plugin edits a bot-sent message, if declared. |
| `delete_message` | Plugin deletes/recalls a bot-sent message, if declared and permissions allow. |
| `forward_message` | Plugin forwards or emulates forwarding a real message, if declared. |
| `get_message` | Plugin retrieves a real message and receives common `MessageReceivedEvent` shape. |
| `get_group_info` | Plugin receives `UserGroup` with ID/name/count where available. |
| `get_group_list` | Plugin receives joined groups/channels list where supported. |
| `get_group_member_list` | Plugin receives list of `UserGroupMember` where supported. |
| `get_group_member_info` | Plugin receives one member with role/display name where available. |
| `set_group_name` | Plugin changes and restores a disposable group name, if declared. |
| `mute_member` | Plugin mutes a disposable target, if declared. |
| `unmute_member` | Plugin unmutes the same target, if declared. |
| `kick_member` | Plugin kicks a disposable target only in destructive test mode, if declared. |
| `leave_group` | Plugin leaves only in destructive test mode and only at the end, if declared. |
| `get_user_info` | Plugin receives common `User` shape. |
| `get_friend_list` | Plugin receives friend/contact list where supported. |
| `approve_friend_request` | Plugin accepts/rejects a disposable friend request, if declared. |
| `approve_group_invite` | Plugin accepts/rejects a disposable group invite, if declared. |
| `upload_file` | Plugin uploads a real small file, if declared. |
| `get_file_url` | Plugin resolves a real file ID to a URL, if declared. |
| `call_platform_api` | Plugin calls every declared platform-specific action with safe parameters. |
Destructive APIs must be opt-in and documented with the exact target used.
The SDK must expose a plugin-side platform API escape hatch for adapter-specific actions. The acceptance plugin should call it from the same EBA event handler that received the real platform event, so the evidence proves both directions of the path:
```text
plugin -> SDK call_platform_api -> LangBot core -> adapter call_platform_api -> platform SDK/API
```
The result must be serialized into JSON-safe values before it is returned to the plugin runtime.
## Platform-Specific API Tests
Every action listed in `manifest.yaml -> spec.platform_specific_apis` must have one acceptance entry:
- `plugin-e2e-ui` or `plugin-e2e-outbound`: called by the plugin against the live/simulator endpoint.
- `plugin-e2e-protocol`: called by the plugin after a protocol-boundary injected event; useful for endpoint-specific simulators but must be labelled.
- `not-supported`: removed from manifest or explained if the platform SDK exposes it but this adapter intentionally does not.
- `blocked`: endpoint did not implement it, permissions missing, or safe fixture unavailable.
Do not leave a platform-specific API in the manifest without a corresponding test record.
## Required Compatibility Tests
Each migrated adapter must also prove:
- Manifest supported events match `adapter.get_supported_events()`.
- Manifest supported APIs match `adapter.get_supported_apis()`.
- Manifest platform-specific actions match `PLATFORM_API_MAP`.
- Legacy `FriendMessage` / `GroupMessage` listeners still work when the core registers them.
- EBA listener dispatch prefers the most specific event class, then `EBAEvent`, then base `Event`.
- Self-message filtering prevents bot echo loops without dropping edit/delete/moderation events needed for API tests.
- `source_platform_object` is present for reply/debug but not required by plugins for common behavior.
## Required Documentation Per Adapter
Each adapter document must include:
- adapter directory and manifest name
- config table
- supported event table with evidence level per event
- supported common API table with evidence level per API
- platform-specific API table with evidence level per action
- receive component table with evidence level per component
- send component table with evidence level per component
- exact test date
- exact platform endpoint or simulator used
- standalone runtime command
- plugin path/name used for testing
- evidence JSONL path
- destructive operations performed or explicitly skipped
- blocked items and reasons
## Acceptance Rule
An adapter can be marked migrated only when:
1. All declared events have `plugin-e2e-ui`, justified `plugin-e2e-protocol`, or `not-supported` evidence.
2. All declared APIs have `plugin-e2e-outbound` or `not-supported` evidence.
3. All platform-supported receive components have `plugin-e2e-ui` evidence; protocol-only receive coverage keeps the status partial.
4. All platform-supported send components have `plugin-e2e-outbound` evidence.
5. Unit tests cover conversion and API-shape boundaries.
6. The adapter document lists every blocked or skipped item honestly.
If any declared capability is only covered by `adapter-live` or `unit`, the adapter status must remain partial.
@@ -1,171 +0,0 @@
# EBA Adapter Acceptance Report
Date: May 10, 2026
Scope:
- `telegram-eba`
- `discord-eba`
- `aiocqhttp-eba`
- `dingtalk-eba`
- `lark-eba`
- `wecom-eba`
- `wecombot-eba`
- `wecomcs-eba`
- `officialaccount-eba`
- `qqofficial-eba`
- `slack-eba`
This report follows `acceptance-checklist.md`. Evidence levels are intentionally strict:
- `plugin-e2e-ui`: real platform or simulator UI event reached LangBot, standalone runtime, and `EBAEventProbe`.
- `plugin-e2e-protocol`: real adapter endpoint event reached LangBot, standalone runtime, and `EBAEventProbe`, but the event was injected at the platform protocol boundary rather than sent through the UI.
- `plugin-e2e-outbound`: the plugin called SDK APIs and the resulting bot message was visible on the platform.
- `unit`: mocked converter/API coverage only.
- `blocked`: not completed, either because the platform/simulator/client could not trigger it or because a safe disposable fixture was unavailable.
- `not-supported`: the platform has no equivalent capability.
## Summary
| Adapter | Status | Honest acceptance summary |
|---------|--------|---------------------------|
| Telegram | Partial EBA acceptance | Real Telegram UI covered private text, group mention text, bot invite, inbound private image/file, outbound component sweep, safe SDK APIs, and safe Telegram platform APIs. Real UI inbound voice/quote was not completed in the latest plugin run. |
| Discord | Partial EBA acceptance | Real Discord UI covered group text, outbound image/file/quote/mention components, safe SDK APIs, and safe Discord platform APIs. Real UI inbound attachment/image/file/reply/mention was not completed. A later UI retry was blocked because the Discord client kept the send button disabled. |
| OneBot v11 / aiocqhttp | Partial EBA acceptance | Matcha UI covered real group text and outbound supported components/APIs. Multi-component inbound `Source/Plain/At/Face/Image/Voice/File/Quote` was verified through the real OneBot reverse WebSocket adapter endpoint, but not through Matcha UI upload/send. Matcha blocks file-send and merged-forward APIs. |
| DingTalk | Partial EBA acceptance | Real DingTalk UI covered private text, emoji-as-text inbound, private inbound image/file, outbound image/file/quote/mention fallback components, safe SDK APIs, and safe DingTalk platform APIs. Real UI inbound voice/quote and group trigger were not completed. |
| Lark / Feishu | Partial EBA acceptance | EBA adapter structure, self-built/store app config, WebSocket/Webhook mode handling, converters, common APIs, platform APIs, and unit tests are in place. One real LangBot organization WebSocket private text event reached `EBAEventProbe`; outbound component sweep was visible in Feishu. Latest real UI image/file sends did not reach local plugin evidence, so media receive remains blocked. |
| WeCom | Partial EBA acceptance | Regular WeCom application-message adapter is split into the EBA directory with manifest, converters, API mixin, platform API map, and unit tests. Private text reached `EBAEventProbe` through standalone runtime and the real WeCom client; safe plugin APIs passed. Real inbound media and broader event coverage remain pending. |
| WeComBot | Partial EBA acceptance | WeCom AI Bot is split into the EBA directory with WebSocket long connection mode and optional webhook mode, EBA message/feedback/platform-specific conversion, cache-backed common APIs, platform API map, unit tests, and a direct live probe. Private text, outbound component sweep, safe common APIs, and all declared WeComBot platform APIs reached `EBAEventProbe`; group, real inbound media, and feedback callback evidence remain pending. |
| WeCom Customer Service | Partial EBA acceptance | WeCom Customer Service is split into the EBA directory with manifest, converters, API mixin, platform API map, unit tests, docs, and a direct live probe scaffold. Real WeChat customer-side UI text reached `EBAEventProbe`; plugin outbound text/image and safe cache-backed common APIs passed. Inbound media and platform-specific API live coverage remain pending; later fallback text sends were blocked by WeCom `95001 send msg count limit`. |
| Official Account | Partial EBA acceptance | WeChat Official Account is split into the EBA directory with manifest, converters, cache-backed safe APIs, platform API map, unit tests, and a direct live probe scaffold. Real WeChat Official Account UI private text reached `EBAEventProbe`; safe cache-backed common APIs and declared platform APIs passed. Proactive outbound `send_message` is not supported because replies must be tied to inbound webhook windows; inbound image/voice live UI evidence remains pending. |
| QQ Official API | Partial EBA acceptance | QQ Official API is split into the EBA directory with manifest, converters, cache-backed safe APIs, platform API map, unit tests, docs, and a direct live probe scaffold. A real WebSocket-mode QQ Official bot reached the LangBot pipeline on `dev.rockchin.top`; reply/outbound evidence is blocked by the test model provider returning `model_not_found` for `deepseek-v3`. |
| Slack | Partial EBA acceptance | Slack is split into the EBA directory with manifest, converters, cache-backed safe APIs, platform API map, unit tests, docs, and a direct live probe scaffold. Real Slack private text reached `EBAEventProbe`; safe common APIs, outbound component fallback sweep, and declared Slack platform APIs passed. Channel mention and real inbound media evidence remain pending. |
Telegram and DingTalk now have real user-side UI image/file upload evidence in plugin JSONL. Discord and aiocqhttp do not yet have real UI inbound image/file evidence.
## Evidence Files
| Adapter | Endpoint | Evidence |
|---------|----------|----------|
| Telegram private | Telegram Lite, `@rockchinq_bot` private chat | `data/temp/telegram-plugin-e2e-rerun.jsonl` |
| Telegram private media | Telegram Lite, `@rockchinq_bot` private chat | `data/temp/telegram-plugin-e2e-media-ui.jsonl` |
| Telegram group | Telegram Lite, `Rock'sBotGroup` | `data/temp/telegram-plugin-e2e-group.jsonl` |
| Discord | Discord client, LangBot server, `#debugging` | `data/temp/discord-plugin-e2e-20260510-final.jsonl` |
| aiocqhttp UI | local Matcha, group `test group` | `data/temp/aiocqhttp-plugin-e2e-20260510-multiformat.jsonl` |
| aiocqhttp protocol | OneBot reverse WebSocket endpoint `127.0.0.1:2280/ws` | `data/temp/aiocqhttp-plugin-e2e-20260510-multiformat.jsonl` |
| DingTalk | DingTalk Mac, `LangBot Team` org private chat | `data/temp/dingtalk-plugin-e2e-20260510-rerun.jsonl` |
| DingTalk private media | DingTalk Mac, `LangBot Team` org private chat | `data/temp/dingtalk-plugin-e2e-media-ui.jsonl` |
| Lark / Feishu unit | local mocked Feishu SDK/client paths | `tests/unit_tests/platform/test_lark_eba_adapter.py` |
| Lark / Feishu partial live | Feishu Mac, LangBot organization `LangBotDev` private chat | `data/temp/lark-plugin-e2e-ws.jsonl` |
| WeCom Customer Service | WeChat customer-side UI, `客服消息 -> 浪波智能客服` on `dev.rockchin.top` | `/home/wgc/LangBotxg/LangBotEbaTest/data/temp/wecomcs_eba_plugin_probe.jsonl` |
| Official Account | WeChat desktop client, subscribed Official Account on `dev.rockchin.top` | `/home/wgc/LangBotxg/LangBotEbaTest/data/temp/officialaccount_eba_plugin_probe.jsonl` |
| QQ Official API unit | local mocked QQ Official client paths | `tests/unit_tests/platform/test_qqofficial_eba_adapter.py` |
| Slack unit | local mocked Slack client paths | `tests/unit_tests/platform/test_slack_eba_adapter.py` |
| Slack private | Slack workspace private DM on `dev.rockchin.top` | `/home/wgc/LangBotxg/LangBotEbaTest/data/temp/slack_eba_plugin_probe.jsonl` |
All plugin runs used SDK standalone runtime ports `5400/5401`, LangBot `--standalone-runtime`, and the real plugin at `langbot-plugin-demo/EBAEventProbe`.
## Unified Shape Verification
All four adapters deliver common SDK entities to plugins before LangBot core/plugin logic handles the event.
| Requirement | Telegram | Discord | aiocqhttp | DingTalk | Lark / Feishu |
|-------------|----------|---------|-----------|----------|---------------|
| `bot_uuid` filled | plugin-e2e | plugin-e2e | plugin-e2e | plugin-e2e | live plugin-e2e pending |
| `adapter_name` filled | `telegram` | `discord` | `aiocqhttp` | `dingtalk` | `lark-eba` in current unit/code; older live text evidence recorded `lark` before the naming fix |
| common `MessageChain` delivered | `Plain`, group `At + Plain`, private `Image`, private `File` | `Source + Plain` | UI `Source + Plain`; protocol `Source + Plain + At + Face + Image + Voice + File + Quote + Plain` | `Source + Plain`, private `Source + Image`, private `Source + File` | live private `Source + Plain`; unit `Source + Plain + At/Image/File`; latest live image/file blocked |
| common user/group entities | plugin-e2e | plugin-e2e | plugin-e2e | plugin-e2e private user; group not completed | live private user; unit private/group |
| raw native object isolation | raw data stays in `source_platform_object` | raw data stays in `source_platform_object` | raw data stays in `source_platform_object` | raw data stays in `source_platform_object` | raw data stays in `source_platform_object` |
## Message Receive Components
| Component | Telegram | Discord | aiocqhttp | DingTalk | Lark / Feishu |
|-----------|----------|---------|-----------|----------|---------------|
| `Source` | design gap: event has message id but chain omits `Source` | plugin-e2e-ui | plugin-e2e-ui/protocol | plugin-e2e-ui | plugin-e2e-ui private text |
| `Plain` | plugin-e2e-ui private/group | plugin-e2e-ui | plugin-e2e-ui/protocol | plugin-e2e-ui | plugin-e2e-ui private text |
| `At` | plugin-e2e-ui group mention | unit; real UI mention not completed in latest run | plugin-e2e-protocol; unit | unit; group trigger not completed | unit; group trigger not completed |
| `AtAll` | not-supported | unit only | unit only | unit/send fallback only | unit only |
| `Image` | plugin-e2e-ui private | converter/unit; real UI attachment not completed | plugin-e2e-protocol, not Matcha UI | plugin-e2e-ui private | unit; real UI image sent but not observed in plugin evidence |
| `Voice` | converter/unit; real UI inbound not completed | not-supported as native voice; audio is attachment/file | plugin-e2e-protocol, not Matcha UI | converter/unit; real UI inbound not completed | unit; real UI inbound not completed |
| `File` | plugin-e2e-ui private | converter/unit; real UI attachment not completed | plugin-e2e-protocol, not Matcha UI | plugin-e2e-ui private | unit; real UI file sent but not observed in plugin evidence |
| `Quote` | converter/unit; real UI reply not completed | unit; real UI reply not completed | plugin-e2e-protocol | converter/unit; real UI quote not completed | unit/API-backed quote lookup; real UI quote not completed |
| `Face` | not-supported as common `Face` | not-supported as common `Face` | plugin-e2e-protocol | UI emoji becomes `Plain` (`[smile]` text), not `Face` | not-supported as common `Face` |
| `Forward` | not-supported inbound | not-supported inbound | unit; Matcha forward UI/action blocked | not-supported inbound | not-supported inbound |
| Mixed chain | group `At + Plain`; media tested as separate messages | not completed inbound | plugin-e2e-protocol | media tested as separate messages; mixed inbound not completed | unit only |
## Message Send Components
| Component | Telegram | Discord | aiocqhttp | DingTalk | Lark / Feishu |
|-----------|----------|---------|-----------|----------|---------------|
| `Plain` | plugin-e2e-outbound | plugin-e2e-outbound | plugin-e2e-outbound | plugin-e2e-outbound | plugin-e2e-outbound |
| `At` | plugin-e2e-outbound equivalent | plugin-e2e-outbound | plugin-e2e-outbound | plugin-e2e-outbound fallback/equivalent | plugin-e2e-outbound |
| `AtAll` | plugin-e2e-outbound fallback | plugin-e2e-outbound | plugin-e2e-outbound | plugin-e2e-outbound fallback | unit; group live not completed |
| `Image` | plugin-e2e-outbound | plugin-e2e-outbound | plugin-e2e-outbound | plugin-e2e-outbound | plugin-e2e-outbound |
| `Voice` | not-supported in current send converter | not-supported as native voice | converter path; not completed against Matcha UI | fallback as file/text depending DingTalk media support | converter path; live not completed |
| `File` | plugin-e2e-outbound | plugin-e2e-outbound | blocked by Matcha endpoint error | plugin-e2e-outbound | plugin-e2e-outbound |
| `Quote` | plugin-e2e-outbound | plugin-e2e-outbound | plugin-e2e-outbound | plugin-e2e-outbound fallback | plugin-e2e-outbound fallback |
| `Face` | not-supported | not-supported | plugin-e2e-outbound attempted in mixed chain | fallback text | not-supported |
| `Forward` | flattened fallback | flattened fallback | blocked by Matcha unsupported action | flattened fallback | plugin-e2e-outbound flattened fallback |
| Mixed chain | plugin-e2e-outbound | plugin-e2e-outbound | plugin-e2e-outbound except blocked file/forward | plugin-e2e-outbound | plugin-e2e-outbound |
## Event Acceptance
| Event category | Telegram | Discord | aiocqhttp | DingTalk |
|----------------|----------|---------|-----------|----------|
| `message.received` | plugin-e2e-ui | plugin-e2e-ui | plugin-e2e-ui and plugin-e2e-protocol | plugin-e2e-ui private |
| `message.edited` | implemented/unit, not plugin-e2e-ui | historical/direct only, not latest plugin-e2e | unit | not declared |
| `message.deleted` | implemented/unit, not plugin-e2e-ui | historical/direct only, not latest plugin-e2e | unit | not declared |
| `message.reaction` | implemented/unit, not plugin-e2e-ui | historical/direct only, not latest plugin-e2e | not-supported in standard OneBot message path | not declared |
| member join/left/ban | implemented/unit or blocked without disposable users | blocked without disposable users | unit; Matcha fixture unavailable | not declared |
| bot invited/removed | invite plugin-e2e-ui for Telegram; removal blocked | invite historical/plugin-series; removal blocked | unit; Matcha fixture unavailable | not declared |
| requests/friend events | not applicable | not applicable | unit; Matcha fixture unavailable | not declared |
| `platform.specific` | implemented; not latest plugin-e2e | not latest plugin-e2e | adapter lifecycle observed; plugin focus was message path | declared for fallback; not reproduced in UI run |
## Common API Acceptance
| API area | Telegram | Discord | aiocqhttp | DingTalk |
|----------|----------|---------|-----------|----------|
| send/reply | plugin-e2e-outbound | plugin-e2e-outbound | plugin-e2e-outbound, with Matcha file/forward gaps | plugin-e2e-outbound |
| edit/delete | historical/direct or unit; destructive/current UI not repeated | historical/direct; destructive/current UI not repeated | unit/destructive blocked | not declared or blocked |
| message lookup | not-supported | not-supported | plugin-e2e | inbound cache-backed where available; limited live coverage |
| group info/member info | plugin-e2e safe subset | plugin-e2e safe subset | plugin-e2e safe subset | private path only; group not completed |
| user/friend info | plugin-e2e where platform allows | plugin-e2e where platform allows | plugin-e2e | plugin-e2e private user |
| moderation/leave | blocked without disposable safe targets | blocked without disposable safe targets | blocked without disposable safe targets | blocked/not declared |
| `get_file_url` | implemented; latest inbound `File` carried downloadable file data in plugin evidence | URL passthrough for attachments; inbound attachment not completed | not portable/endpoint-dependent | implemented through DingTalk media API; latest inbound `File` carried a platform file URL |
| `call_platform_api` | plugin-e2e safe actions | plugin-e2e safe actions | plugin-e2e safe actions, Matcha gaps documented | plugin-e2e safe `check_access_token` |
## Platform-Specific API Acceptance
| Adapter | plugin-e2e verified | Blocked or not reproduced |
|---------|---------------------|---------------------------|
| Telegram | safe chat/admin/member count/chat-action actions | mutating actions and callback-only actions were not repeated |
| Discord | safe channel/guild/role/typing actions | mutating pin/reaction/invite actions were not repeated in the latest plugin run; inbound attachment paths not completed |
| aiocqhttp | safe OneBot actions such as status/version/can-send checks | `get_group_honor_info` unsupported by Matcha; admin/card/title/ban/record/file/forward require better endpoint fixtures |
| DingTalk | `check_access_token`; real inbound file produced a file URL in the common `File` component | separate media-download replay APIs and group actions need a working follow-up fixture |
## SDK API Acceptance
`EBAEventProbe` exercised the standalone runtime path for:
- bot discovery and bot info lookup
- send message
- component sweep where enabled
- platform API sweep where enabled
- plugin storage
- workspace storage
- plugin/command/tool/knowledge-base list APIs
The probe logs set `ok=true` when the sweep completed with only expected unsupported/blocked items. Individual call details are stored in the JSONL evidence files.
## Residual Risks And Required Follow-Up
- Discord still requires real UI inbound image/file upload evidence before it can be called media-complete.
- aiocqhttp has rich inbound component evidence only at the OneBot reverse WebSocket boundary; Matcha UI did not provide image/file upload coverage.
- DingTalk group trigger remains unclosed; current evidence is private chat only.
- Lark / Feishu requires a clean follow-up live pass: the latest LangBot organization WebSocket run connected, but UI-sent text/image/file after the loop-scheduling fix did not append plugin events.
- Discord UI retry on May 10, 2026 was blocked by the client keeping the send button disabled even after text was entered.
- Destructive moderation and leave APIs are intentionally blocked until disposable users/groups are available.
## Conclusion
The EBA conversion path is implemented and partially proven for the migrated adapters. Telegram and DingTalk now have real UI private-chat image/file inbound evidence. Discord, aiocqhttp, and Lark / Feishu still have explicit UI-level media gaps, so the overall adapter set remains partial acceptance rather than production-complete media acceptance.
@@ -1,162 +0,0 @@
# OneBot v11 / aiocqhttp EBA Adapter
## Status
OneBot v11 has been migrated to the EBA adapter directory:
```text
src/langbot/pkg/platform/adapters/aiocqhttp/
├── adapter.py
├── api_impl.py
├── event_converter.py
├── manifest.yaml
├── message_converter.py
├── platform_api.py
├── types.py
└── onebot.svg
```
The EBA adapter is registered as `aiocqhttp-eba`. The legacy adapter remains at `src/langbot/pkg/platform/sources/aiocqhttp.py`.
## Configuration
| Field | Required | Default | Description |
|-------|----------|---------|-------------|
| `host` | Yes | `0.0.0.0` | Host for the reverse WebSocket server that the OneBot endpoint connects to. |
| `port` | Yes | `2280` | Reverse WebSocket listen port. |
| `access-token` | No | `""` | OneBot access token, if the endpoint is configured to use one. |
## Events
The adapter declares these EBA events:
- `message.received`
- `message.deleted`
- `group.member_joined`
- `group.member_left`
- `group.member_banned`
- `friend.request_received`
- `friend.added`
- `bot.invited_to_group`
- `bot.removed_from_group`
- `bot.muted`
- `bot.unmuted`
- `platform.specific`
`platform.specific` is used for OneBot notice/request/meta events that do not yet have a common EBA event type, such as group admin changes, group file uploads, pokes, honor changes, and group join requests from non-bot users.
## Common APIs
| API | Status | Notes |
|-----|--------|-------|
| `send_message` | Supported | Supports private and group text, mentions, images, voice, files, faces, and flattened forwards. Group merged forwards are sent through OneBot forward APIs when possible. |
| `reply_message` | Supported | Uses the original OneBot event and can prepend a reply segment. |
| `edit_message` | Not supported | OneBot v11 has no standard message edit action. |
| `delete_message` | Supported | Uses `delete_msg`; permission depends on endpoint and group role. |
| `forward_message` | Supported | Emulates forward by fetching the source message with `get_msg` and sending its content to the target chat. |
| `get_message` | Supported | Uses `get_msg` and converts the response into `MessageReceivedEvent`. |
| `get_group_info` | Supported | Uses `get_group_info`. |
| `get_group_list` | Supported | Uses `get_group_list`. |
| `get_group_member_list` | Supported | Uses `get_group_member_list`. |
| `get_group_member_info` | Supported | Uses `get_group_member_info`. |
| `set_group_name` | Supported | Uses `set_group_name`; may be unsupported by mock endpoints. |
| `get_user_info` | Supported | Uses `get_stranger_info`. |
| `get_friend_list` | Supported | Uses `get_friend_list`. |
| `approve_friend_request` | Supported | Uses `set_friend_add_request`. |
| `approve_group_invite` | Supported | Uses `set_group_add_request` with `sub_type=invite`. |
| `upload_file` | Not supported | OneBot v11 has endpoint-specific file upload extensions but no portable standalone upload action. |
| `get_file_url` | Not supported | OneBot v11 file URL resolution is endpoint-specific. Use `call_platform_api("get_image")`, `get_record`, or endpoint extensions when available. |
| `mute_member` | Supported | Uses `set_group_ban`. |
| `unmute_member` | Supported | Uses `set_group_ban` with duration `0`. |
| `kick_member` | Supported | Destructive; test only with disposable members. |
| `leave_group` | Supported | Destructive; should run last in live tests. |
| `call_platform_api` | Supported | See below. |
## Platform-Specific APIs
`call_platform_api(action, params)` supports:
- `get_login_info`
- `get_status`
- `get_version_info`
- `get_group_honor_info`
- `set_group_card`
- `set_group_special_title`
- `set_group_admin`
- `set_group_whole_ban`
- `send_group_forward_msg`
- `get_forward_msg`
- `get_record`
- `get_image`
- `can_send_image`
- `can_send_record`
## Message Conversion Notes
Incoming OneBot segments are converted into common `MessageChain` components before LangBot core/plugin dispatch:
- `text` -> `Plain`
- `at` -> `At` / `AtAll`
- `image` -> `Image` or `Face` for OneBot emoji-package images
- `record` -> `Voice`
- `file` -> `File`
- `reply` -> `Quote`
- `face`, `rps`, `dice` -> `Face`
- unsupported segments -> `Unknown`
Outgoing `MessageChain` components are converted back into `aiocqhttp.Message` segments. Base64 media strings are normalized to OneBot `base64://...` format.
## Live Test Record
The direct live probe is:
```bash
PYTHONPATH=/Users/qinjunyan/code/projects/langbot/langbot-plugin-sdk/src \
uv run python tests/e2e/live_aiocqhttp_eba_probe.py --host 127.0.0.1 --port 2280
```
It starts the reverse WebSocket adapter directly, records observed EBA events to `data/temp/aiocqhttp_eba_live_probe.jsonl`, waits for a real Matcha or OneBot message, then tries reply/send/get/delete/group/user/platform API calls as far as the endpoint supports them.
Verified on May 10, 2026 with local Matcha connected to `ws://127.0.0.1:2280/ws`:
- Real inbound group message converted to `MessageReceivedEvent`.
- Real lifecycle connection converted to `PlatformSpecificEvent`.
- Real reply API succeeded and rendered a quoted bot reply in Matcha.
- Real proactive send API succeeded and rendered a bot group message in Matcha.
- Real outgoing component sweep succeeded for text, `At`, `AtAll`, `Face`, and base64 `Image`.
- Real `get_message`, `get_group_info`, `get_login_info`, `get_status`, `get_version_info`, `can_send_image`, and `can_send_record` calls succeeded against Matcha.
- Unit conversion and API-shape tests passed for `Plain`, `At`, `AtAll`, `Image`, `Voice`, `File`, `Quote`, `Face`, `rps`, `dice`, `Forward`, `Unknown`, private/group message events, delete notices, group join/leave/ban notices, bot mute notices, friend requests, group invites, friend added notices, dispatch specificity, send, reply, delete, forward, get message, group APIs, user APIs, request approval APIs, moderation APIs, leave group, unsupported file APIs, and all declared `call_platform_api` actions.
Skipped or residual live-test items:
- `edit_message`: not implemented because OneBot v11 has no standard edit action.
- `upload_file` and `get_file_url`: not implemented as common APIs because portable OneBot v11 file upload/download URL semantics are endpoint-specific.
- `kick_member` and `leave_group`: destructive; run only with explicit `--destructive` and disposable Matcha/OneBot state.
- `group.info_updated`, message reactions, and message edits are not declared because OneBot v11 does not provide standard equivalents for them.
- Matcha returned `ActionFailed` for outgoing `File` segment rendering and did not support merged-forward actions in this run. The adapter keeps the conversion/API implementations because they are valid OneBot/NapCat-style capabilities, but the Matcha live probe records them as skipped.
- Matcha returned an empty `get_group_member_list` for the test group, so `get_group_member_info`, mute/unmute, kick, and leave were covered by unit/API-shape tests only in this run.
## Standalone Runtime Plugin E2E Record
Verified on May 10, 2026 with `EBAEventProbe`, SDK standalone runtime, LangBot `--standalone-runtime`, local Matcha, and group `测试群`.
Evidence:
- Plugin JSONL: `data/temp/aiocqhttp-plugin-e2e-20260510-multiformat.jsonl`
Observed and verified:
- A real Matcha group message reached the plugin as `MessageReceived` with `bot_uuid=eba-aiocqhttp-matcha`, `adapter_name=aiocqhttp`, common `Source`/`Plain` message components, common sender, and common group identifiers.
- A protocol-level OneBot reverse WebSocket event reached the plugin as `MessageReceived` with a mixed common chain: `Source`, `Plain`, `At`, `Face`, `Image`, `Voice`, `File`, `Quote`, and trailing `Plain`. This proves the real adapter + LangBot + standalone runtime + plugin path for mixed inbound OneBot payloads, but it was not sent through Matcha UI.
- SDK API calls succeeded: `get_langbot_version`, `get_bots`, `get_bot_info`, `send_message`, plugin storage, workspace storage, `list_plugins_manifest`, `list_commands`, `list_tools`, and `list_knowledge_bases`.
- Outbound component sweep succeeded for plain text plus `At`/`Face`, `AtAll`, base64 `Image`, and quoted reply.
- Common APIs succeeded through the plugin path: `get_message`, `get_user_info`, `get_friend_list`, `get_group_info`, `get_group_list`, `get_group_member_list`, and `get_group_member_info`.
- Safe OneBot platform APIs succeeded through `call_platform_api`: `get_login_info`, `get_status`, `get_version_info`, `can_send_image`, and `can_send_record`.
Documented Matcha limits in this E2E run:
- Matcha UI did not provide a completed image/file upload/send path for inbound media. The rich inbound media evidence is `plugin-e2e-protocol`, not UI-level media upload evidence.
- Outbound `File` failed in Matcha even after the adapter emitted an official `file` segment shape.
- Outbound `Forward` failed because Matcha returned unsupported action for merged-forward.
- `get_group_honor_info` failed because Matcha returned unsupported action.
- Destructive/admin APIs such as mute, unmute, kick, leave, group rename, card/title/admin/whole-ban changes, and request approvals were not run without disposable fixtures.
@@ -1,115 +0,0 @@
# DingTalk EBA Adapter Migration Record
Status: migrated with partial plugin E2E evidence.
Adapter directory: `src/langbot/pkg/platform/adapters/dingtalk/`
## What Changed
The DingTalk adapter now has an Event-Based Agents adapter package with:
- `manifest.yaml` for adapter metadata, configuration, events, common APIs, and platform-specific APIs.
- `adapter.py` for DingTalk client startup, native callback handling, legacy compatibility, and EBA dispatch.
- `event_converter.py` for native DingTalk events and card callbacks to common EBA events.
- `message_converter.py` for DingTalk message payloads to/from common `MessageChain` components.
- `api_impl.py` for common EBA API implementations.
- `platform_api.py` for DingTalk-specific `call_platform_api` actions.
The legacy DingTalk HTTP client now returns successful JSON response bodies from proactive send methods and raises with response details on non-200 responses.
## Configuration
| Field | Required | Notes |
|-------|----------|-------|
| `client-id` | yes | DingTalk robot/client identifier. |
| `client-secret` | yes | DingTalk client secret. |
| `robot-code` | yes | Robot code used for send APIs. |
| `robot-name` | no | Used for bot mention/self filtering and display. |
| `encrypt-key` | no | DingTalk callback encryption key when configured. |
| `verification-token` | no | DingTalk callback verification token when configured. |
## Supported Events
| Event | Support | Evidence |
|-------|---------|----------|
| `message.received` | implemented | `plugin-e2e-ui` private text and emoji-as-text. |
| `feedback.received` | unit covered | DingTalk card callback actions with feedback-like values (`like`, `dislike`, `cancel`, or `1`/`2`/`3`) map to `FeedbackReceivedEvent`. Other card actions remain `platform.specific`. |
| `platform.specific` | implemented | Non-feedback card callbacks and unmapped callback/message shapes are emitted as structured platform-specific events. |
## Receive Components
| Component | Support | Evidence |
|-----------|---------|----------|
| `Source` | supported | `plugin-e2e-ui` private message. |
| `Plain` | supported | `plugin-e2e-ui` private text. DingTalk emoji currently arrives as plain text such as `[smile]`. |
| `At` | converter path | Group trigger was not completed in the latest run. |
| `AtAll` | fallback/send-side only | Not completed inbound. |
| `Image` | supported | Real DingTalk Mac private-chat image upload reached the plugin as common `Image`. |
| `Voice` | converter path | Real UI inbound voice was not completed. |
| `File` | supported | Real DingTalk Mac private-chat file upload reached the plugin as common `File`. |
| `Quote` | converter path | Real UI inbound quote was not completed. |
| `Face` | not native common mapping | DingTalk emoji was observed as `Plain`, not `Face`. |
| `Forward` | not-supported inbound | DingTalk does not expose a portable structured forward event in this adapter. |
## Send Components
| Component | Support | Evidence |
|-----------|---------|----------|
| `Plain` | supported | `plugin-e2e-outbound`. |
| `At` | supported or text fallback | `plugin-e2e-outbound`. |
| `AtAll` | fallback | `plugin-e2e-outbound`. |
| `Image` | supported | `plugin-e2e-outbound`. |
| `File` | supported | `plugin-e2e-outbound`. |
| `Quote` | fallback | `plugin-e2e-outbound`. |
| `Face` | fallback | `plugin-e2e-outbound` as text fallback. |
| `Forward` | flattened fallback | `plugin-e2e-outbound`. |
| `Voice` | fallback/endpoint-dependent | Not separately verified as a native DingTalk voice send. |
## Common APIs
| API | Support | Notes |
|-----|---------|-------|
| `send_message` | supported | Verified through `EBAEventProbe`. |
| `reply_message` | supported | Verified through quoted/fallback send path. |
| `get_message` | cache-backed | Requires the message to have been observed by this adapter process. |
| `get_group_info` | cache-backed/API-backed where available | Group path not completed in latest UI run. |
| `get_group_list` | supported where DingTalk API allows | Limited live coverage. |
| `get_group_member_info` | supported where DingTalk API allows | Limited live coverage. |
| `get_user_info` | supported | Private sender path verified. |
| `get_friend_list` | limited | DingTalk does not expose a portable friend-list equivalent. |
| `get_file_url` | supported with media/file identifiers | Real inbound file yielded a platform file URL in the converted `File` component. |
| `call_platform_api` | supported | Safe action `check_access_token` verified. |
## Platform-Specific APIs
| Action | Support | Evidence |
|--------|---------|----------|
| `check_access_token` | supported | `plugin-e2e`. |
| `refresh_access_token` | supported | Implemented; not separately reproduced in the latest plugin run. |
| `get_file_url` | supported | Real inbound file yielded a platform file URL in the converted `File` component. |
| `get_audio_base64` | supported | Needs real inbound audio/media ID. |
| `download_image_base64` | supported | Real inbound image reached the plugin as `Image`; separate image-download API replay was not completed. |
## End-to-End Evidence
Evidence files:
- Text/API/component JSONL: `data/temp/dingtalk-plugin-e2e-20260510-rerun.jsonl`
- Real UI inbound media JSONL: `data/temp/dingtalk-plugin-e2e-media-ui.jsonl`
Verified:
- DingTalk Mac private chat in the `LangBot Team` organization produced `MessageReceived` through LangBot standalone runtime and `EBAEventProbe`.
- The common chain was `Source + Plain` for normal text.
- DingTalk emoji was received as `Source + Plain`, not common `Face`.
- Real DingTalk Mac private-chat image upload was received as `Source + Image`.
- Real DingTalk Mac private-chat file upload was received as `Source + File`.
- The plugin sent outbound text, mention/fallback, image, quote/fallback, file, and forward/fallback messages visible in DingTalk.
- The plugin called safe SDK and DingTalk platform APIs.
Not completed:
- Real UI inbound voice.
- Real UI inbound quote.
- Group trigger with a real robot mention.
- Destructive or organization-mutating APIs.
-146
View File
@@ -1,146 +0,0 @@
# Discord EBA Adapter
## Status
Discord has been migrated from the legacy source adapter:
```text
src/langbot/pkg/platform/sources/discord.py
src/langbot/pkg/platform/sources/discord.yaml
```
EBA adapter directory:
```text
src/langbot/pkg/platform/adapters/discord/
├── adapter.py
├── api_impl.py
├── event_converter.py
├── manifest.yaml
├── message_converter.py
├── platform_api.py
├── types.py
└── voice.py
```
The adapter is registered as `discord-eba`.
## Configuration
| Field | Required | Default | Description |
|-------|----------|---------|-------------|
| `client_id` | Yes | `""` | Discord application client ID. |
| `token` | Yes | `""` | Discord bot token. |
The bot needs gateway permissions and intents for the target test server. Message Content intent is required for message bodies, Server Members intent is required for member APIs/events, and reaction events require the Reactions intent and channel permissions.
## Events
Discord declares these EBA events:
- `message.received`
- `message.edited`
- `message.deleted`
- `message.reaction`
- `group.member_joined`
- `group.member_left`
- `bot.invited_to_group`
- `bot.removed_from_group`
- `platform.specific`
Discord-specific events that do not map cleanly to common events should be surfaced as `platform.specific`.
## Common APIs
| API | Status | Notes |
|-----|-----------------|-------|
| `send_message` | Supported | Supports text, image, file, and mixed message chains through Discord messages and attachments. |
| `reply_message` | Supported | Uses Discord message references when replying to a received EBA message event. |
| `edit_message` | Supported | Bot can edit its own messages. File edits are implemented by clearing old attachments and sending replacement files when needed. |
| `delete_message` | Supported | Requires message management permissions for non-bot messages. |
| `forward_message` | Emulated | Discord has no native forward API; the adapter copies content and attachments. |
| `get_group_info` | Supported | Maps Discord guild metadata to EBA group info. |
| `get_group_member_list` | Supported | Requires member cache or the Server Members intent/fetch permission. |
| `get_group_member_info` | Supported | Maps Discord roles/permissions into EBA member roles. |
| `get_user_info` | Supported | Uses Discord user fetch/cache. |
| `upload_file` | Not supported | Discord uploads files as message attachments; standalone upload raises `NotSupportedError`. |
| `get_file_url` | Supported | Discord attachment URLs are already downloadable URLs, so the adapter returns the input URL. |
| `mute_member` | Supported where possible | Uses Discord timeout API and requires guild moderation permission. |
| `unmute_member` | Supported where possible | Clears timeout and requires guild moderation permission. |
| `kick_member` | Supported | Destructive; test only with a disposable account/bot. |
| `leave_group` | Supported | Bot leaves a guild; destructive and should run last. |
| `call_platform_api` | Supported | Discord-specific actions live here. |
## Platform-Specific APIs
`call_platform_api(action, params)` supports:
- `get_channel`
- `get_guild`
- `get_guild_channels`
- `get_guild_roles`
- `create_invite`
- `pin_message`
- `unpin_message`
- `add_reaction`
- `remove_reaction`
- `typing`
Voice helpers are intentionally kept Discord-specific:
- `join_voice_channel`
- `leave_voice_channel`
- `get_voice_connection_status`
- `list_active_voice_connections`
- `get_voice_channel_info`
## Live Test Record
The live probe is:
```bash
uv run python tests/e2e/live_discord_eba_probe.py --help
```
Verified on May 7, 2026 with a newly created Discord application/bot named `LangBot EBA Test 0507`, the LangBot Discord server, and the `#🐞-debugging` channel:
- SDK standalone runtime started with WebSocket control/debug ports, and the `EBAEventProbe` plugin connected through `lbp run`.
- Plugin runtime received real Discord events through LangBot: `BotInvitedToGroup`, `MessageReceived`, `MessageReactionReceived` add/remove, `MessageEdited`, and `MessageDeleted`.
- Plugin runtime API calls succeeded through the standalone runtime: `get_langbot_version`, `get_bots`, `get_bot_info`, `send_message`, plugin storage APIs, workspace storage APIs, `list_plugins_manifest`, `list_commands`, `list_tools`, and `list_knowledge_bases`.
- Direct live adapter probe observed `message.received`, `message.edited`, `message.deleted`, and `bot.removed_from_group`.
- Message APIs verified: send, reply, edit, delete, forward, text/image/file mixed message chains.
- User and guild APIs verified: `get_user_info`, `get_group_info`, `get_group_member_list`, `get_group_member_info`.
- Platform-specific APIs verified: `get_channel`, `get_guild`, `get_guild_channels`, `get_guild_roles`, `create_invite`, `typing`, `pin_message`, `unpin_message`, `add_reaction`, `remove_reaction`.
- Unsupported API behavior verified: `upload_file` raises `NotSupportedError`.
- Destructive API verified at the end: `leave_group`, which emitted `bot.removed_from_group`.
Not verified in the shared LangBot server live run: `mute_member`, `unmute_member`, and `kick_member`, because the run did not use a disposable target member. They are implemented through Discord timeout/kick APIs and should only be exercised against a disposable account or bot.
The test fixed one real test-fixture issue: `EBAEventProbe` previously assumed `get_bots()` returned UUID strings. The current standalone runtime returns bot dictionaries, so the probe now selects an enabled bot dictionary and passes its `uuid` to `get_bot_info` and `send_message`. The probe also now subscribes to `MessageDeleted`.
## Standalone Runtime Plugin E2E Record
Verified again on May 10, 2026 with SDK standalone runtime, LangBot `--standalone-runtime`, Discord web client, the LangBot server, and `#🐞-debugging`.
Evidence:
- Main plugin JSONL: `data/temp/discord-plugin-e2e-20260510-final.jsonl`
- LangBot runtime log: `data/temp/discord-langbot-e2e-20260510-rerun.log`
Observed and verified:
- A newly invited Discord bot connected to the LangBot server and received a real web-client message in `#🐞-debugging`.
- `MessageReceived` reached the plugin with `bot_uuid=eba-discord-live`, `adapter_name=discord`, common `Source`/`Plain` message components, common `User`, and common `UserGroup` for the guild.
- SDK API calls succeeded: `get_langbot_version`, `get_bots`, `get_bot_info`, `send_message`, plugin storage, workspace storage, `list_plugins_manifest`, `list_commands`, `list_tools`, and `list_knowledge_bases`.
- Outbound component sweep succeeded: plain text plus user mention, `AtAll`/`@everyone`, base64 image, quoted reply, file attachment, and flattened forward fallback.
- Common APIs succeeded: `get_user_info`, `get_group_info`, `get_group_member_list`, and `get_group_member_info`.
- Discord platform APIs succeeded through `call_platform_api`: `get_channel`, `typing`, `get_guild`, `get_guild_channels`, and `get_guild_roles`.
Documented limits in this E2E run:
- Real Discord UI inbound attachment/image/file, reply/quote, and fresh mention-chain messages were not completed in the plugin E2E evidence. Outbound image/file attachments from the bot do not prove inbound attachment conversion.
- A later May 10 UI retry could write text into the Discord message box, but the client kept the send button disabled and did not send the message, so it produced no new plugin evidence.
- `get_message`, `get_friend_list`, and `get_group_list` are not supported by this Discord adapter.
- Destructive moderation and guild-leave APIs were not repeated against the shared LangBot server.
- Native Discord voice is not represented as common `Voice`; audio-like payloads are treated as file attachments.
- `create_invite`, pin/unpin, and reaction mutation were covered by prior direct live probes but were not repeated by the final plugin run to avoid extra shared-server side effects.
-108
View File
@@ -1,108 +0,0 @@
# KOOK EBA Adapter
## Status
KOOK has been migrated to the EBA adapter directory:
```text
src/langbot/pkg/platform/adapters/kook/
├── adapter.py
├── api_impl.py
├── event_converter.py
├── manifest.yaml
├── message_converter.py
├── platform_api.py
└── types.py
```
The adapter is registered as `kook-eba`.
## Configuration
| Field | Required | Default | Description |
|-------|----------|---------|-------------|
| `token` | Yes | `""` | KOOK bot token. |
| `enable-stream-reply` | Yes | `false` | Reserved for shared platform configuration compatibility. |
## Events
| Event | Evidence | Notes |
|-------|----------|-------|
| `message.received` | `plugin-e2e-ui` | Real KOOK UI channel message reached `EBAEventProbe` as `MessageReceivedEvent`. |
| `platform.specific` | `plugin-e2e-ui` | KOOK gateway event without a common EBA mapping reached `EBAEventProbe` as `PlatformSpecificEventReceived`. |
## Common APIs
| API | Evidence | Notes |
|-----|----------|-------|
| `send_message` | `plugin-e2e-outbound` | Probe plugin sent channel messages through SDK `send_message`; KOOK returned message IDs. |
| `reply_message` | `unit` | Supports `reply_msg_id` and optional quoted replies when the source message ID is available. |
| `get_message` | `plugin-e2e-outbound` | Probe plugin fetched the cached triggering message. |
| `get_group_info` | `plugin-e2e-outbound` | Probe plugin received cached KOOK channel info. |
| `get_group_list` | `plugin-e2e-outbound` | Probe plugin received cached channel/group entities observed by the adapter. |
| `get_group_member_info` | `plugin-e2e-outbound` | Probe plugin received cached sender info as a group member. |
| `get_user_info` | `plugin-e2e-outbound` | Probe plugin received cached sender user info. |
| `get_friend_list` | `plugin-e2e-outbound` | Probe plugin received cached users. |
| `upload_file` | `unit` | Uses KOOK `asset/create` and returns URL/ID. |
| `get_file_url` | `unit` | KOOK media IDs are URL-like in the adapter path; returns the ID unchanged. |
| `delete_message` | `unit` | Calls KOOK delete endpoints. Live permission verification is still required. |
| `forward_message` | `plugin-e2e-outbound` | Probe plugin sent flattened forward content through SDK `send_message`. |
| `call_platform_api` | `plugin-e2e-outbound` | Probe plugin called safe KOOK platform-specific APIs through SDK `call_platform_api`. |
## Platform-Specific APIs
| Action | Evidence | Notes |
|--------|----------|-------|
| `get_current_user` | `plugin-e2e-outbound` | Probe plugin called `user/me`. |
| `get_user` | `plugin-e2e-outbound` | Probe plugin called `user/view` for the triggering sender. |
| `get_channel` | `plugin-e2e-outbound` | Probe plugin called `channel/view` for the triggering channel. |
| `get_guild` | `plugin-e2e-outbound` | Probe plugin called `guild/view`; gateway URLs redact token query values. |
| `get_gateway` | `plugin-e2e-outbound` | Probe plugin called `gateway/index`; returned token query values are redacted. |
| `send_direct_message` | `unit` | Calls `direct-message/create`. |
## Components
| Component | Receive Evidence | Send Evidence | Notes |
|-----------|------------------|---------------|-------|
| `Source` | `plugin-e2e-ui` | N/A | KOOK message ID and timestamp are preserved. |
| `Plain` | `plugin-e2e-ui` | `plugin-e2e-outbound` | Text and KMarkdown are represented as plain common text. |
| `At` | `plugin-e2e-ui` | `plugin-e2e-outbound` | KOOK `(met)<id>(met)` mentions map to common `At`. |
| `AtAll` | `unit` | `plugin-e2e-outbound` | KOOK `(met)all(met)` maps to common `AtAll`; real inbound UI AtAll was not tested. |
| `Image` | `unit` | `unit` | URL/image ID based path only; live rendering still needs verification. |
| `Voice` | `unit` | `unit` | URL based path only; live rendering still needs verification. |
| `File` | `unit` | `unit` | URL based path only; upload API is exposed separately. |
| `Forward` | `unit` | `unit` | Outbound forwards are flattened; inbound structured forwards are not exposed by current legacy implementation. |
| `Unknown` | `unit` | N/A | Unsupported KOOK message types become `Unknown` or `PlatformSpecificEvent`. |
## Acceptance Record
Test date: June 4, 2026.
Plugin E2E verified on June 4, 2026 with `EBAEventProbe`, SDK standalone runtime, KOOK WebSocket adapter, and a real KOOK channel UI message.
Evidence:
- JSONL: `data/temp/kook_eba_plugin_probe.jsonl`
- Plugin log: `data/logs/eba-probe-kook.log`
Observed and verified:
- A real KOOK UI channel message reached the plugin as `MessageReceived` with `bot_uuid=7ab5b065-6e4e-4def-95f0-3c265366e26f`, `adapter_name=kook`, common sender/group/chat fields, and common `MessageChain` components.
- KOOK gateway-specific event reached the plugin as `PlatformSpecificEventReceived`.
- Probe plugin called SDK `send_message`; KOOK returned message IDs for text, At, AtAll, image URL/base64 fallback path, quote fallback, file fallback, and flattened forward cases.
- Probe plugin called common API methods through the SDK path: `get_message`, `get_user_info`, `get_friend_list`, `get_group_info`, `get_group_list`, and `get_group_member_info`.
- Probe plugin called safe KOOK platform-specific APIs through SDK `call_platform_api`: `get_current_user`, `get_user`, `get_channel`, `get_gateway`, and `get_guild`.
Run:
```bash
uv run pytest tests/unit_tests/platform/test_kook_eba_adapter.py
git diff --check
```
Blocked or partial items:
- `plugin-e2e-ui` inbound coverage for image, file, voice, AtAll, quote, and forward.
- `plugin-e2e-outbound` visual verification in KOOK UI for image/file/voice rendering. KOOK returned message IDs, but UI inspection was not performed in this run.
- `reply_message` and `delete_message` live permission verification.
- Destructive or permission-sensitive APIs were not declared beyond delete; KOOK mute/kick/leave remain explicit `NotSupportedError` paths until a safe fixture is available.
-135
View File
@@ -1,135 +0,0 @@
# Lark / Feishu EBA Adapter Migration Record
Status: migrated with unit coverage and partial live plugin E2E. WebSocket text reached the standalone runtime once in the LangBot organization test app, but the latest real UI image/file inbound attempts did not reach the local adapter log, so media receive is not release-complete yet.
Adapter directory: `src/langbot/pkg/platform/adapters/lark/`
## What Changed
The Lark/Feishu adapter now has an Event-Based Agents adapter package with:
- `manifest.yaml` for adapter metadata, configuration, events, common APIs, platform-specific APIs, app type, and communication mode.
- `adapter.py` for self-built/store app token handling, WebSocket long connection startup, Webhook callback handling, card feedback, streaming-card replies, and EBA dispatch.
- `event_converter.py` for native Feishu events to common EBA events.
- `message_converter.py` for Feishu text/post/image/file/audio payloads to/from common `MessageChain` components.
- `api_impl.py` for common EBA API implementations.
- `platform_api.py` for Feishu-specific `call_platform_api` actions.
The legacy `lark` adapter remains available while the EBA adapter is registered separately as `lark-eba`.
## Configuration
| Field | Required | Notes |
|-------|----------|-------|
| `app_id` | yes | Feishu/Lark application App ID. |
| `app_secret` | yes | Feishu/Lark application App Secret. |
| `bot_name` | yes | Must match the bot name so group mentions can be recognized. |
| `enable-webhook` | yes | `false` uses WebSocket long connection; `true` uses Request URL/Webhook callbacks. |
| `webhook_url` | no | Generated callback URL for Webhook mode. |
| `encrypt-key` | no | Webhook decrypt key when event encryption is enabled. |
| `enable-stream-reply` | yes | Enables streaming replies through an updating Feishu card. |
| `app_type` | no | `self` for self-built apps; `isv` for store apps. |
| `bot_added_welcome` | no | Optional group welcome message sent after bot-added events. |
## Application And Communication Modes
| Mode | Support | Implementation |
|------|---------|----------------|
| Self-built application | implemented | Uses standard app credentials and tenant token behavior from the Feishu SDK client. |
| Store application | implemented | Builds an ISV client, requests app tickets, and resolves app/tenant access tokens with per-tenant caching. |
| WebSocket long connection | implemented | Registers `im.message.receive_v1` and card-action callbacks through `lark_oapi.ws.Client`. |
| Webhook Request URL | implemented | Handles URL verification, encrypted payloads, message events, app-ticket events, bot-added events, and card-action feedback. |
## Supported Events
| Event | Support | Evidence |
|-------|---------|----------|
| `message.received` | implemented | Unit coverage for private and group native events to common EBA events. |
| `bot.invited_to_group` | implemented | Webhook bot-added event maps to common bot invite event and optional welcome send. |
| `platform.specific` | implemented | Unknown callback events are preserved as `platform.specific`. |
| `FeedbackEvent` | compatibility event | Card button feedback is still dispatched through the existing SDK `FeedbackEvent` type. |
## Receive Components
| Component | Support | Evidence |
|-----------|---------|----------|
| `Source` | supported | Unit coverage; live private text evidence. |
| `Plain` | supported | Text and post payloads convert to common text; live private text evidence. |
| `At` | supported | Feishu mentions map to common `At` with user ID and display name. |
| `AtAll` | supported | `user_id=all` maps to common `AtAll`. |
| `Image` | supported | Image payloads download through message resource API and map to common `Image`; real UI image send attempted, but not observed in local plugin evidence yet. |
| `Voice` | supported | Audio payloads download through message resource API and map to common `Voice`. |
| `File` | supported | File payloads download through message resource API and map to common `File`; real UI file send attempted, but not observed in local plugin evidence yet. |
| `Quote` | supported | Parent/thread reply lookup maps quoted content into common `Quote`. |
| `Face` | not native common mapping | Feishu emoji/stickers are not exposed as a portable common `Face` component here. |
| `Forward` | not-supported inbound | Feishu does not expose a portable structured forward event in this adapter. |
## Send Components
| Component | Support | Evidence |
|-----------|---------|----------|
| `Plain` | supported | Unit coverage; sends Feishu `text`. |
| `At` | supported | Unit coverage; sends Feishu `post` at element. |
| `AtAll` | supported | Unit coverage; sends Feishu `post` at-all element. |
| `Image` | supported | Uploads image resource and sends Feishu `image`. |
| `Voice` | supported | Uploads OPUS/audio resource and sends Feishu `audio`. |
| `File` | supported | Uploads file resource and sends Feishu `file`. |
| `Quote` | supported/fallback | Sends quote marker plus origin content. |
| `Face` | not-supported | No portable send mapping. |
| `Forward` | flattened fallback | Flattens forward nodes into text/media messages. |
## Common APIs
| API | Support | Notes |
|-----|---------|-------|
| `send_message` | supported | Supports private/open_id and group/chat_id targets; live plugin outbound component sweep produced visible Feishu messages. |
| `reply_message` | supported | Replies to the source Feishu message; fixed to recover the native Feishu message ID from legacy-wrapped source events. |
| `get_message` | cache-backed/API-backed | Returns cached inbound event where possible and converts uncached Feishu message API items into common `MessageReceivedEvent`. |
| `get_group_info` | supported | Uses cached group or Feishu chat metadata. |
| `get_group_member_info` | limited | Uses cached user data when available. |
| `get_user_info` | limited | Uses cached user data when available. |
| `get_file_url` | limited | Returns `file://` paths from downloaded inbound resources; remote Feishu resource download uses platform-specific API params. |
| `call_platform_api` | supported | See below. |
## Platform-Specific APIs
| Action | Support | Evidence |
|--------|---------|----------|
| `check_tenant_access_token` | supported | Unit coverage. |
| `refresh_app_access_token` | supported | Store-app token path implemented. |
| `refresh_tenant_access_token` | supported | Store-app tenant token path implemented. |
| `get_chat` | supported | Feishu chat metadata API wrapper. |
| `get_message` | supported | Feishu message API wrapper with JSON-safe return values for plugin calls. |
| `get_message_resource` | supported | Feishu message resource download wrapper. |
## End-to-End Evidence
Current code-level evidence:
- `tests/unit_tests/platform/test_lark_eba_adapter.py`
- `PYTHONPATH=../langbot-plugin-sdk/src uv run pytest tests/unit_tests/platform/test_lark_eba_adapter.py -q`
Live evidence collected on May 11, 2026:
- Standalone runtime: `uv run lbp rt --ws-control-port 5400 --ws-debug-port 5401 --skip-deps-check`
- LangBot: `uv run main.py --standalone-runtime --debug`
- Plugin: `LangBot__EBAEventProbe`
- Feishu org/app: LangBot organization, `LangBotDev` private chat.
- Observed plugin JSONL: one private `MessageReceived` event with `Source + Plain`; plugin API probe then exercised bot discovery, bot info, `send_message`, outbound component sweep, storage/list APIs, and safe platform API calls.
- Real UI sends attempted after the fixes: private text, local file, and image/video image upload. These appeared in the Feishu client but did not append new `EBAEventProbe` records in the local JSONL during this run.
- Fixes from live testing: reply path now extracts the native Feishu `message_id` from legacy-wrapped source events; WebSocket callbacks are scheduled onto the adapter event loop instead of assuming the SDK callback has a running asyncio loop; platform API results are converted to JSON-safe values.
Live E2E items still required before marking release-complete:
- WebSocket self-built app in LangBot organization: repeat private text after callback-loop fix, plus private image/file/audio and group mention message received by `EBAEventProbe`.
- Webhook self-built app in LangBot organization: URL verification plus text/image/file message received by `EBAEventProbe`.
- Store app token path: at least token acquisition/tenant-token safe API through `call_platform_api`; full message E2E if a LangBot organization store-app fixture is available.
- Outbound component sweep: text, mention, at-all, image, file, voice where Feishu accepts the fixture, quote/fallback, and forward/fallback.
- Safe platform API sweep: token check, chat metadata, message lookup, and message resource download using real inbound IDs.
## Known Limits
- Store-app live E2E requires a real ISV app ticket/tenant installation fixture.
- Current LangBot organization WebSocket run connected successfully but did not deliver the latest UI-sent image/file attempts to local plugin evidence; this blocks release-complete media acceptance.
- Feishu native emoji/sticker semantics are not represented as common `Face`.
- Destructive org or chat mutations are not declared in this adapter.
@@ -1,101 +0,0 @@
# OfficialAccount EBA Adapter
Adapter directory: `src/langbot/pkg/platform/adapters/officialaccount/`
Manifest name: `officialaccount-eba`
Status: partial migration. Unit/API-shape coverage is present, and private text `plugin-e2e-ui` plus safe API evidence has been verified against the `dev.rockchin.top` Official Account fixture. Proactive outbound `send_message` remains not supported by this adapter because WeChat Official Account replies must be tied to inbound webhook windows.
## Config
| Field | Required | Notes |
| --- | --- | --- |
| `webhook_url` | no | Generated by LangBot and copied into the Official Account callback settings. |
| `token` | yes | WeChat callback token. |
| `EncodingAESKey` | yes | WeChat message encryption key. |
| `AppID` | yes | Official Account app ID. |
| `AppSecret` | yes | Official Account app secret. |
| `Mode` | yes | `drop` waits for an in-callback reply; `passive` returns the loading text first and queues the answer for the user's next message. |
| `LoadingMessage` | no | Only used by `passive` mode. |
| `api_base_url` | no | Optional API base URL for proxy deployments. |
## Events
| Event | Evidence | Notes |
| --- | --- | --- |
| `message.received` | plugin-e2e-ui, unit | Text UI message verified through WeChat Official Account on `dev.rockchin.top`; image and voice webhook payloads are covered by unit tests. |
| `platform.specific` | unit | Subscribe/unsubscribe/menu/etc. native events are emitted as structured `PlatformSpecificEvent`. |
## Common APIs
| API | Evidence | Notes |
| --- | --- | --- |
| `reply_message` | unit | Queues/passively returns text through the inbound webhook source event. |
| `get_message` | plugin-e2e-ui, unit | Cached inbound message retrieved by `EBAEventProbe` platform API sweep. |
| `get_user_info` | plugin-e2e-ui, unit | Cached inbound sender retrieved by `EBAEventProbe` platform API sweep. |
| `get_friend_list` | plugin-e2e-ui, unit | Cached inbound sender list retrieved by `EBAEventProbe` platform API sweep. |
| `call_platform_api` | plugin-e2e-ui, unit | Safe diagnostic actions verified through `get_mode` and `get_cached_response_status`. |
| `send_message` | not-supported | Official Account customer-service proactive messaging is not implemented by the existing SDK adapter; only webhook reply is supported here. |
## Platform APIs
| Action | Evidence | Notes |
| --- | --- | --- |
| `get_mode` | plugin-e2e-ui, unit | Returned `{"mode": "drop", "longer_response": false}` in live probe. |
| `get_cached_response_status` | plugin-e2e-ui, unit | Returned `{"pending": false}` in live probe. |
## Components
| Receive Component | Evidence | Notes |
| --- | --- | --- |
| `Source` | plugin-e2e-ui, unit | Uses `MsgId` and `CreateTime`; live UI text message included `Source`. |
| `Plain` | plugin-e2e-ui, unit | Live UI text message mapped to `Plain`. |
| `Image` | unit | `PicUrl` and `MediaId` map to common `Image`. |
| `Voice` | unit | `MediaId` maps to common `Voice`. |
| `Unknown` | unit | Unsupported message/event types do not crash. |
| `At`, `AtAll`, `File`, `Quote`, `Face`, `Forward`, mixed chain | not-supported | WeChat Official Account inbound webhook payloads used by the current SDK do not expose these as common structured components. |
| Send Component | Evidence | Notes |
| --- | --- | --- |
| `Plain` | unit | Sent as webhook reply text. |
| `Image`, `Voice`, `File`, `Quote`, `At`, `AtAll`, `Face`, `Forward`, mixed chain | not-supported | Existing SDK reply path is text XML only; non-text components degrade to readable placeholders in tests and are not declared as supported outbound components. |
## Verification Record
Test date: 2026-05-28
Endpoint/simulator: `dev.rockchin.top` with WeChat desktop client and a real subscribed Official Account conversation. The running EBA test stack used SDK standalone runtime ports `5400/5401`, LangBot from `/home/wgc/LangBotxg/LangBotEbaTest`, and `EBAEventProbe`.
Verified UI message: `EBA officialaccount single probe 2026-05-28 16:53`
Observed event/API evidence:
- `MessageReceived`: `bot_uuid=d7c46880-a9f8-431a-9172-5d3e0d663dbc`, `adapter_name=officialaccount-eba`, `chat_type=private`, `chat_id=ovH9L7OW6hNpWZWvp_NMmypVh26w`, `message_chain=[Source, Plain]`.
- Common safe APIs through probe platform sweep: `get_message`, `get_user_info`, `get_friend_list`.
- Platform APIs through `call_platform_api`: `get_mode`, `get_cached_response_status`.
- `send_message` and outbound component sweep returned explicit `NotSupportedError: send_message:official_account_requires_inbound_webhook_reply`, as expected for this adapter.
Standalone runtime command:
```bash
cd langbot-plugin-sdk
uv run python -m langbot_plugin.cli.__init__ rt --debug-only --ws-control-port 5400 --ws-debug-port 5401 --skip-deps-check
```
Probe plugin: `data/plugins/LangBot__EBAEventProbe` when live credentials are available.
Adapter live probe:
```bash
uv run python -m py_compile tests/e2e/live_officialaccount_eba_probe.py
OFFICIALACCOUNT_TOKEN=... OFFICIALACCOUNT_ENCODING_AES_KEY=... OFFICIALACCOUNT_APP_SECRET=... OFFICIALACCOUNT_APP_ID=... uv run python tests/e2e/live_officialaccount_eba_probe.py
```
Evidence JSONL path: `/home/wgc/LangBotxg/LangBotEbaTest/data/temp/officialaccount_eba_plugin_probe.jsonl` for plugin E2E, or `data/temp/officialaccount_eba_probe.jsonl` for direct adapter live probe.
Destructive operations: none.
Blocked items:
- `plugin-e2e-outbound`: proactive `send_message` is not supported for this adapter; Official Account responses must be produced through the inbound webhook reply window.
- Inbound image and voice live UI evidence remains pending; webhook conversion is covered by unit tests.
@@ -1,119 +0,0 @@
# QQOfficial EBA Adapter
Adapter directory: `src/langbot/pkg/platform/adapters/qqofficial/`
Manifest name: `qqofficial-eba`
Status: partial migration. The EBA adapter structure, manifest, converters, cache-backed safe APIs, platform API map, unit tests, and direct live probe scaffold are in place. A real QQ Official WebSocket bot on `dev.rockchin.top` received an inbound user message and drove LangBot into the normal pipeline path; the response path was blocked by the test environment model service returning `model_not_found` for `deepseek-v3`.
## Config
| Field | Required | Notes |
| --- | --- | --- |
| `appid` | yes | QQ Official app ID. |
| `secret` | yes | QQ Official app secret. |
| `token` | yes | QQ Official callback token. |
| `enable-webhook` | yes | Uses LangBot unified webhook when true; otherwise uses the QQ WebSocket gateway. |
| `enable-stream-reply` | yes | Enables C2C streaming replies when supported by the QQ Official endpoint. |
| `webhook_url` | no | Generated by LangBot and copied into the QQ Official callback settings in webhook mode. |
## Events
| Event | Evidence | Notes |
| --- | --- | --- |
| `message.received` | adapter-live, unit | `C2C_MESSAGE_CREATE`, `DIRECT_MESSAGE_CREATE`, `GROUP_AT_MESSAGE_CREATE`, and `AT_MESSAGE_CREATE` map to common `MessageReceivedEvent`. A real WebSocket-mode QQ Official bot reached the LangBot pipeline on `dev.rockchin.top`; plugin JSONL evidence remains pending. |
| `message.reaction` | unit | `MESSAGE_REACTION_ADD` and `MESSAGE_REACTION_REMOVE` map to common `MessageReactionEvent`. Live gateway evidence is pending. |
| `group.member_joined` | unit | `GUILD_MEMBER_ADD` and `GROUP_MEMBER_ADD` map to common `MemberJoinedEvent` when the gateway payload carries a group/guild/channel ID and member openid. |
| `group.member_left` | unit | `GUILD_MEMBER_REMOVE` and `GROUP_MEMBER_REMOVE` map to common `MemberLeftEvent`. Live gateway evidence is pending. |
| `bot.invited_to_group` | unit | `GUILD_CREATE` and `GROUP_ADD_ROBOT` map to common `BotInvitedToGroupEvent`. |
| `bot.removed_from_group` | unit | `GUILD_DELETE` and `GROUP_DEL_ROBOT` map to common `BotRemovedFromGroupEvent`. |
| `platform.specific` | unit | Unmapped gateway events are emitted as structured `PlatformSpecificEvent`; live evidence is pending. |
## Common APIs
| API | Evidence | Notes |
| --- | --- | --- |
| `send_message` | unit, blocked | Sends private C2C, group, and text-only channel messages through the existing QQ Official client. Live outbound UI verification is pending because the test pipeline failed before producing a bot response. |
| `reply_message` | unit, blocked | Replies using the source `QQOfficialEvent` message ID when available. Live reply was blocked by the test environment model service returning `model_not_found`. |
| `get_message` | unit | Returns cached inbound `MessageReceivedEvent`. |
| `get_user_info` | unit | Returns cached inbound sender. |
| `get_friend_list` | unit | Returns cached private senders. |
| `get_group_info` | unit | Returns cached group/channel metadata from inbound events. |
| `get_group_member_info` | unit | Returns cached group sender as a common member. |
| `get_group_member_list` | unit | Returns cached group members observed by the adapter. |
| `call_platform_api` | unit, blocked | Safe diagnostic actions are implemented; live calls are pending credentials. |
## Platform APIs
| Action | Evidence | Notes |
| --- | --- | --- |
| `check_access_token` | unit, blocked | Calls the existing client token check. |
| `refresh_access_token` | unit, blocked | Forces token refresh. |
| `get_gateway_url` | unit, blocked | Fetches the WebSocket gateway URL. |
| `get_mode` | unit | Returns webhook and stream-reply mode. |
## Components
| Receive Component | Evidence | Notes |
| --- | --- | --- |
| `Source` | unit | Uses QQ message/event IDs and timestamp. |
| `Plain` | unit | Preserves text content. |
| `At` | unit | Group and channel mention events insert an adapter bot mention marker. |
| `Image` | unit | QQ image attachment URL is converted to common `Image`; falls back to URL if download fails. |
| `Unknown` | unit | Unsupported/empty native payloads become `Unknown`. |
| `Voice`, `File`, `Quote`, `Face`, `Forward`, mixed chain | blocked | Current native parser only exposes text and image attachments; live endpoint behavior still needs verification. |
| Send Component | Evidence | Notes |
| --- | --- | --- |
| `Plain` | unit, blocked | Sends through private, group, or channel text APIs. |
| `At`, `AtAll` | unit, blocked | Converted to readable mention text. |
| `Image` | unit, blocked | Sends through the QQ Official rich media upload/send path for C2C and group targets. |
| `Voice` | unit, blocked | Sends through the QQ Official rich media upload/send path for C2C and group targets. |
| `File` | unit, blocked | Sends through the QQ Official rich media upload/send path for C2C and group targets. |
| `Quote`, `Forward`, mixed chain | unit, blocked | Flattened to ordered send payloads where possible. |
| `Face` | not-supported | No common QQ Official face mapping is implemented. |
## Verification Record
Test date: 2026-06-02
Endpoint/simulator: `dev.rockchin.top` with a real QQ Official WebSocket bot (`qqofficial-eba`, bot UUID `80a5560b-52b1-40e7-b7d6-4a2341eb4780`) and LangBot running from `/home/wgc/LangBotxg/LangBotEbaTest`.
Observed evidence:
- The QQ Official WebSocket bot was enabled with `enable-webhook=false`.
- A real user message reached LangBot and entered the standard pipeline path.
- The response path stopped at the model layer with `model_not_found` for `deepseek-v3`; this is a model/provider configuration issue, not an adapter conversion failure.
- `qq-webhook.langbot.dev` was temporarily routed through Caddy to `127.0.0.1:5301` for webhook checks, but the observed EBA bot used WebSocket mode.
Standalone runtime command:
```bash
cd langbot-plugin-sdk
uv run python -m langbot_plugin.cli.__init__ rt --debug-only --ws-control-port 5400 --ws-debug-port 5401 --skip-deps-check
```
Probe plugin: `data/plugins/LangBot__EBAEventProbe` when live credentials are available.
Adapter live probe:
```bash
uv run python -m py_compile tests/e2e/live_qqofficial_eba_probe.py
QQOFFICIAL_APPID=... QQOFFICIAL_SECRET=... QQOFFICIAL_TOKEN=... uv run python tests/e2e/live_qqofficial_eba_probe.py
```
Webhook-mode probe:
```bash
QQOFFICIAL_APPID=... QQOFFICIAL_SECRET=... QQOFFICIAL_TOKEN=... uv run python tests/e2e/live_qqofficial_eba_probe.py --webhook --host 0.0.0.0 --port 5312
```
Evidence JSONL path: `data/temp/qqofficial_eba_probe.jsonl` for direct adapter live probe; plugin E2E evidence should use `data/temp/qqofficial_eba_plugin_probe.jsonl`.
Destructive operations: none implemented.
Blocked items:
- `plugin-e2e-ui`: standalone probe plugin JSONL evidence is still pending; the observed live run reached LangBot core/pipeline but was not recorded by the EBA probe plugin.
- `plugin-e2e-outbound`: waiting for visible QQ client verification of plugin `send_message`/`reply_message` output after a working model/provider is configured.
- Inbound non-text media and platform lifecycle events require endpoint evidence before they can be marked complete.
-84
View File
@@ -1,84 +0,0 @@
# Slack EBA Adapter
## Structure
Slack is migrated into `src/langbot/pkg/platform/adapters/slack/` with the standard EBA adapter layout:
- `adapter.py` owns lifecycle, listener dispatch, unified webhook handling, outbound send/reply, and event caches.
- `event_converter.py` maps Slack `im` and `app_mention` channel events to `message.received`.
- `message_converter.py` maps common `MessageChain` components to Slack text fallback and maps inbound Slack text/image payloads back to EBA components.
- `api_impl.py` provides cache-backed common read APIs.
- `platform_api.py` declares safe Slack-specific API actions.
- `manifest.yaml` declares `slack-eba`.
The legacy `src/langbot/pkg/platform/sources/slack.py` adapter is kept unchanged.
## Configuration
| Field | Required | Notes |
|-------|----------|-------|
| `webhook_url` | No | Generated by LangBot. Paste it into Slack Event Subscriptions. |
| `bot_token` | Yes | Slack bot token, usually `xoxb-...`. |
| `signing_secret` | Yes | Slack app signing secret. |
## Events
| Event | Notes |
|-------|-------|
| `message.received` | Emitted for private `im` messages and channel `app_mention` events. Channel messages are mapped to group chats. |
| `platform.specific` | Reserved for Slack event types that are not converted into common message events. |
## Common APIs
Required:
- `send_message`
- `reply_message`
Optional:
- `get_message`
- `get_user_info`
- `get_friend_list`
- `get_group_info`
- `get_group_list`
- `get_group_member_list`
- `get_group_member_info`
- `call_platform_api`
Cache-backed APIs are only available after the relevant inbound event has been observed.
## Platform APIs
| Action | Notes |
|--------|-------|
| `get_mode` | Returns webhook mode and configured bot account id. |
| `auth_test` | Calls Slack `auth.test` with the configured bot token. |
## Known Limits
- Slack file/image outbound is currently represented as text fallback because the existing Slack SDK wrapper only exposes `chat_postMessage`.
- Inbound channel coverage follows the legacy adapter behavior: only `app_mention` events are treated as group messages.
- Real live testing requires a public callback URL configured in Slack Event Subscriptions.
## Verification
Local mocked unit coverage validates manifest parity, event conversion, legacy listener compatibility, cache-backed APIs, send/reply routing, and declared platform APIs.
Plugin E2E evidence was captured on June 2, 2026 against `dev.rockchin.top` with Slack private DM input and `EBAEventProbe` through the standalone runtime.
Evidence file: `/home/wgc/LangBotxg/LangBotEbaTest/data/temp/slack_eba_plugin_probe.jsonl`.
Observed:
- Real Slack private text produced `MessageReceived` with `adapter_name=slack-eba`, `Source + Plain`, private chat type, and filled `bot_uuid`.
- Safe common APIs passed: `get_message`, `get_user_info`, `get_friend_list`.
- Outbound component fallback sweep passed through `send_message`: plain/at/face, image, quote, file, and forward.
- Declared Slack platform APIs passed: `get_mode`, `auth_test`.
Still pending:
- Channel `app_mention` plugin E2E.
- Real inbound Slack file/image UI evidence.
Live probe scaffold: `tests/e2e/live_slack_eba_probe.py`.
@@ -1,139 +0,0 @@
# Telegram EBA Adapter
## Status
Telegram has been migrated to the EBA adapter directory:
```text
src/langbot/pkg/platform/adapters/telegram/
├── adapter.py
├── api_impl.py
├── event_converter.py
├── manifest.yaml
├── message_converter.py
├── platform_api.py
└── types.py
```
The adapter is registered as `telegram-eba`.
## Configuration
| Field | Required | Default | Description |
|-------|----------|---------|-------------|
| `token` | Yes | `""` | Telegram Bot API token from BotFather. |
| `markdown_card` | No | `true` | Whether to render Markdown card style replies. |
| `enable-stream-reply` | Yes | `false` | Whether to use Telegram streaming reply mode. |
## Events
Telegram declares these EBA events:
- `message.received`
- `message.edited`
- `message.reaction`
- `group.member_joined`
- `group.member_left`
- `group.member_banned`
- `bot.invited_to_group`
- `bot.removed_from_group`
- `bot.muted`
- `bot.unmuted`
- `platform.specific`
`platform.specific` is currently used for Telegram-only callback and chat-member update payloads that do not yet have a more specific common event type.
## Common APIs
| API | Status | Notes |
|-----|--------|-------|
| `send_message` | Supported | Supports text, image, file, and mixed message chains. |
| `reply_message` | Supported | Supports quoted replies through the original message event. |
| `edit_message` | Supported | Uses Telegram message editing APIs. |
| `delete_message` | Supported | Deletes messages where bot permissions allow it. |
| `forward_message` | Supported | Forwards a message between Telegram chats. |
| `get_group_info` | Supported | Uses Telegram chat metadata. |
| `get_group_member_list` | Supported | Telegram only exposes administrators through the Bot API; this returns the available member set. |
| `get_group_member_info` | Supported | Maps Telegram member status to EBA member roles. |
| `get_user_info` | Supported | Uses Telegram `get_chat` for user chat metadata. |
| `upload_file` | Not supported | Telegram has no standalone upload endpoint; files are uploaded as part of messages. The adapter raises `NotSupportedError`. |
| `get_file_url` | Supported | Returns the Bot API file URL. Test output redacts the bot token. |
| `mute_member` | Supported | Requires a supergroup and bot moderation permission. |
| `unmute_member` | Supported | Uses current `telegram.ChatPermissions` fields. |
| `kick_member` | Supported | Destructive; should only be run against disposable users/bots in tests. |
| `leave_group` | Supported | Destructive; should run at the end of a live test. |
| `call_platform_api` | Supported | See below. |
## Platform-Specific APIs
`call_platform_api(action, params)` supports:
- `pin_message`
- `unpin_message`
- `unpin_all_messages`
- `get_chat_administrators`
- `set_chat_title`
- `set_chat_description`
- `get_chat_member_count`
- `send_chat_action`
- `create_chat_invite_link`
- `answer_callback_query`
## Live Test Record
The live probe is:
```bash
uv run python tests/e2e/live_telegram_eba_probe.py --help
```
It supports private chat tests, group/supergroup tests, moderation tests, destructive tests, and a callback-only mode.
Verified on May 7, 2026:
- Private chat message APIs: send, reply, edit, delete, forward.
- Private chat media APIs: image/file sending and `get_file_url`.
- User API: `get_user_info`.
- Supergroup APIs: group info, member list, member info, administrators, member count, invite link.
- Supergroup mutation APIs: pin, unpin, unpin all, set title, restore title, set description, restore description.
- Moderation APIs: mute and unmute against a non-owner target bot.
- Destructive APIs: kick a disposable target bot, then make the test bot leave the test group.
- Event conversion observed for `message.received`, `group.member_banned`, `group.member_left`, `bot.removed_from_group`, and Telegram-specific chat-member updates.
The test fixed one real compatibility issue: `unmute_member` previously used Telegram's removed `can_send_media_messages` permission field. It now uses the split media permission fields required by current `python-telegram-bot`.
## Standalone Runtime Plugin E2E Record
Verified on May 10, 2026 with `EBAEventProbe`, SDK standalone runtime, Telegram Lite, `@rockchinq_bot`, and `Rock'sBotGroup`.
Evidence:
- Private chat JSONL: `data/temp/telegram-plugin-e2e-rerun.jsonl`
- Group chat JSONL: `data/temp/telegram-plugin-e2e-group.jsonl`
- Private media JSONL: `data/temp/telegram-plugin-e2e-media-ui.jsonl`
Observed and verified:
- `MessageReceived` reached the plugin with `bot_uuid=eba-telegram-live`, `adapter_name=telegram`, common sender/chat fields, and common `MessageChain` content.
- `BotInvitedToGroup` reached the plugin after adding the bot to `Rock'sBotGroup`.
- SDK API calls succeeded: `get_langbot_version`, `get_bots`, `get_bot_info`, `send_message`, plugin storage, workspace storage, `list_plugins_manifest`, `list_commands`, `list_tools`, and `list_knowledge_bases`.
- Outbound component sweep succeeded in private and group chats: plain text, mention text/equivalent, base64 image, quoted reply, file/document, and flattened forward fallback. Group mode also covered `AtAll` fallback behavior.
- Real Telegram Lite private-chat inbound media was verified through the plugin path: a sent document arrived as common `File`, and a sent photo arrived as common `Image`.
- Telegram platform API sweep succeeded for safe group actions: `get_chat_administrators`, `get_chat_member_count`, and `send_chat_action`.
- Common group/user APIs succeeded in group mode: `get_user_info`, `get_group_info`, `get_group_member_list`, and `get_group_member_info`.
Documented limits in this E2E run:
- Real Telegram UI inbound voice, sticker/emoji-as-common-component, and reply/quote messages were not completed in the plugin E2E evidence.
- `get_message`, `get_friend_list`, and `get_group_list` are not supported by this Telegram adapter.
- Mutating/destructive Telegram-specific actions such as pin/unpin, title/description changes, invite-link creation, moderation, kick, and leave were not repeated in the plugin run. They remain opt-in live-probe cases.
- Telegram does not expose a portable common `Face` component for native sticker/emoji semantics in the current adapter.
## Notes for Future Adapters
Telegram is the reference implementation for:
- Keeping platform-specific actions behind `call_platform_api`.
- Treating unsupported common APIs as explicit `NotSupportedError`.
- Marking destructive live test operations behind CLI flags.
- Redacting access tokens from live probe output.
-130
View File
@@ -1,130 +0,0 @@
# WeCom EBA Adapter
## Status
WeCom application messages now have an EBA adapter directory:
```text
src/langbot/pkg/platform/adapters/wecom/
├── adapter.py
├── api_impl.py
├── event_converter.py
├── manifest.yaml
├── message_converter.py
├── platform_api.py
└── types.py
```
The adapter is registered as `wecom-eba`.
This record covers the regular WeCom application-message adapter. WeCom AI Bot (`wecombot-eba`) uses a different protocol flow and is documented separately in `wecombot.md`. WeCom Customer Service (`wecomcs`) remains a separate follow-up migration.
## Configuration
| Field | Required | Default | Description |
|-------|----------|---------|-------------|
| `webhook_url` | No | `""` | Unified webhook URL copied into the WeCom application callback settings. |
| `corpid` | Yes | `""` | WeCom corporate ID. |
| `secret` | Yes | `""` | WeCom application secret. |
| `token` | Yes | `""` | WeCom callback token. |
| `EncodingAESKey` | Yes | `""` | WeCom callback encryption key. |
| `contacts_secret` | No | `""` | Contacts secret for contact-list based helper APIs. |
| `api_base_url` | No | `https://qyapi.weixin.qq.com/cgi-bin` | WeCom API base URL, overrideable for proxy/private-network deployments. |
## Events
WeCom declares these EBA events:
- `message.received`
- `platform.specific`
`message.received` currently covers text and image application callbacks. Other WeCom callback types are surfaced as `platform.specific` so plugins can inspect the raw structured payload without crashing the common message path.
## Common APIs
| API | Status | Notes |
|-----|--------|-------|
| `send_message` | Supported | Private/person target only. `target_id` must be `user_id|agent_id`. Supports text, image, voice, file, flattened forward, and quote fallback. |
| `reply_message` | Supported | Replies to the original WeCom sender and application agent from `source_platform_object`. |
| `get_message` | Supported from cache | Returns cached inbound `MessageReceivedEvent` by message ID. |
| `get_user_info` | Supported | Uses cached event users first, then WeCom `user/get`. |
| `get_friend_list` | Partial | Returns users seen by this adapter instance. Full contacts listing is not declared as common coverage. |
| `call_platform_api` | Supported | See below. |
| `edit_message` | Not supported | WeCom application messages do not expose a general edit endpoint for sent messages. |
| `delete_message` | Not supported | WeCom application messages do not expose a general delete endpoint for sent messages. |
| `get_group_info` / member APIs | Not supported | Regular WeCom application callbacks handled here are private user messages, not group-chat bot messages. |
| `upload_file` / `get_file_url` | Not supported as common APIs | WeCom media upload is used internally while sending image/voice/file components; no portable standalone common file URL is exposed. |
## Platform-Specific APIs
`call_platform_api(action, params)` supports:
- `check_access_token`
- `refresh_access_token`
- `get_user_info`
- `send_to_all`
`send_to_all` requires a configured `contacts_secret` with suitable contact visibility and should be treated as a broad-send operation in live testing.
## Unit Verification
Covered by:
```bash
uv run pytest tests/unit_tests/platform/test_wecom_eba_adapter.py
```
The unit tests cover:
- Manifest events/APIs/platform actions match adapter declarations.
- Outbound component conversion for text, image, voice, file, quote fallback, and byte-safe text splitting.
- Text callback conversion to `MessageReceivedEvent`.
- Legacy `FriendMessage` compatibility.
- EBA listener dispatch and inbound message/user cache.
- `send_message`, `reply_message`, and safe platform API dispatch against a mocked WeCom client.
## Standalone Runtime Plugin E2E Record
Verified on May 27, 2026 with `EBAEventProbe`, SDK standalone runtime, LangBot core, and a real WeCom desktop client against the server test environment.
```bash
cd langbot-plugin-sdk
uv run python -m langbot_plugin.cli.__init__ rt --debug-only --ws-control-port 5400 --ws-debug-port 5401 --skip-deps-check
cd LangBot
uv run main.py --standalone-runtime
cd data/plugins/LangBot__EBAEventProbe
EBA_PROBE_API=1 EBA_PROBE_COMPONENT_SWEEP=1 EBA_PROBE_PLATFORM_API=1 \
uv --project /absolute/path/to/langbot-plugin-sdk run python -m langbot_plugin.cli.__init__ run
```
Evidence:
- JSONL: `data/temp/wecom_eba_plugin_probe.jsonl`
- Bot: `wecom-eba`
- Client: real WeCom desktop client
- Environment: `dev.rockchin.top` test server
Observed and verified:
- A real private WeCom user message reached the plugin as `MessageReceived` with `adapter_name=wecom-eba`, common sender/chat fields, and `Source + Plain`.
- SDK API calls succeeded through the standalone runtime, including `get_langbot_version`, `get_bots`, `get_bot_info`, `send_message`, plugin/workspace storage, and manifest/list APIs.
- Safe adapter API checks succeeded through the plugin path for cached message/user data and declared safe platform API actions.
Still required for stricter acceptance:
- Send a private image and confirm common `Image` reaches the plugin.
- Have the plugin call `send_message` and `reply_message` for text and one media component, then verify the WeCom client receives the bot output.
- Exercise `send_to_all` only with a disposable visible-contact scope.
- Trigger one non-text/image callback, if available, and confirm it becomes `PlatformSpecificEventReceived`.
## Current Acceptance
Current status is **partial EBA acceptance**.
Blocked items:
- Real inbound image/voice/file evidence was not completed in this run.
- Inbound voice/file callback parsing is not present in the legacy `WecomClient.get_message()` path, so the EBA adapter does not claim those receive components yet.
- Group/member/moderation APIs do not apply to this regular WeCom application-message adapter.
@@ -1,148 +0,0 @@
# WeComBot EBA Adapter
## Status
WeCom AI Bot now has an EBA adapter directory:
```text
src/langbot/pkg/platform/adapters/wecombot/
├── adapter.py
├── api_impl.py
├── event_converter.py
├── manifest.yaml
├── message_converter.py
├── platform_api.py
└── types.py
```
The adapter is registered as `wecombot-eba`.
This is separate from regular WeCom internal applications (`wecom-eba`). WeComBot supports WebSocket long connection mode, which does not require a webhook URL. Webhook mode remains available when `enable-webhook=true`.
## Configuration
| Field | Required | Default | Description |
|-------|----------|---------|-------------|
| `BotId` | Yes for WebSocket mode | `""` | WeCom AI Bot ID. |
| `robot_name` | Yes | `""` | Bot display name used to strip bot mentions from incoming group text. |
| `enable-webhook` | Yes | `false` | `false` uses WebSocket long connection mode; `true` uses webhook callback mode. |
| `webhook_url` | No | `""` | Unified webhook URL, only needed when webhook mode is enabled. |
| `Secret` | Yes for WebSocket mode | `""` | WeCom AI Bot secret for long connection mode. |
| `Corpid` | Yes for webhook mode | `""` | WeCom corporate ID for webhook callback mode. |
| `Token` | Yes for webhook mode | `""` | WeCom callback token. |
| `EncodingAESKey` | Yes for webhook mode; optional for WebSocket media decrypt | `""` | Message encryption/decryption key. |
| `enable-stream-reply` | No | `true` | Enables WeComBot streaming replies. |
## Events
WeComBot declares these EBA events:
- `message.received`
- `feedback.received`
- `platform.specific`
`message.received` covers private and group messages from the WeComBot SDK. `feedback.received` covers WeComBot like/dislike feedback callbacks. Native SDK events without a common EBA equivalent are emitted as `platform.specific`.
## Common APIs
| API | Status | Notes |
|-----|--------|-------|
| `send_message` | Supported in WebSocket mode | Sends proactive markdown/text to a person or group chat ID. Webhook mode raises `NotSupportedError` because the platform callback flow has no proactive send path here. |
| `reply_message` | Supported | Replies through native `req_id` in WebSocket mode or stream finalization/cache in webhook mode. |
| `get_message` | Supported from cache | Returns cached inbound `MessageReceivedEvent` by message ID. |
| `get_user_info` | Supported from cache | WeComBot events carry user info; no full user lookup endpoint is declared. |
| `get_friend_list` | Partial | Returns users observed by this adapter instance. |
| `get_group_info` | Supported from cache | Returns groups observed from inbound group messages. |
| `get_group_member_info` | Supported from cache | Returns observed sender/group-member pairs. |
| `get_group_member_list` | Partial | Returns observed members for the cached group only. |
| `call_platform_api` | Supported | See below. |
| `edit_message` / `delete_message` / `forward_message` | Not supported | WeComBot does not expose portable common APIs for these operations in the current SDK wrapper. |
| `upload_file` / `get_file_url` | Not supported as common APIs | Media is represented inside messages; no portable standalone file upload/URL API is declared. |
| moderation / leave APIs | Not supported | WeComBot does not expose equivalent common moderation operations through this adapter. |
## Platform-Specific APIs
`call_platform_api(action, params)` supports:
- `is_websocket_mode`
- `get_stream_session_status`
- `send_markdown`
`send_markdown` is only available in WebSocket mode.
## Unit Verification
Covered by:
```bash
PYTHONPATH=/Users/wangqiang/code/python/langbot-plugin-sdk/src uv run pytest tests/unit_tests/platform/test_wecombot_eba_adapter.py
```
The unit tests cover:
- Manifest events/APIs/platform actions match adapter declarations.
- Outbound common components flatten to WeComBot markdown/text.
- Private and group native events become `MessageReceivedEvent`.
- Inbound image, file, voice, and quote components map to common `MessageChain`.
- Legacy `FriendMessage`/`GroupMessage` compatibility.
- EBA listener dispatch, message/user/group/member cache, reply, send, streaming chunk, feedback, and platform API calls.
## Live Probe
The direct adapter probe is:
```bash
PYTHONPATH=/absolute/path/to/langbot-plugin-sdk/src uv run python tests/e2e/live_wecombot_eba_probe.py --help
```
Default mode is WebSocket long connection and requires:
- `WECOMBOT_BOT_ID`
- `WECOMBOT_SECRET`
- `WECOMBOT_ROBOT_NAME`
- optional `WECOMBOT_ENCODING_AES_KEY`
Webhook mode uses `--webhook` and requires:
- `WECOMBOT_TOKEN`
- `WECOMBOT_ENCODING_AES_KEY`
- `WECOMBOT_CORPID`
The probe writes JSONL evidence to `data/temp/wecombot_eba_live_probe.jsonl`, waits for a real WeComBot message, records common EBA event fields and message components, then runs safe cached/common/platform API checks.
## Standalone Runtime Plugin E2E Record
Verified on May 27, 2026 with `EBAEventProbe`, SDK standalone runtime, LangBot core, and the real WeCom desktop client in a WeCom AI Bot private chat.
Evidence:
- JSONL: `data/temp/wecombot_eba_plugin_probe.jsonl`
- Bot UUID: `9f5d4125-7b6d-4c98-8ca2-111111111111`
- Adapter: `wecombot-eba`
- Client: real WeCom desktop client, private `LangBot` BOT chat
- Mode: WebSocket long connection (`enable-webhook=false`)
Observed and verified:
- A real user-side message reached the plugin as `MessageReceived` with `adapter_name=wecombot-eba`, common sender/chat fields, and `Source + Plain`.
- SDK API calls succeeded through the standalone runtime: `get_langbot_version`, `get_bots`, `get_bot_info`, `send_message`, plugin/workspace storage, manifest/list APIs, and safe cached common platform APIs.
- Outbound component sweep was visible in the WeCom client and returned `errcode=0`: plain/mention/face fallback, base64 image marker, quote fallback, file marker, and flattened forward fallback.
- Declared WeComBot platform APIs succeeded through `plugin.call_platform_api`: `is_websocket_mode`, `get_stream_session_status`, and `send_markdown`.
- The `send_markdown` platform API produced visible bot output in the WeCom client.
Not completed:
- Clicking the visible WeCom AI feedback button did not produce a `FeedbackReceived` JSONL entry in this run, so `feedback.received` remains unverified at plugin E2E level.
- Group chat inbound and group cache/member coverage still need a real group-side trigger.
- Real inbound image/file/voice from the WeCom client was not exercised.
## Current Acceptance
Current status is **partial EBA acceptance**.
Blocked or limited items:
- `feedback.received` is implemented and unit-covered, but real plugin E2E feedback evidence was not observed from the desktop client click.
- Outbound image/voice/file are flattened as textual markers because the WeComBot SDK reply/proactive path used here is markdown/text oriented.
- Group member APIs are cache-backed and only know members observed in received messages.
- Destructive or moderation APIs are not declared because the current WeComBot protocol surface does not provide safe common equivalents.
-161
View File
@@ -1,161 +0,0 @@
# WeCom Customer Service EBA Adapter
## Status
WeCom Customer Service now has an EBA adapter directory:
```text
src/langbot/pkg/platform/adapters/wecomcs/
├── adapter.py
├── api_impl.py
├── event_converter.py
├── manifest.yaml
├── message_converter.py
├── platform_api.py
└── types.py
```
The adapter is registered as `wecomcs-eba`. It is separate from regular WeCom application messages (`wecom-eba`) and WeCom AI Bot (`wecombot-eba`).
## Configuration
| Field | Required | Default | Description |
|-------|----------|---------|-------------|
| `webhook_url` | No | `""` | Unified webhook URL copied into the WeCom Customer Service callback settings. |
| `corpid` | Yes | `""` | WeCom corporate ID. |
| `secret` | Yes | `""` | Customer Service secret used for access tokens. |
| `token` | Yes | `""` | Customer Service callback token. |
| `EncodingAESKey` | Yes | `""` | Customer Service callback encryption key. |
| `api_base_url` | No | `https://qyapi.weixin.qq.com/cgi-bin` | WeCom API base URL, overrideable for proxy/private-network deployments. |
## Events
| Event | Status | Notes |
|-------|--------|-------|
| `message.received` | Plugin E2E UI covered for text | Text, image, file, and voice payloads convert to common EBA message components in unit tests. Real WeChat customer-side UI text reached `EBAEventProbe` on May 27, 2026. |
| `platform.specific` | Unit covered | Non-message or unknown Customer Service payloads become structured `PlatformSpecificEvent` records. |
## Common APIs
| API | Status | Notes |
|-----|--------|-------|
| `send_message` | Plugin E2E outbound covered | Private/person target only. `target_id` must be `external_userid|open_kfid`. Text and image are implemented; voice/file are explicitly unsupported. |
| `reply_message` | Plugin E2E partial | Replies through Customer Service `kf/send_msg` using the original `source_platform_object`. The pipeline reply path reached the send API, but the dev account later hit WeCom `95001 send msg count limit`. |
| `get_message` | Plugin E2E covered from cache | Returns cached inbound `MessageReceivedEvent` by message ID. |
| `get_user_info` | Plugin E2E covered | Uses cached event users first, then Customer Service `customer/batchget`. |
| `get_friend_list` | Plugin E2E covered, partial | Returns customer users seen by this adapter instance. |
| `call_platform_api` | Unit covered | See platform-specific APIs below. |
| `edit_message` / `delete_message` | Not supported | WeCom Customer Service does not expose a general edit/delete endpoint for bot-sent messages in this adapter. |
| Group/member/moderation APIs | Not supported | Customer Service conversations handled here are private customer sessions, not group chats. |
| `upload_file` / `get_file_url` | Not supported | Media upload is used internally for outbound image; no portable file URL common API is exposed. |
## Platform-Specific APIs
| Action | Status | Notes |
|--------|--------|-------|
| `check_access_token` | Unit covered | Checks whether the current access token is present. |
| `refresh_access_token` | Unit covered | Refreshes the Customer Service access token. |
| `get_customer_info` | Unit covered | Calls Customer Service customer lookup by `external_userid`. |
## Message Components
Receive:
| Component | Status | Notes |
|-----------|--------|-------|
| `Source` | Unit covered | Uses Customer Service `msgid` and `send_time`. |
| `Plain` | Unit covered | Text payload content is preserved. |
| `Image` | Unit covered | Uses the base64 data URL produced by the existing SDK image download path. |
| `Voice` | Unit covered | Maps exposed voice media ID to common `Voice.voice_id`; live UI evidence pending. |
| `File` | Unit covered | Maps exposed file media ID/name/size to common `File`; live UI evidence pending. |
| `Quote`, `At`, `AtAll`, `Face`, `Forward` | Not supported inbound | The current Customer Service SDK event model does not expose these as structured inbound fields. |
| `Unknown` | Unit covered | Unsupported message types become `Unknown` in message conversion or `platform.specific` at event level. |
Send:
| Component | Status | Notes |
|-----------|--------|-------|
| `Plain` | Plugin E2E outbound covered | Sends through `kf/send_msg` text. |
| `Image` | Plugin E2E outbound covered | Uploads media as WeCom image media and sends through `kf/send_msg` image. |
| `Quote`, `At`, `AtAll`, `Forward` | Unit covered fallback, live partially blocked | Flattened to text where possible. In the May 27 sweep, later text sends hit WeCom `95001 send msg count limit` after the successful text/image sends. |
| `Voice`, `File`, `Face` | Not supported | The adapter raises `NotSupportedError`; no tested Customer Service send path is implemented. |
## Unit Verification
Covered by:
```bash
PYTHONPATH=/Users/wangqiang/code/python/langbot-plugin-sdk/src uv run pytest tests/unit_tests/platform/test_wecomcs_eba_adapter.py
```
Result on May 27, 2026: `10 passed`.
The local `PYTHONPATH` is required in this workspace because the installed SDK package in the LangBot venv does not contain the newer `langbot_plugin.api.entities.builtin.platform.errors` module; the existing EBA adapter tests need the same SDK override.
## Live Probe
Auxiliary direct adapter probe:
```bash
PYTHONPATH=/path/to/langbot-plugin-sdk/src uv run python -m py_compile tests/e2e/live_wecomcs_eba_probe.py
WECOMCS_CORPID=... \
WECOMCS_SECRET=... \
WECOMCS_TOKEN=... \
WECOMCS_ENCODING_AES_KEY=... \
PYTHONPATH=/path/to/langbot-plugin-sdk/src \
uv run python tests/e2e/live_wecomcs_eba_probe.py \
--path /wecomcs/callback \
--log data/temp/wecomcs_eba_live_probe.jsonl
```
This probe is diagnostic only. Final EBA acceptance still requires the standalone SDK runtime plus `EBAEventProbe` plugin path.
## Standalone Runtime Plugin E2E Record
Completed partial plugin E2E on May 27, 2026 against `dev.rockchin.top` and the WeChat customer-side UI entry `微信 -> 客服消息 -> 浪波智能客服`.
Evidence:
- Server JSONL: `/home/wgc/LangBotxg/LangBotEbaTest/data/temp/wecomcs_eba_plugin_probe.jsonl`
- Trigger text: `EBA wecomcs dedupe probe 2026-05-27`
- `bot_uuid`: `cc810d2c-91f3-4f92-8f27-e1bf9f7b6cb4`
- `adapter_name`: `wecomcs-eba`
- Observed common event: `MessageReceived`, `event.type=message.received`
- Observed message chain: `Source + Plain`
- Observed chat: `chat_type=private`, `chat_id=external_userid|open_kfid`
- Observed sender: customer `User` with nickname/avatar from Customer Service lookup
- Plugin API probe: `send_message`, `get_message`, `get_user_info`, `get_friend_list`, plugin/workspace storage, and manifest/list APIs succeeded
- Component sweep: outbound `Plain` and `Image` succeeded; `Face` and `File` returned explicit `NotSupportedError`; later quote/forward fallback sends were blocked by WeCom `95001 send msg count limit`
Command shape used:
```bash
cd langbot-plugin-sdk
uv run python -m langbot_plugin.cli.__init__ rt --debug-only --ws-control-port 5400 --ws-debug-port 5401 --skip-deps-check
cd LangBot
PYTHONPATH=/absolute/path/to/langbot-plugin-sdk/src uv run main.py --standalone-runtime
cd data/plugins/LangBot__EBAEventProbe
DEBUG_RUNTIME_WS_URL=ws://127.0.0.1:5401/plugin/ws \
EBA_PROBE_LOG=/absolute/path/to/LangBot/data/temp/wecomcs_eba_plugin_probe.jsonl \
EBA_PROBE_API=1 \
EBA_PROBE_COMPONENT_SWEEP=1 \
EBA_PROBE_PLATFORM_API=1 \
uv --project /absolute/path/to/langbot-plugin-sdk run python -m langbot_plugin.cli.__init__ run
```
Required real UI trigger: send a Customer Service message from the WeCom/WeChat customer-side UI to the configured `dev.rockchin.top` Customer Service account.
## Current Acceptance
Current status is **partial EBA acceptance**.
Blocked or pending items:
- Inbound UI media (`Image`, `Voice`, `File`) was not sent from the real WeChat customer UI during this run, so receive-side media remains unit-covered only.
- Pipeline auto-reply reached `kf/send_msg`, but the test account hit WeCom `95001 send msg count limit` after successful plugin outbound text/image sends. This is recorded as an account/platform rate-limit block, not a conversion or API-shape failure.
- The current `EBAEventProbe` run did not call the adapter-specific `call_platform_api` actions (`check_access_token`, `refresh_access_token`, `get_customer_info`); the platform API map remains unit-covered.
- Inbound voice/file depends on whether the real Customer Service callback plus `sync_msg` endpoint returns those fields in the shape the local SDK models.
- Group, member, edit, delete, moderation, and standalone file URL APIs are intentionally not declared because this Customer Service protocol path does not provide tested common equivalents.
@@ -0,0 +1,114 @@
# LangBot Cloud 24 小时资源 Soak 门禁
`scripts/cloud_runtime_soak.py` 是生产候选拓扑的最终资源稳定性门禁。它不替代单元测试、历史 churn 探针或 nsjail 隔离测试;它把以下三类证据按同一时间轴采集并给出可机读的 pass/fail:
- Core、Plugin Runtime 和 Box Runtime 的 HTTP liveness/readiness。
- 三个 Python 进程的 event-loop recent max/p95 调度延迟。
- Linux `/proc` 进程树的 current RSS、累计 CPU、线程、文件描述符和子进程数。
- cgroup v2 的 `memory.current/peak/events`、swap、CPU usage/throttling、PID current/events 和实际硬限制。
生产批准必须使用 cgroup 证据。`--pid` 只适合本地诊断,因为进程指标无法证明 OOM kill、PID limit 或 CPU throttling。
## 运行位置
建议把采集器放在独立的 node agent 或监控 sidecar 中,并只读挂载三个目标容器的 cgroup 路径。不要把采集器和样本文件放进被测容器自己的 cgroup/数据卷,否则采集器的 CPU、内存和 page cache 会污染目标数据。
Kubernetes/containerd 生成的 cgroup 路径不是稳定 API。每次生产候选部署都必须从实际 pod/container ID 解析,不能从 pod 名猜路径。传入的每个目录都必须至少可读:
- `memory.current``memory.events`
- `cpu.stat``cpu.max`
- `pids.current``pids.events`
最终门禁应加 `--require-hard-limits`。该选项要求每个目标 cgroup 都能观察到有限的 CPU quota、memory、swap 和 PID 上限;任一值为 `max` 都失败。
## 标准 24 小时命令
```bash
uv run python scripts/cloud_runtime_soak.py \
--duration 24h \
--startup-grace 5m \
--sample-interval 15s \
--cooldown 30m \
--analysis-window 30m \
--http-timeout 5s \
--max-memory-growth-mib 64 \
--max-memory-slope-mib-per-hour 32 \
--max-tail-cpu-cores 0.5 \
--max-throttled-period-ratio 0.25 \
--max-event-loop-lag-ms 1000 \
--max-event-loop-p95-lag-ms 250 \
--require-hard-limits \
--endpoint core=http://langbot:5300/healthz \
--endpoint plugin=http://langbot-plugin-runtime:5400/healthz \
--endpoint box=http://langbot-box:5410/readyz \
--cgroup core=/host-cgroup/CURRENT_CORE_CONTAINER \
--cgroup plugin=/host-cgroup/CURRENT_PLUGIN_RUNTIME_CONTAINER \
--cgroup box=/host-cgroup/CURRENT_BOX_RUNTIME_CONTAINER \
--samples-file artifacts/cloud-soak-samples.jsonl \
--report-file artifacts/cloud-soak-report.json \
--workload uv run python tests/load/cloud_candidate_workload.py
```
`--duration` 是包含启动观察、负载和冷却期的最大墙钟时间。工作负载必须在截止时间前退出并至少留出 30 分钟冷却;否则门禁会终止负载并失败。若负载由外部系统控制,可以省略 `--workload`,但必须保证最后 `--analysis-window` 完全无测试流量,该窗口才可解释为空闲尾段。
工作负载命令的 stdout/stderr 会转发到采集器 stderr,不会混入 stdout 的最终 JSON 报告。命令以独立 process group 启动;超时或中断时整组收到 TERM,10 秒后仍未退出则收到 KILL。
凭据只能通过 workload 进程环境或 secret mount 注入,不能放在命令参数中。最终报告只记录可执行文件名和参数个数,不保存参数正文;采集器也拒绝带 userinfo、query 或 fragment 的健康 URL。
## 必须覆盖的负载
同一候选版本至少要覆盖:
1. 大批 Workspace 注册、成员邀请、登录和 entitlement 刷新。
2. Plugin installation reconcile、依赖准备、正常调用、进程崩溃与重启。
3. Dashboard/Embed/平台 WebSocket 建连、突发消息和批量断连。
4. Box session、文件同步、并发 exec、managed-process 输出和清理。
5. PostgreSQL pool 接近容量、事务超时和恢复。
6. Core、Plugin Runtime、Box 分别收到 SIGTERM 后的优雅重启。
工作负载不能把 API 过载拒绝当作成功吞掉。默认情况下,Core health 中 blocking executor 的 global/scope rejection counter 只要增长,门禁即失败;只有专门验证“过载会正确返回 429”的独立测试才可以使用 `--allow-rejections`,该次运行不能作为生产批准证据。
## 判定规则
整个有效观察期内出现以下任一情况即失败:
- 健康接口请求失败、非 2xx、Core `code != 0`,或 Box `ready=false`
- `memory.events.high/max/oom/oom_kill/oom_group_kill` 增长。
- `pids.events.max` 增长。
- cgroup 单调计数器回退,表示目标很可能发生了未记录的重启或 cgroup 替换。
- CPU throttled-period ratio 超过配置阈值。
- 任一健康采样窗口的 event-loop recent max 超过 1 秒,或冷却尾段 recent p95 超过 250 ms。
- 健康接口缺少 event-loop monitor、monitor 未持续运行,或其 sample counter 回退。
- blocking executor rejection counter 增长。
- Plugin Runtime restart circuit 的累计打开次数增长。
- Core 目录 active Workspace、最近 snapshot/delta Workspace 或 membership 基数
超过各自配置上限,或 PostgreSQL `checked_out` 超过配置 pool 容量;相关 current/max
指标只出现一半或 max 非法也失败。
负载结束后的冷却尾段还必须满足:
- `memory.current`/RSS 的稳健首尾增长和线性斜率不能同时超过阈值。
- 平均 CPU 核数不超过 `--max-tail-cpu-cores`
- event-loop recent p95 不超过 `--max-event-loop-p95-lag-ms`
- blocking executor `pending` 至少回到过零;不能整个尾段持续积压。
- Plugin Runtime restart coordinator 的 active launch、half-open probe 和
circuit open remaining time 必须回到零,`gate_waiters` 必须至少归零一次。
- Core 的 MCP projection retirement queue/worker 和 message aggregation
buffer/scope 必须至少归零一次。
- telemetry、QueryPool、MCP host/dispatch、Box creating/closing/background 等临时 gauge 不能继续增长。
内存判定要求“增长量”和“斜率”同时越界,避免几 MiB allocator/page-cache 噪声在短窗口被外推成很大的每小时斜率。最终报告仍保留实际增长与斜率,人工审查时不能只看 verdict。
## 产物与退出码
- `--samples-file`:逐样本 JSONL,写入后立即 flush,供时序图和故障定位。
- `--report-file`:最终汇总、阈值、资源硬限制、OOM/PID/throttle delta、尾段斜率和 workload 状态。
- stdout:与 report 文件相同的最终 JSONworkload 日志只写 stderr。
退出码:
- `0`:全部门禁通过。
- `1`:采样完成但资源门禁失败。
- `2`CLI 参数或目标配置错误。
必须保存原始 JSONL、最终报告、三个镜像 digest、LangBot/SDK commit、生产配置摘要和工作负载版本。滚动更新、节点迁移或镜像变化后,旧报告不能继续作为新候选版本的批准证据。
@@ -0,0 +1,273 @@
# Cloud v2 仍待验证事项
状态:`NOT APPROVED FOR SAAS ACTIVATION`
更新日期:2026-07-29
本文是 Cloud v2 首期上线前的剩余验证清单。它只记录尚不能由当前代码审查、
单元测试、集成测试、合成容量探针或短时 Linux 容器实验替代的证据。这里的项目
不属于 2026-07-29 代码与本地测试资源审查的完成条件,也不会让该审查持续保持未完成;
它们只在准备最终 SaaS 激活时重新进入验收范围。
相关文档:
- [多租户架构决策](./pending-architecture-decisions.md)
- [实现决策记录](./implementation-decisions.md)
- [Runtime 资源安全审查](./runtime-resource-audit-2026-07-28.md)
- [24 小时资源 Soak 门禁](./cloud-runtime-soak-gate.md)
## 1. 当前已形成的交付基线
- LangBot Core 全量 `2855 passed, 33 skipped`Plugin SDK 全量
`1328 passed`,闭源适配器 `40 passed`,Space Go 全量测试通过;三仓格式、
静态检查和 `git diff --check` 已通过。
- Plugin Runtime 和 Box Runtime 的公开健康接口、event-loop lag 与有界
blocking executor 指标已经过真实进程短时验证。
- 仓库 Dockerfile 构建的 Linux/cgroup v2 短时探针已证明 CPU、memory、
swap 和 PID 限制代码路径可工作。
- PostgreSQL 16 + RLS 的 1,000 Workspace 真实启动测试,以及 5,000
Workspace 三代替换合成探针已通过。
- Core 已精确钉住 Plugin SDK 提交
`1d65ed301a6afc52150a998043f73cd6032c8162`。最终验证必须使用包含该提交的
Core、Plugin Runtime 和 Box Runtime 镜像,不能混用旧 SDK。
- 独立资源复核已经移除 Cloud MCP 每会话 5 秒查询执行绑定的轮询,改由签名目录
投影提交后向一个合并回收任务发布代次变化;工具与资源调用前后仍使用数据库
execution fence。Plugin restart 冷却等待者、MCP 投影回收、消息聚合 buffer/scope
均已纳入健康快照和 soak 归零门禁。
- 单实例目录现在有一致的操作容量契约:Space 在注册事务内通过 PostgreSQL
advisory lock 串行执行 active Workspace check-and-createSpace 全量快照只返回
active Workspace,并在查询阶段限制 Workspace/membership 数量;闭源适配器限制
解压后的 HTTP 响应字节和签名目录基数;Core 在持有目录投影行锁的事务内再次
COUNT active Workspace,超限时整批回滚且不推进 cursor。任何一层都不截断权威数据。
- Core Cloud PostgreSQL pool 的 `pool_size + max_overflow` 有绝对上限 100
Cloud runtime 连接默认强制 60 秒 statement/idle-transaction timeout 和 5 秒
lock timeoutpool 使用量、超时累计数与目录 active/max 基数进入 `/healthz`
Box Runtime 的 session、process、admission record、RPC 文件和 completed retention
配置也有不可被实例配置放大的绝对上限。
- Space 的 concurrent registration 容量准入已在一次性 PostgreSQL 16 上真实执行:
两个 Account 同时争用最后一个 Workspace 槽位时,精确一个事务成功、一个事务
得到 capacity error,最终 active Workspace 数为 1。active-only snapshot 与
archived delta tombstone 的同一真实 PostgreSQL 集成流程也通过。
- Core Cloud manager 已连接一次性 PostgreSQL 16,并从 `pg_settings` 读回
`statement_timeout=60000ms``lock_timeout=5000ms`
`idle_in_transaction_session_timeout=60000ms`;测试结束后引擎已显式 dispose。
- 独立异常路径复核已补齐 HTTPX 超限/取消时的底层流关闭;Monitoring 查询、导出和
detail 物化量均有实例上限与绝对上限,detail 统计使用数据库聚合。Token statistics
不再拉取全部历史 LLM call 在 Python 中分桶,而由 PostgreSQL/SQLite 聚合并只返回
有界的最新时间桶和模型分组,截断状态在响应中显式可见。邀请、Monitoring 和 Storage
周期清理已合并为一个先等待首个 interval 的调度器,同一周期只进行一次 Workspace
discovery;数据库删除批次和本地/S3 文件候选也有每轮硬上限。
以上结果是进入生产候选验证的前提,不是 SaaS 上线批准。
## 2. 尚有实现前置条件的阻断项
以下项目不是“再跑一次测试”即可关闭。必须先完成实现,再执行对应验收。
| 编号 | 阻断项 | 完成实现后的最低验收证据 |
| --- | --- | --- |
| B-01 | Cloud 插件缺少生产 egress policy | 证明插件只能访问允许的公网目标,不能访问 Core/Box/数据库、其他内部服务、loopback、link-local 或云 metadata endpoint |
| B-02 | Plugin installation 与 Box Workspace/Skill/root/tmp/home 缺少真实的 byte 和 inode 硬配额 provider | 在写入边界原子拒绝超额;并发写入、重启和配额耗尽后不能越界,也不能用目录扫描或事后清理冒充硬配额 |
| B-03 | 普通业务写入尚未具备贯穿 commit 的 generation-aware fence、同事务 business outbox,以及 generation cutover 后稳定的 durable-object 引用 | 在旧 generation 与新 generation 并发、事务提交竞态和重复投递下,旧 owner 不产生业务写入或外部副作用,outbox 可幂等恢复 |
任一 B 类项目未关闭时,不得把 24 小时 soak 的通过结果解释为可以上线。
## 3. 最终部署环境验证
### V-01Plugin Runtime 与 Box 的 Linux 隔离
必须在最终 Cloud Pod security context、容器 runtime 和 cgroup 拓扑中验证,
不能使用开发机或权限不同的一次性容器替代。
验证内容:
1. nsjail 可以建立 mount、PID、IPC、UTS namespace 和 private `/proc`
2. delegated cgroup v2 对每个插件进程和 sandbox 强制 CPU、memory、
`memory.swap.max=0` 和 PID 上限。
3. open files、process 和单文件大小 rlimit 生效。
4. 插件不能枚举、读取或 signal Runtime 及其他 installation 的进程,
不能读取其他 installation 的 home/tmp/data,也不能修改共享只读
artifact/environment。
5. 超额只杀死或拒绝当前 installation/sandboxRuntime、其他租户和健康接口
继续工作。
6. 进程退出、取消、超时、generation 切换和 Pod SIGTERM 后,cgroup、nsjail
目录、子进程和文件描述符均被回收。
7. 硬限制或 namespace 能力缺失时 readiness 失败关闭,不允许降级成普通进程。
通过证据必须包含实际容器安全配置、cgroup 文件值、探针原始输出和失败注入结果。
### V-02Box 持久卷与硬存储配额
在 B-02 的 quota provider 实现后,必须验证:
1. Core 与 Box Runtime 通过随机 marker challenge 证明使用同一共享持久卷。
2. Workspace、Skill store、ephemeral root/tmp/home 的 byte 和 inode quota
都在写入点生效。
3. 并发写、压缩包展开、文件同步、重启恢复和删除重建不能绕过配额。
4. 配额耗尽只影响目标 Workspace,其他 Workspace 仍能执行。
5. 任一硬存储能力缺失时 `/readyz` 返回非 2xx,Pod 不进入就绪流量。
### V-03PostgreSQL 与 pgvector 生产边界
必须在最终 PostgreSQL endpoint、凭据和网络策略下验证:
1. migrator 与 runtime 使用不同 roleruntime role 无 superuser、
`BYPASSRLS`、DDL、对象所有权、role membership 或额外 schema 权限。
2. runtime credential 只能连接目标 business database。需要专用
cluster/endpoint,或经过测试的 HBA/proxy policydatabase 内 catalog
audit 本身不能证明这一点。
3. release migration Job 的 advisory lock、失败重试、回滚和精确 Alembic head
校验有效;应用启动角色不能执行 migration 或其他 DDL。
4. 若使用 PgBouncertransaction pooling、异常回滚和连接复用不会残留
tenant context。
5. 故意遗漏应用层 Workspace filter 时,RLS 仍阻止跨租户读写。
6. 两个 Workspace 使用相同 `vector_id`、猜测其他 Workspace ID、后台任务和
连接复用时,pgvector CRUD 均不能越权。
7. dimension mismatch、extension/schema/ACL drift 或 runtime audit 失败时,
Core 启动失败且不回退到其他向量后端。
### V-04:最终镜像和配置一致性
生产候选验证前必须固定:
- Core、Plugin Runtime、Box Runtime 的不可变镜像 digest
- LangBot 与 SDK commit
- `data/config.yaml` 的非敏感摘要和所有环境变量覆写;
- Space 的 `CLOUD_V2_MAX_DIRECTORY_WORKSPACES` 必须与 Core
`cloud.directory.max_active_workspaces`
`cloud.directory.max_snapshot_workspaces` 一致;Space 的 membership 上限必须与
Core `cloud.directory.max_snapshot_memberships` 一致;
- Core 的 PostgreSQL pool、statement/lock/idle-transaction timeout,以及 Plugin
Runtime/Box Runtime 的全部实例级资源上限;
- PostgreSQL migration revision
- Cloud Adapter、Space control plane 和 workload 的版本。
滚动更新、节点迁移、配置变化或任一镜像 digest 变化后,旧验证报告失效。
## 4. 多租户行为与故障注入
### V-05:目录、entitlement 与 generation
在真实 Space control plane、闭源 Cloud Adapter 和 Core 之间验证:
1. 注册自动创建个人 Workspace、owner membership、Free subscription、
entitlement snapshot 和 outbox 事件,且重试不重复创建。
2. 邀请、成员变更、套餐变更和 Workspace 撤销只影响目标 Workspace。
3. 全量快照与增量事件覆盖乱序、重复、断点续传、缺页、签名错误、
high-water gap、snapshot coverage 和消费者重启。
4. Core、Plugin Runtime 或 Box 重启后,权威 desired state 可恢复;
本地进程表和缓存不是唯一真相。
5. generation/revision 切换期间,旧 callback、RPC、WebSocket、Box relay、
plugin worker 和缓存写入全部失败关闭。
6. 为未来多副本预留的 replica-local cursor 语义通过故障注入:
一个副本追平不能使另一个副本跳过本地 cache 刷新。
7. 同时使大量 plugin worker 因系统性故障退出,证明 restart launch 全局并发受限、
失败阈值触发 Runtime circuit、冷却后只有一个 half-open probe,且 probe 未稳定前
其他 installation 不会继续重启;冷却计时器/状态等待者数量不得超过全局 restart
并发,取消 probe 不得把 circuit 永久卡在 half-open24 小时门禁必须把 circuit
打开或 `gate_waiters` 未归零判为失败。
8. 使用大量空闲 remote MCP session 做 generation 切换,证明目录投影只创建一个
合并回收任务,不产生每 session 周期数据库查询、计时器或同时唤醒;旧 session
最终关闭,`mcp_projection_retirements`
`mcp_projection_reconcile_active` 在冷却期归零。
### V-06:套餐、Box 与 stdio MCP
1. Free/非 Pro Workspace 不会自动获得 managed sandbox。
2. 合资格 Workspace 最多只有一个持久 `global` sandbox,且新增 Workspace
不创建专属 Runtime、Pod、PVC、database、schema、role 或连接池。
3. Cloud 即使 `box.enabled=true`,也不能 create/update/test/start stdio MCP
旧记录和直接 API 调用同样失败关闭,且不创建 `mcp-shared` session。
4. OSS 默认仍是单 Workspace、多用户,stdio MCP 保持兼容,多租户能力不会被
未签名配置或普通环境变量开启。
### V-07:跨租户安全回归
至少使用两个恶意测试 Workspace 验证:
- 同 digest 插件只共享只读代码和依赖;进程、secret、日志及所有可写目录隔离。
- 同 author/name/version 但 digest 不同的 artifact 不共享目录。
- Plugin Host API、Box RPC、对象 key、WebSocket、RAG、storage、model/session
cache 和平台回调不能接受调用方伪造的 Workspace scope。
- 撤销 entitlement、删除 installation 或 generation 切换后,已有长连接和
in-flight 请求不能继续访问旧权限。
## 5. 生产候选容量与 24 小时门禁
### V-08:真实容量曲线
现有 fake adapter/requester/Plugin handler 探针不能替代真实容量数据。必须使用
计划上线的平台 SDK、外部 HTTP/WebSocket 连接池、真实插件进程、真实
PostgreSQL/pgvector 和代表性 Workspace 配置分布,测量:
- 空 Workspace、活跃 Workspace、每个启用插件和每个 Pro sandbox 的边际
RSS、线程、文件描述符、连接和 PostgreSQL pool 成本;
- 启动、目录重放、批量 reconcile 和故障恢复的耗时与峰值;
- remote MCP 数量增加及目录 generation 批量切换时的数据库 QPS、回收队列和
event-loop lag,确认不存在与 session 数量成比例的空闲轮询;
- 在最大 retention/backlog 和并发 Dashboard 请求下执行不带时间范围的 Monitoring
overview/token statistics,验证 SQL 分桶、statement timeout、响应截断和 cleanup
追赶不会形成 PostgreSQL CPU 尖峰或 Core RSS 增长;
- 单实例可批准的 Workspace、活跃 Bot、plugin worker 和 sandbox 上限。
容量上限必须写入生产配置与告警,不能只保留在测试报告中。
代码中的默认值和绝对上限只是失控配置的最后防线,不等于生产容量结论。V-08 必须
根据最终镜像的真实曲线把 Space 与 Core 的匹配上限调到已验证容量以内;如果最终
批准值高于当前默认 1,000 active Workspace,必须重新执行目录启动、故障恢复和
24 小时门禁。
### V-0924 小时资源 soak
使用 [标准 24 小时命令](./cloud-runtime-soak-gate.md#标准-24-小时命令),并强制
`--require-hard-limits`。工作负载至少覆盖:
1. 注册、邀请、登录和 entitlement 刷新;
2. plugin reconcile、依赖准备、调用、崩溃与重启;
3. Dashboard/Embed/平台 WebSocket 建连、突发消息和断连;HTTP Bot 覆盖
高基数 session/idempotency、硬容量拒绝、空闲回收及 callback 堵塞;
remote MCP 覆盖大量空闲连接、批量 generation 切换和合并回收;
4. Box session、文件同步、并发 exec、输出与清理;
5. PostgreSQL pool 接近容量、事务超时和恢复;
6. Core、Plugin Runtime、Box 分别 SIGTERM 和恢复。
最后至少保留 30 分钟无测试流量冷却。任一健康失败、OOM/memory pressure、
PID limit、blocking executor rejection、超阈值 CPU throttling/event-loop lag、
目录 `active_workspaces > max_active_workspaces`、数据库 pool 使用量超过配置容量、
冷却尾段内存持续增长,或 Plugin restart `gate_waiters`、MCP 投影回收、
消息聚合 buffer/scope 等临时 gauge 不回落都判为失败。
标准 soak 工具已自动比较目录 active/最近批次与各自配置上限,并比较 PostgreSQL
`checked_out` 与配置 pool 容量;名为 `core` 的标准 endpoint 缺少任一容量指标、
current/max 只出现一半、数值非法或任一样本越界都会直接失败。
必须归档:
- 原始 `cloud-soak-samples.jsonl`
- 最终 `cloud-soak-report.json`
- 三个镜像 digest、Core/SDK commit
- 生产配置摘要、数据库 migration revision 和 workload 版本;
- 故障注入时间线及关联日志/trace。
## 6. 本轮不作为验收条件的后续事项
以下能力已明确暂缓,不能混入当前验证结果,也不能以“尚未验证”为理由临时发明方案:
- Workspace export、释放、delete、单 Workspace restore 和在线迁移;
- Workspace 级 BYOK E2B WebUI 配置;
- 多 Core/Plugin Runtime/Box replica 的 lease store 与调度实现;
- PostgreSQL 多 shard、dedicated shard 和跨地域部署;
- 多 CloudInstance、Cell Router 或 Workspace Placement。
这些事项需要后续单独决策。首期实现仍需保留稳定 UUID、generation fence、
幂等事件和无副本地址泄漏的协议边界。
## 7. 关闭规则
每个 B/V 项只能通过以下方式关闭:
1. 记录被测 commit、镜像 digest、配置摘要和环境拓扑;
2. 保存可复现命令、原始输出和失败注入证据;
3. 由报告明确给出 pass/fail,不能只依赖日志中“看起来正常”;
4. 任一生产候选输入变化后,重跑受影响的验证。
在 B-01 至 B-03 全部实现,且 V-01 至 V-09 均有当前生产候选版本的通过证据前,
Cloud v2 状态保持 `NOT APPROVED FOR SAAS ACTIVATION`
@@ -0,0 +1,292 @@
# Multi-tenant implementation checklist
This checklist turns the Workspace architecture into implementation and
verification gates. Exact commands and observed results are recorded in the
[verification report](./verification-report.md).
## Scope guard
- [x] LangBot uses branch feat/multi-tenants.
- [x] langbot-plugin-sdk uses branch feat/multi-tenants.
- [x] langbot-space implements the greenfield Cloud v2 modular-monolith control plane without extending the legacy per-account Pod topology; the old Pod UI remains available when Cloud v2 is disabled and only retained Pods appear in the v2 view.
- [x] Unrelated untracked files in either repository remain untouched.
- [x] Open-source startup cannot enable SaaS multi-workspace through edition flags or unsigned configuration.
## SaaS activation gates
These items intentionally remain incomplete. Some require additional Core
transaction/cutover primitives and others require the closed Control Plane or
deployment. The feature branch delivers the Core isolation kernel, not the
closed SaaS product or a production Cloud v2 deployment. Checked implementation
items later in this document do not supersede these gates.
- [x] The closed Control Plane owns the global Account, Workspace, Membership, and Invitation directory.
- [ ] The closed Control Plane execution-ownership module issues monotonic generations and owner leases for projected Workspaces.
- [x] Core verifies a signed `InstanceManifest` before the closed bootstrap can inject `CloudWorkspacePolicy`.
- [ ] Tenant database writes hold a generation-aware shared transaction fence through commit, while execution-owner cutovers take the exclusive fence.
- [ ] Business writes and non-transactional side effects use a generation-stamped outbox or equivalent publish fence.
- [ ] Durable object references survive an execution-generation change through stable published keys or an explicitly atomic key/reference migration.
- [ ] The SaaS runtime pools enforce tenant-safe egress and SSRF controls for Webhooks, providers, MCP servers, and every tenant-configurable outbound URL.
- [ ] Entitlement checks, usage aggregation, and subscription lifecycle are implemented in the closed Control Plane; production activation still requires provider callback amount/currency/session/expiry binding inside the locked fulfillment transaction.
- [ ] Account registration persists a `new_api.provision_account` outbox item with the Account and personal Workspace, and an in-process reconciler provisions New API idempotently after commit.
- [ ] EPay and Stripe callbacks bind provider identity, amount, currency, channel/session, payment status, and expiry to the locked order before entitlement fulfillment.
- [ ] OAuth state and directory projection use an atomic shared store suitable for horizontally scaled SaaS services.
- [ ] A greenfield Cloud v2 deployment is designed and validated independently of the legacy Space deployment scheme.
- [ ] The Plugin Runtime shared profile refuses to run without delegated cgroup v2 CPU, memory-plus-swap, and PID limits, all verified in a real Linux container; production tenant-safe egress remains incomplete.
- [x] The Plugin Runtime Supervisor automatically restores an unexpectedly exited enabled worker with bounded per-installation backoff.
- [ ] Jitter, a global restart concurrency limit, and a Runtime-level circuit breaker prevent a systemic failure from creating a cross-tenant restart storm.
- [ ] Until authenticated Runtime takeover or an owner lease/fence exists, the M0 deployment rolls Core and Plugin Runtime together and forbids an independent Core-only rollout.
- [ ] Plugin installation data has an operator-owned hard disk quota provider that atomically rejects writes over the limit; directory scans are not accepted as enforcement.
- [ ] The Box deployment provides an operator-owned quota provider that proves hard byte and inode limits for Workspace, Skill, root, tmp, and home storage.
- [ ] Core and Box Runtime mount the same durable volume and pass the authenticated marker challenge during startup and reconnect.
- [ ] Production provisions distinct migrator/runtime credentials and runs the implemented same-host/port/database release command as a one-shot Job, with tested orchestration retry, backup, and rollback procedures.
- [ ] Production PostgreSQL uses a dedicated cluster/endpoint, or a tested HBA/proxy policy proves the cluster-wide runtime credential can connect only to the target business database.
- [ ] Any future direct-migrator/pooler-runtime endpoint split is admitted only by a migrator-owned, runtime-read-only database cluster identity that the runtime role cannot spoof.
- [ ] Legacy pgvector migration failure and retry integration paths prove exact source-table RLS/FORCE restoration; the non-superuser, non-`BYPASSRLS` success path is already covered below.
- [ ] Multi-workspace is enabled in SaaS only after all closed Control Plane, deployment, and security gates pass.
## 1. Persistence foundation
### Account and directory
- [x] User has a stable, unique account UUID and explicit status.
- [x] Existing email and password behavior remains compatible during migration.
- [x] Workspace table represents the instance-local tenant.
- [x] WorkspaceMembership has a unique Workspace and Account pair.
- [x] WorkspaceInvitation stores only a token hash and supports expiry, revoke, and one-time accept.
- [x] WorkspaceExecutionState stores generation, state, source, and write fence.
- [x] OSS initialization creates exactly one Workspace and one owner membership atomically.
- [x] OSS refuses a second Workspace while allowing multiple members.
### Migration
- [x] Alembic migration upgrades SQLite.
- [x] Alembic migration upgrades PostgreSQL.
- [x] Existing first user becomes owner of the default Workspace.
- [x] Existing tenant resources are backfilled with the default Workspace UUID.
- [x] SQLite destructive boundaries create verified, revision-aware backups and atomically restore after failure.
- [x] Migration can resume safely after interruption.
- [x] New installs and upgraded installs produce the same tenancy-kernel schema.
- [x] The first Cloud release pins migrator and runtime sessions to `public` with `current_schemas(false)` containing only that business schema; runtime-role/database `search_path` overrides are rejected.
- [x] Cloud sessions require `session_replication_role=origin`, `row_security=on`, and `lo_compat_privileges=off`; every persistent `pg_db_role_setting` applicable to the runtime role or current business database is rejected.
- [x] The release migrator grants the runtime role exact business-table DML, `alembic_version` read-only access, and business-sequence `USAGE/SELECT`, with no `WITH GRANT OPTION` or non-business object grants.
- [x] Every Cloud runtime startup revalidates the login role, current user, schema, effective/direct ACLs, ownership, memberships in all directions, column ACLs, routines, extensions, foreign objects, parameter ACLs, and other-schema access before serving traffic.
- [x] The business database requires `vector`, permits only `plpgsql`/`vector` extensions, forbids runtime extension ownership, and contains no foreign data wrapper, foreign server, or user mapping.
- [x] The runtime role and `PUBLIC` have no explicit routine or parameter ACL; the runtime owns no routine and cannot effectively execute any `SECURITY DEFINER` routine, including extension-owned routines.
- [x] PostgreSQL's default `PUBLIC TEMP` is documented and tested as a dedicated-business-database v1 compatibility exception; the migrator never grants `TEMP` directly to the runtime role.
- [x] Legacy pgvector migration succeeds as a non-superuser, non-`BYPASSRLS` source-table owner and restores mixed source-table RLS/FORCE states exactly.
- [ ] Legacy pgvector migration still needs explicit failure-and-retry integration coverage before SaaS activation.
### Runtime transaction enforcement
- [x] Each tenant UoW owns one task, root transaction, database bind, and transaction-local scope.
- [x] A scoped Session and every captured bound method become permanently unusable when the owning UoW exits.
- [x] Public transaction/session control, raw/textual SQL, connection/bind escape, nested transactions, execution/loader options, foreign binds, live results, unapproved functions/operators/casts/types, `INSERT FROM SELECT`, hidden `ON CONFLICT` and batch-value expressions, forced-unquoted identifiers, and custom AST/compiler nodes fail closed and make the UoW rollback-only.
- [x] ORM `SessionEvents` fail before a registered callback can receive the synchronous Session or transaction connection; rollback cleanup cannot execute the rejected listener.
- [x] ORM flush, implicit autoflush, and commit reject SQL expressions assigned to mapped attributes before compilation.
- [x] Tenant relationship loading uses eager loading or explicit async `refresh`; synchronous object-session access and `AsyncAttrs.awaitable_attrs` are not supported tenant APIs.
- [x] The UoW guard is documented as a trusted-Core misuse boundary rather than an in-process Python sandbox; mapped metadata/compiler registration is trusted, plugins remain out of process, and SQLAlchemy upgrades must rerun the private-container regression suite.
## 2. Authentication and authorization
### Identity
- [x] JWT sub uses account UUID, with a bounded compatibility path for legacy email tokens.
- [x] Disabled or deleted accounts cannot authenticate.
- [x] Local password and Space-linked account flows support more than one local Account.
- [x] Public registration closes after initialization by default.
- [x] Invitation registration works without requiring SMTP.
- [x] An unknown Space OAuth subject cannot claim an existing Account by email; explicit account-bound binding is required.
### Request context
- [x] PrincipalContext identifies Account, API Key, or trusted runtime principal.
- [x] WorkspaceContext contains Workspace, Membership, role, permissions, and revision.
- [x] RequestContext contains instance UUID, Workspace context, auth type, request ID, and generation.
- [x] ExecutionContext propagates Workspace and generation to runtime work.
- [x] SaaS-style requests never fall back to the first or most recent Workspace.
- [x] OSS may resolve the single Workspace when the selector is omitted.
- [x] Account-token bootstrap can list only the authenticated Account's active memberships before a Workspace selector exists.
### Fixed RBAC
- [x] owner, admin, developer, operator, and viewer permissions match the architecture matrix.
- [x] Invitation cannot grant owner.
- [x] The last owner cannot be removed or demoted.
- [x] Cross-Workspace resources return 404.
- [x] Same-Workspace permission failures return 403.
## 3. Workspace and member APIs
- [x] GET /api/v1/workspaces returns the OSS singleton Workspace.
- [x] POST /api/v1/workspaces returns edition_limit in OSS.
- [x] Current Workspace endpoint returns the authenticated Membership.
- [x] Member list is permission scoped.
- [x] Invitation create, revoke, inspect, and accept are atomic.
- [x] Member role update and removal enforce owner rules.
- [x] Invitation tokens travel in a request body and are redacted from logs.
- [x] Relevant MCP tools and in-repo skills are updated with the same contract.
## 4. Tenant-scoped persistence and services
Each row type must have a non-null Workspace UUID, scoped indexes, scoped uniqueness, and scoped CRUD tests.
- [x] Bots and bot admins.
- [x] Legacy pipelines and pipeline run records.
- [x] Model providers.
- [x] LLM models.
- [x] Embedding models.
- [x] Rerank models.
- [x] Plugin installations, settings, and configuration.
- [x] MCP servers and resource preferences.
- [x] Knowledge bases, files, and chunks.
- [x] Vector collections and handles.
- [x] Monitoring messages, calls, sessions, errors, embeddings, and feedback.
- [x] API keys and scopes.
- [x] Webhooks and public route resolution.
- [x] Binary storage and Workspace storage.
- [x] Workspace metadata, separated from system metadata.
### Service and API rules
- [x] Every tenant Service receives RequestContext or an explicit Workspace UUID.
- [x] No tenant Service treats context None as global access.
- [x] Every applicable get, list, create, update, delete, copy, export, and bulk operation is scoped.
- [x] Parent-child references use the same Workspace.
- [x] API Key authentication derives Workspace from the key, not a header.
- [x] Webhook and Bot public routes derive Workspace from a trusted resource.
- [x] Background jobs carry Workspace and generation explicitly.
## 5. Runtime isolation
### Core runtime
- [x] RuntimeBot carries Workspace UUID and execution generation (currently stored in the compatibility field `placement_generation`).
- [x] RuntimePipeline carries Workspace UUID and execution generation (currently stored in the compatibility field `placement_generation`).
- [x] Query and Event carry Workspace UUID without making it an authorization source.
- [x] Session key includes Workspace UUID, Bot UUID, launcher type, and launcher ID.
- [x] QueryPool and manager indexes cannot collide across Workspaces.
- [x] Query and aggregation cache keys and locks include Workspace UUID.
- [x] Runtime transports, cached results, object operations, and long-lived tasks revalidate WorkspaceExecutionState generation at side-effect boundaries.
- [ ] Ordinary tenant database writes hold the generation fence in the same transaction until commit; this remains a SaaS activation gate.
### Plugin
- [x] Plugin installation and configuration are Workspace scoped.
- [x] Runtime control actions carry trusted Workspace binding and execution generation (wire-compatible as `placement_generation`).
- [x] The Plugin Runtime supervisor is instance-scoped and intentionally serves multiple Workspaces.
- [x] Every plugin process is bound to exactly one Workspace, installation, generation, revision, and verified artifact digest.
- [x] Same-digest plugin code may be cached once, while worker processes and writable data remain isolated.
- [x] Same-digest plugin dependencies are prepared once in a Runtime-owned immutable environment and mounted read-only into each isolated worker; dependency failure is surfaced before launch and recorded per installation without blocking other desired-state recovery.
- [x] Host API derives Workspace from the connection, installation, and trusted action context, not plugin input.
- [x] Plugin get_bots, models, tools, vector, RAG, configuration, and messaging calls are scoped.
- [x] Plugin Workspace storage no longer uses owner default.
- [x] Plugin page APIs check Membership and installation ownership.
- [x] Local plugin launches use short-lived, one-use registration capabilities bound to manifest identity.
### MCP, RAG, and Box
- [x] MCP runtime key contains instance UUID, Workspace UUID, execution generation, and server UUID.
- [x] Same-named MCP servers in two Workspaces do not share sessions.
- [x] Pipeline cannot reference another Workspace's MCP resource.
- [x] RAG collection names and handles are server-derived and Workspace scoped.
- [x] Legacy global vector migration is available only to the local OSS singleton Workspace.
- [x] Object storage paths include instance, Workspace, and execution generation for the fixed-generation OSS runtime.
- [x] Object storage revalidates generation before touching a provider or resolving an opaque key.
- [ ] Cloud cutover uses generation-scoped staging plus stable published object references, rather than making the staging generation the durable identity.
- [x] Box persistent and ephemeral namespaces include the required instance, Workspace, and generation scope.
- [x] Same-named Box sessions and processes cannot collide across Workspaces or execution generations.
- [x] Box relay and process I/O reject or retire stale generations.
- [x] External paths and privileged mounts cannot be supplied by an untrusted plugin.
- [x] Cloud attachment host I/O uses query UUIDs and link-free dirfd operations with bounded inode traversal.
- [x] Cloud Skill package paths are Runtime-owned, Workspace-scoped, read-only mounts; Python env/cache stays tenant-writable.
- [x] Skill ZIP preview/install rejects path escape, links, non-regular files, duplicate entries, excessive compression ratio, entry count, per-file size, and total size.
- [x] Cloud Box code paths and automated tests require the authenticated marker challenge before startup or reconnect can proceed.
- [x] Cloud Box readiness fails until hard Workspace, Skill, ephemeral-storage, and inode quota capabilities are available.
## 6. SDK and protocol
- [x] Public Query, Event, Session, and context entities carry backward-compatible Workspace data.
- [x] Action RPC request models carry trusted Workspace binding where required.
- [x] Action enums and callers remain consistent.
- [x] Old plugins continue to deserialize compatible events.
- [x] Plugins cannot select an arbitrary Workspace through a Host API argument.
- [x] Runtime storage uses the bound Workspace UUID.
- [x] SDK API tests pass.
- [x] Runtime tests pass.
- [x] Action consistency script passes.
## 7. Frontend
- [x] Every browser tenant API request carries the current Workspace selector after bootstrap.
- [x] OSS automatically selects the singleton Workspace.
- [x] OSS does not show Create Workspace or a misleading switcher.
- [x] Workspace settings show current Workspace information.
- [x] Members page lists roles and permissions.
- [x] Invitation creation shows a one-time link when SMTP is unavailable.
- [x] Invitation acceptance supports a signed-out user flow.
- [x] Role controls are hidden or disabled consistently with backend permissions.
- [x] Switching accounts clears stale Workspace query cache and local state.
- [x] User-facing strings support en_US, zh_Hans, and ja_JP.
## 8. Automated verification
### Persistence and authorization
- [x] SQLite fresh install.
- [x] SQLite upgrade from pre-tenant schema, including verified failure recovery.
- [x] PostgreSQL fresh install.
- [x] PostgreSQL upgrade from pre-tenant schema.
- [x] All fixed roles have positive and negative permission-matrix tests.
- [x] Concurrent invitation acceptance creates one Membership.
- [x] Concurrent owner changes never leave zero owners.
### Cross-tenant isolation
- [x] Two Workspaces are created through a test-only policy.
- [x] Applicable resource operations and parent-child references have cross-Workspace negative coverage.
- [x] Resource UUID guessing cannot cross Workspace.
- [x] API Key cannot cross Workspace.
- [x] Plugin cannot enumerate or invoke another Workspace's resources.
- [x] Sessions, caches, locks, MCP, RAG, Box, storage, and monitoring do not collide.
- [x] Background jobs cannot execute without an explicit Workspace and execution generation.
### Security and revocation
- [x] Space login and binding use purpose-bound, one-time opaque OAuth state; caller-supplied state is rejected.
- [x] OAuth redirects trust only server-configured WebUI or webhook origins, never request `Host` or `Origin` headers.
- [x] Dashboard WebSockets revalidate authentication, Membership, resource, permission, and generation per message.
- [x] Public embed WebSockets re-resolve Bot availability and execution binding per message.
- [x] Runtime, storage, Plugin Runtime, MCP, RAG, and Box reject a stale execution generation.
- [x] Unhandled API and webhook failures return a generic error plus request ID without exception text.
- [x] URL user information and sensitive query parameters are redacted before configuration is serialized or logged.
### Regression
- [x] LangBot unit tests pass.
- [x] LangBot integration tests pass.
- [x] Frontend lint completes without errors and the production build passes.
- [x] SDK focused and full relevant tests pass.
- [x] LangBot is pinned to the exact pushed SDK commit and cross-repo tests pass against that revision.
## 9. Real browser E2E
- [x] Start from a clean local data directory.
- [x] First user initializes the singleton Workspace as owner.
- [x] Owner creates an invitation link.
- [x] A second signed-out browser identity accepts the invitation and registers.
- [x] owner, admin, developer, operator, and viewer UI permissions match backend enforcement.
- [x] Direct API calls cannot bypass hidden controls.
- [x] Account switch does not expose prior account or Workspace data.
- [x] Refresh and a new browser tab recover the correct Workspace safely.
- [x] OSS rejects a second Workspace with `edition_limit`; same-name and same-identifier isolation is covered by the test-only multi-Workspace policy because OSS deliberately has no multi-Workspace browser surface.
- [x] Explicit error states are visible for expired, revoked, reused, and email-mismatched invitations.
## 10. Completion evidence
- [x] LangBot and SDK branch refs are recorded in the verification report.
- [x] Space contains the closed adapter package and Cloud v2 control plane, billing, migration, and Workspace UI changes; unrelated pre-existing files remain unstaged.
- [x] Migration output is captured for SQLite and PostgreSQL.
- [x] Test commands and results are recorded.
- [x] Browser E2E actions and observed results are recorded.
- [x] No remaining tenant table, global Service query, owner default, or unscoped runtime key is found by the final audit.
@@ -0,0 +1,305 @@
# Multi-tenant implementation decisions
This log records implementation choices made while delivering the Workspace architecture. It is intended to make trade-offs auditable without interrupting implementation for routine decisions.
> Architecture decisions, activation gates, and still-open follow-ups are tracked in
> [pending-architecture-decisions.md](./pending-architecture-decisions.md). Sections marked as decided there are authoritative;
> this file records the concrete implementation choices and compatibility names used to realize them.
## 2026-07-18
### OSS remains a singleton Workspace with multiple Accounts
- Decision: Community builds create exactly one Workspace per LangBot instance and allow multiple Accounts through invitations.
- Reason: This preserves a simple self-hosted deployment while making authorization and ownership explicit. Creating a second Workspace is an edition error, not a hidden fallback.
- SaaS boundary: Multi-Workspace directory, execution ownership, entitlement, and billing are the responsibility of a separate closed SaaS Control Plane. Core consumes a validated projection and remains the final isolation and authorization enforcement point; it does not become the SaaS system of record or billing engine.
- Deployment boundary: Cloud v2 is a greenfield deployment design. The previous per-account instance/pod scheme is not migrated or extended and remains available only for existing subscriptions. Existing OAuth, marketplace, and payment rails are reused through explicit adapters where they still fit; new Workspace, subscription, entitlement, directory, and usage modules live alongside the legacy Pod flow in `langbot-space`.
### Workspace selection is trusted only after authentication
- Decision: Browser requests carry `X-Workspace-Id`, but the server resolves it against the authenticated Account membership. API keys, public Bot routes, webhooks, jobs, and plugin calls derive Workspace from their trusted owning resource or binding instead of trusting the header.
- Reason: A selector is routing input, not authorization evidence.
- Compatibility: Community builds may select the singleton Workspace when the header is omitted. A multi-Workspace-capable build must reject an omitted selector.
### Stable Account UUID is the token subject
- Decision: New JWTs use the stable Account UUID as `sub`; a bounded compatibility path accepts legacy email-subject tokens and rotates them when checked.
- Reason: Email can change and therefore cannot be a durable authorization identity.
### Fixed roles are authoritative in Core
- Decision: `owner`, `admin`, `developer`, `operator`, and `viewer` map to a fixed permission matrix in LangBot Core. The last owner cannot be removed or demoted, and invitations cannot create an owner directly.
- Reason: Core must remain the final authorization boundary in both OSS and SaaS deployments.
### Cross-Workspace access is indistinguishable from absence
- Decision: Resource lookups always include Workspace UUID. A guessed UUID belonging to another Workspace returns 404; a visible resource with insufficient same-Workspace permission returns 403.
- Reason: This avoids leaking resource existence across tenants while preserving actionable same-tenant errors.
### Plugin Runtime is shared; every plugin process is single-Workspace
- Decision: One instance-scoped Plugin Runtime control plane serves all Workspaces in the logical LangBot instance. Each running plugin installation has its own nsjail worker with an immutable binding containing `instance_uuid`, `workspace_uuid`, `execution_generation` (stored as the compatibility field `placement_generation` until the schema rename), `installation_uuid`, `runtime_revision`, and verified artifact digest; enabled-resident is the desired semantic. A worker never routes actions for another Workspace or installation, and plugin-supplied scope fields are stripped.
- Isolation: Plugin code is mounted read-only. Home, tmp, and data paths are installation-scoped; process, file-descriptor, file-size, CPU, memory, and PID limits come only from `data/config.yaml` (including native environment overrides), never from a plugin manifest. Cloud requires nsjail and delegated cgroup v2 hard limits or fails closed.
- Cost boundary: Identical verified package bytes share one digest-addressed code cache. A dependency environment is keyed by the artifact and requirements digests, Python ABI, Runtime version, and installer schema, then atomically published read-only for reuse. Installations and processes are not merged, even for the same plugin and version. Registration creates database desired state only; a worker is launched only for an enabled installation.
- Recovery: PostgreSQL installation desired state and durable binary storage are authoritative. Runtime reconnect performs an instance-wide full reconciliation, removes stale workers, and can replay a verified package after Runtime-local cache loss. Dependency preparation failure is recorded per installation with `dependency_prepare_failed`; it prevents that worker launch without blocking recovery of other desired installations, and the same revision can be retried. The installation Supervisor now restores an unexpectedly exited enabled worker through a completion callback with bounded exponential backoff. Jitter, global restart concurrency limits, and a Runtime-wide circuit breaker are still required to prove that an infrastructure-wide failure cannot create a cross-tenant restart storm.
- Compatibility: Older SDK payloads and legacy `data/plugins` remain an OSS-only bridge. Shared mode requires complete bindings and rejects incomplete context.
- Reason: Sharing the supervisor and immutable code cache removes per-Workspace service cost without turning an untrusted plugin process into a cross-tenant router.
### Invitation delivery does not require SMTP
- Decision: Core returns an invitation secret once for copy-and-share, persists only its hash, and supports expiry, revocation, and one-time acceptance.
- Reason: Self-hosted OSS must support adding users without an email service while avoiding recoverable invitation secrets at rest.
- Browser handling: The copyable invitation URL carries the secret in its fragment, which browsers do not send in HTTP requests or Referer headers. The acceptance page immediately removes the fragment and keeps the secret only in `sessionStorage` until login or acceptance completes; it is never placed in a path, query string, analytics event, or persistent local storage.
### Schema rollout is additive before enforcement
- Decision: Add Account/Workspace directory tables first, then add non-null Workspace ownership to every tenant resource with a deterministic default-Workspace backfill. Runtime and service enforcement is enabled only with matching migration and isolation tests.
- Reason: A Workspace column alone is not isolation, and enforcing queries before data backfill would break upgraded installations.
### Login capability discovery is instance-scoped, not account-scoped
- Decision: The unauthenticated login bootstrap endpoint reports only which login mechanisms the instance supports. It does not inspect the first Account or expose whether that Account has a password. Both password and Space OAuth entry points are available on a multi-user instance; the submitted identity determines which mechanism is valid.
- Reason: This avoids projecting the original owner's authentication type onto invited users and removes a public Account-state disclosure.
### Space OAuth identity does not choose a SaaS Workspace
- Decision: Space OAuth tokens remain Account credentials. In OSS singleton mode, an OAuth refresh may update the singleton Workspace's Space provider only when that Account's role can manage provider secrets. In SaaS multi-Workspace mode an OAuth callback without an authenticated Workspace selector never guesses which Workspace to mutate; explicit Workspace configuration or the closed control plane owns that linkage.
- Reason: An Account may belong to several Workspaces, and authentication must not silently mutate a shared tenant secret.
### SaaS execution state is a validated Core projection
- Decision: Core can resolve both local and `cloud_projection` Workspaces, but only from an explicit Workspace UUID and an active, unfenced `WorkspaceExecutionState` for the current instance and matching source. OSS-only bootstrap paths additionally require `source=local`.
- Reason: The closed control plane owns execution ownership and generation decisions, while Core remains the enforcement point for instance binding, generation, and write fences.
### API-key secrets are one-time and Workspace-bound
- Decision: Database API keys persist only a globally unique SHA-256 hash, an opaque UUID, one Workspace UUID, explicit fixed-permission scopes, status, expiry, creator, and last-used time. The raw secret is returned once. Authentication derives Workspace and generation from the key record and ignores Workspace selectors. Legacy plaintext keys are hashed during migration and receive a compatibility `*` scope. The plaintext config key works only for the OSS singleton Workspace and is disabled in multi-Workspace mode.
- Reason: A bearer key is an identity and routing credential, not merely a password layered on top of caller-controlled tenant selection.
### MCP tools inherit the authenticated API-key context
- Decision: The MCP ASGI mount authenticates the API key once, binds an immutable per-request `RequestContext`, and every tool checks a fixed permission before calling tenant services with that same context.
- Reason: Authenticating the transport without propagating Workspace identity into tool calls would leave the direct service path globally scoped.
### Unreleased SDK protocol is pinned reproducibly without publishing
- Decision: The SDK tenancy protocol is versioned as 0.4.18. This task does not create a GitHub release or publish PyPI because the user authorized pushing code, not a package release. After the SDK feature branch is final, LangBot's feature branch temporarily pins the exact pushed SDK Git commit. Before merging to master, the release gate is to publish `langbot-plugin==0.4.18` and replace the Git pin with the registry pin.
- Reason: The current registry release does not contain the complete tenant action context and shared Runtime hardening. An exact Git commit is reproducible and keeps the feature branch testable without expanding release authority.
### Cloud directory writes stay outside Core
- Decision: The open-source Core startup always installs `SingleWorkspacePolicy`, creates or repairs one local Workspace, and permits local membership/invitation workflows. Changing mutable configuration such as `system.edition` cannot activate multi-Workspace routing. The future closed Cloud bootstrap will install `CloudWorkspacePolicy` only after verifying a signed `InstanceManifest`; that policy requires an explicit projected Workspace selector, does not create Workspaces, and rejects invitation or membership mutations with `control_plane_required`; member reads use the versioned local projection.
- Ownership split: The closed Control Plane owns the global Account/Workspace/Membership/Invitation directory, execution ownership and generation, entitlements, subscription state, usage aggregation, and billing decisions. Core owns request authorization, resource scoping, execution-generation validation, and fail-closed enforcement. Provisioning and invoice computation do not belong in open-source Core.
- Reason: The closed control plane is authoritative for SaaS Account, Workspace, Membership, and Invitation state. Allowing Core to mutate the same directory would create split-brain ownership and would make an ownerless compatibility Workspace a dangerous fallback.
- Release gate: Multi-Workspace activation is deliberately unavailable in the open-source bootstrap. Production Cloud v2 must implement the signed `InstanceManifest` verifier and closed bootstrap described in the architecture document before it can inject `CloudWorkspacePolicy`; `edition=cloud`, an environment variable, or any unsigned local configuration is never a valid activation credential.
### Workspace bootstrap is reactive and ordered before browser resource calls
- Decision: The web application blocks Workspace-owned pages until Account and current Workspace bootstrap completes. A `useSyncExternalStore` Workspace store publishes permission changes to React consumers; direct mutation-only routes and controls are hidden or disabled when the fixed role lacks the required permission.
- Reason: Mutating a module-level variable after the initial React render did not reliably re-render permission controls, and mounting resource pages before the selector was established could issue tenant requests without `X-Workspace-Id`.
### JWTs are bound to one LangBot instance
- Decision: New Core JWTs require `iss=langbot-core`, an audience derived from the immutable instance UUID, and an expiry. Legacy community tokens are accepted only when they have the historical issuer, carry no audience, and the active policy is the OSS singleton policy.
- Reason: A token issued by one instance must not authenticate against another instance that happens to share a secret, and a compatibility decoder must not become an alternate path around the SaaS trust boundary.
### Runtime control transports authenticate before protocol dispatch
- Decision: External Plugin Runtime and Box WebSocket control channels require independent strong shared secrets in handshake headers. Locally managed child processes receive ephemeral secrets through their environment; secrets are not placed in URLs, process arguments, request payloads, or logs. Box additionally binds the first authenticated control channel to one trusted instance. Plugin Runtime debug and control credentials remain separate.
- Reason: Workspace context inside an RPC payload is not trustworthy until the transport peer itself is authenticated. Separating control and debug credentials also limits accidental privilege reuse.
- Deployment consequence: Docker Compose and Kubernetes wire one shared secret to each host/runtime pair. An empty external-runtime secret fails startup instead of silently exposing an unauthenticated socket.
### Dashboard WebSocket sessions are tenant runtime objects
- Decision: A dashboard WebSocket sends an authentication frame immediately after upgrade. The server validates Account, Membership, permission, Pipeline ownership, instance, Workspace, and execution generation before registering the connection. Connection indexes, sessions, broadcasts, attachments, and resets include the complete execution scope.
- Reason: Browser WebSocket APIs cannot attach the normal authorization headers, and a process-global `pipeline_uuid` or `session_type` index can collide across Workspaces.
### Read permissions never imply secret permissions
- Decision: `resource.view` responses recursively redact Bot, Plugin, MCP, and provider credentials. Provider secrets require `provider_secret.manage`; Bot and Plugin configuration writes require `resource.manage`. Masked Plugin values can be round-tripped by a manager without overwriting the stored secret. Plugin Runtime debug credentials require `resource.manage`, not the operator-only `runtime.operate` permission.
- Reason: A multi-user Workspace needs useful viewer access without turning every visible configuration endpoint into credential export. Plugin debug attachment can register executable code and is therefore a resource-management operation.
### Temporary credential exchanges are bound to their initiator
- Decision: Lark, Weixin, DingTalk, WeComBot, and QQOfficial one-click registration sessions require `resource.manage` and store the initiating instance, Workspace, execution generation, and principal. Status and cancellation by any other scope return the same 404 as an unknown session.
- Reason: Random session IDs reduce guessing probability but do not authorize access to credentials returned by a completed exchange.
### Uploaded images and documents use different storage capabilities
- Decision: Browser images use the scoped `upload_image` owner type and may be resolved only through the opaque public-image route. RAG documents use `upload_document` and can be read, sized, or deleted only by an exact instance, Workspace, generation, and owner-type match. Legacy `upload` objects are cleanup-only.
- Reason: Treating every upload as a public image made a leaked document key sufficient to bypass authenticated RAG access.
## 2026-07-19
### Space OAuth state is server-issued and single-use
- Decision: Core issues an opaque, cryptographically random OAuth state for each Space login or Account-binding attempt, stores only its digest, and consumes it exactly once within a short expiry. Login and binding states are different capabilities; a binding state is additionally bound to the authenticated Account. Caller-supplied state, including a LangBot JWT, is rejected.
- Redirect boundary: Callback redirects are accepted only for the known callback path and an origin declared by the server-side `api.webui_url` or `api.webhook_prefix`. Request `Host` and `Origin` headers never expand this allowlist.
- Current deployment: The OSS state store is bounded and process-local, so a Core restart safely invalidates outstanding attempts. A horizontally scaled SaaS deployment must move this exchange to an atomic, shared Control Plane store before enabling the closed Cloud bootstrap.
- Reason: OAuth state is a narrow, one-time CSRF and flow-binding capability. Reusing a bearer JWT or trusting caller-controlled Host or Origin data would turn an authorization redirect into an Account-token theft or open-redirect primitive.
### OAuth provider subjects, not email addresses, bind Accounts
- Decision: A known Space `account_uuid` may refresh the credentials of its already-bound local Account. An unknown provider subject that presents an email belonging to an existing Account is rejected, even when the normalized emails match. The Account owner must authenticate locally and use the one-time, account-bound binding flow.
- Reason: Email is contact and display data, not a stable federated identity key. Email-only auto-linking would let provider verification drift or identity reassignment become a local Account takeover.
### Workspace discovery is an account-only bootstrap capability
- Decision: `ACCOUNT_TOKEN` validates the active Account JWT but intentionally cannot resolve a Workspace, receive `RequestContext`, or declare Workspace permissions. Its narrow bootstrap endpoint returns only active Workspace memberships belonging to that Account and never chooses the first Workspace when several exist. All tenant resource routes still require the explicit selector in multi-Workspace mode.
- Reason: Requiring a Workspace header to discover the Account's Workspaces creates an authentication deadlock; allowing the bootstrap route to perform tenant actions would create an authorization bypass. Separating the two capabilities resolves the cycle without weakening tenant routes.
### SQLite tenancy migrations have a verified recovery boundary
- Decision: Before each destructive tenant-schema boundary, a file-backed SQLite installation creates an online-consistent backup with its source and target revisions, runs `PRAGMA quick_check`, writes a durable manifest, and fsyncs restrictive-permission files and directories. A failed boundary disposes the engine, removes stale journal sidecars, atomically restores the verified source revision, and verifies the restored database before startup continues.
- Compatibility: In-memory SQLite cannot provide this recovery guarantee and is rejected for destructive production migration boundaries; it remains usable in tests that create the final schema directly.
- Reason: SQLite batch table rebuilds can leave an installation between schemas if a process or migration fails. A verified pre-boundary image makes retry behavior recoverable instead of merely idempotent in the happy path.
### Execution generation is an execution revocation capability
- Decision: RuntimeBot, RuntimePipeline, background tasks, object storage, Plugin Runtime, MCP, RAG, and Box operations carry the complete instance, Workspace, and execution-generation scope. The current schema and wire compatibility field remains `placement_generation` until a coordinated rename. They revalidate the active execution binding before accessing a provider or transport; long-running calls validate again before accepting results. A stale generation is fenced before it can read, write, or reuse a cached object.
- Plugin boundary: Each locally launched plugin receives a short-lived, one-use registration capability bound to the expected manifest identity and execution scope. The production child environment does not inherit the reusable debug credential, and Host APIs derive scope from the trusted connection and action context.
- Box boundary: Persistent skill content remains Workspace-scoped, while session/process state and relay requests also include execution generation. A generation change retires matching live sessions and closes a stale relay before further stdin, stdout, or file operations.
- Transaction boundary: Request admission and runtime side effects are fenced in this branch, but ordinary tenant database mutations do not yet hold a generation-aware lock through commit. The closed Cloud bootstrap must remain disabled until Core provides the shared-write/exclusive-cutover transaction primitive and a generation-stamped outbox (or an equivalent atomic publish fence). The OSS singleton policy has a fixed local generation and cannot trigger an execution-owner cutover.
- Durable-object boundary: Current opaque storage keys include generation and therefore fail closed after a generation change. That is safe for OSS's fixed generation, but a Cloud cutover must not strand durable KB files, images, or plugin references. Cloud v2 must publish stable final object identities from generation-scoped staging, or perform an atomic object-and-reference migration before activating the new generation.
- Reason: Workspace UUID prevents cross-tenant collisions, but it cannot revoke work after execution ownership changes or is fenced. Execution generation is the monotonic revocation value that makes old runtimes unusable; it does not express membership in a product-level deployment entity.
### Long-lived WebSockets continuously revalidate authority
- Decision: Dashboard WebSockets re-authenticate the Account, Membership, permission, resource ownership, instance, Workspace, and execution generation for every inbound message, not only during the initial frame. A changed role, removed Membership, or fenced execution binding takes effect without waiting for reconnect.
- Public embed boundary: The embed connection re-resolves its Bot before every message and rejects a Bot that was disabled, deleted, moved, or rebound. The public connection may identify a Bot, but it cannot make the initial Bot object an indefinite authorization capability.
- Reason: Authorization and resource state can change while a socket remains open. Connection-time validation alone leaves a revocation gap.
### Legacy vector migration is an OSS-local compatibility path
- Decision: Status, backup, execute, dismiss, and background entry points for legacy global vector collections require an active local Workspace binding under `SingleWorkspacePolicy`. A `cloud_projection` Workspace cannot observe or migrate the old global collection, even when it carries a legacy marker.
- Reason: The legacy collection predates tenant ownership. Treating it as a SaaS fallback would expose one installation's historical vectors to an arbitrary projected Workspace.
### External errors and persisted URLs are redacted centrally
- Decision: Unhandled HTTP and webhook failures return a stable `internal_error` response and request ID, and expose that ID in `X-Request-Id`; the detailed exception is retained only in server logs correlated by the same ID. Explicit domain and validation errors keep their documented status and code.
- Secret boundary: Shared sanitization removes URL user information and masks sensitive query parameters before provider or MCP configuration is serialized, logged, or shown to a reader. Masked placeholders can be round-tripped by an authorized manager without replacing the stored secret.
- Reason: Tenant isolation is incomplete if framework exceptions, connection URLs, or configuration reads can export credentials across otherwise authorized interfaces.
### Box is one shared control plane with one admitted sandbox per Workspace
- Decision: The logical instance has one shared Box Runtime control plane, implemented by one Runtime replica in M0. A closed entitlement adapter projects generic `managed_sandbox` capability and `managed_sandbox_sessions` limit; Core and Runtime never branch on a plan name. An eligible Workspace receives at most one persistent logical `global` session, while each ordinary command remains a one-shot nsjail process. Managed processes and network are disabled in the first Cloud release.
- Storage boundary: Core and Runtime prove they see the same durable volume with an authenticated random-marker challenge. Attachments use opaque query UUID directories and link-free dirfd operations. Skill packages remain in the Runtime-owned Workspace store and enter a sandbox only as a read-only logical-name mount; Python environments and caches live in the tenant's writable Workspace.
- Resource boundary: Cloud readiness requires cgroup v2 plus hard byte and inode limits for Workspace files, Skill storage, root, tmp, and home. The existing directory scan is only a compatibility soft check. Plain nsjail reports these storage capabilities as unavailable, so Cloud Box intentionally cannot start until the greenfield deployment supplies and verifies a real quota provider.
- Archive boundary: Skill ZIP processing is bounded by compressed input, entry count, per-entry size, total uncompressed size, and compression ratio, and rejects links, non-regular entries, duplicates, and path escape before streaming extraction.
- Reason: A shared supervisor removes per-Workspace services and idle control-plane cost, but storage and process admission must still fail closed at the untrusted execution boundary.
### Cloud business data and vectors share one PostgreSQL schema
- Decision: SaaS uses one PostgreSQL business database and shared schema. Every tenant row has an explicit Workspace key; application scope is the first boundary and precise `ENABLE` plus `FORCE ROW LEVEL SECURITY` policies are the second. The Cloud runtime role must be non-owner and have neither superuser nor `BYPASSRLS`.
- Vector boundary: pgvector is the Cloud default in the same business database. Vectors use `(workspace_uuid, knowledge_base_uuid, vector_id)` identity, an untyped vector column with explicit checked dimension, and release-created partial expression indexes for the enabled dimensions. Cloud never falls back to Chroma or performs vector DDL at runtime.
- Transaction boundary: A tenant UoW binds `SET LOCAL` and SQL to one transaction. Long-running pipeline and streaming MCP execution carry a trusted transaction-free tenant scope; each database helper opens a short scoped transaction, avoiding a held pool connection during LLM or network waits. Detached tasks start only after commit and create their own short UoW; rollback cancels them.
- Schema boundary: The first release has exactly one business schema, `public`. Both migrator and runtime sessions must report `current_schema() = 'public'` and `current_schemas(false) = ARRAY['public']`; the runtime role and business database must not carry a `search_path` override. Runtime startup validates this before using the prepared schema and reruns the complete catalog and privilege validation on every process start; it never runs DDL.
- Session boundary: Both Cloud modes require `session_replication_role = 'origin'`, `row_security = 'on'`, and `lo_compat_privileges = 'off'`. Every persistent setting applicable to the runtime role or current business database in `pg_db_role_setting` is rejected, even if its present value appears safe; tenant context remains transaction-local application state rather than a persistent role/database override.
- Grant boundary: The migrator grants the runtime role direct `CONNECT` on the dedicated business database and `USAGE` on `public`; exact `SELECT, INSERT, UPDATE, DELETE` on every allowlisted business table; `SELECT` only on `alembic_version`; and exact `USAGE, SELECT` on business-owned sequences. It grants neither `CREATE`, `TRUNCATE`, `REFERENCES`, `TRIGGER`, sequence `UPDATE`, nor any privilege with `WITH GRANT OPTION`, and grants nothing on other relations or schemas.
- Role boundary: The runtime identity is a `LOGIN` role with no superuser, `BYPASSRLS`, `CREATEDB`, `CREATEROLE`, or replication attribute; no role membership in any direction, including acting as grantor; no ownership of the business database, `public` schema, relations, sequences, routines, or extensions; no column ACLs; and no use, create, or ownership in another non-system schema. Neither the runtime role nor `PUBLIC` may have an explicit routine or parameter ACL, and the runtime role may not effectively execute any `SECURITY DEFINER` routine, including an extension-owned one. PostgreSQL's default `TEMP` privilege inherited from `PUBLIC` is an explicit first-release compatibility decision for this dedicated business database, not a direct runtime-role grant.
- Catalog boundary: The business database must contain `vector` and may contain no extension other than `plpgsql` and `vector`; the runtime role owns neither. It contains no foreign data wrapper, foreign server, or user mapping. These checks remove catalog-level escape paths without forbidding the ordinary implicit execution of non-`SECURITY DEFINER` built-in routines.
- Migration boundary: In the first release, the migrator and runtime URLs must name the same normalized PostgreSQL host, port, and database while using different roles. The migrator owns the application schema, establishes the exact allowlist above, and validates both required access and every prohibited escalation path before releasing the advisory lock. An exact Alembic head, RLS checks, and pgvector table/index/constraint validation remain mandatory; concurrent Jobs fail explicitly and are retried by orchestration.
- Deployment boundary: PostgreSQL roles are cluster-wide, while the in-database audit proves only the target business database contract. SaaS production must therefore use a dedicated PostgreSQL cluster or endpoint that exposes only this business database to the runtime credential, or enforce and test an HBA/proxy policy proving that the credential cannot connect to any other database. This external connectivity proof is still an incomplete SaaS activation gate.
- Endpoint evolution: A future deployment may use a direct endpoint for migrations and a pooler endpoint for runtime traffic. That topology may relax literal host/port equality only after both endpoints are proven to reach the same database through a database-internal, migrator-owned cluster identity that the runtime role can read but cannot create, alter, or spoof.
- Legacy pgvector boundary: Revision 0013 records the exact `ENABLE` and `FORCE ROW LEVEL SECURITY` state of each RLS-protected source table, temporarily suspends those source policies as their table owner inside the migration transaction, and restores every table to its recorded state in `finally`. The migrator does not require superuser or `BYPASSRLS` for this data move.
- Activation gate: The shared schema, pgvector adapter, and database-local runtime audit are implemented, but the external cluster/endpoint or HBA/proxy connectivity proof remains deployment work. Ordinary business writes also do not yet hold a generation-aware fence through commit; a generation-stamped outbox (or equivalent atomic publish fence) and stable durable-object references across generation cutover remain required before SaaS activation.
- Reason: Sharing one database and pool keeps marginal Workspace cost low, while transaction-local context and RLS prevent that shared storage from becoming shared authority.
### stdio MCP has an independent deployment gate
- Decision: `mcp.stdio.enabled` is independent of Box availability and entitlement. OSS defaults it on for compatibility; Cloud requires it off at bootstrap and enforces the same gate on create, update, test, startup loading, and final runtime execution.
- Reason: Treating Box availability as stdio permission would silently create another persistent `mcp-shared` sandbox for each Workspace and bypass the one-sandbox subscription and cost boundary.
## 2026-07-20
### The tenant UoW owns its task, root transaction, bind, and scope
- Decision: A tenant UoW creates one task-owned `TenantScopedAsyncSession` and one root transaction. Public commit, rollback, close, connection, bind, nested-transaction, synchronous-Session, live-streaming, raw SQL, public execution options, and public `set_config` paths fail closed and mark the transaction rollback-only. ORM objects cannot expose a usable synchronous Session, captured methods cannot run in child tasks, an explicit foreign bind is rejected, and a captured Session is permanently retired when its UoW exits rather than being reset for reuse. Tenant scope is installed only through a private UoW capability; pgvector index-plan `SET LOCAL`/`EXPLAIN` diagnostics use a test/operator connection rather than the business Session API.
- SQL boundary: Public UoW calls accept only structured SQLAlchemy query and DML trees. `TextClause`, literal SQL columns, textual labels, prefixes/suffixes/hints, statement execution options, `VALUES` roots, `INSERT FROM SELECT`, `EXTRACT`, literal-execute parameters, unknown/custom AST nodes, forced-unquoted identifiers, named `ON CONFLICT` constraints, unknown dialect post-values clauses, and untrusted casts/types fail closed. PostgreSQL/SQLite `ON CONFLICT DO UPDATE` and batch-insert containers are traversed explicitly because SQLAlchemy's standard visitor omits their executable values. Function classes are exactly allowlisted as `count`, `coalesce`, `sum`, `now`, `length`, and `nullif`; the only custom operator/cast admitted is the validated pgvector cosine operator and `Vector` cast.
- Legacy migration boundary: The local-only RAG backup restore uses explicit table and column objects, never raw SQL, but deliberately leaves legacy values untyped. This preserves SQLite's string-valued `DATETIME` rows and both the historical PostgreSQL `TEXT` and fresh-schema `JSON` settings columns while keeping every value bound rather than interpolated.
- ORM boundary: SQLAlchemy `SessionEvents` are unsupported on a tenant-scoped Session. If a listener is registered before or during a UoW, the operation fails before the callback executes and cleanup proceeds against an empty dispatch surface. Public `get`, `get_one`, `refresh`, and `merge` reject caller-supplied loader, bind, lock, shard, and execution options. Flush, implicit autoflush, and commit reject a SQL expression assigned to a mapped attribute before it can reach the compiler. Tenant code uses the async Session directly; relationships use eager loading or explicit `await session.refresh(entity, [attribute])`. LangBot's persistence base does not expose `AsyncAttrs.awaitable_attrs` as a supported tenant API.
- Compiler trust boundary: This guard prevents accidental scope/transaction escape by trusted LangBot Core code; it is not an in-process Python sandbox. Registered SQLAlchemy compilers and mapped schema metadata are trusted boot-time code. The fail-closed traversal of dialect containers necessarily covers SQLAlchemy private fields, so dependency upgrades require the regression suite and remain pinned until verified. Untrusted plugins cannot import or call this Session because they remain isolated in Plugin Runtime child processes.
- Result boundary: A caught database or boundary failure rolls back the root transaction and cancels after-commit work. Buffered results contain only already-authorized rows and no live connection; live database results cannot escape the UoW operation.
- Reason: `SET LOCAL` plus RLS protects a tenant only while every statement stays on the same owned connection and callers cannot end the transaction, replace the GUC, recover the synchronous proxy, or route a statement through another bind.
### Parallel request work re-enters tenant scope explicitly
- Decision: Child coroutines created by request-level `asyncio.gather` open their own explicit transaction-free tenant scope before calling persistence-backed Plugin, MCP, or Skill operations. They never inherit the parent's active database Session.
- Reason: Python copies ContextVars into child tasks, but SQLAlchemy Sessions are not task-safe and task identity is part of the tenant UoW boundary.
### Ordinary monitoring is readable; audit and export remain privileged
- Decision: Workspace monitoring dashboards, Bot logs, sessions, messages, calls, errors, and feedback require `resource.view`. Monitoring export requires `data.export`; system/runtime audit logs keep `audit.view`. Frontend tabs and controls use the same split.
- Reason: A Viewer needs useful read-only product observability, while bulk data extraction and privileged runtime/system logs are separate capabilities.
### Invitation failures survive login without contradictory success state
- Decision: Invitation terminal codes map to stable browser states. A login-mediated email mismatch preserves the fragment-captured secret only in session storage, returns to the acceptance page with the stable error code, and suppresses the generic login-success toast. A transient acceptance failure retains the authenticated session and offers the same one-time invitation for retry. Invalid Bearer tokens on acceptance map to the normal authentication error instead of an internal failure.
- Reason: Invitation acceptance crosses signed-out and authenticated states; losing or masking the domain error makes recovery ambiguous and can present contradictory UI feedback.
### Shared Plugin Runtime starts only from verified desired state
- Decision: SDK shared mode waits for immutable runtime configuration before inspecting plugin state, never scans or launches legacy `data/plugins`, and rejects legacy install/restart/delete/upgrade control actions. Worker RPC files use installation-private directories with aggregate size enforcement, and resident nsjail workers explicitly disable the default 600-second wall-time limit.
- Remaining gate: completion-callback recovery, jittered per-installation backoff, globally bounded restart launch admission, Runtime-level circuit breaking, one half-open probe, and worker ready timeout are implemented. Production cross-tenant fault injection must still prove the restart-storm controls. Hard installation disk quota and production egress policy remain Cloud activation requirements; Linux nsjail/cgroup CPU, memory-plus-swap, PID, namespace, and cgroup-reaping behavior have real-container evidence.
- Reason: A shared supervisor reduces per-Workspace services only if legacy global paths, writable transfer state, and lifecycle defaults cannot bypass installation isolation.
## 2026-07-24
### Cloud v2 remains a modular monolith with one closed adapter
- Decision: Workspace directory, plans, subscriptions, payment fulfillment, entitlements, usage, and signed runtime feeds are implemented as modules in the existing Space service and PostgreSQL database. The only separately packaged closed component is the thin Core bootstrap/control-plane adapter; it verifies signed data and owns no business state.
- Compatibility: The legacy Pod card and fulfillment path remain visible only to accounts with old subscriptions. New purchases create Workspace subscriptions and never provision a per-user LangBot Pod.
- Reason: A separate tenancy or billing service would add deployment, queue, network, and consistency cost without providing an isolation boundary. The signed adapter keeps SaaS logic closed while Core retains the ORM, RLS, authorization, and runtime enforcement boundary.
### Registration creates data, not infrastructure
- Decision: Account registration and personal Workspace creation share one transaction and include an active owner membership, Free subscription, entitlement snapshot, and outbox notifications. Repair is idempotent for older accounts. Empty Workspace creation starts no Plugin worker, Box sandbox, database, queue, bucket, or tenant service.
- Activation: This automatic Cloud v2 transaction is enabled only when Space has an explicit valid `CLOUD_V2_INSTANCE_UUID`. A legacy marketplace-only Space deployment with no Cloud v2 instance configured keeps its existing Account registration path; Cloud v2 internal endpoints still fail closed instead of inventing an instance identity.
- Billing boundary: Core and runtimes consume only generic capability and numeric-limit entitlements. They never branch on `free` or `pro`; plan names, prices, payment providers, and fulfillment stay in Space.
- Reason: This keeps the marginal cost of a new user close to a few PostgreSQL rows while preserving an immediate, usable Workspace.
### Directory bootstrap is full; steady-state projection is per Workspace
- Decision: Space generates the initial signed directory snapshot and its outbox high-water mark in one read-only PostgreSQL `REPEATABLE READ` transaction. After bootstrap, each event page names affected Workspaces and Core fetches one signed `directory.delta` for that set. Missing requested Workspaces are tombstones; unrelated Workspaces are untouched.
- Cursor boundary: A delta carries no event cursor. Each signed event page carries the transaction-consistent current high-water mark, while Core advances only to the final event it actually consumed, even when the authoritative delta already contains a later revision. Event payload Workspace and revision fields must exactly match the signed event envelope; readiness is renewed only after the replica-local cursor reaches the signed high-water and the shared projection is not ahead.
- Replica boundary: Every Core replica owns a process-local consumer cursor because its verified entitlement cache is process-local. Replicas share the PostgreSQL projection high-water mark and inbox. The state separately records which cursor range was atomically subsumed by a full snapshot, so a lagging replica can add missing receipts within that coverage and refresh its local caches without repeating tenant mutations; a missing receipt beyond snapshot coverage fails closed.
- Reason: A shared consumer cursor would let one replica starve another replica's local cache, while a full snapshot on every change would make steady-state cost grow with every registered Workspace.
### Cloud OAuth can authenticate only an existing projected Account
- Decision: In multi-Workspace Cloud mode, Space OAuth may refresh tokens only for an active `cloud_projection` Account whose Space subject UUID and normalized email exactly match. It cannot create an Account, relink by email, mutate directory identity, or choose a Workspace.
- Redirect boundary: When Cloud v2 configures its Core public URL, Space issues authorization codes only to that exact callback origin or the explicitly retained legacy managed-Pod domain. Community installations remain dynamic OAuth clients because they cannot pre-register with the public Space service: the consent screen displays their hostname, and only HTTPS or loopback HTTP with the fixed callback path is accepted. Remote HTTP, userinfo, fragments, arbitrary callback paths, and unrecognized query parameters always fail closed.
- Reason: Space is the SaaS identity and directory authority, but Core remains the authentication enforcement point. Projected-only matching avoids split-brain identities; redirect allowlisting prevents bearer authorization-code exfiltration.
## 2026-07-29
### Directory capacity is one instance-level admission contract
- Decision: Space and Core share an explicitly configured operational ceiling for active
Workspace and snapshot membership cardinality. Space serializes new personal-Workspace
creation across replicas with a PostgreSQL transaction advisory lock and rejects the
registration transaction before the limit is crossed. Core independently counts the
projected active Workspace set while holding the per-instance directory projection row
lock. Exceeding any limit rolls back the complete projection and does not advance its
cursor; neither side truncates authoritative data.
- Snapshot boundary: A full snapshot is current desired state and contains active
Workspaces only. Archived Workspace revisions are retained as bounded, targeted deltas
so Core can validate and apply monotonic tombstones without every bootstrap carrying
unbounded history.
- Memory boundary: Space bounds database result cardinality before signing. The closed
adapter bounds decompressed response bytes before JSON/JWS parsing and validates
Workspace/membership cardinality before entitlement-cache fan-out. Core schema models
have absolute list ceilings, validate duplicates in one pass, and bulk-read Accounts in
bounded chunks rather than issuing two serial queries per Account. Manifest and
entitlement responses have smaller endpoint-specific byte ceilings; entitlement refresh
validates and releases batches of at most 16 raw responses instead of retaining the
entire directory fan-out.
- Operations boundary: Core health exposes aggregate active/max directory cardinality and
PostgreSQL pool occupancy/timeouts. The default active Workspace ceiling is 1,000, with
a hard code ceiling of 5,000; this is a safety stop, not a production capacity claim.
Space and Core configuration must match, and the approved production value comes from
the real V-08 capacity curve plus the V-09 24-hour soak.
- Reason: One LangBot instance should admit new tenants at the cheapest data-only boundary,
while still preventing legitimate registration growth or a malformed control-plane
response from causing an unbounded startup allocation, database connection storm, or
CPU spike.
@@ -0,0 +1,525 @@
# Cloud v2 多租户架构决策与待决策项
状态:`DECIDED — Core isolation kernel implemented; SaaS activation gates remain`
创建日期:2026-07-19
最近更新:2026-07-24
本文记录 Cloud v2 多租户架构中已经确认的首期决策、明确淘汰的方案和仍需在后续阶段决定的扩展项。
本文同时记录实现状态。“实现完成”仅指开源 Core/SDK 的隔离内核和 fail-closed 门禁,
不表示闭源 Control Plane、计费或 Cloud v2 部署已经可上线。最终实现选择同步记录在
[implementation-decisions.md](./implementation-decisions.md),剩余发布门禁记录在实施清单和验证报告。
## 0. 已确认的 SaaS 拓扑前提
1. SaaS 只有一个逻辑 LangBot 实例,全部 Workspace 都是该实例内的租户。
2. 产品和领域模型中不引入 Cell 内多个 CloudInstance、Workspace Placement 或 Workspace 到 CloudInstance 的路由。
3. 当前不实现分布式,但同一个逻辑实例未来可以运行多个 Core、Plugin Runtime 和 Box Runtime replica
也可以增加 PostgreSQL shard;这些只是内部实现,不成为新的租户或产品实体。
4. 所有副本共享稳定的 `instance_uuid``replica_id``worker_id` 和进程地址是短期运行身份,不能写进业务资源的永久主键。
5. `workspace_uuid` 始终是数据、任务和运行时的租户键,也是未来内部路由与分片的候选键。
6. generation/epoch 的语义是执行所有权、故障转移和任务撤销,不代表 Workspace 在多个 CloudInstance 之间 Placement。
7. 注册 Account 时自动创建 Workspace,但只新增目录记录和业务行,不创建租户专属部署、数据库、队列或 Runtime。
8. OSS 仍是单租户 LangBot 实例,但允许该 Workspace 内存在多个用户;只有 SaaS 开启多 Workspace 租户模式。
“单个 LangBot 实例”表示单个逻辑服务和安全域,不等于永远只有一个 OS 进程或一个 Kubernetes replica。
当前代码字段 `placement_generation` 在完成架构迁移前继续兼容,目标语义和候选命名是 `execution_generation`
### 0.1 本轮确认的首期决策
| 编号 | 结论 | 首期状态 |
| ----- | --------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------- |
| D-001 | 一个共享 Plugin Runtime 控制面;每个运行中的 plugin installation 独占一个 nsjail 子进程;只有 digest 相同且已验证的代码 artifact 可以只读共享 | `IMPLEMENTED — egress/disk-quota pending; restart-storm fault injection pending` |
| D-002 | 一个共享 Box RuntimeCloud 固定使用 nsjail;符合套餐的 Workspace 最多一个持久 `global` 逻辑 sandbox,普通执行按需启动 nsjail 进程 | `IMPLEMENTED FAIL-CLOSED — hard filesystem quota provider pending` |
| D-003 | SaaS 业务数据使用 PostgreSQL shared schema、应用层作用域和 RLS 双重隔离;pgvector 使用同一 PostgreSQL,作为 SaaS 默认向量后端 | `PARTIALLY IMPLEMENTED — transaction/outbox/deployment gates remain` |
| D-004 | stdio MCP 与 Box availability 解耦;Cloud v2 首期强制关闭 stdio MCP,避免为每个 Workspace 创建额外的 `mcp-shared` persistent sandbox | `IMPLEMENTED` |
| D-005 | 目录启动使用事务一致的全量快照,运行时按事件涉及的 Workspace 拉取增量;每个 Core replica 独立消费事件,共享 PostgreSQL 投影和 inbox | `IMPLEMENTED — production fault injection pending` |
Workspace 的具体创建、释放、数据导出和单 Workspace 恢复机制不在本轮决定;本文只保证这些后续能力不会改变稳定的
`workspace_uuid`,也不会要求重建租户专属部署。
## 1. 本轮重构的最高目标
> 共享可信控制面和基础设施池,隔离不可信执行单元;减少独立部署、扩缩容和运维组件,使新增 Account 或 Workspace 的静态成本接近零。
这里的“减少组件”指减少独立 Deployment、Service、数据库、消息系统和租户专属常驻控制面,
不是通过合并安全边界来减少必要的隔离进程。
统一评估原则:
1. 注册 Account、自动创建空 Workspace 时,不启动 Plugin worker 或 Box sandbox。
2. 启用插件后,每个 installation 的常驻成本来自其独立安全边界;首次使用托管 sandbox 后,符合套餐的 Workspace 才承担一个持久逻辑 session 的成本。
3. 可信 supervisor、artifact cache、数据库连接池和 Runtime 容量可以多租户共享。
4. 一个不可信插件进程不能服务多个 installation;一个 sandbox/session 不能服务多个 Workspace。
5. 默认使用共享 Runtime;dedicated 只作为未来高隔离、大客户或合规资源等级,不建立第二套外部协议。
6. 没有明确容量证据前,不新增 Kafka、Redis、Runtime 专用数据库、Box 专用数据库或租户级调度服务。
7. 多租户隔离必须覆盖身份、路由、存储、缓存、日志、配额、撤销和故障恢复,不能只给请求增加 `workspace_uuid`
## 2. 首期部署形态与未来演进
```mermaid
flowchart LR
Traffic["SaaS traffic"] --> Core["One logical LangBot instance<br/>1 Core replica in MVP"]
Core --> PluginRuntime["Shared Plugin Runtime<br/>trusted supervisor"]
Core --> BoxRuntime["Shared Box Runtime<br/>nsjail backend"]
PluginRuntime --> PA["Workspace A / installation 1<br/>isolated nsjail process"]
PluginRuntime --> PB["Workspace B / installation 2<br/>isolated nsjail process"]
BoxRuntime --> BA["Workspace A<br/>one persistent global logical session"]
BoxRuntime --> BB["Workspace B<br/>one persistent global logical session"]
Core --> PG["Shared PostgreSQL business schema<br/>RLS + pgvector"]
```
| 档位 | 内部部署形态 | 新 Workspace 静态成本 | 启用条件 |
| -------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------- | ------------------------ |
| M0. 单副本 MVP | 一个 Core、一个共享 Plugin Runtime、一个共享 Box Runtime、一个 PostgreSQL business database;插件按启用状态运行,托管 sandbox 按首次使用与 entitlement 创建 | 只新增 Workspace 业务行 | 当前已确认目标 |
| M1. 同逻辑实例内部横向扩展 | Core、Plugin Runtime、Box Runtime 按容量增加 replica;运行所有权由内部 lease 和 generation fence 决定;PostgreSQL 可增加共享 shard | 不创建 Workspace 专属部署 | 出现容量或可用性证据后 |
| M2. Dedicated 资源档位 | 特定 workload 使用独享 worker pool、sandbox class 或 PostgreSQL shard,但沿用相同身份、协议、schema 和控制面 | 仅由购买 dedicated 的客户承担 | 合规、数据驻留或超大负载 |
M1 是 M0 的透明扩容,M2 是相同架构下的资源等级;两者都不是新的 LangBot 实例、Cell 或 CloudInstance。
外部 API 只认识稳定的 `instance_uuid``workspace_uuid`,不认识 replica、worker、pool 或 shard。
Plugin Runtime 与 Core 在 M0 使用独立容器和 security contextCore 不能继承 Plugin Runtime 所需的
nsjail/cgroup 权限。当前 Runtime 在进程生命周期内绑定首次认证的 `runtime_id`,因此 M0 必须把 Core 与
Plugin Runtime 放在同一 rollout/restart unit 中协调重启;在实现受认证 takeover 或 owner lease/fencing 前,
不能单独滚动 Core 并让它接管仍存活的 Runtime。Box Runtime 同样使用独立进程身份和安全配置,
不与 Plugin Runtime 合并成一个高权限进程。
## 3. D-001Plugin Runtime 多租户控制面
状态:`IMPLEMENTED — Cloud egress/disk-quota pending; restart-storm fault injection pending`
### 3.1 已实现的基础
- Plugin Runtime 控制连接只绑定稳定实例身份;一个逻辑共享控制面通过完整 installation binding 管理多个 WorkspaceM0 由一个 Supervisor replica 承担。
- 每个运行中的 installation 使用独立 nsjail workerenabled-resident 是 desired semantics。代码只读,home/tmp/data 私有,shared profile 不读取 artifact `.env`
- 实例级 `PluginWorkerPolicy` 由 Core 的 `data/config.yaml` 下发,支持原生环境变量覆写;manifest 不能覆盖。
- `installation_uuid``artifact_digest``runtime_revision` 已持久化并进入 desired-state、注册、Host API 和 generation/revision fence。
- 已验证 `.lbpkg` 先进入 Workspace-scoped durable binary storageRuntime 本地缓存丢失后可由 Core replay。
- 相同 digest 的代码和 Runtime 准备的只读依赖环境可以共享,但 worker、运行时写入、配置和数据不合并;
dependency preparation 在启动 worker 前完成,失败会进入明确的 installation failed 状态。
- Cloud shared profile 强制 Linux nsjail`plugin.worker.require_hard_limits=true` 时 cgroup v2 delegation 不可用会启动失败。
### 3.2 已确认的不变量
1. 每个运行中的 plugin installation 独占一个 worker process tree,任何时刻都不能与其他 installation 共用;
停用或删除的 installation 可以没有进程。
2. 插件进程只绑定一个
`(instance_uuid, workspace_uuid, execution_generation, installation_uuid, runtime_revision, artifact_digest)`,且运行期间不可重绑。
3. 插件不能通过 payload、Host API 参数、环境变量或重连选择 Workspace。
4. 插件进程的 home、tmp、可写数据、secret、进程视图和配额必须按 installation 隔离。
5. 只有 `artifact_digest` 相同且完整性已验证的代码文件和依赖环境可以只读共享;
同名同版本但 digest 不同的 artifact 不能共享。配置、持久数据和运行进程不能共享。
6. generation、installation revision 或 capability 被撤销后,旧进程必须失去 Host API 和副作用权限。
7. Supervisor 不在自身解释器中加载第三方插件代码。
### 3.3 首期执行模型
- 整个 SaaS 实例共享一个可信 Plugin Runtime 逻辑控制面,M0 运行一个 Supervisor replica;新 Workspace 不创建专属 Runtime、连接、卷或进程。
- Supervisor 的控制连接只绑定稳定 `instance_uuid` 和短期 Runtime identity,不绑定某个 Workspace。
每条 installation desired-state 命令都携带并验证完整的 installation binding;每个 worker action context 在注册后永久绑定该 tuple。
- 安装并启用插件后,Supervisor 在自己的 Runtime 容器内直接启动一个 nsjail 子进程;
不再为每个插件创建 nested container、Pod、sidecar 或租户级 Runtime service。
- desired semantics 要求 enabled installation 保持 resident,不做 idle eviction;停用、删除、revision/generation 变化或 entitlement 撤销时停止并按需重建。
Supervisor 通过 completion callback、带 jitter 的有界指数 backoff 恢复意外退出的 worker。所有 restart launch 共用实例级并发槽;
在配置的失败窗口达到阈值后打开 Runtime 级 circuit breaker,冷却后只允许一个 half-open probe。
probe 必须完成初始化并持续稳定一个窗口后才能恢复其他 installation;未在 30 秒内 ready 的子进程会被取消回收。
生产候选环境仍需执行跨租户系统性故障注入,证明熔断、恢复和告警符合预期。
- 子进程使用一次性 registration capability 向 Supervisor 注册;capability 由可信 desired state 派生并绑定完整 installation tuple
不是插件直接建立 Core Host connection,也不能只绑定 author/name/path。Supervisor/Core 据此注入 tenant context
丢弃插件 payload 中自带的 scope 字段。
- Supervisor 的进程表、nsjail root/tmp 和 artifact cache 都是可重建运行态;PostgreSQL 中的 installation desired state 才是权威业务状态。
- M0 不增加 Runtime 专用数据库、Redis、Kafka、scheduler 或 artifact serviceCore 重连后向 Supervisor replay desired state。
### 3.4 nsjail 和文件边界
首期目标目录模型:
```text
data/plugin-runtime/
├── artifacts/sha256/<artifact_digest>/code/ # digest 校验后只读共享
├── environments/sha256/<environment_digest>/ # 原子发布、只读共享依赖环境
└── installations/<installation_uuid>/
├── home/ # 私有可写
├── tmp/ # 私有可写、可清理
└── data/ # 私有持久数据
```
- artifact 只有在内容摘要和完整性校验一致时才允许共享,不能只凭 author/name/version 复用目录;
cache 可接受的签名/来源、撤销和 GC 规范属于后续发布规则,不改变本轮基于已验证 digest 的只读共享边界。
- artifact 与按环境摘要构建的共享依赖环境以只读 mount 进入 nsjailinstallation 的 home/tmp/data 使用独立可写 mount。
环境摘要包含 artifact、requirements、Python ABI、Runtime 版本和 installer schema。依赖只能从已验证 artifact 的 PEP 508 声明构建,
index/trusted-host 只由实例配置控制;构建在独立 nsjail 的临时路径中完成并在成功后原子发布,失败或并发安装不能留下可见半成品。
- 插件 cwd 可以是其私有 mount namespace 内的只读 `/plugin`,不要求为每个 installation 复制代码;
必须私有的是 home/tmp/data 等所有可写路径。
- nsjail 必须启用 mount、PID、IPC、UTS 和 private `/proc` 等必要 namespace,插件不能枚举或 signal 其他插件及 Runtime 进程,
不能读取 Runtime 文件系统、宿主机路径、其他 installation 目录或平台 metadata endpoint。
- 公开 SaaS 禁止从插件 artifact 自动加载 `.env`。secret 只能由可信控制面按 installation 注入,且不能进入共享 artifact/cache。
- 插件需要外网时使用受控 egress;不得通过共享 host network 访问 Core loopback、Box Runtime、数据库或其他内部服务。
- Cloud 部署必须提供可用的 cgroup v2 delegation 和所需 namespace 权限;如果硬 CPU/内存/PID 限制不可用,
Plugin Runtime readiness 必须失败,不能只记录告警后降级为普通子进程。
### 3.5 统一资源上限与配置
首期资源规格完全由 LangBot 实例配置决定,manifest 不能声明、放宽或覆盖资源。以下数值是建议默认值,
最终仍由同一实例的 `data/config.yaml` 统一配置:
```yaml
plugin:
worker:
max_cpus: 1.0
max_memory_mb: 512
max_pids: 128
max_open_files: 256
max_file_size_mb: 512
require_hard_limits: true # Cloud; OSS defaults false
```
配置文件路径为 `data/config.yaml`,沿用现有原生环境变量覆写:
- `PLUGIN__WORKER__MAX_CPUS`
- `PLUGIN__WORKER__MAX_MEMORY_MB`
- `PLUGIN__WORKER__MAX_PIDS`
- `PLUGIN__WORKER__MAX_OPEN_FILES`
- `PLUGIN__WORKER__MAX_FILE_SIZE_MB`
- `PLUGIN__WORKER__MAX_CONCURRENT_RESTARTS`
- `PLUGIN__WORKER__RESTART_FAILURE_THRESHOLD`
- `PLUGIN__WORKER__RESTART_FAILURE_WINDOW_SECONDS`
- `PLUGIN__WORKER__RESTART_CIRCUIT_OPEN_SECONDS`
- `PLUGIN__WORKER__REQUIRE_HARD_LIMITS`
Core 启动时校验配置并通过现有 `SET_RUNTIME_CONFIG` 下发不可变 `PluginWorkerPolicy`
Runtime 不读取另一份环境变量配置,避免两个配置源不一致。CPU、内存和 PID 使用 cgroup 硬限制,
open files/file size 使用 rlimit。Cloud deployment profile 固定使用 nsjail,不能通过插件 manifest 或 SaaS 环境变量降级为普通进程。
installation data 的总空间硬配额需要 filesystem project quota 或独立 quota volume,不能用目录扫描伪装成硬限制;
该字段在选定可原子拒绝写入的存储机制前不进入首期配置。
### 3.6 淘汰与暂缓方案
| 状态 | 方案 | 结论 |
| ---------- | -------------------------------------------------- | -------------------------------------------------------- |
| 淘汰 | 每 Workspace 一个 Plugin Runtime | 部署、连接和固定内存随 Workspace 线性增长 |
| 淘汰 | 一个插件进程服务多个 Workspace/installation | 全局状态、本地文件和依赖无法形成可信租户边界 |
| 淘汰 | 同 Workspace 多插件合并到一个 worker | 与“每 installation 独立进程”冲突,扩大故障和权限边界 |
| 淘汰 | manifest 自行声明 CPU、内存或更高限额 | 首期统一执行实例级最大值 |
| MVP 不引入 | Runtime 专用数据库、Redis、Kafka 或独立 scheduler | 当前无容量证据,会增加组件和运维面 |
| 后续演进 | 多 Supervisor replica、owner lease、dedicated pool | 保留接口,达到容量或可用性阈值后再决定具体存储与调度方式 |
架构扩展项包括:Core/Supervisor 是否共置、artifact/venv cache 的签名/来源/撤销/GC 规范、installation data hard-quota provider、
v1 connection 的兼容期限,以及进入多 replica 后的 lease TTL、fencing token 和 owner 转移顺序。
这些不改变“每个运行中的 installation 一个隔离进程”的首期边界。
### 3.7 验收条件
- 两个 Workspace 安装 digest 相同且已验证的 artifact 时,共享目录仍为只读,进程、配置、data、home、tmp、日志和 Host API 完全隔离;
同名同版本但 digest 不同的 artifact 绝不共享目录。
- 插件不能读取其他 installation 文件、枚举或 signal 其他进程,也不能修改共享代码/依赖目录。
- CPU、内存、PID、open files 和单文件上限在真实 nsjail/cgroup 环境中生效;超额只终止或拒绝对应 installation。
- 修改 manifest 不能改变任何资源上限。
- installation data 的总空间硬配额在写入边界原子拒绝超额,并证明目录扫描不是生产 enforcement。
- 旧 generation/revision 的回调、消息、副作用和存储访问全部失败关闭。
- Runtime 重启能从业务 desired state 恢复,不依赖本地进程表作为权威真相。
- 意外退出的 enabled worker 由 completion callback 触发带有界 backoff 的自动恢复;连续失败只影响对应 installation,不能形成跨租户重启风暴。
- requirements 中存在 Runtime 基础镜像未预装的包时,Supervisor 仍能先完成共享依赖环境准备再启动 worker;
安装失败不会留下持续重启的半启动进程,也不会影响同 digest 已就绪环境的其他 installation。
## 4. D-002Box 多租户控制面和套餐边界
状态:`IMPLEMENTED FAIL-CLOSED — production quota provider pending`
### 4.1 已实现的基础
- 共享 Box 控制连接可服务多个 Workspace;所有操作绑定 instance、Workspace 和 generationRuntime namespace 由可信 context 派生。
- 短期 `SandboxAdmissionGrant`、revision tombstone 和原子 session admission 强制每个合资格 Workspace 最多一个 `global` persistent sessionmanaged process 固定为零。
- Core 与 Runtime 使用认证 host-control challenge 校验同一个 durable volume,而不是比较路径字符串;不一致时启动和重连失败。
- Cloud skill 只传逻辑名称;Runtime 从 Workspace-scoped store 解析只读包路径,Python env/cache 写入租户自己的 `/workspace/.skill-envs`
- ZIP 安装限制压缩输入、条目、单项、总解压量和压缩比,采用流式解压并拒绝 link、非普通文件、重复项和路径逃逸。
- 附件 host path 使用 query UUID 和 dirfd/openat/O_NOFOLLOWCloud replica 启动不再全局清理其他请求目录,遍历和删除有 inode 预算。
- grant-enforced readiness 强制 cgroup、namespace、mount、共享卷、Workspace hard quota、Skill hard quota、ephemeral storage 和 inode quota 全部被证明。
普通 nsjail backend 对尚未实现的硬磁盘能力明确返回 false,因此当前 Cloud Box 会按设计拒绝启动,直到新部署提供真实 quota provider。
### 4.2 首期套餐与 entitlement 模型
- 闭源订阅管理/Control Plane 负责把套餐映射为版本化 entitlementCore 和 Box Runtime 不硬编码 `plan == pro`
- 首期复用 Cloud Control Plane(可结合现有 Space 的订阅模块)承载闭源套餐、计费和 entitlement 投影,
不再拆一个独立 billing/tenant microservice;开源 Core 只实现通用 capability 和数值限额。
- 首期套餐投影为:Pro 的 `managed_sandbox_sessions = 1`,其他套餐为 `0`。建议 capability 形态:
```json
{
"features": {
"managed_sandbox": true,
"external_sandbox": false,
"mcp_stdio": false
},
"limits": {
"managed_sandbox_sessions": 1
}
}
```
- `box.enabled` 只表示当前 LangBot 实例是否部署了 Box Runtime,不能替代 Workspace entitlement。
- 工具发现层根据 entitlement 隐藏/禁用托管 sandbox。Core 校验 Control Plane 的 entitlement 后,
通过受认证控制连接向 Box Runtime 下发短期 `SandboxAdmissionGrant`,绑定
`instance_uuid + workspace_uuid + execution_generation + entitlement_revision + expires_at + max_sessions + max_managed_processes`
Runtime 只验证和执行该内部 grant,不理解 Pro 等套餐名称,也不相信业务调用方提交的 plan、session ID 或 host path。
- entitlement 缺失、过期或无法验证时失败关闭。并发创建必须用原子 admission 保证同一 Workspace 永远不超过一个 managed session。
- entitlement 被撤销后停止 managed process 并关闭逻辑 sessionWorkspace 数据保留/删除策略随未来 Workspace 释放机制一并决定。
### 4.3 Cloud nsjail 执行模型
- 整个逻辑 SaaS 实例共享一个 Box Runtime 逻辑控制面,M0 运行一个 Runtime replica;不创建每 Workspace Box service、worker pool、PVC、bucket、scheduler、Redis 或 Box 数据库。
- Cloud 显式固定 `box.backend: nsjail`。sandbox 直接作为 Box Runtime 容器内的 nsjail 子进程运行,
不创建 nested Docker container、独立 Pod、microVM 或 warm pool,也不挂宿主机 `docker.sock`
- 符合 entitlement 的 Workspace 首次使用时懒创建一个逻辑 session,内部固定 ID 为 `global`,并强制 `persistent=True`
外部调用方不能选择或覆盖 session ID、persistence、host path 或 backend。
- “全局 sandbox 一直存活”在当前机制中的精确定义是:每个合资格 Workspace 最多一个稳定的 `global` 逻辑 session
它不被 TTL reaper 回收,其 `/workspace` 持久保存;普通命令仍按需启动并退出 nsjail 进程,不能承诺一个空闲 OS 进程永久驻留。
- Box Runtime 重启后,旧进程、attach token、root/tmp/home 和内存 session 状态失效;下一次使用时以相同 Workspace namespace 懒重建
`global` session。持久 `/workspace` 必须继续存在,旧 generation 权限必须失败关闭。
- 共享 Box Runtime 采用单 owner 的 M0 实现;未来多 replica 才引入 session owner lease 和跨 replica 路由,
但 session handle 永远不包含 replica 地址。
- Cloud 首期强制 `network=off`,调用方和 WebUI 不能覆盖。当前 `network=on` 会关闭 nsjail 的独立 network namespace
不能用于共享 SaaS。未来如需联网,必须先实现每 session 独立 netns 和受控 egress,再单独开放。
- Cloud 首期禁止 `START_MANAGED_PROCESS``SandboxAdmissionGrant.max_managed_processes` 固定为 `0`
普通 exec 在同一 Workspace 的 `global` session 内串行执行。未来开放 resident process 前必须增加数量和聚合 CPU/内存上限。
### 4.4 文件与资源边界
- 文件机制沿用当前 nsjail 方案:Box-owned durable volume 上的 Workspace 目录只 bind mount 到对应租户的 `/workspace`
不在首期新增对象存储双向同步服务或文件服务。
- `/workspace` 的持久性来自独立 durable host path,而不是 `persistent=True`;后者只禁止 TTL/普通 shutdown 回收逻辑 session。
Box Runtime 容器必须挂载持久卷,Workspace 数据不能只放在容器可写层。root/tmp/home 可以在 Runtime 重启时丢失。
- Cloud MVP 要求 Core 与 Box Runtime 以相同路径挂载同一持久卷并沿用直接文件读写;现有 exec/base64 fallback 的单文件上限
低于正常附件上限,不能当作等价 Cloud 文件机制。无法共享路径时必须先扩展传输协议,否则 deployment readiness 失败。
- 现有执行前后目录扫描只能提供软检查,不能阻止单次命令写满共享卷。生产 Cloud 的 Workspace 总空间上限必须由
Box-owned volume 的 filesystem project quota/subvolume quota 在写入点原子执行;Box Runtime 负责设置和验证,不能因 Core 看不到路径而跳过。
- Cloud 的 nsjail CPU、内存、PID、单文件和总空间限制使用运维配置统一设置;套餐只决定 session 数量,
不允许 Workspace 放宽 sandbox 上限。
- Box Runtime 必须在 cgroup v2 hard limit、namespace 和 mount 条件满足后才通过 readiness;不能在共享 SaaS 中告警后降级运行。
- 同 Workspace 的 `global` session 可以复用持久 `/workspace`,但不同 Workspace 即使使用相同文件名、进程名或逻辑 session ID,
物理 namespace、路径、进程和 capability 也必须完全隔离;Cloud 首期没有可暴露端口或共享网络 namespace。
### 4.5 非 Pro 和未来外部 E2B
- 非 Pro Workspace 在首期没有 Cloud managed sandbox,直接调用内部 API 也必须被拒绝。
- 未来允许 Workspace 在 WebUI 配置自己购买的远程 E2B endpoint/template/secret;它属于 tenant-owned external sandbox
不消耗 Cloud 的 `managed_sandbox_sessions` 配额,也不能读取其他 Workspace 的凭证。
- 当前 Box Runtime 只有实例级全局 backend 和 E2B credentialWebUI 也没有 Workspace 级配置,因此 BYOK E2B 明确不在首期实现。
- 未来实现时在共享 Box Runtime 内增加按可信 Workspace context 选择 backend/provider 的 registry
仍不创建租户专属 Box 控制面或新协议。
### 4.6 淘汰与暂缓方案
| 状态 | 方案 | 结论 |
| ---------- | ------------------------------------------------ | -------------------------------------------------------- |
| 淘汰 | 每 Workspace 一个 Box service | 组件和空闲成本随 Workspace 线性增长 |
| 淘汰 | 多 Workspace 共享一个活 sandbox/session | 不能承载不可信代码 |
| 淘汰为 MVP | Docker、独立 Pod、microVM 或 warm pool | Cloud v2 首期固定使用 Runtime 容器内 nsjail |
| 淘汰为 MVP | 非 Pro 使用 Cloud managed sandbox | 首期数值 entitlement 为 0 |
| 后续演进 | 多 Box Runtime replica、dedicated pool、BYOK E2B | 保留 provider/ownership 接口,有真实容量或产品需求后实现 |
### 4.7 验收条件
- Pro entitlement 首次使用时懒创建一个 persistent `global` session;重复和并发请求都不能产生第二个 session。
- 非 Pro、entitlement 缺失/过期及伪造 plan 的 API 直调全部失败关闭。
- TTL 不回收 persistent sessionRuntime 重启后进程和临时目录失效,但 `/workspace` 保留并能在下一次使用时安全重建。
- 两个 Workspace 的文件、进程、session、attach token 和 generation 完全隔离;network/managed-process 请求在首期失败关闭。
- Core 与 Box Runtime 通过随机 marker challenge 证明同一共享持久卷;只配置相同路径字符串不算通过。
- Workspace、Skill store、ephemeral root/tmp/home 的 byte quota 与 inode quota 在写入点真实生效;现有目录扫描不被当作硬配额。
- cgroup 或任一硬存储能力不可用时 Cloud Box Runtime readiness 失败;普通 nsjail 因此不会被误当成 production-ready provider。
## 5. D-003SaaS PostgreSQL 与 pgvector
状态:`PARTIALLY IMPLEMENTED — shared schema/pgvector complete; SaaS transaction and deployment gates remain`
当前分支已实现 PostgreSQL shared schema、transaction-local scope、FORCE RLS、Cloud runtime 非 DDL 模式、
同业务数据库 pgvector、显式向量维度和 tenant-scoped vector 主键。一次性 migrator 使用独立凭据、advisory lock
负责建立并校验 runtime role 的最小权限,并完成全量 schema 验证;
普通业务写入贯穿 commit 的 generation-aware fence、与外部副作用同事务的 outbox,以及 generation cutover 后稳定的 durable object 引用尚未实现。
这些 Core 事务原语与生产 Job、凭据发放、备份和回滚流程,以及 runtime credential 的跨 database 连接隔离证明一起,
都是 Cloud v2 的 SaaS activation gate。
### 5.1 已确认的数据库边界
- PostgreSQL 是 SaaS 的业务数据库,不把它扩展成通用 Runtime coordinator、Box session directory 或新控制面数据库。
- M0 使用一个 PostgreSQL business database/shared schema 承载全部 Workspace。创建 Workspace 不创建 database、schema、role 或专属连接池。
- 首版 migrator URL 和 runtime URL 必须归一化到同一 host、port 和 database,但必须使用不同 role。
未来如果 migrator 使用 direct endpoint、runtime 使用 pooler endpoint,只能在两个端点都验证同一个数据库内部 cluster identity 后放宽 host/port 相等。
该 identity 由 migrator 所有并固定,runtime role 只能读取,不能创建、修改或伪造。
- 首版唯一业务 schema 固定为 `public`。migrator 和 runtime 连接都必须满足
`current_schema() = 'public'``current_schemas(false) = ARRAY['public']`;禁止 runtime role 级和 business database 级 `search_path` 覆写。
- migrator 和 runtime session 的安全值固定为 `session_replication_role=origin``row_security=on`
`lo_compat_privileges=off`。runtime role 或当前 business database 作用域内只要存在任意 `pg_db_role_setting` 持久化设置就失败关闭,
即使该设置当前看似等于安全值也不接受;tenant context 只能通过事务内 `SET LOCAL` 建立。
- 业务行显式携带 `workspace_uuid`Repository/Service 的应用层 scope 是第一道边界,PostgreSQL RLS 是第二道边界。
- runtime role 的直接 ACL 固定为:business database 的 `CONNECT``public``USAGE`、全部 allowlisted business table 的
`SELECT/INSERT/UPDATE/DELETE``alembic_version` 的只读 `SELECT`,以及业务表自有 sequence 的 `USAGE/SELECT`
不授予 database/schema `CREATE`、table `TRUNCATE/REFERENCES/TRIGGER`、sequence `UPDATE`、其他对象权限或任何 `WITH GRANT OPTION`
- runtime role 必须是 `LOGIN`,但不得具有 superuser、`BYPASSRLS``CREATEDB``CREATEROLE` 或 replication 属性;
不得在 role membership 中以 granted role、member 或 grantor 任一方向出现;不得拥有 database、schema、table、view、sequence、routine 或 extension
不得持有 column ACL,也不得使用、创建或拥有其他非系统 schema。
- business database 必须安装 `vector`,且 extension catalog 只允许 `plpgsql``vector`;不得存在 FDW、foreign server 或 user mapping。
runtime role 和 `PUBLIC` 都不得有显式 routine ACL 或 parameter `SET/ALTER SYSTEM` ACLruntime role 不得有效执行任何
`SECURITY DEFINER` routine,包括被 allowlisted extension 收编的 routine。普通非 `SECURITY DEFINER` 内建函数的隐式执行权限不在此禁令内。
- PostgreSQL 新 database 默认向 `PUBLIC` 提供的 `TEMP` 是首版在专用业务 database 上明确接受的兼容性决定,
不是 migrator 对 runtime role 的直接 grant;首版不得据此把业务 database 与不受信任工作负载混用。
migrator 在释放 advisory lock 前建立并校验上述精确 allowlistCloud runtime 每次启动都必须重新完成 schema、身份、有效权限和 catalog 负向校验,发现 drift 立即失败关闭。
- PostgreSQL role 是 cluster-wide identity,当前 database 内的 catalog audit 不能证明同一 credential 无法连接 cluster 中的其他 database。
SaaS 生产环境必须使用仅向该 credential 暴露目标 business database 的专用 PostgreSQL cluster/endpoint
或通过已验证的 HBA/proxy policy 证明该 credential 只能连接目标 business database;这项部署隔离仍是未完成的 activation gate。
- 关键租户表使用 `FORCE ROW LEVEL SECURITY`migration/repair/audit 使用独立受控 migrator role。
- 每个租户事务通过 `SET LOCAL` 设置 tenant context,并由统一 `TenantUnitOfWork` 保证设置 context 和业务查询使用同一事务/连接。
禁止使用连接级 session variable 或 `search_path`,避免连接池、PgBouncer、异常回滚和后台任务串租户。
- 一个 `TenantUnitOfWork` 只能访问一个 Workspace。业务写入与对应 business outbox 在同一事务中提交;
写入可以校验由执行层传入的 generation/fencing token,但 Runtime owner、lease 和 Box session directory 不由业务 PostgreSQL 承担。
- SaaS schema、extension 和 policy 只由 release migration job 创建;应用启动角色不执行 `CREATE EXTENSION``create_all` 或自动 migration。
- OSS 继续默认 SQLite,并保留自托管 PostgreSQL 选项;Cloud RLS 约束不让 OSS 强制依赖 PostgreSQL。
### 5.2 pgvector 首期方案
- SaaS 使用 pgvector 作为默认向量数据库,并与业务表使用同一个 PostgreSQL cluster/database
vector schema 可以使用独立 adapter、受控 role 和有上限的 pool,但不新增 Chroma、Milvus 或独立向量数据库服务。
- OSS 默认仍是 SQLite + Chroma,用户可以显式选择 pgvectorCloud 配置 pgvector 失败时必须启动失败,不能静默回退到 Chroma。
- 向量表必须显式保存 `workspace_uuid``knowledge_base_uuid`,并至少以
`(workspace_uuid, knowledge_base_uuid, vector_id)` 建立唯一键/主键和查询条件;服务端生成的 collection name/hash 不是安全边界。
- pgvector adapter 复用相同的 tenant-context/RLS 契约,每次向量操作在自己的事务中执行 `SET LOCAL`
是否复用普通业务 UoW、role 或 connection pool 由实现决定,但 adapter 不能丢弃 tenant metadata。
- `vector(1536)` 不能继续作为无条件硬编码。首期使用无 typmod 的 `vector` 列和显式 `embedding_dimension`
`CHECK (vector_dims(embedding) = embedding_dimension)` 校验;release migration 为允许的维度创建带 dimension predicate 的 expression/partial ANN index。
知识库/model 元数据必须选择已启用维度,写入和查询 mismatch 或未启用维度时失败关闭,不能截断、补齐、退化为无界扫描或换后端。
- `vector` extension、表、索引和 RLS 由 release migration 创建。应用进程不在启动时执行 DDL。
- 0013 如需搬迁 legacy pgvector 数据,migrator 作为源表 owner 先记录每个受保护源表的 `ENABLE/FORCE RLS` 状态,
仅在同一 migration transaction 内临时暂停 RLS,并在 `finally` 中精确恢复各表原状态。
该流程不依赖 superuser 或 `BYPASSRLS`,也不允许在迁移事务外留下已禁用的 RLS。
### 5.3 候选拓扑与未来演进
| 状态 | 方案 | 结论 |
| -------- | ----------------------------------------- | --------------------------------------------------------------------------- |
| 首期决定 | P0. shared database/shared schema | 一个 pool、一套 migration;应用 scope + RLS |
| 后续演进 | P1. 多 shared database shard | 每个 shard 仍承载多个 Workspace,并使用相同 schema;有容量/地域证据后再设计 |
| 后续例外 | P2. dedicated shard | 只作为合规、驻留或超大 workload 的资源等级,不建立第二套代码路径 |
| 淘汰 | schema/database per Workspace | catalog、pool、migration、备份成本随 Workspace 线性增长 |
| 淘汰 | database/schema per replica/Cell/Instance | 把业务数据拓扑错误绑定到计算副本或已删除的产品实体 |
M0 不提前增加始终返回 `primary` 的 resolver、shard router 或 shard binding。
P1 的 resolver、映射、在线迁移、连接池预算、shard-affine replica 和 dedicated shard 细节等到出现容量、地域或合规需求时再设计。
在此之前,direct endpoint 与 pooler endpoint 分离只能通过数据库内部、runtime 不可伪造的 cluster identity 开启,不使用 DNS 名、数据库名或配置声明代替。
### 5.4 备份与生命周期边界
- PostgreSQL PITR 是 database/cluster 级恢复手段,不等同于单 Workspace 恢复。
- Workspace 创建、释放、export、delete、单 Workspace restore 和在线迁移机制本轮暂缓,后续单独决策;
首期不以尚未设计的 export 能力作为数据库架构验收条件。
- 除关系业务数据和 pgvector 的向量/检索字段外,大对象、插件 artifact 和 sandbox 文件仍存放在对象存储或 Runtime 持久卷;
PostgreSQL 保存相应业务元数据和稳定引用。
- tenant-visible usage/billing 业务行可以进入 PostgreSQL;基础设施 log/metric/trace 不进入业务数据库。
高增长业务表在有数据量证据后再决定 retention、时间分区或分析存储。
### 5.5 验收条件
- 故意遗漏应用层 Workspace filter 时,RLS 仍阻止跨租户读写。
- 连接池/事务池复用、异常回滚、并发请求和后台任务不会残留 tenant context;如部署 PgBouncer,也必须覆盖 transaction pooling。
- migration 对 shared schema 只执行一次,不产生 Workspace 级 schema drift;应用启动角色不能执行 DDL。
- 首版拒绝 host、port 或 database 不同的 migrator/runtime URL,并拒绝相同 rolemigrator/runtime 都只解析到 `public`
migrator 在迁移后完成精确 table/sequence/`alembic_version` ACL grant 和正反向 role 校验,runtime 每次启动重新校验。
- runtime role 没有任一方向的 role membership、`WITH GRANT OPTION``search_path` 覆写、对象所有权、其他 schema 访问或非业务对象权限;
专用业务 database 上可继承 PostgreSQL 默认 `PUBLIC TEMP`,但 runtime role 没有直接 `TEMP` ACL。
- migrator/runtime session GUC 保持 `session_replication_role=origin``row_security=on``lo_compat_privileges=off`
runtime role/当前 database 没有任何 `pg_db_role_setting`extension 仅为 `plpgsql/vector` 且 runtime 不拥有 extension
database 中没有 FDW/server/user mapping、runtime 或 `PUBLIC` 显式 routine/parameter ACL、runtime-owned routine 或 runtime 可执行的 `SECURITY DEFINER` routine。
- 生产 deployment 证明 cluster-wide runtime credential 只能连接目标 business database;专用 cluster/endpoint 或 HBA/proxy 隔离未经验证前不得启用 SaaS。
- legacy pgvector 搬迁在非 superuser、非 `BYPASSRLS` 的 table-owner migrator 下可成功,成功、异常和重试路径都精确恢复所有源表的 RLS/FORCE 状态。
- 业务写入和对应 business outbox 在同一事务内具备可证明的提交顺序;外部 generation/fencing token 校验失败时不产生写入。
- pgvector 使用真实 PostgreSQL 集成测试覆盖:两个 Workspace 使用相同 `vector_id`、猜测其他 Workspace ID、
故意遗漏 scope、连接复用、CRUD 和后台任务,全部不能越权。
- embedding dimension 不匹配或 pgvector extension 不可用时失败关闭,不回退到其他向量后端。
- 新建 Workspace 只新增目录与业务行,不创建 database、schema、role 或专属连接池。
## 6. D-004stdio MCP 独立开关
状态:`IMPLEMENTED`
### 6.1 已修复的原问题
- 修复前,stdio MCP 的启用条件只检查 transport 为 `stdio` 且 Box available,没有独立 feature gate。
- 修复前,所有 stdio MCP 使用固定的 `mcp-shared` 逻辑 session,并强制 `persistent=True`
- 在该旧逻辑下,如果多租户 Cloud 只通过 `box.enabled` 开放能力,每个配置 stdio MCP 的 Workspace 都会额外保留一个 persistent sandbox
绕过“每 Workspace 最多一个 managed `global` sandbox”的成本和套餐边界。
### 6.2 首期决定
新增独立实例配置:
```yaml
mcp:
stdio:
enabled: true
```
- OSS 默认 `true`,保持当前本地部署兼容;Cloud v2 通过 `MCP__STDIO__ENABLED=false` 强制关闭。
- 该开关与 `box.enabled``managed_sandbox` entitlement 和 sandbox session 数量相互独立,不能从任一条件推导。
- 后续如开放给特定套餐,可在实例开关之上再叠加 Workspace capability;实例开关为 `false` 时任何 entitlement 都不能绕过。
- HTTP/SSE/其他远程 MCP transport 不受此开关影响。
### 6.3 强制检查点
开关必须同时覆盖:
1. MCP create
2. MCP update 到 stdio
3. transient connection test
4. Core 启动时加载已有 stdio 配置;
5. RuntimeMCPSession/loader 的最终执行门禁;
6. WebUI transport selector 和错误提示。
不能只在 WebUI 隐藏选项。Cloud 配置关闭时,已有 stdio 记录保留但不自动启动,并返回明确的 feature-disabled 错误;
最终 gate 必须位于 Box 分支和 legacy host-stdio 分支之前,不能误报为 `box_unavailable`
也不得创建 `mcp-shared` session 或 stdio 子进程。
### 6.4 验收条件
- Cloud 即使 `box.enabled=true` 且 Workspace 拥有一个 managed sandbox,也无法 create/update/test/start 任何 stdio MCP。
- 直接调用 API、重放旧配置和启动 bootstrap 都失败关闭,且不会产生 `mcp-shared` session、nsjail 进程或额外配额占用。
- OSS 默认行为保持兼容;HTTP/SSE MCP 正常工作。
## 7. 五项决策之间的关系
五项决策共同遵循:
> 多租户共享可信控制面、连接池、只读 artifact 和基础容量;租户独占不可信执行进程、sandbox、secret、可写文件和数据作用域。
- Plugin Runtime 通过“共享 Supervisor + 每 installation 独立 nsjail 进程”降低控制面数量,同时保留进程级租户隔离。
- Box 通过“共享 Runtime + 每个合资格 Workspace 一个持久逻辑 session + one-shot nsjail exec”避免每租户部署服务和空闲容器。
- stdio MCP 独立关闭,防止从 Box availability 隐式产生第二套 persistent sandbox。
- PostgreSQL 和 pgvector 共享数据库组件,但使用显式 tenant key、应用层 scope 和 RLS 防止共享存储变成共享权限。
- 订阅管理只在闭源 Control Plane 维护套餐与计费规则,并向开源 Core 投影签名/版本化 entitlement
Core/Runtime 执行通用 capability 和数值限额,不复制套餐名称或计费逻辑。
- 目录同步只在启动或恢复时读取全量快照;常态变更按 Workspace 聚合为签名增量,新增租户不会使每次目录事件退化为全实例重投影。
## 8. 当前不做分布式时仍保留的能力
1. 所有运行时协议继续携带稳定 `instance_uuid``workspace_uuid` 和 execution generation;不能依赖进程地址表达身份。
2. Core、Plugin Runtime 和 Box Runtime 的本地进程表不能成为 durable desired state 或撤销状态的唯一真相。
3. 创建、重试、回调、outbox 和 worker 注册使用稳定 idempotency key;重复投递不能产生第二个 owner 或副作用。
4. Plugin installation 和 Box session 使用稳定 owner abstraction;启用第二个 replica 前再实现带 expiry、CAS 和 fencing token 的 lease。
具体 lease store 后续决定,不预设复用业务 PostgreSQL,更不因此新增 Runtime/Box 数据库。
5. Repository/UoW 不允许无边界跨 Workspace 事务;`workspace_uuid` 从第一天就是内部路由与分片候选键。
6. schema migration、任务扫描、监控聚合和运维接口不能假设永远只有一个 Core 进程。
7. 外部 API 不暴露 replica、worker 或 shard 标识;未来扩容不改变 Workspace URL、UUID 或客户端协议。
8. 只有出现容量、可用性、地域或合规需求时才增加 replica/shard;预留协议不等于现在部署额外组件。
9. 每个 Core replica 保存自己的事件消费 cursor,因为 entitlement cache 是进程本地状态;PostgreSQL 中的 projection high-water mark
、snapshot coverage 和 inbox 仍由所有 replica 共享,用于幂等投影和冲突检测。事件页携带签名 high-water,副本追平前不能续期
ready;不能让一个 replica 的共享 cursor 使其他 replica 跳过本地 cache 刷新。
## 9. 本轮明确不做的事情
- 不合并 plugin installation 进程,即使插件和版本完全相同;只允许共享摘要校验后的只读代码/依赖文件。
- 不允许插件 manifest 声明或覆盖 CPU、内存、PID、文件或存储上限。
- 不为每个 Workspace 创建 Plugin Runtime、Box service、database、schema、role、bucket、PVC 或消息队列。
- 不在 Cloud v2 首期使用 Docker sandbox、microVM、warm pool 或非 Pro managed sandbox。
- 不在首期实现 Workspace 级 BYOK E2B WebUI 配置。
- 不在 Cloud v2 首期支持 stdio MCP,也不让 Box availability 隐式开启它。
- 不把业务 PostgreSQL 用作未决定的 Runtime/Box 通用协调数据库,不在缺少容量证据时引入 Redis、Kafka 或新 scheduler。
- 不在本轮实现 Workspace export、释放、单租户恢复或在线迁移;具体生命周期另行决策。
- 不实现多个 CloudInstance、Workspace Placement 或 Cell Router;未来分布式只作为单逻辑实例内部的副本和分片能力。
- 不修改旧 Space 部署模型;Cloud v2 继续按绿地方案设计。
@@ -0,0 +1,348 @@
# LangBot Cloud Runtime 资源安全审查
日期:2026-07-28 至 2026-07-29
审查分支:
- LangBot`feat/multi-tenants`,审查起点 `32abbb636f4455e965141d8d209b359dbfbb5aae`
- Plugin SDK`feat/multi-tenants`,审查起点 `0cddf3c2bea5939c67b71e488a719e9903c28d17`
## 结论
本轮已覆盖 LangBot Core、Plugin Runtime 和 Box Runtime 的主要常驻对象、后台任务、队列、网络客户端、进程生命周期及数据库连接池。本轮定位到的攻击者可控或历史累积状态均已补充容量、超时、淘汰或确定性清理边界;修正后的高基数探针没有观察到随历史请求继续增长的活跃缓存。按本轮“代码审查 + 本地可重复测试”的验收口径,审查已经完成,未发现仍未处理的严重内存泄漏或 CPU 抢占路径。该结论不等于证明任意生产负载下不存在资源问题。
最终 Cloud 拓扑和生产环境不是本轮完成条件。代码级审查、跨仓全量测试和仓库 Dockerfile 构建的 Linux/cgroup v2 探针已经通过,但当前状态仍不能单独作为 Cloud 生产激活批准。完整的环境侧剩余清单见
[Cloud v2 仍待验证事项](./cloud-v2-pending-verification.md)。其中与本轮资源审查直接相关、上线前还必须完成的项目包括:
1. 在最终 Cloud 部署权限和 cgroup 拓扑下重复 nsjail、namespace 和 delegated cgroup v2 的 CPU、内存、swap、PID、文件句柄验证。本轮一次性 Linux 容器已经证明代码路径可工作,但普通容器和仅 `--privileged` 的 private cgroup namespace 都不满足条件。
2. 为 Cloud Box 提供并验证硬文件系统 quota provider。普通 nsjail bind mount 不能证明总字节数和 inode 硬配额,当前严格 readiness 按设计会失败关闭。
3. 使用最终生产配置分布继续做容量测试,并据此确定单实例 Workspace placement 上限。本轮真实 PostgreSQL 16 + RLS 启动测试已经覆盖 1,000 个各带 Provider、三类 Model、Bot、Pipeline、KnowledgeBase、MCP 和 Plugin setting 的 Workspace,启动加载耗时和 SQL 次数保持线性;5,000 Workspace 的合成三代替换探针也证明旧运行时会释放。仓库已新增可同时采集 Core/Plugin/Box HTTP、进程树和 cgroup v2 的 24 小时门禁工具,并在受 CPU、memory、swap、PID 硬限制的 Linux 容器中完成短时自检;但最终生产候选拓扑的 24 小时运行仍未执行。测试中的 fake adapter/requester/Plugin handler 仍不能替代真实平台 SDK、外部连接池和插件进程的容量数据;合法活跃租户本身仍会线性占用内存。
SDK 已先行发布到分支提交 `1d65ed301a6afc52150a998043f73cd6032c8162`,本提交集中的 LangBot
`pyproject.toml``uv.lock` 已精确钉住该提交。最终镜像仍需按待验证清单记录并核对实际安装版本。
## 覆盖范围
### LangBot
- 启动、停机、全局任务管理和运行时配置。
- PostgreSQL、Tenant UoW、RLS、迁移和共享 pgvector。
- HTTP、MCP、WebSocket、上传下载、S3/本地存储、维护任务和异步用户任务。
- QueryPool、Pipeline Controller、会话/对话、限流、第三方 Agent/LLM runner 和同步 SDK 桥接。
- PlatformManager 以及 DingTalk、QQ、Lark、WeCom、WeComCS、WeChatPad、LINE、Kook、Satori、OpenClaw Weixin、Telegram、Discord、Matrix、HTTP/WebSocket 等适配器。
- Plugin Runtime connector、插件包校验、Marketplace 下载、pip 安装输出和 desired-state reconcile。
- Box connector、admission、session/process 生命周期和 RPC 文件。
- RAG、向量后端、Skill、Storage、Telemetry 和日志缓存。
### Plugin SDK
- stdio/WebSocket transport、请求 waiter、action task 和文件传输。
- Runtime control handler、Workspace/generation fence 和 EventContext。
- 插件 artifact、Marketplace、pip 依赖安装、共享依赖环境、installation desired state、Supervisor 和 worker launcher。
- nsjail 参数、cgroup/rlimit、进程注册 capability 和 Runtime shutdown。
- Box admission、generation fence、session/process/reaper、RPC/relay WebSocket 和 nsjail backend。
## 主要修复
### 确定性生命周期
- 修复 `Application.shutdown()` 使用 `contextlib.suppress` 却未导入 `contextlib` 的问题。原行为会在实际资源关闭分支直接 `NameError`,阻断后续 Plugin、Box、HTTP 和数据库释放。
- `make_app()` 在任一启动 stage 或初始化失败时会关闭尚未返回给 `main()` 的半构建 ApplicationTelemetry、Box、Tool、Platform、Vector 和 HTTP manager 在初始化前即挂到 Application,避免初始化中途失败后清理器无法发现已经创建的连接、会话、子进程或后台任务。
- MCP streamable-HTTP session manager 现在随 Application shutdown 显式退出。
- MCP loader 按 `(instance, workspace, generation)` 管理 host task 和 session;任务完成会从注册表移除,代次推进会取消旧 host task 并关闭旧 sessionreload/shutdown 会先清空现有运行时,避免 completed task 和旧代连接永久驻留。
- Platform bot reload/remove/shutdown 统一串行化,旧 bot、代理、adapter 任务和进程会先停止再从注册表移除。
- Model provider requester 新增异步关闭契约;provider reload/remove、Workspace generation 替换、全量 reload 和 Application shutdown 都会确定性关闭旧 requester,允许第三方 requester 安全持有自己的 HTTP client 或连接池。
- Plugin Runtime、Box Runtime、stdio transport、adapter 连接和共享 HTTP client 均补齐 close/cancel/await。
- HTTPX 有界流在超限异常或消费者取消时会立即关闭底层响应流;原来的 response hook
只在正常读完后由 HTTPX 自动关闭,持久客户端反复收到超大响应时可能积累未释放连接。
已消费响应在超限分支也会先 `aclose()` 再传播错误。
- `Application.dispose()` 只允许一个可追踪 shutdown task;重复的信号、窗口关闭或调用方清理不会铺开多个并行停机流程。
- Lark、微信、钉钉、企业微信和 QQ Official 的凭证交换后台任务统一进入 Application TaskManager,受全局/单 Workspace admission 约束并随应用停机取消;容量满时关闭尚未调度的 coroutine 并返回 429,不留下游离 task。
- `TaskCapacityError` 已下沉到无 Application/controller 依赖的纯错误模块。原来的 HTTP 过载异常路径会在特定冷启动导入顺序下触发 TaskManager/controller 循环导入,把应返回的 429 变成框架 500。
- S3 storage provider 在初始化失败或 Application shutdown 时关闭 botocore HTTP connection poolStorage manager 在 provider 初始化前即挂到 Application,避免 bucket probe 失败后遗留 client。
- 修复 Coze runner 每次请求创建 `aiohttp.ClientSession` 却不关闭的问题;现在 runner 的 `aclose()` 会确定性关闭底层 API client。
- LINE SDK client 现在随 adapter 停止关闭;Plugin Runtime shutdown 回调只创建一个可追踪、可等待的后台任务,重复回调不会累积清理 task。
- SDK `lbp publish` 现在用上下文管理器关闭插件包上传文件;原实现的成功、API 错误和 HTTP 错误返回路径都会遗留文件句柄。
- Plugin worker、Box Docker/CLI backend、Box nsjail backend 和 nsjail 依赖安装在调用方取消时会 terminate/kill、读取管道并 `wait()` 回收子进程;Windows worker 的原生进程路径也进入相同的 `finally` 清理契约,避免取消安装、停机或超时后留下孤儿进程。
- Box 服务入口现在从 Runtime initialize、aiohttp app/runner setup、端口 bind 到主循环共用一个外层清理边界;建站期间的非 `OSError` 也会关闭 Runtime 和 reaper。WebSocket 控制模式端口绑定失败会退出并交给编排器重启,不再在没有任何可用 RPC/health 端口时永久等待;stdio 模式仍允许仅 relay 绑定失败后继续控制通道。
- Plugin artifact 解压、installation staging/activation/rollback/delete、共享依赖环境和 nsjail session 目录的关键文件系统变更均移出事件循环。已经开始的原子变更在取消时会先等待线程结束,再回滚临时目录、目标目录或旧 supervisor,不会让后台线程继续修改一个调用方已经认为清理完毕的路径。
- Core 和 SDK 的阻塞 executor 新增独立的有界清理入口。普通工作在容量耗尽时仍快速拒绝;已经拥有资源的 close/unlink/rmtree/进程回收则等待有限 worker 槽位并保证原子操作结束后再传播取消,避免过载恰好导致清理任务被拒绝。SeekDB、Milvus、S3、LINE、WeChatPad、MCP staging 和 TBox 临时文件等关键路径已接入。
### 有界队列、缓存和历史状态
- QueryPool 同时限制全局与单 Workspace 的 queued/running query;调度后不再被过载淘汰;历史 scope counter 有上限。
- Session、Conversation、WebSocket connection/proxy/message、rate-limit identity、task record/log、telemetry task、vector handle 和 adapter 私有队列均有容量或 LRU/TTL。
- SessionManager 现在维护 Workspace 二级索引和带 revision 校验的最小过期堆。新会话只扫描目标 Workspace 的有界会话集,TTL 回收只消费已过期堆前缀,全局 idle 淘汰使用最小堆;高频命中产生的旧堆项按活跃会话的有界倍数压缩。原实现会在每个攻击者可制造的新 launcher id 上扫描并排序实例全部会话。
- SDK 的 EventContext 和依赖准备锁使用 weak referencegeneration、admission、installation、capability 和 completed-process 状态有上限。
- Plugin restart circuit 打开期间只有 `max_concurrent_restarts` 个 supervisor
能持有冷却计时器/状态等待,其他 installation 睡眠在同一个 semaphore FIFO
probe 状态变更在调用者取消时仍会完成,避免 half-open 永久占用。
- Box nsjail 启动时只扫描一次 `/proc`,并流式删除遗留 session 目录;不再为每个
遗留目录重复扫描全部进程或先把全部目录物化到内存。
- Box Skill discovery、目录列表和列表正文分别限制扫描 entry、package、返回 entry
与累计文本字节;BFS 使用 deque,拒绝 inode 洪泛导致的 O(N²) 或无界列表。
- Core message aggregation 使用 `(instance, workspace, generation)` O(1) scope
counter 做准入,不再为每个新 launcher 扫描全实例 buffer。
- Cloud remote MCP 的 idle execution fence 不再由每个 session 每 5 秒查询
Workspace/ExecutionState;签名目录投影事务提交后将失效 scope 合并进一个有界
cleanup worker。实际工具/资源调用前后仍保留数据库强校验。
- 空 Workspace 不再预分配 Model generation scope、Plugin installation set 或 Box generation event;只有 Workspace 实际拥有对应运行时资源或等待任务时才创建这些对象。
- Runtime RPC 文件同时限制单文件字节数和单连接未消费文件数量,连接关闭时清理连接拥有的临时文件。
- Box Runtime 维护实例与 Workspace 到活跃 session 的二级索引;创建、删除、过期、撤销和 shutdown 共用同一清理路径,避免每个租户 RPC 都扫描实例中的全部 session。
- Box Runtime 另外只索引“可过期 session”和“持有 managed process 的 session”。Cloud 的持久 `global` sandbox 不进入 TTL 索引,managed process 被禁用时进程索引为空;session 创建、周期 reaper、状态和 `/healthz` 因此不会随全部持久租户数线性扫描。测试把总 session 字典替换为禁止迭代的映射,仍能完成第二个持久 session 创建、reap、status 和 health。
- Core MCP loader 同样维护 Workspace/generation 到 session、host task 的二级索引;请求、代次回收和动态配置不再扫描实例中的全部租户 MCP session,已完成 task 的 done callback 会同步移除所有索引。
- Box admission 过期回收使用带 revision/generation 校验的最小堆,只访问已到期记录;重复续期产生的旧堆项会被忽略,堆大小超过活跃 grant 的有界倍数时压缩,不再在每次 RPC 上全表扫描所有租户 grant。
- 旧 QQ message ID/object cache 和 stdio MCP Workspace copy lock 不再随历史请求无限增长。
- LLM/Agent runner 的单次生成结果默认限制为 1 MiB,流式传输限制单事件 1 MiB、单请求累计 16 MiB,并限制最多 100,000 个流式事件,避免上游异常响应无限占用内存或 CPU。
- Marketplace JSON 限制为 1 MiB、插件包限制为 64 MiBpip stdout/stderr 各最多保留 1 MiB,超出部分继续 drain 但不驻留内存。
- Plugin Runtime stdio/WebSocket 协议除 16 MiB 消息字节上限外,新增最多 4,096 个入站和出站碎片的对象数量上限;大消息的 UTF-8 编码、分片、拼接、JSON 编解码和 Pydantic 验证均在线程执行。WebSocket receive 异常使用稳定的 `ConnectionClosed` 类型,不再因库顶层未导出 `exceptions` 属性而在错误路径二次失败。
- HTTPX response hook 和 aiohttp 有界读取统一把第三方响应限制为 10 MiB;JSON 解析及诊断文本转换在线程执行,错误正文最多保留 4 KiB,避免大 JSON 在共享事件循环集中解析。
- 图片、data URI 和平台媒体默认限制为 10 MiB,Base64 在解码前先校验编码长度;Plugin binary storage 默认单值 10 MiB,并设置不可由错误配置绕过的 64 MiB 绝对上限。
- Skill 文本单文件限制为 1 MiBPlugin UI 文件限制为 4 MiBhost edit 文件限制为 1 MiBBox Skill ZIP、插件 artifact 和 GitHub Skill archive 同时限制条目数、单文件、解压总量和压缩比。
- SDK E2B 文件同步限制为最多 2,048 个目录项、1,024 个文件、单文件 10 MiB、总计 50 MiB;同步文件 IO 从事件循环移到线程。
- Dify 待提交表单、用户 Space OAuth state 和 Cloud launch JTI replay cache 改为带 revision 校验的最小过期堆;Space credits 使用按时间有序的 LRU/TTL 队列。原实现会在每次攻击者可触发的请求上扫描整个历史缓存并在满容量时再次线性寻找最旧项;现在过期回收为摊销 `O(log N)` 或仅消费已过期前缀,旧堆项会忽略并按活跃状态的有界倍数压缩。
- Cloud launch JTI cache 达到 4,096 个仍有效 token 时失败关闭,不再为了接纳新 token 淘汰仍有效的 replay 记录;否则攻击者可以在容量满后重放被提前遗忘的合法签名。
- Entitlement resolver 现在跟随 Cloud directory 的权威 Workspace 活跃集合。全量目录投影会丢弃已 fenced/removed Workspace 的历史 entitlement snapshotdelta 批量更新不会对每个变化重复扫描;provider 请求进行中发生目录撤销时,返回前的第二次 active fence 会阻止旧结果重新写回缓存。
- Cloud directory 的签名响应、Workspace、membership 和实例 active Workspace
均新增可配置操作上限与绝对上限。闭源适配器按流读取响应并在 JSON/JWS 解析前
拒绝超过 32 MiB 默认值的解压后正文;Manifest/entitlement/event endpoint 再
分别限制为 256 KiB、256 KiB 和 2 MiB。entitlement 刷新最多并发并驻留 16 个
原始响应,逐批验证成小型 snapshot 后释放;Core 在目录投影行锁保护的事务内检查最终
active 数量,超限时回滚 Workspace、Account、membership、inbox 和 cursor,不会
截断权威数据或让并发副本各自越过最后一个容量槽。
- Space 全量目录只投影 active Workspace,历史 archived Workspace 只在最多 100 个
目标的签名 delta 中作为 tombstone 返回。注册创建新个人 Workspace 前通过
PostgreSQL transaction advisory lock 串行执行全局 active 数量准入;达到上限
返回 503,避免多个 Space 副本同时观察到最后一个空位。
- Monitoring 分页、offset、CSV export 和 session/message detail 均在 service
边界执行实例配置上限与不可放大的绝对上限;detail 的完整统计改为 SQL aggregate
只物化有界的 tool/LLM/error 明细并显式返回 `detail_truncated`。默认分页 1,000、
export 10,000、detail 2,000,绝对上限分别为 5,000、50,000、10,000。
- Token statistics 的时间序列不再把筛选范围内的全部 LLM call 拉回 Python 分桶;
PostgreSQL 使用 `date_trunc`、SQLite 使用 `strftime` 在数据库中聚合,并只返回
最近 1,000 个时间桶(绝对上限 10,000)。模型分组复用分页上限并在 SQL 中按 token
排序、限制;两类结果都返回显式的 `*_truncated` 标志。
- Monitoring 过期数据每表每轮默认最多删除 4 个批次、绝对最多 100 个批次;本地/S3
过期上传文件候选和每轮删除默认最多 1,000、绝对最多 10,000。单个历史数据量异常的
Workspace 不再能让一次维护循环无限物化候选或持续清空全部 backlog。
- Workspace webhook 数量默认限制为 16、绝对限制为 64;管理查询和运行时 fan-out
都只物化有界结果。实例同时发送的 webhook 请求默认限制为 16、绝对限制为 128;
满载时直接跳过未获准目的地,不创建一批等待 semaphore 的 task。取消调用时会取消并
await 已创建的所有请求任务,归还实例槽位。
- Local/S3 Storage 的对象读取在实际 IO 中只读取 `limit + 1` 字节,S3 body 在成功、
超限和异常分支都会关闭;默认单对象 10 MiB、绝对上限 64 MiB。所有 scoped load
以及 WebSocket attachment 都经过同一边界,写入也不能产生当前实例无法安全读取的对象。
- Valkey Search 的批量删除改为固定页流式搜索、删除并累计计数,不再把全部匹配 key
保留在 Python 列表;每次删除后从 offset 0 继续,避免结果集缩短造成跳项,并设置
1,000 轮绝对终止条件。
### CPU 和事件循环保护
- 修复旧 QQ `repeat_seed('')` 空输入无限循环。
- ZIP 校验/重打包、PIL、Base64、AES、JSON 解析、fsync probe、插件 artifact/依赖文件、Skill、S3、本地存储和维护目录扫描从事件循环移到线程。
- 公开 Slack、QQ Official、HTTP Bot、公众号、WeCom/WeComCS 回调体显式限制为 1 MiB;JSON/XML 解码移出共享事件循环。QQ、DingTalk、Satori、WeCom AI 和 WeChatPad 网关帧同样设置 1 MiB 上限或在解码前拒绝超限消息;KOOK zlib 数据使用 10 MiB 解压后硬上限,阻断小压缩包制造的大内存解压。
- HTTP Bot 的幂等键和 outbound session 现在均在写入前执行硬容量 admission;满额时只按固定 64 项预算检查最旧记录,不能先超限后整体清空,也不再在每条回复上扫描和排序全部 session。已有 session 继续 O(1) 访问,新 session 在没有可安全回收的空闲记录时失败关闭。
- Dashboard、Embed 和 Plugin Runtime 的协议 JSON 编解码在线程执行;Dashboard/Embed 在接收端提交 terminal error 后为发送端保留有界 drain 窗口并使用内部 sentinel 唤醒,不会因“任一方向结束即取消”在撤权错误帧发出前关闭连接。
- 租户配置的敏感词、内容忽略和群响应正则统一使用声明为直接依赖的 `regex` 引擎:最多 64 个 pattern、单 pattern 1,024 字符、输入 1 MiB、单次总匹配 CPU 预算 50 ms,并在线程中执行。超时、非法正则和替换放大均失败关闭;灾难性 `(a+)+$` 回归在 1 ms 测试预算内被中断。
- 原生 `read/write/edit/glob/grep` 文件工具移出事件循环并继承 Workspace 阻塞预算。目录列举、递归 walk、grep 文件/总字符、单行、pattern、结果和 regex CPU 均有硬上限;glob 只用固定大小最小堆保留最新 100 项,不再先把全部命中路径驻留内存。Box 内执行的 glob/grep 脚本同样限制命中集合、扫描量和正则时间。
- Dify、DingTalk、QQ、WeCom 等客户端复用连接并在生命周期结束时关闭,响应体和下载有字节上限。
- DashScope、TBox 等同步第三方 SDK 的调用和生成器迭代改为在线程执行;单个同步生成器最多消费 100,000 个事件。
- Dashboard 和 Embed WebSocket 改为任一收发 task 结束即取消并等待另一方向,避免发送端退出后接收 task 永久阻塞;两方向 task 同时继承从认证结果或 RuntimeBot 得到的可信 Workspace 阻塞预算。
- Plugin installation 生命周期全局串行化;不同租户的依赖 pip/nsjail 准备不会在安装高峰并发抢占 CPU。
- Plugin installation 的意外退出除了每 installation 的 jittered exponential backoff,还经过 Runtime 全局 restart launch 并发槽和失败窗口熔断。
熔断冷却后只允许一个 half-open probeprobe 必须完成初始化并持续稳定后才恢复其他 installation。未在 30 秒内 ready 的 worker
会被取消回收。健康指标输出 active launch、窗口失败数、circuit 状态和累计打开次数,24 小时门禁把 circuit 打开或尾段启动槽不归零判为失败。
- S3 同步 SDK 使用线程执行,并通过实例级 semaphore 限制并发;默认 `storage.s3.max_concurrency=16`,可通过实例配置和环境变量覆写。
- Box 子进程 stderr 以 64 KiB 块读取,日志最多每秒输出 4 个摘录并汇总抑制数量,避免无换行或刷屏输出制造无界缓冲与日志放大。
- Plugin worker 日志单行最多保留 64 KiBBox managed-process stdout relay 以固定 64 KiB 块读取,不再依赖换行符,避免超长无换行输出触发 `StreamReader` limit 或堵塞子进程。
- Box generation fence 的代次更新改为只访问目标 Workspace 的 event 和 active-task 二级索引。原实现每次更新都会遍历全部 Workspace 的 fence/task 记录,10,000 个 Workspace 的第二阶段更新会退化为 O(N²) 并在 40 秒后仍未完成;修正后包括其他 SDK 高基数负载和本轮协议 offload 在内的当前完整双阶段探针耗时 `11.270s`
- Box session 枚举、旧 generation 回收和 admission 计数均通过 Workspace 索引执行;admission 过期回收通过最小堆执行,不再在每次 RPC 上产生 O(实例总 session/grant 数) 的扫描。
- Model、Pipeline、RAG 和 Platform manager 均维护 Workspace 到运行时 key 的二级索引。Workspace generation 更新只清理目标 Workspace 的缓存和运行时,不再扫描实例内所有租户的 provider/model、pipeline、knowledge runtime 或 bot;回归测试使用禁止全局迭代的映射验证该边界。
- Cloud heartbeat 直接读取已加载且有容量边界的 Pipeline、MCP、KnowledgeBase 和 Bot registry 计数,不再为每个活跃 Workspace 依次打开 Tenant UoW、执行四类 COUNT 查询;这消除了租户数增长后每日周期性形成的串行 SQL/CPU 尖峰。OSS 模式仍保留数据库统计语义。
- 邀请、Monitoring 和 Storage 的三个周期清理 task 合并为一个
`resource-maintenance` 调度器。调度器先等待首个 interval,不与启动加载争抢资源;
同一到期周期只执行一次 active Workspace discovery,然后按 Workspace 串行运行
有界 job,单 Workspace 失败不跳过其他 Workspace。默认相同的一小时周期由此从
三次全租户发现和三个同时唤醒的任务收敛为一次发现和一个任务。
- Cloud 启动阶段先生成一份经过部署适配器和目录投影校验的 Workspace binding 快照,Model、Platform、Pipeline、RAG 和 Plugin 初始化共用该快照,初始化完成后立即释放;避免启动期间为每个 manager 重复执行整批租户发现和投影校验。
- Platform、Pipeline 和 RAG 的资源加载在使用已验证启动快照时不再为每个 Bot/Pipeline/KnowledgeBase 重新查询同一个 execution binding;常规请求和动态更新路径仍保留数据库 generation fence。
- MCP 初始 host 和 shutdown burst 由实例级 semaphore/批次限制;默认 `mcp.lifecycle_concurrency=16`,支持 `MCP__LIFECYCLE_CONCURRENCY` 覆写并硬性限制最大 128。初始加载不再先为每个 server 创建一个等待 semaphore 的 task,而是由一个可取消 dispatcher 每批最多物化 `lifecycle_concurrency` 个子 task;同时去掉了 ORM server/config 的双份临时列表,避免大量租户启动时集中占用 CPU、内存、socket 和文件句柄。
- Core、Plugin Runtime、Box Runtime 和独立 Plugin worker 的默认 `asyncio.to_thread()` executor 统一改为硬有界线程池。默认最多同时运行 8 个阻塞调用、排队 128 个,达到容量后立即抛出 admission 错误,不再使用 Python 默认 `ThreadPoolExecutor` 的无界工作队列保留任意数量的请求对象、Future 和闭包。每个可信 Workspace 的 running + queued 默认再限制为 4,并强制配置值不超过 worker 数的一半,避免单个租户先提交一整批同步工作占满全部 worker/FIFO 队列。Core 使用 `system.blocking_executor.max_workers/max_pending/max_inflight_per_scope`,原生支持对应的 `SYSTEM__BLOCKING_EXECUTOR__*` 覆写;SDK 进程使用 `LANGBOT_BLOCKING_EXECUTOR_MAX_*`,并分别限制全局最大值为 64/4096。
- Core、Plugin Runtime 和 Box Runtime 各自运行固定 1 秒间隔、仅保留最近 120 个样本的 event-loop lag monitor;健康快照输出 current/recent max/recent p95/进程期最大延迟和累计样本数。Plugin Runtime 的两个 WebSocket 端口现在都在免认证 `/healthz` 返回同一份无凭据、无租户标识的聚合 JSON,Box `/readyz` 同时附带资源快照。24 小时门禁默认拒绝缺失或停止的 monitor、超过 1 秒的 recent max、超过 250 ms 的尾段 recent p95 及 sample counter 回退。
- Workspace 阻塞预算由服务端认证后的 `RequestContext`、公开 bot 的 RuntimeBot、公开对象 key 中经 binding fence 验证的 Workspace、Platform/TaskManager 的 ExecutionContext,以及 SDK 入站 ActionContext 建立,不接受调用方伪造的租户 header。公开 webhook、公开对象下载、Dashboard/Embed WebSocket、普通 HTTP handler、Platform adapter 和 detached tenant task 均已覆盖。容量拒绝在 Core HTTP 路径返回稳定的 429health/debug counter 分开报告 global 与 scope rejection。
- Argon2 密码 hash/verify 只允许一个实例级在途操作,额外并发立即返回容量错误而不是在 asyncio semaphore 中无限积累等待请求;该 CPU/内存密集工作同时使用独立的 `system:authentication` 阻塞作用域。Cloud 本身仍禁用本地密码登录。
- WeCom 扩展 API 的无限客户端超时改为 120 秒;平台 webhook 的 AES、媒体 Base64 与同步 SDK 调用均移出共享事件循环。
- 长文本转图片限制为 100,000 字符、256 行、800 万 RGBA 像素和 10 MiB 输出;
超限时回退到 forward message。数字边界查找从重复 `count/find/sort` 改成线性扫描,
PIL image 使用显式关闭,压缩步长为零时也能终止。
- Core 在每次 quota-enforced Box exec 前后遍历 Workspace 时使用非递归 DFS,并在
超过字节 quota 或默认 100,000/绝对 1,000,000 个目录项后立即停止;目录项洪泛
失败关闭,不再重复完整扫描 inode bomb。远程 outbox fallback 同时限制扫描项、
文件数、单文件和总字节,Python project manifest 使用分块 hash 并限制单文件 10 MiB。
### 插件和 Box 资源隔离
- Plugin worker 数量受 `max_workers``max_total_cpus / max_cpus``max_total_memory_mb / max_memory_mb` 的最小值约束。
- Shared profile 强制 Linux 和 nsjailCloud 强制 `plugin.worker.require_hard_limits=true`cgroup v2 delegation 不可用时拒绝启动。
- 每个 worker 下发 CPU、memory、swap、PID cgroup 限制,以及 process、open-file、file-size rlimit;插件 manifest 不能提高限制。
- Box nsjail 的 cgroup v2 路径现在同时设置 `memory.max``memory.swap.max=0`。修复前,48 MiB 沙盒可以把强制提交的 128 MiB 页面换出并正常退出,形成宿主 swap 抢占;修复后同一探针以 exit 137 被 cgroup 杀死。
- 仓库 Docker Compose/Kubernetes 示例显式下发 Core、Plugin Runtime 和 Box Runtime 的 blocking executor 上限;Kubernetes Box readiness probe 从仅报告进程存活的 `/healthz` 改为 `/readyz`,使 backend 或 managed-mode 隔离检查失败时不会把 Pod 加入就绪流量。
- 相同 digest 的已验证代码和依赖环境可只读共享,每个 installation 的 home/tmp/data 和进程独立。
- SDK 在发布共享依赖环境前最多校验 100,000 个目录项和 2 GiB 常规文件元数据总量;
超限的 staging tree 会被原子清理而不会进入 worker。`requirements.txt` 和插件
`manifest.yaml` 都使用 `limit + 1` 有界读取,manifest 额外限制为 1 MiB。
- Box session、managed process、completed process、admission record 和 RPC 文件均有实例级上限;Cloud entitlement 仍限制每个合资格 Workspace 一个 `global` session、零 managed process。
- Box Runtime 对上述实例级配置再增加不可放大的硬上限:session 5,000、managed
process 1,024、completed process 10,000、admission record 250,000、RPC 单文件
100 MiB、completed retention 86,400 秒。初始化与远程 INIT 对错误类型、负数和
超上限均失败关闭,错误动态更新不会留下部分生效的 limit。
### PostgreSQL
- Cloud 强制 PostgreSQL 业务库、共享 pgvector 和允许的固定向量维度。
- pgvector Cloud 模式复用业务数据库的同一个 AsyncEngine,不创建第二个连接池。
- `database.postgresql` 新增并校验 `pool_size``max_overflow``pool_timeout_seconds``pool_recycle_seconds`;默认最大连接数为 `10 + 10`
- `pool_size + max_overflow` 的绝对上限为 100timeout/recycle 也有绝对上限;
Cloud runtime 的 asyncpg 连接默认设置 60 秒 statement timeout、5 秒 lock
timeout 和 60 秒 idle-in-transaction timeout,并分别限制最大 300/60/300 秒。
一次性 release migration 不继承这些短 runtime timeout。
- `/healthz` 输出 pool 配置容量、checked-in/out、overflow、pool admission timeout
累计数和 SQL timeout 配置;目录同时输出 active/max 与最近批次
Workspace/membership 数,供生产 soak 和告警核对。
- Application shutdown 显式 dispose 业务引擎;standalone pgvector 仅关闭自己拥有的引擎。
- PersistenceManager 提供统一异步 shutdown;Cloud 常驻进程的启动失败、正常停机和一次性 release migration 的成功/异常路径都会释放数据库引擎。真实 PG catalog 测试还覆盖了“入口已经关闭后测试再次复用 manager 会重开 pool”的第二生命周期,严格资源告警模式下无 asyncpg socket/transport 遗留。
## 本轮采用的默认决策
- 优先 fail closed 或淘汰最老的 idle cache,不允许攻击者控制的历史 key 无限驻留。
- 插件依赖准备选择实例级串行化,以稳定 CPU/磁盘峰值;代价是批量安装耗时增加。
- PostgreSQL 使用一个显式有界共享连接池;未拆分 pgvector pool。
- 单实例目录的默认 active/full-snapshot Workspace 上限均为 1,000membership
上限为 20,000,签名响应上限为 32 MiB;绝对上限分别为 5,000、100,000 和
64 MiB。Space 和 Core 必须配置为同一操作上限,生产值只能根据 V-08 容量曲线
向下调整或在重跑全部门禁后提高。
- 第三方 runner 采用 1 MiB 单结果、16 MiB 单流总量和 100,000 个同步/异步事件的统一实例级安全上限;超限请求失败关闭。
- S3 默认允许 16 个并发阻塞调用,最大配置值 128;在没有独立 worker service 的前提下限制线程池排队和上游连接压力。
- MCP 生命周期默认并发 16、最大 128;该限制统一约束实例启动时的 session host 峰值和 shutdown 批次,不允许租户配置单独放大。
- Core 与 SDK 各进程的通用阻塞 executor 默认使用 8 个 worker、128 个 pending 槽位、每 Workspace 4 个在途槽位;它是实例/进程级共享背压,不由 Workspace 或插件 manifest 调高,单 Workspace 配置硬性不得超过 worker 的一半。生产值应按容器 CPU 和上游阻塞时延校准,不能把 pending 当吞吐配置无限放大。
- 插件包下载上限 64 MiBpip stdout/stderr 保留上限各 1 MiB;这不会限制安装进程实际输出,只限制父进程内存中的诊断副本。
- 通用远程响应和媒体默认上限 10 MiB;错误诊断正文只保留 4 KiB。Plugin binary storage 默认 10 MiB、绝对上限 64 MiBSkill 文本、Plugin UI 和 host edit 分别限制为 1 MiB、4 MiB 和 1 MiB。
- Storage scoped object 默认读写上限 10 MiB、代码绝对上限 64 MiBWebhook 默认每
Workspace 16 个、实例 16 个同时出站请求,代码绝对上限分别为 64 和 128。Box
Workspace quota 扫描默认最多访问 100,000 个目录项、绝对最多 1,000,000 个。
- SDK 共享依赖环境在发布前最多接受 100,000 个条目、2 GiB 常规文件元数据总量;
artifact manifest 与 requirements 各最多 1 MiB。这些是 Runtime 控制面在启动
worker 前的保护,不替代最终文件系统的 byte/inode 硬配额。
- Monitoring 查询上限由 `monitoring.query_limits` 配置并支持原生环境变量覆写,但始终
受代码绝对上限约束;cleanup 的每表批次数和 Storage 每轮文件数同样采用实例配置加
绝对上限。时间序列默认/绝对上限为 1,000/10,000 个数据库聚合桶,模型分组复用分页
上限。提高这些值必须计入 V-08/V-09 的数据库 CPU 与 Core RSS 容量曲线。
- Managed-process relay 保留 stdout 的原始换行,并按 64 KiB WebSocket frame 分块;不再承诺“一行对应一个 frame”。这是为无换行输出提供确定内存边界所需的协议收敛。
- 本轮没有把 Pipeline、Model、KnowledgeBase 等合法租户资源改成 lazy runtime。该改动会改变启动和请求语义,留到 Workspace placement/释放机制一起设计。
- 本轮没有为普通 nsjail 声称伪硬盘配额;严格 Cloud readiness 保持失败关闭。
## 验证结果
| 验证项 | 结果 |
| --- | --- |
| LangBot Ruff + `git diff --check` | 通过 |
| Plugin SDK Ruff + `git diff --check` | 通过 |
| LangBot 全量测试(含 unit/integration/Box/E2E | `2855 passed, 33 skipped` |
| Plugin SDK 全量测试 | `1328 passed` |
| Space Go 全量测试与闭源 Cloud Adapter 测试 | Go `go test ./...` 通过;Adapter `40 passed` |
| Space PostgreSQL 16 Cloud v2 目录与并发容量准入 | 通过;两个注册并发争用最后一个槽位时 `1 success / 1 capacity rejection / 1 active Workspace` |
| Core PostgreSQL 16 Cloud runtime server timeout | 真实连接从 `pg_settings` 读回 `60000ms / 5000ms / 60000ms` 的 statement/lock/idle-transaction timeout,并显式 dispose |
| 真实 PostgreSQL 16 + pgvector 迁移/RLS/发布测试(严格资源告警) | `22 passed` |
| 真实 PostgreSQL 16 + RLS populated Cloud 启动容量 | 500 Workspace `6.178s / CPU 3.026s`;当前 1,000 Workspace 复跑 `12.109s / CPU 5.967s` |
| 较早 Core Dockerfile Linux 镜像构建与 `regex` 导入 | 通过,image SHA `8893a14053df`;该镜像使用旧 SDK pin,已失效,最终候选必须重建 |
| `ResourceWarning` + `PytestUnraisableExceptionWarning` 全量门禁 | Core 与 SDK 均通过,并已固化到 pytest 配置 |
| Plugin SDK Box 专项测试(含全局扫描回归保护) | `669 passed` |
| Docker Compose 渲染、Compose/Kubernetes YAML 解析与 diff 检查 | 通过 |
| Cloud soak 门禁解析/采样/判定单元测试 | `27 passed` |
| Core/Plugin SDK event-loop monitor 专项测试 | 两仓各 `7 passed`,包含真实 50 ms scheduler stall |
| Cloud soak Linux 硬限制短时自检 | 通过;CPU `0.5`、memory+swap `256 MiB`、PID `128` 均从 cgroup v2 读回,冷却尾段 verdict `pass` |
| Core 双阶段历史 churn 资源探针(使用当前本地 SDK 分支) | audit 通过,`12.559s` |
| Core 5,000 个 populated Workspace 三代容量探针(使用当前本地 SDK 分支) | 当前复跑通过,最大替换耗时比 `1.405` |
| Plugin SDK 双阶段资源探针 | audit 通过,`11.270s` |
两个仓库新增了可重复执行的历史 churn 探针,Core 另有 populated Workspace 三代替换探针:
```bash
# LangBot Core
PYTHONPATH=../langbot-plugin-sdk/src uv run python scripts/runtime_resource_probe.py --scale audit --json
# LangBot Core5,000 个带代表性资源的 Workspace
PYTHONPATH=../langbot-plugin-sdk/src uv run python scripts/workspace_runtime_capacity_probe.py --scale audit --json
# LangBot Core:真实 PostgreSQL 16 + RLS populated Workspace 启动
TEST_POSTGRES_URL=postgresql+asyncpg://... \
LANGBOT_PG_CAPACITY_WORKSPACES=1000 \
uv run pytest \
tests/integration/persistence/test_migrations_postgres.py::TestPostgreSQLTenantRuntime::test_populated_cloud_startup_is_linear_and_task_bounded \
-q -W error::ResourceWarning --log-cli-level=INFO
# langbot-plugin-sdk
uv run python scripts/runtime_resource_probe.py --scale audit --json
```
Core audit 每个阶段执行 10,000 个空 Workspace 的真实 Model/Plugin manager 加载与 reconcile、25,000 次 Query、2,500 次 session churn、10,000 个限流身份、5,000 个 task 和 2,500 次 WebSocket churn。第一、第二阶段的保留状态完全一致:
- 20,000 个历史空 WorkspaceModel scope/provider/LLM、Plugin Workspace set/installation 均为 `0`
- 50,000 个历史 Query:活跃 query cache `0`,历史 scope counter `100`
- 5,000 个会话身份:session cache `200`
- 20,000 个限流身份:rate-limit container `10,000`
- 10,000 个历史 tasktask record `200`
- 5,000 次 WebSocket churnconversation 与 stream index 均为 `200`
- event-loop task、线程和文件描述符保持 `1 / 1 / 6`;使用当前本地 SDK 分支的复跑中,第二阶段相对第一阶段 RSS 增长 `2,228,224 bytes`、tracemalloc current 增长 `344,669 bytes`,总耗时 `12.559s`。Session 淘汰改为 Workspace 索引和最小堆后,同一 audit 工作量相对此前 `16.150s` 明显下降。
Populated Workspace audit 为 5,000 个 Workspace 各加载一个 Provider、LLM、Embedding、Rerank、Pipeline、Bot、KnowledgeBase 和 MCP session,然后全部推进两个 generation
- 三个阶段的活跃 provider/model、pipeline、bot、knowledge 和 MCP registry 均精确维持 `5,000`,不存在按历史 generation 增长。
- 到第三阶段,前两代的 requester、Bot adapter 和 MCP session 各 `10,000` 个全部收到确定性关闭;weak reference 断言旧代对象可被回收。
- event-loop task、线程和文件描述符保持 `1 / 1 / 6`;使用远端精确钉住 SDK 的当前复跑中,第三阶段相对第二阶段 RSS 增长 `1,245,184 bytes`tracemalloc current 仅增长 `2,061 bytes`
- 初始/第一次替换/第二次替换分别耗时 `1.893s / 2.549s / 2.659s`,最大替换耗时比为 `1.405`,未随历史代次出现 CPU 退化。
- macOS RSS sample 从初始的 `154,648,576` 增至第一阶段 `368,181,248`、第二阶段 `389,087,232` 和第三阶段 `390,332,416 bytes`;第二次替换只比第一次替换增加约 1.19 MiB,但“合法活跃租户资源的线性容量”仍必须作为 placement 容量输入。这里使用轻量 fake adapter/requester,不应把第一阶段约 204 MiB 增量外推为生产每租户成本。
Plugin SDK audit 每个阶段执行 25,000 次 loopback RPC、5,000 次安装 binding 激活/撤销、10,000 个 Workspace generation 更新和 2,500 次带 Workspace 上下文的 Box session 创建/删除。第一、第二阶段的保留状态完全一致:
- RPC waiter、stream queue、action task 和活跃 installation binding 均为 `0`
- installation watermark 为有界的 `5,000`Workspace generation record 为有界的 `10,000`,没有等待者时 generation event 为 `0`
- generation active task/index、Box session、Box Workspace session index、creating/closing/background task 和 session lock 均为 `0`
- event-loop task 和文件描述符保持 `1 / 7`;当前复跑第二阶段相对第一阶段 RSS peak 增长 `2,637,824 bytes`、tracemalloc current 增长 `289,746 bytes`,总耗时 `11.270s`。耗时增加来自本轮把大协议消息的 JSON/Pydantic、UTF-8 编码、分片和拼接移入有界线程池;结构状态和第二阶段 tracemalloc 增量保持平稳。
第二轮反向静态审查另外枚举了 Core 的 50 个显式 task 创建点和 204 个线程、阻塞调用及子进程调用点,以及 SDK 的 28 个显式 task 创建点和 62 个线程、阻塞调用及子进程调用点。第三轮独立复核继续从高基数定时器、目录遍历、准入全表扫描和取消竞态反推,新增关闭了 Plugin restart 冷却唤醒群、MCP idle 数据库轮询、nsjail orphan 的 O(session × process) 启动扫描、message aggregation 的 O(buffer) 准入及 Skill inode/文本列表边界。显式 task 均具有持有者、完成回调或 `finally` 回收路径;所有生产入口在第一次 `asyncio.to_thread()` 前安装有界默认 executor。Core、Plugin Runtime 和 Box 的公开 `/healthz`Box `/readyz` 亦同)会输出各自的 aggregate runtime/resource counter 和 event-loop lag,供 soak 对比活跃量、pending、累计 capacity rejection 与调度延迟;不输出 debug key、控制 token、租户或插件身份。Plugin Runtime 的授权 debug info 复用同一资源快照,避免公开/私有指标语义漂移。
真实 PostgreSQL populated 启动门禁会先通过 release migration 创建最新 schema,再用无 `BYPASSRLS` 的临时 Cloud Runtime 角色启动。每个 Workspace 都含九类代表性资源,测试会走实际的 instance discovery、tenant UoW、启动 binding 快照和 Model/Platform/Pipeline/RAG/MCP/Plugin 加载路径:
- 500 Workspace:启动加载 `6.178s`,进程 CPU `3.026s`
- 当前 1,000 Workspace 复跑:启动加载 `12.109s`,进程 CPU `5.967s`;相对此前 500 Workspace 的墙钟比为 `1.960`
- `model_providers``llm_models``embedding_models``rerank_models``bots``legacy_pipelines``knowledge_bases``mcp_servers``plugin_settings` 九张表的 SELECT 次数均精确等于 Workspace 数,没有重复的全租户发现或超线性资源扫描。
- MCP host dispatcher、host task 和临时 Runtime 角色/asyncpg 连接在测试结束后均清空;严格 `ResourceWarning` 模式通过。
探针要求第二阶段的结构状态与第一阶段精确相等,并对第二阶段 RSS 与 tracemalloc 增长设置失败阈值。macOS 的 RSS 来源是 `getrusage` peak,因此这里验证的是峰值增量边界而非“当前 RSS 回落”;最终 Linux 24 小时 soak 仍需采集 current RSS/PSS 和 cgroup `memory.current`
LangBot 全量测试的 33 个 skip 中,22 个是默认全量运行未提供 PostgreSQL/pgvector 而跳过的集成用例,10 个是未提供 Valkey,另 1 个是可选环境的 collection skip;真实 PostgreSQL 相关路径已由上表单独运行覆盖。Plugin SDK 的 26 个 warning 为现有 Pydantic v2 deprecation 与 aiohttp AppKey 建议;没有失败、未关闭资源或资源上限降级。Core 当前全量产生 194 个既有第三方/兼容性 warning;`ResourceWarning``PytestUnraisableExceptionWarning` 仍由 pytest 配置提升为错误,本轮没有此类泄漏告警。
Linux Runtime 探针使用上述镜像并只读挂载本地最新 SDK 源码:
- 普通容器:nsjail binary 可执行,但 namespace、mount、network 与 cgroup v2 检查均为 `false`,严格 readiness 按预期失败关闭。
- `--privileged` + private cgroup namespacenamespace、mount、network 通过,但 cgroup v2 delegation 为 `false`,仍按预期不能进入 Cloud ready。
- 一次性容器内建立可写 delegated cgroup 子树后:Plugin 与 Box cgroup 探针均为 `true`nsjail namespace、mount、network 和 cgroup v2 均通过;硬文件系统与 inode quota 继续报告 `false`
- `cpus=0.1` 的 1.0 秒 process-CPU busy loop 实际耗时 `9.13s``memory_mb=48` 下逐页提交 128 MiB 以 exit `137` 终止;`pids_limit=8` 下批量 fork 返回 `EAGAIN`。这些结果验证了 CPU、memory+swap 和 PID 的实际内核执行路径。
- 新增 `scripts/cloud_runtime_soak.py` 后,在同一 Linux 镜像的独立容器中设置 `--cpus 0.5 --memory 256m --memory-swap 256m --pids-limit 128`,工具从目标 cgroup 读回 quota `50000/100000 usec`、memory `268435456 bytes`、swap `0 bytes` 和 PID `128`。最终复跑中,32 MiB 子负载退出后的 4 秒冷却尾段 `memory.current` 稳健增长和斜率均为 `0`,平均 CPU `0.00132 cores`OOM、memory pressure、PID max 和 throttle delta 均为 `0`,最终 verdict 为 `pass`。这只是采集器/判定器自检,不替代最终 24 小时生产候选运行。
- 本地实际启动 Plugin Runtime 后,控制端口与 debug 端口的公开 `/healthz` 均返回相同聚合 JSONevent-loop monitor 为 running,且正文不含 debug key。采集器显式绕过进程级 HTTP proxy 后,对控制端口执行 6 秒短时 endpoint gate:无失败,观测到的 recent max/p95 均为 `2.233 ms`verdict 为 `pass`
- 本地实际启动 Box Runtime(未创建 sandbox session)后,`/healthz``/readyz` 均返回 event-loop、blocking executor、session/process/task 聚合快照;monitor 为 running、样本持续增长,两个端点观测到的 recent max 均为 `2.265 ms`。SIGINT 后 aiohttp、Runtime、reaper 与 monitor 走统一清理路径并正常退出。
## 上线配置与监控门禁
最终 24 小时命令、运行位置、阈值语义、负载矩阵和产物要求见 [LangBot Cloud 24 小时资源 Soak 门禁](./cloud-runtime-soak-gate.md)。该工具默认把任一健康失败、OOM/memory pressure、PID limit、CPU throttling 超阈值、blocking executor rejection、冷却尾段内存持续增长或空闲 CPU 过高判为失败;生产运行必须使用 `--require-hard-limits`
至少需要监控并告警:
- Core/Plugin Runtime/Box Runtime 的 RSS、CPU throttling、OOM、PID 数和 event-loop lag。
- 各进程 blocking executor 的 running、pending、inflight、active scopes、`global_rejected_total``scope_rejected_total`pending 持续不归零或 rejection 增长都应告警。
- QueryPool、WebSocket、session、task、plugin worker、Box session 的当前量、容量拒绝和淘汰计数。
- Plugin crash/restart 频率、dependency prepare 耗时和失败率。
- PostgreSQL pool checked-out/overflow/wait timeout、事务耗时和连接错误。
- 临时文件、artifact、dependency environment、Box Workspace volume 的字节数和 inode。
生产 soak 应覆盖租户突发登录、批量插件 reconcile、插件崩溃重启、WebSocket 断连、Box 并发执行、PG pool 饱和和应用 SIGTERM;持续运行至少 24 小时,并验证负载停止后 RSS、task、socket、文件和子进程数量回到稳定基线。
+271
View File
@@ -0,0 +1,271 @@
# Cloud v2 multi-tenant verification report
Date: 2026-07-24
Status: `FOUNDATION AND CONTROL PLANE VERIFIED — PRODUCTION ACTIVATION REMAINS GATED`
This report records the implementation and verification evidence for one
logical LangBot instance serving multiple Workspace tenants. It covers the
open-source Core, the shared Plugin/Box runtimes, the closed Space adapter, and
the Space Cloud v2 modular-monolith control plane. It does not claim that the
production Cloud deployment may enable `CLOUD_V2_ENABLED` yet.
## Repository refs and scope
- LangBot Core: branch `feat/multi-tenants`, commit
`e8a09b7537ef285a967f24add05fdb9bb557b97e`
- langbot-plugin-sdk: branch `feat/multi-tenants`, head
`ca545d079ca1657a5d4efb4e31bfeafe1a374a46`
- langbot-space: branch `feat/cloud-v2-control-plane`, head
`ce41ff370e94a405f70e2fbb2f99b0946e0e0387`
- Closed Core adapter: `langbot-space/cloud-adapter`
- SDK protocol/package version: `0.4.18`
Core pins SDK commit `e7d946af4a6b1494fbe74627c1815ace19ac8991`;
the SDK branch head adds CI-only follow-up. Cloud v2 is a greenfield
multi-Workspace deployment and does not provision one Pod, database, queue, or
Runtime per tenant. Legacy Space Pods remain a compatibility surface only.
## Implemented product and security boundaries
### Workspace and identity
- SaaS has one stable `instance_uuid`; `workspace_uuid` is the tenant key.
- Registering an Account creates its personal Workspace and owner Membership in
the same PostgreSQL transaction.
- A Workspace may contain multiple users with owner, admin, and member roles.
Invitation acceptance is token-hash based, email-bound, single-use,
concurrency-safe, and member-limit checked while locked.
- Community remains exactly one local Workspace while supporting multiple
Accounts and fixed RBAC roles. Installing the closed adapter is required to
inject the SaaS policy; configuration alone cannot activate it.
- Disabled or deleted Accounts cannot use password/code/OAuth login, sessions,
refresh/access tokens, or personal access tokens.
### Closed Space control plane
- `CLOUD_V2_ENABLED` defaults to false. Invalid or incomplete configuration
fails startup, and disabled endpoints return 503.
- Space owns Workspace/Membership/Invitation, versioned plans, subscriptions,
entitlements, usage events, directory outbox, and signed instance manifests.
- Free and Pro are Workspace plans. Pro projects one managed global sandbox;
Free projects none. Both force stdio MCP off in the first release.
- Subscription changes use Workspace advisory locking, expected entitlement
revision, one-live-order uniqueness, idempotent usage ingestion, period-end
downgrade, renewal, and lazy expiry settlement.
- Account provisioning and quota projection into New API use PostgreSQL
transactional outboxes, replica-safe claims, retry/backoff, revision fencing,
and database-enforced identity/user/token ownership. No additional service is
introduced.
- EPay and Stripe callbacks are bound to the locked order's provider identity,
amount, currency, channel/session, expiry, and provider transaction ID.
Replays are idempotent, EPay is CNY-only, credential rotation is supported by
encrypted per-order snapshots, and permanent fulfillment conflicts are
retained for reconciliation.
- Directory snapshot and per-Workspace delta reads use repeatable-read,
read-only transactions with a high-water cursor. Core stores a replica-local
consumer cursor and shared PostgreSQL projection state, snapshot coverage,
and inbox rows.
- The `/cloud` page selects a Workspace and shows its independent subscription,
entitlement, limits, and usage. When Cloud v2 is disabled or the backend does
not expose the feature field, the complete legacy Welcome/Pod UI is used.
### PostgreSQL and pgvector
- SaaS business data uses one PostgreSQL shared schema with application scope
plus forced RLS. Cloud directory writes are separated from local tenant
writes.
- Projected Account and Membership revisions are monotonic. Tombstones remove
memberships, and stale revisions cannot resurrect them.
- pgvector shares the business PostgreSQL database and remains
Workspace-scoped.
- Space runs versioned SQL migrations before application seeds. A fresh
database, a partial pre-migration database, and an existing-baseline path
converge without startup-time `AutoMigrate`.
### Shared Plugin Runtime
- One instance-scoped trusted Supervisor serves multiple Workspaces.
- Every enabled installation has its own nsjail process bound to
`(instance, workspace, execution generation, installation, runtime revision,
artifact digest)`.
- Verified same-digest code and dependency files are mounted read-only and may
be shared; home, tmp, data, process namespace, registration capability, and
cgroup are private.
- Instance configuration owns CPU, memory, PID, open-file, and file-size
limits. Plugin manifests cannot increase them.
- Memory includes swap: nsjail receives `memory.max` and `swap.max=0`.
- Unexpected worker exit is recovered by a completion callback with bounded
per-installation exponential backoff. Remove, reconcile, Runtime shutdown,
and container SIGTERM perform a graceful-to-SIGKILL bounded reap.
### Shared Box Runtime and MCP
- One shared Box control plane serves Workspace-bound logical sessions.
- Cloud grants allow at most one persistent `global` session for an entitled
Workspace and no managed processes in the first release.
- Cloud is fixed to nsjail and network-off. Core and Box prove the shared
durable Workspace mount with an authenticated marker challenge.
- stdio MCP is independently gated and forced off for Cloud v2.
- The current nsjail Box backend does not provide hard byte/inode quotas, so
Cloud readiness correctly fails closed instead of silently using a soft
directory scan.
## Automated verification
### LangBot Core
```text
uv run --no-sync pytest -q
2590 passed, 32 skipped, 177 warnings
real PostgreSQL migration, pgvector, and release-migrator suites
21 passed, 11 warnings
uv run --no-sync ruff check .
uv lock --check
git diff --check
passed
```
The full suite ran without the closed adapter installed, proving the open-source
single-Workspace/multi-user path remains standalone. Focused closed-adapter,
directory projection, runtime connector, Box cleanup, and configuration suites
also passed with the adapter installed.
### Plugin SDK and real Linux runtime
```text
SDK full suite
1226 passed, 22 existing warnings
Ruff check and format check
git diff --check
passed
```
A privileged Linux test container with host cgroup namespace ran one shared
Runtime and two Workspace installations:
- both workers referenced the same artifact inode;
- home, tmp, and data inodes were distinct;
- each plugin saw only PID 1 in its private PID namespace;
- a tampered binding and an unknown installation were rejected;
- the control token was absent from worker environments;
- cgroups were distinct with `memory.max=134217728`, `memory.swap.max=0`,
`pids.max=32`, and `cpu.max=500000 1000000`;
- touching 256 MiB exited with code 137 without swap growth;
- the 32nd fork failed with `EAGAIN`;
- reconcile and container SIGTERM removed the worker cgroups.
The same run started the Runtime from a non-root working directory, covering
absolute nsjail mount-source normalization.
### Space backend, adapter, and frontend
```text
MIGRATIONS_TEST_DSN=... MIGRATIONS_TEST_DSN_FRESH=... \
go test -count=1 ./...
go vet ./...
passed against PostgreSQL 16
fresh PostgreSQL app startup, partial-baseline migration,
Cloud v2 migration rerun, and control-plane integration
passed
closed adapter pytest and Ruff
passed
pnpm exec tsc --noEmit
pnpm check:i18n
pnpm check:cloud-checkout-currency
passed; 7 checkout/currency cases
```
The PostgreSQL checks started from an empty database and verified all 34
registered migrations in order, Cloud v2 Free/Pro seeds, legacy plan seeds,
Cloud columns/indexes, payment callback constraints, New API outbox/ownership
constraints, and repeatable reruns.
## Cross-service and browser E2E
### Signed Space-to-Core directory projection
Using an isolated Space PostgreSQL database and a migrated Core PostgreSQL
database:
1. Space issued a signed manifest for the fixed LangBot instance.
2. Space returned a directory snapshot at cursor 5 containing two Workspaces,
two projected Accounts, and owner/admin memberships.
3. Core verified the signature, instance, release, validity window, and
capability before injecting the Cloud Workspace policy.
4. Core stored both Workspaces, all active Memberships, both projected
Accounts, snapshot coverage, inbox entries, and cursor 5.
5. Account-field-only projection revisions and Workspace directory revisions
remained independent, and cross-Workspace Account conflicts failed closed.
### Real browser Cloud v2 flow
A real local browser operated the Space frontend and backend:
1. An existing legacy-Pod owner logged in and saw the automatically created
personal Workspace on Free.
2. The page showed Free and Pro Workspace plans while retaining the legacy Pro
instance card, its Online state, URL, version, billing period, and actions.
3. Annual Pro checkout through the configured EPay/Alipay rail was re-quoted
from the displayed USD plan price to `¥490.00 CNY`.
4. The browser reached the EPay gateway with `money=490.00`; no USD amount was
sent through the CNY-only rail.
5. A valid signed `TRADE_SUCCESS` callback returned `success`. An exact replay
also returned `success`, leaving one successful order and one Pro Workspace
subscription.
6. Refreshing payment state cleared the pending indicator. The page then showed
the Pro annual period and managed-sandbox entitlement while the legacy Pod
remained Online.
The browser run used the real Space UI and HTTP handlers. Its disposable local
development harness added a same-origin Next.js rewrite only in the temporary
worktree; that harness change was removed after the run.
The legacy feature-flag branch was then covered by API/static checks and the
production build: false or missing `cloud_v2_enabled` renders the original
Welcome/Pod client; a failed web-config request renders an explicit retry
instead of guessing a deployment mode.
### Deliberate Core startup failure
Core completed signed manifest verification and directory projection, then
stopped at the Box readiness gate because the current nsjail backend cannot
prove hard Workspace byte/inode quota enforcement. Connector shutdown and
reconnect tasks were cleanly reaped; no event-loop or never-awaited coroutine
warning remained.
This is a successful fail-closed acceptance result, not a passing production
Cloud boot.
## Remaining production activation gates
Cloud v2 must remain disabled until these gates are closed:
- Box provides and proves hard byte and inode quotas for Workspace, Skill,
root, home, and tmp storage.
- Plugin installation writable data receives an operator-owned hard total
disk quota.
- Plugin and future networked Box workloads have tenant-safe egress/SSRF
policy.
- Plugin Runtime adds jitter, global restart concurrency limiting, and a
Runtime-level circuit breaker, then passes systemic-failure injection.
- M0 rolls Core and Plugin Runtime together until authenticated Runtime takeover
or an owner lease/fencing protocol exists.
- Payment operations add scheduled reconciliation and alerting for stale
`processing` orders and persisted permanent fulfillment conflicts.
- Cloud v2 subscription service periods are stored immutably and included in
recognized-revenue reporting.
- Production migration Job, backup/rollback, PostgreSQL credential/network
boundaries, horizontal-replica fault injection, Workspace release/export,
deletion, and restore semantics are completed.
These gates intentionally add no tenant-specific service. They are implemented
inside the existing Space, Core, Plugin Runtime, Box Runtime, and PostgreSQL
components to preserve the architecture goal: near-zero static cost for a new
Workspace.
@@ -0,0 +1,941 @@
# LangBot Workspace 多用户与 SaaS 多租户架构
状态:`ARCHITECTURE BASELINE — isolation kernel implemented; SaaS activation gates remain`
本文描述 Cloud v2 的目标架构和安全边界。详细的 Runtime、Box、PostgreSQL、pgvector 与 stdio MCP 决策以
[pending-architecture-decisions.md](./pending-architecture-decisions.md) 为权威来源;已经落地的实现选择记录在
[implementation-decisions.md](./implementation-decisions.md)。
“隔离内核已实现”仅表示开源 Core/SDK 已具备多租户数据和运行时隔离所需的基础能力,
不表示闭源控制面、计费、生产部署或 Cloud v2 已经可以上线。
## 1. 架构决策摘要
Cloud v2 采用以下模型:
> SaaS 对外只有一个逻辑 LangBot 实例,全部 Workspace 都是该实例内的租户;
> 开源 Core 提供完整隔离内核,闭源 Cloud Control Plane 管理 SaaS 目录、订阅、权益和计费。
核心决策如下:
1. `Workspace` 是数据、成员、权限、用量和不可信执行的租户边界,不是一个 Pod、namespace、数据库或独立 LangBot 部署。
2. SaaS 注册 Account 时自动创建个人 Workspace;这只新增目录与业务记录,不创建租户专属服务、数据库、队列或 Runtime。
3. OSS 每个 LangBot 实例只能存在一个 Workspace,但该 Workspace 可以有多个 Account、邀请和固定角色。
4. SaaS 才允许一个 Account 拥有或加入多个 Workspace,并在 WebUI 中切换当前 Workspace。
5. MVP 可以各运行一个 Core、Plugin Runtime 和 Box Runtime 进程;未来增加副本或 PostgreSQL shard 仍属于同一个逻辑实例的内部扩展,不改变产品模型和外部 API。
6. 一个共享 Plugin Runtime 控制面管理所有 Workspace,但每个运行中的 plugin installation 独占一个 nsjail 进程;enabled-resident 是 desired semantics,只读代码和依赖可按已验证摘要共享。
7. 一个共享 Box Runtime 管理所有 Workspace;首期符合 entitlement 的 Workspace 最多拥有一个持久 `global` 逻辑 sandbox,实际命令继续以 nsjail 子进程执行。
8. SaaS 业务数据使用 PostgreSQL shared schema、应用层 scope 与 RLS 双重隔离;pgvector 位于同一个业务数据库并作为 SaaS 默认向量后端。
9. stdio MCP 有独立实例开关,Cloud v2 首期强制关闭,不能由 Box availability 或套餐能力隐式开启。
10. 闭源 Control Plane 可以作为模块化单体复用现有账户、支付和运营能力,但历史 Cloud 的租户专属部署模型不进入新架构。
11. Workspace 创建、释放、export、单 Workspace restore 和在线迁移的具体流程仍待后续决策。
本轮重构的最高目标是:
> 共享可信控制面和基础设施池,隔离不可信执行单元;减少独立部署和常驻组件,使新增 Account 或空 Workspace 的静态成本接近零。
减少组件数量不意味着合并安全边界。插件进程、sandbox、secret、可写文件和租户数据仍必须严格隔离。
## 2. 范围与非目标
### 2.1 本方案覆盖
- OSS 单 Workspace 多用户、邀请和固定 RBAC。
- SaaS 多 Workspace 账户、成员和 Workspace 切换模型。
- HTTP、WebSocket、API Key、Bot、Webhook、后台任务和内部调用的可信 Workspace 上下文。
- Bot、Pipeline、Provider、Knowledge、Plugin、MCP、RAG、Session、Storage 和 Monitoring 的租户隔离。
- Plugin Runtime 与 Box Runtime 的共享控制面和进程级隔离。
- SaaS PostgreSQL shared schema、RLS 与 pgvector 边界。
- 开源 Core 与闭源 Control Plane 的职责、协议和故障边界。
- 当前单副本运行和未来同一逻辑实例内横向扩展的兼容约束。
- 分阶段实施、激活门禁和验收策略。
### 2.2 本方案不覆盖
- 兼容或原地升级历史 Cloud 的租户专属部署方案。
- 为每个 Workspace 创建独立服务、数据库、schema、role、bucket、PVC、队列或 Runtime。
- 当前阶段实现多副本调度、跨地域 active-active 或 PostgreSQL 在线分片迁移。
- 第一版自定义角色、SAML、SCIM 或企业离线授权。
- 第一版 Workspace 级 BYOK E2B WebUI 配置。
- Cloud v2 首期 stdio MCP。
- Workspace export、释放、单租户恢复和在线迁移的具体产品流程。
历史客户数据、账户和财务记录如需迁移,应单独立项;旧部署拓扑不作为本架构的设计约束。
## 3. 术语与不变量
### 3.1 术语
| 术语 | 定义 |
| --- | --- |
| Account | 登录主体。OSS 中是实例本地账户;SaaS 中是全局账户 |
| Workspace | 逻辑 LangBot 实例内的租户,是资源、成员、权限、用量和不可信执行的首要边界 |
| Membership | Account 与 Workspace 的关系,包含固定角色、状态和权限版本 |
| Invitation | 邀请一个 Account 或邮箱加入 Workspace 的一次性凭证 |
| Logical Instance | 对外唯一的 LangBot 服务与安全域,拥有稳定 `instance_uuid`,不等同于某个进程或 Pod |
| Replica | Core、Plugin Runtime 或 Box Runtime 的短期内部运行副本,不是产品实体 |
| Execution Generation | Workspace 执行所有权和撤销的单调代数,用于隔离旧任务、旧连接和故障转移 |
| Billing Account | SaaS 付款主体,可以为一个或多个 Workspace 付费 |
| Entitlement | Control Plane 签发、Core 与 Runtime 本地执行的功能和数值额度快照 |
| Cloud Control Plane | 闭源 SaaS 控制面,管理全局身份、Workspace 目录、订阅、权益、计费和生命周期 |
| LangBot Core | 开源数据面,执行 Bot、Pipeline、Plugin、MCP、RAG 等业务并实施最终授权与隔离 |
当前代码中的 `placement_generation` 字段在迁移完成前保留兼容;其架构语义和目标命名均为
`execution_generation`,不表达 Workspace 属于某个产品级部署单元。
### 3.2 必须始终成立的不变量
1. SaaS 只有一个稳定 `instance_uuid`;所有副本共享该身份。
2. `replica_id``worker_id`、Pod 名称、进程地址和数据库连接地址都是短期运行信息,不能进入业务资源的永久主键或外部 URL。
3. `workspace_uuid` 是租户数据、任务、缓存、文件、日志、用量和运行时隔离的稳定键,也是未来内部路由与分片的候选键。
4. OSS 一个实例最多一个 WorkspaceSaaS 才能激活多个 Workspace。
5. 一个 Workspace 可以有多个 Account;一个 SaaS Account 可以加入多个 Workspace。
6. 所有租户业务资源都具有非空 `workspace_uuid`,并使用 `(workspace_uuid, resource_uuid)` 定位。
7. Workspace 选择器只是路由输入,不是授权凭证;服务端必须重新验证 Account、Membership、资源所有权和权限。
8. API Key、Bot、Webhook、后台任务、Plugin 与 Box 调用从可信所有权或绑定派生 Workspace,不能信任调用方自报 scope。
9. SaaS 缺少有效 Workspace 上下文时必须失败关闭,不能回退到第一个、最近或 OSS 默认 Workspace。
10. Core 是资源访问、运行时授权和 entitlement 执行的最后一道边界;Control Plane 不同步代理每条消息或普通资源请求。
11. 一个不可信插件进程只能属于一个 installation;一个 sandbox/session 只能属于一个 Workspace。
12. execution generation 失效后,旧任务、连接、回调和副作用必须被拒绝。
13. 本地进程表、缓存和临时目录都可重建,不能成为 desired state、撤销状态或业务数据的唯一真相。
14. 创建空 Workspace 不启动插件 worker、sandbox 或租户专属常驻组件。
15. 未来横向扩展不能改变 Workspace UUID、外部 API、权限模型或隔离语义。
## 4. 产品与部署模型
### 4.1 SaaS 逻辑拓扑
```mermaid
flowchart LR
User["Browser / API / Bot traffic"] --> Edge["SaaS Edge"]
User --> CP["Closed Cloud Control Plane<br/>directory + subscription + billing"]
Edge --> Core["One logical LangBot instance<br/>Core replica pool; MVP = 1"]
CP -->|"signed manifest, directory projection,<br/>entitlement and desired state"| Core
Core -->|"usage outbox and observed state"| CP
Core --> PG["Shared PostgreSQL business database<br/>RLS + pgvector"]
Core --> PluginRT["Shared Plugin Runtime<br/>trusted supervisor"]
Core --> BoxRT["Shared Box Runtime<br/>trusted supervisor"]
PluginRT --> PluginA["Workspace A installation<br/>isolated nsjail process"]
PluginRT --> PluginB["Workspace B installation<br/>isolated nsjail process"]
BoxRT --> SandboxA["Workspace A<br/>persistent global logical sandbox"]
BoxRT --> SandboxB["Workspace B<br/>persistent global logical sandbox"]
Core --> ObjectStore["Shared durable object storage<br/>Workspace-scoped keys"]
```
这里的“一个逻辑实例”是一个服务、安全域和稳定身份,不是“永远只有一个 OS 进程”。
MVP 不实现分布式,但从第一天保留内部扩展所需的身份、幂等、generation 和 owner 抽象。
### 4.2 容量演进
| 阶段 | 内部部署形态 | 新 Workspace 静态成本 | 启用条件 |
| --- | --- | --- | --- |
| M0 单副本 MVP | 一个 Core、一个共享 Plugin Runtime、一个共享 Box Runtime、一个 PostgreSQL business database | 只新增目录和业务行 | 当前目标 |
| M1 同逻辑实例横向扩展 | 按容量增加 Core/Runtime 副本;使用 owner lease、fencing 和 generationPostgreSQL 可增加 shared shard | 不创建 Workspace 专属部署 | 出现容量或可用性证据后 |
| M2 Dedicated 资源等级 | 特定 workload 使用独享 worker pool、sandbox class 或 database shard,但沿用相同身份、协议和 schema | 仅购买该等级的客户承担 | 合规、驻留或超大负载需求 |
M1 是 M0 的透明扩容,M2 是相同架构下的资源等级。外部 API 只认识稳定的
`instance_uuid``workspace_uuid`,不认识 replica、worker、pool 或 shard。
### 4.3 当前不做分布式时必须预留的能力
1. 运行时协议携带稳定 `instance_uuid``workspace_uuid``execution_generation`,不依赖进程地址表达身份。
2. Plugin installation 和 Box session 使用稳定 owner 抽象;启用第二个副本前再实现带 expiry、CAS 和 fencing token 的 lease。
3. 创建、重试、回调、worker 注册和 outbox 使用稳定 idempotency key,重复投递不能产生第二个 owner 或副作用。
4. Repository/UoW 不允许无边界跨 Workspace 事务;`workspace_uuid` 可直接作为未来 shard key。
5. schema migration、后台任务扫描、监控聚合和运维接口不能假设永远只有一个 Core 进程。
6. Runtime 重启通过 durable desired state reconciliation 恢复,不依赖原进程或本地 cache。
7. 只有出现容量、可用性、地域或合规证据后才增加副本、lease store 或 shard router;预留协议不等于提前部署组件。
### 4.4 组件边界
- Core、Plugin Runtime 和 Box Runtime 必须保持独立进程身份、容器和 security context。M0 中 Core 与 Plugin Runtime
需要处于同一 rollout/restart unit;在实现受认证 takeover 或 owner lease/fencing 前,Core 不能单独重启后接管仍存活的 Runtime。
- Core 不能继承 nsjail、cgroup 或 mount namespace 所需的高权限。
- Plugin Runtime 与 Box Runtime 不合并为一个高权限进程。
- MVP 不新增 Runtime 专用数据库、Box 专用数据库、Kafka、Redis、租户级 scheduler 或 artifact service。
- 可信 supervisor、数据库连接池、只读 artifact cache 和基础容量可以多租户共享。
## 5. OSS 与 SaaS 产品行为
### 5.1 能力矩阵
| 能力 | OSS | SaaS |
| --- | --- | --- |
| Workspace 数量 | 实例固定一个 | Account 可拥有或加入多个,受 ProductPolicy 约束 |
| Workspace 成员 | 多用户 | 多用户,受 entitlement 约束 |
| 邀请成员 | 支持 | 支持 |
| 固定 RBAC | 支持 | 支持 |
| 自定义角色 | 不支持 | 后续商业能力 |
| Workspace 创建 | 首次初始化创建唯一 Workspace | 注册自动创建个人 Workspace;后续创建受 ProductPolicy 约束 |
| Workspace 切换 | 无需展示 | 支持 |
| 订阅与计费 | 无远端依赖 | 闭源 Control Plane 管理 |
| 租户隔离 | 完整实现 | 完整实现 |
OSS edition policy 应表达为:
```text
workspace_limit = 1
members_enabled = true
invitations_enabled = true
fixed_rbac_enabled = true
multi_workspace_enabled = false
```
不能用 `member_limit = 1`、关闭邀请或移除 RBAC 来实现单租户限制。
### 5.2 OSS 初始化和邀请
首次初始化在一个事务中完成:
1. 创建本地 Account。
2. 创建实例唯一 Workspace。
3. 创建 owner Membership。
4. 创建默认 Pipeline、metadata 等 Workspace 初始资源。
5. 标记实例初始化完成。
初始化后默认关闭公开注册。后续用户由 owner/admin 创建一次性 Invitation,注册或登录后接受邀请并加入唯一 Workspace。
OSS 后续注册不创建第二个 Workspace。未配置 SMTP 时,系统返回只展示一次的邀请链接供管理员通过可信渠道发送。
### 5.3 SaaS 注册和邀请
普通注册由 Control Plane 通过幂等工作流完成:
1. 创建或确认全局 Account 与 AuthIdentity。
2. 创建 personal Workspace 和 owner Membership。
3. 创建初始 Subscription/Entitlement 投影。
4. 完成 verified email、速率限制和基础风控。
5. 将 Account、Workspace 和 Membership 投影到 Core。
6. Core 达到要求的目录 revision 后返回可访问 route。
注册只创建逻辑记录,不启动 Runtime 或租户专属基础设施。
通过邀请注册的新用户也创建自己的 personal Workspace,同时加入受邀 Workspace;已注册用户接受邀请时只新增目标 Membership。
个人 Workspace 与团队 Workspace 的付费关系必须由 ProductPolicy 明确,不允许代码根据名称或创建路径隐式推断。
### 5.4 Invitation 安全规则
- token 使用至少 256-bit 加密安全随机数,数据库只保存 hash。
- token 具有 `expires_at``accepted_at``revoked_at`,只能使用一次。
- Membership 创建与 token 消费在同一事务中提交。
- Invitation 不能授予 ownerowner 转移使用独立流程。
- SaaS 接受邀请时必须验证目标邮箱;OAuth 邮箱相同不能跳过 token 和显式确认。
- Workspace 必须始终至少有一个 active owneradmin 不能移除或降级 owner。
- 浏览器邀请链接把 secret 放在 URL fragment 中,页面读取后立即清除 fragment,并只短期保存在 `sessionStorage`
### 5.5 固定 RBAC
Core 权威定义 `owner``admin``developer``operator``viewer` 固定角色。
权限按能力划分,例如资源查看、资源管理、运行操作、成员管理、provider secret 管理、审计查看和数据导出。
规则:
- 普通资源可见性不自动授予 secret 可见性。
- 跨 Workspace 猜测资源 UUID 返回 404,不泄露存在性。
- 同 Workspace 资源存在但缺少权限时返回 403。
- 最后一个 owner 不能被删除或降级。
- 前端隐藏或禁用无权限入口只改善体验;后端仍必须执行所有授权检查。
## 6. 开源与闭源职责边界
### 6.1 LangBot Core OSS
Core 负责:
- 本地 Account、Workspace、Membership 和 OSS Invitation。
- 固定 RBAC 与单 Workspace edition policy。
- 业务资源及其 Workspace scope。
- HTTP、WebSocket、后台任务和运行时请求上下文。
- Plugin、MCP、RAG、Box、Session、Storage 和 Monitoring 隔离。
- SaaS Account/Workspace/Membership 的版本化执行投影。
- InstanceManifest、EntitlementSnapshot 和 Runtime 控制通道验证。
- 通用 capability 与数值 quota enforcement。
- UsageEvent/business outbox 和基础安全审计。
Core 是 Bot、Pipeline、Model、Knowledge、Plugin installation、MCP configuration 和 Monitoring 数据的权威来源,
也是每个业务和运行时请求的最终授权边界。
### 6.2 Closed Cloud Control Plane
Control Plane 负责:
- SaaS 全局 Account、AuthIdentity、Session、OIDC 和后续 SSO。
- SaaS Workspace、Membership 和 Invitation 的权威目录。
- Workspace 创建、暂停、归档和删除工作流。
- BillingAccount、Product、PlanVersion、Price、Subscription、Invoice、Refund 和 provider event。
- Entitlement 计算、签名与版本。
- Usage ledger、聚合、额度和欠费策略。
- 实例 manifest、release、capacity、内部 desired state 和 observed state。
- SaaS 运营后台、平台角色和高级审计。
首期不把这些职责拆成多个租户、计费和调度微服务。推荐以一个独立于 Core 的闭源模块化单体承载,
并通过模块边界复用已有账户、OAuth、支付、邮件和运营能力。历史 Cloud 的租户专属部署代码不复用。
Control Plane 不保存 Bot、Pipeline、Model 或 Knowledge 等业务内容,也不代理普通消息执行。
### 6.3 SaaS Adapter
Core 中只保留薄的协议适配层:
- 验证 InstanceManifest、Account token 和 JWKS。
- 消费 DirectoryEvent 并写入本地投影。
- 缓存并验证 EntitlementSnapshot。
- 将 UsageEvent 写入 durable outbox。
- 接收 execution desired state 并上报 observed state。
适配层不得 monkey patch ORM、绕过 Core 权限检查或在普通资源请求中同步调用 Control Plane。
### 6.4 Source of Truth
| 数据 | OSS | SaaS |
| --- | --- | --- |
| Account、Workspace、Membership | Core 本地数据库 | Control Plane 权威,Core 保存版本化投影 |
| Invitation | Core 本地数据库 | Control Plane 权威,不向 Core 投影 pending secret |
| Bot、Pipeline、Model、KB、Plugin、MCP | Core | Core |
| Subscription、Payment、Invoice、Usage ledger | 无远端依赖 | Control Plane |
| Feature 和 quota | 本地 edition policy | Control Plane 签发,Core/Runtime 验证执行 |
| Execution generation | OSS 固定本地值 | Control Plane desired stateCore 执行 |
| 运行时授权 | Core | Core 根据本地投影和 entitlement 执行 |
SaaS 不维护两套可写目录。Control Plane 是目录权威写模型;Core 只保存带 revision 的执行投影。
## 7. 控制面协议
### 7.1 InstanceManifest
仅设置 `system.edition=cloud`、环境变量或前端 feature flag 不得启用 SaaS 多 Workspace。
Cloud bootstrap 必须验证由预置根信任签名的 InstanceManifest,并据此安装闭源 Workspace policy。
Manifest 至少绑定:
```text
iss, aud, sub, jti, iat, nbf, exp
instance_uuid
release
capabilities
tenant_isolation_version
execution_generation
delegated issuers and keyset revision
```
签名错误、audience 不匹配、过期、generation 回滚或信任链缺失时必须失败关闭,不能降级为 OSS 默认 Workspace。
### 7.2 DirectoryEvent 与目录新鲜度
Control Plane 通过 transactional outbox 发布 Account、Workspace 和 Membership 的版本化事件。
Core 使用 inbox 按 `event_id` 去重,以 aggregate revision 拒绝旧写,并追踪连续应用水位。启动时读取一个 PostgreSQL
`REPEATABLE READ` 事务内生成的签名全量 snapshot;运行时先消费携带当前 high-water 的签名事件页,再只请求该页涉及的 Workspace 签名增量。
增量响应不携带新的事件 cursor,因此即使其内容已包含并发提交的后续 revision,也不能跳过尚未消费的事件。
要求:
- 事件和 batch 经过实例绑定的强认证与签名。
- 重复、乱序、延迟、断流和全量 replay 都安全。
- 删除使用 tombstone。
- 新实例先导入带 high watermark 的 snapshot,再消费增量。
- 常态目录更新成本与本页发生变化的 Workspace 数量相关,不得为每个 `directory.changed` 重新读取和投影全部 Workspace。
- 每个 Core replica 独立保存进程内消费 cursor,以确保各自的 entitlement cache 都看到事件;共享 PostgreSQL 保存投影
high-water mark、全量 snapshot coverage 和 inbox。同一事件被多个 replica 消费时,第二个 replica 验证已有 receipt
snapshot coverage 内缺少的 receipt 可以补写,coverage 之外缺失则失败关闭。只有本地 cursor 追平签名 high-water 后才续期 ready。
- projection 未就绪或落后于授权 lease 要求时,交互与自动化请求按策略失败关闭。
- SaaS pending Invitation、email 和 token hash 不进入 Core 投影。
MVP 可采用一个共享、原子且可恢复的 Control Plane store;未来多副本不能继续使用进程内状态承担一次性 token 或目录水位。
### 7.3 EntitlementSnapshot
Entitlement 使用版本化签名快照,至少绑定:
```text
instance_uuid
workspace_uuid
plan_revision
entitlement_revision
status
features
limits
nbf, exp, grace_until
```
Core 校验 issuer、audience、subject、instance、revision、时间和签名;旧 revision 不覆盖新快照。
套餐名称和价格规则只存在于闭源 Control PlaneCore 与 Runtime 只理解通用 capability 和数值限额。
Control Plane 故障时,已缓存且仍有效的快照可继续执行;过期后只能进入明确、有限的 grace 模式或失败关闭。
### 7.4 UsageEvent 与 outbox
用量事件 append-only、至少一次投递,Control Plane 按 `event_id` 去重。事件至少包含:
```text
event_id
instance_uuid
workspace_uuid
execution_generation
meter
quantity_integer
unit
source
occurred_at
entitlement_revision
schema_version
```
Core 不计算账单金额,也不在普通请求中同步扣费。业务写入与相应 business outbox 必须在同一事务中提交;
generation-aware write fence 与 outbox 原子性尚是 SaaS 激活门禁。
### 7.5 Desired state 与 observed state
闭源控制面发布版本化的 release、capacity 和 execution desired stateCore/Runtime 幂等 reconcile 并上报 observed state。
desired state 只描述同一逻辑实例内部的执行所有权和容量,不产生新的产品级实例或租户实体。
Workspace 安全状态由 directory revision 决定,订阅状态由 entitlement revision 决定,执行撤销由
execution generation 决定。三者取最严格有效状态,但任何通道都不能修改另一个通道的权威字段。
## 8. 身份、鉴权与请求上下文
### 8.1 上下文模型
租户业务入口统一解析不可变的 `RequestContext`
```python
@dataclass(frozen=True)
class RequestContext:
instance_uuid: str
workspace_uuid: str
execution_generation: int
principal_type: str
principal_uuid: str
permissions: frozenset[str]
auth_method: str
entitlement_revision: int | None
request_id: str
```
不同入口的 Workspace 来源:
| 入口 | Workspace 来源 |
| --- | --- |
| Browser Account token | `X-Workspace-Id` 只作候选;服务端校验 Membership |
| API Key | key 记录绑定的 Workspace,忽略 caller selector |
| Public Bot / Webhook | Bot 或 webhook route 的可信所有权 |
| Background job | durable payload 中的完整 scope,执行前重新验证 generation |
| Plugin Host API | 认证控制连接和 immutable action context |
| Box operation | 已验证 entitlement、admission grant 和 Runtime namespace |
| System operation | 显式、最小能力的 SystemContext,禁止隐式全局上下文 |
禁止从模块全局变量、进程默认 Workspace、请求 payload 或“第一个 Workspace”推断 scope。
### 8.2 Account token 与 Workspace discovery
- 新 JWT 使用稳定 Account UUID 作为 `sub`,并绑定 issuer、当前 `instance_uuid` audience 和 expiry。
- 账户级 Workspace discovery 是一个窄 bootstrap capability,只列出该 Account 的 active Membership,不能执行租户业务。
- multi-Workspace 模式下,tenant route 缺少 selector 必须拒绝;OSS singleton 模式可由 policy 选择唯一 Workspace。
- Account token 不直接证明任一 Workspace 权限;Membership 必须在服务端解析并验证状态与 revision。
### 8.3 API Key、WebSocket 与长任务
- API Key 只持久化 hashraw secret 仅返回一次;记录绑定 Workspace、固定 scopes、状态、expiry 和 creator。
- Dashboard WebSocket 在升级后认证,并在每条入站消息前重新验证 Account、Membership、权限、资源所有权和 generation。
- 长时间 LLM、MCP、Plugin 或 Box 调用在产生副作用或接受结果前再次校验 execution generation。
- 临时凭证交换绑定发起者、Workspace、instance 和 generation;其他 scope 查询返回与不存在相同的 404。
### 8.4 错误语义
| 场景 | 语义 |
| --- | --- |
| 未认证或 token 无效 | 401 |
| 同 Workspace 资源存在但权限不足 | 403 |
| 资源不存在或属于其他 Workspace | 404 |
| edition / entitlement / quota 禁止 | 稳定领域错误码,不伪装为 500 |
| execution generation 过期 | fail closed,并停止旧运行态 |
| 未处理异常 | 稳定 `internal_error` + request ID;细节只进入服务端日志 |
## 9. Core 数据模型
### 9.1 Account、Workspace 与 Membership
核心实体至少包含:
```text
Account
uuid
email_normalized
display_name
status
auth bindings
Workspace
uuid
name
status
source: local | cloud_projection
directory_revision
WorkspaceExecutionState
workspace_uuid
instance_uuid
execution_generation
status
write_fenced_at
revision
WorkspaceMembership
workspace_uuid
account_uuid
role
status
directory_revision
```
约束:
- Membership 对 `(workspace_uuid, account_uuid)` 唯一。
- Workspace 的 source 不允许通过可变本地配置从 local 升级成 cloud projection。
- Cloud projection 只有在 manifest、instance binding、目录 revision 和 execution state 均有效时才可路由。
- OSS bootstrap 只创建或修复 local singleton Workspace。
### 9.2 Invitation
OSS Invitation 存在 Core 本地数据库;SaaS Invitation 只存在于闭源目录。
```text
WorkspaceInvitation
uuid
workspace_uuid
email_normalized
role
token_hash
expires_at
accepted_at
revoked_at
created_by
```
数据库约束必须保证同一 Workspace 与邮箱只有一个有效邀请,并保证 token hash 全局唯一。
### 9.3 业务资源
所有租户资源显式包含 `workspace_uuid`,包括但不限于:
- Bot、Pipeline、Provider、Model、Knowledge Base 和 vector record。
- Plugin installation、MCP configuration、API Key 和 webhook binding。
- Query、Message、Session、Monitoring、Usage 和 AuditEvent。
- Upload、ObjectRef、Skill、Runtime desired state 和 temporary credential session。
唯一键、索引、缓存 key、object key、日志维度和幂等键都必须包含 Workspace scope。
服务层不得暴露可绕过 Workspace 条件的普通 `get(id)``list()``delete(id)`
### 9.4 防御性约束
- tenant table 的 `workspace_uuid` 非空并有外键。
- SaaS PostgreSQL 关键表启用并强制 RLS。
- 需要全局唯一的 opaque token 使用 hash 唯一索引,不依赖 Workspace 内唯一。
- owner 保底、Membership revision、invitation one-shot 等规则同时由 service 和数据库事务保护。
- 任何跨 Workspace 运维操作必须走显式受审计的 system capability,不得复用普通 repository。
## 10. PostgreSQL、pgvector 与存储
### 10.1 数据库边界
- OSS 继续默认 SQLite,并可显式选择自托管 PostgreSQL。
- SaaS 使用一个 PostgreSQL business database、一个 `public` shared schema 和共享连接池。
- 创建 Workspace 不创建 database、schema、role 或专属连接池。
- 每个 tenant transaction 使用 `SET LOCAL` 建立 scope,并由统一 TenantUnitOfWork 保证 context 与 SQL 使用同一事务和连接。
- 应用层 Workspace scope 是第一道边界,`ENABLE` + `FORCE ROW LEVEL SECURITY` 是第二道边界。
- runtime role 必须是非 owner、最小权限、无 superuser、无 `BYPASSRLS`、无 role membership 和跨 schema 权限。
- schema、extension、policy 和 ACL 只由独立 release migrator 创建与验证;Cloud runtime 不执行 DDL。
- PostgreSQL 仅承载业务数据和 pgvector,不成为 Plugin/Box 通用协调数据库、进程目录或新的控制面数据库。
首期 migrator 和 runtime URL 必须连接同一个 host、port、database,但使用不同 role。
生产部署还必须证明 runtime credential 无法连接 PostgreSQL 集群中的其他 database;专用 endpoint 或经验证的 HBA/proxy 隔离仍是激活门禁。
### 10.2 Transaction 与后台任务
- 一个 TenantUnitOfWork 只绑定一个 Workspace、一个 execution generation 和一个事务所有者任务。
- 子任务不能继承并提交、回滚或关闭父任务的 tenant session。
- 长时间 LLM 或网络等待不持有数据库连接;每次数据库 helper 打开短事务。
- detached task 只在父事务提交后启动,并自行建立新 scope;父事务回滚时取消待启动任务。
- generation-aware write fence 必须保持到 commit,并与 business outbox 原子提交;该能力完成前不得激活 SaaS 写流量。
### 10.3 pgvector
- SaaS 默认使用同一业务 PostgreSQL 中的 pgvector,不静默回退到 Chroma。
- 向量身份至少为 `(workspace_uuid, knowledge_base_uuid, vector_id)`
- 向量操作使用相同 tenant context 与 RLS 契约。
- embedding 维度显式存储和校验;不匹配时失败关闭,不截断、补齐或改用无界扫描。
- extension、表、constraint 和 ANN index 由 release migration 创建。
- OSS 默认仍可使用 SQLite + Chroma;选择 pgvector 时遵守相同 scope。
### 10.4 Object storage
- 大对象、plugin artifact、upload、knowledge 文件和 sandbox 文件不作为 PostgreSQL blob 存储。
- durable object key 和 metadata 都包含 Workspace scope;临时 staging 可包含 generation,但稳定业务引用不能因未来 generation 切换而永久失效。
- 现有 generation-scoped opaque key 在固定 generation 的 OSS 中安全,但 Cloud cutover 前必须实现稳定 final identity 或原子引用迁移。
- public image 与 private document 使用不同 capability;不能把通用 upload key 当作公开读取凭证。
## 11. Plugin Runtime
### 11.1 共享 supervisor、独立 worker
整个逻辑实例共享一个可信 Plugin Runtime 逻辑控制面;M0 由一个 supervisor replica 承担。新 Workspace 不创建专属 Runtime、连接、卷或进程。
每个运行中的 plugin installation 独占一个 nsjail worker process treeenabled-resident 是 desired semantics。worker 运行期间永久绑定:
```text
instance_uuid
workspace_uuid
execution_generation
installation_uuid
runtime_revision
artifact_digest
```
插件不能通过 payload、Host API 参数、环境变量或重连改变该绑定。Supervisor 不在自身解释器中加载第三方插件代码。
停用、删除、revision/generation 变化或 entitlement 撤销时,旧 worker 必须停止并失去 Host API 权限。
### 11.2 文件和进程边界
```text
data/plugin-runtime/
├── artifacts/sha256/<artifact_digest>/code/ # 已验证、只读共享
├── environments/sha256/<environment_digest>/ # 原子发布、只读共享
└── installations/<installation_uuid>/
├── home/ # 私有可写
├── tmp/ # 私有可写
└── data/ # 私有持久数据
```
- 同插件同版本只有在 package digest 完全相同且完整性已验证时才共享只读代码。
- dependency environment key 包含 artifact/requirements digest、Python ABI、Runtime version 和 installer schema。
- installation 进程、配置、secret、home、tmp、data 和日志永不合并。
- namespace、private `/proc`、mount、PID、IPC、UTS、cgroup 与 rlimit 阻止读取其他文件、枚举或 signal 其他进程。
- Cloud 不从 artifact 自动加载 `.env`;secret 只由可信控制面按 installation 注入。
- 插件 egress 必须阻止访问 Core loopback、Box Runtime、数据库和平台 metadata endpoint。
### 11.3 统一资源上限
资源限制只来自实例级 `data/config.yaml`,并支持现有环境变量覆写;plugin manifest 不能声明、放宽或覆盖。
```yaml
plugin:
worker:
max_cpus: 1.0
max_memory_mb: 512
max_pids: 128
max_open_files: 256
max_file_size_mb: 512
require_hard_limits: true
```
CPU、内存和 PID 使用 cgroup 硬限制,open files 和单文件大小使用 rlimit。
Cloud deployment profile 强制 nsjail;硬限制不可用时 readiness 失败,不能降级为普通子进程。
installation 总磁盘配额需要可原子拒绝写入的 quota provider,不能以目录扫描冒充硬限制。
### 11.4 Desired state 与恢复
- PostgreSQL 中的 installation desired state 与 durable binary storage 是权威状态。
- Runtime 本地进程表、nsjail 目录、artifact/venv cache 都可重建。
- Runtime 重连执行实例范围 full reconciliation,清理 stale worker 并恢复 enabled installation。
- dependency preparation 失败记录在对应 installation,不启动半就绪 worker,也不阻塞其他 installation。
- desired semantics 要求 enabled installation 常驻,不做 idle eviction;是否按负载回收以后再决定。
- 当前 Supervisor 已在意外退出时通过 completion callback 和有界指数 backoff 恢复 enabled worker。
Cloud 激活前仍需加入 jitter、全局重启并发上限和 Runtime 级 circuit breaker,并验证系统性故障不会形成跨租户重启风暴。
真实 Linux/nsjail/cgroup 与受控 egress 的 Cloud 部署验证尚未完成,是生产激活门禁。
## 12. Box Runtime 与 stdio MCP
### 12.1 共享 Box 控制面
整个逻辑实例共享一个可信 Box Runtime 逻辑控制面;M0 由一个 Runtime replica 承担。Core 与 Runtime 控制通道绑定稳定 instance identity
每个 operation 绑定 `workspace_uuid``execution_generation`、session revision 和短期 admission grant。
首期 entitlement 模型:
```json
{
"features": {
"managed_sandbox": true,
"external_sandbox": false
},
"limits": {
"managed_sandbox_sessions": 1
}
}
```
闭源订阅模块把套餐映射为该通用 capabilityCore 与 Runtime 不判断 `plan == pro`
预期 Pro 得到 `managed_sandbox_sessions = 1`,其他套餐为 `0`
### 12.2 Sandbox 模型
- 合资格 Workspace 首次使用时懒创建一个持久 `global` 逻辑 session。
- `global` 表示 Workspace 内默认逻辑 sandbox,不表示跨 Workspace 共享。
- session TTL 不自动回收;Runtime 重启后进程和临时目录失效,但 `/workspace` 持久数据保留。
- 每次普通命令在 Box Runtime 容器内启动一个 one-shot nsjail 子进程。
- 首期禁止 managed background process 和 network,避免 session 被当成常驻共享主机。
- Core 与 Runtime 通过认证 random-marker challenge 证明看到同一 durable volume,不能只比较路径字符串。
- 文件同步、attachment 和 skill mount 沿用现有 nsjail 机制,但所有 host path 解析必须由可信 Workspace context 派生并防止 symlink/path escape。
Cloud readiness 必须证明 cgroup、namespace、mount、Workspace/Skill/ephemeral byte quota 和 inode quota 均为硬限制。
当前普通 nsjail backend 不具备全部硬磁盘能力,因此 Cloud Box 应失败关闭,直到绿地部署提供并验证真实 quota provider
不能把软目录扫描写成“生产已就绪”。
### 12.3 外部 E2B
非 Pro 用户后续可在 WebUI 配置 Workspace 自有的远程 E2B sandbox。该功能尚未实现,首期不纳入。
未来 credential 必须属于 Workspace、加密存储且读取受 secret 权限保护,不消耗 Cloud managed sandbox 配额。
### 12.4 stdio MCP 独立开关
```yaml
mcp:
stdio:
enabled: true
```
- OSS 默认 `true` 保持兼容。
- Cloud v2 通过 `MCP__STDIO__ENABLED=false` 强制关闭。
- 该 gate 独立于 `box.enabled`、managed sandbox entitlement 和 session quota。
- gate 同时覆盖 create、update、test、bootstrap load 和最终 Runtime execution。
- 已有 stdio 配置在 gate 关闭时保留但不启动,并返回明确的 feature-disabled 错误。
- HTTP/SSE 等远程 MCP transport 不受影响。
## 13. HTTP API 与 WebUI
### 13.1 Core API
OSS 与 SaaS 执行面共用通用 Workspace API
```text
GET /api/v1/workspaces
GET /api/v1/workspaces/{workspace_uuid}
GET /api/v1/workspaces/{workspace_uuid}/members
POST /api/v1/workspaces/{workspace_uuid}/invitations
PATCH /api/v1/workspaces/{workspace_uuid}/members/{account_uuid}
DELETE /api/v1/workspaces/{workspace_uuid}/members/{account_uuid}
```
Cloud policy 下,目录 mutation 由闭源 Control Plane 负责;Core 对本地创建、邀请和成员修改返回稳定的
`control_plane_required`,只提供执行投影的安全读取。
所有 tenant resource route 必须经过统一 decorator/middleware
1. 认证 principal。
2. 解析可信 Workspace。
3. 校验 Workspace/ExecutionState。
4. 校验 Membership 或资源绑定。
5. 校验 permission 和 entitlement。
6. 创建 RequestContext 与 TenantUnitOfWork。
### 13.2 SaaS Control Plane API
SaaS 产品 API 包含:
```text
POST /cloud/workspaces
GET /cloud/workspaces
POST /cloud/workspaces/{workspace_uuid}/invitations
POST /cloud/invitations/{token}/accept
GET /cloud/workspaces/{workspace_uuid}/subscription
POST /cloud/workspaces/{workspace_uuid}/checkout
GET /cloud/workspaces/{workspace_uuid}/usage
```
这些 API 管理目录、产品和计费,不直接操作 Bot/Pipeline 等 Core 业务资源。
### 13.3 WebUI
OSS
- 首次注册进入唯一 Workspace。
- owner/admin 可邀请成员并管理固定角色。
- 不展示 Workspace 切换器和创建第二 Workspace 的入口。
SaaS
- 登录先获取 Account 级 Workspace 列表,再显式选择当前 Workspace。
- 当前 Workspace UUID 保存在受控客户端状态中;所有 tenant request 自动附带 selector。
- 切换 Account 或 Workspace 时清理缓存、WebSocket、上传、表单、错误和 optimistic state,不能显示前一租户数据。
- 页面 refresh、新 tab 和邀请跳转恢复同一个经过授权的 Workspace;失效 Membership 不回退到其他 Workspace。
- UI 权限变化必须响应式更新,但 API 仍是最终授权边界。
## 14. 故障、安全与降级
### 14.1 Fail-closed 场景
以下情况必须拒绝新的租户业务和副作用:
- Cloud manifest 缺失、签名失败、audience 错误或回滚。
- Account token、Membership、Workspace status 或 execution generation 无效。
- 目录投影未就绪或落后于有效 lease 要求。
- Entitlement 缺失、过期且不在明确 grace 范围内。
- Runtime 控制通道认证失败或实例绑定不一致。
- Plugin nsjail/cgroup hard limit 在 Cloud profile 下不可用。
- Box 的任一硬存储或 namespace capability 无法证明。
- PostgreSQL RLS、runtime role、schema、catalog 或 endpoint 隔离校验失败。
- stdio MCP 在 Cloud profile 下被尝试启用。
不能把上述错误静默降级为 OSS singleton、普通子进程、Chroma、软 quota 或 caller-supplied Workspace。
### 14.2 撤销语义
- Membership 删除或降权必须影响下一次 HTTP 请求,并使长连接在下一条消息前重新授权。
- Workspace 暂停禁止新交互、自动化工作负载和新副作用;恢复只允许当前 generation。
- entitlement 到期按 capability 明确停止新创建或新执行,不隐式删除已有数据。
- generation 变化使旧 worker、session、callback、cached runtime object 和 outbox publisher 失效。
- 控制面暂时不可达时,只能在有效签名快照和本地投影允许的范围内继续;过期后失败关闭。
### 14.3 安全清单
- 所有 identifier 使用不可猜 UUID,但不把随机性当成授权。
- 所有 token/secret 只存 hash 或加密值,raw secret 一次展示。
- 日志、trace、metric、cache 和 object key 都包含 Workspace 维度并过滤 secret。
- Provider、Bot、Plugin、MCP 配置的 read response 递归遮蔽 credential。
- Runtime control、debug、registration 和 attachment capability 分离,不能复用万能 secret。
- untrusted code 不访问 Core loopback、数据库、其他 Runtime、宿主文件系统或 metadata endpoint。
- bulk operation、后台扫描和 monitoring 聚合使用显式 tenant/system capability。
- 所有跨 Workspace 运维操作记录 principal、reason、scope、request ID 和结果。
## 15. 实现状态与 SaaS 激活门禁
### 15.1 已实现的隔离内核
当前分支已经实现或具备基础的部分包括:
- OSS singleton Workspace、多 Account、Invitation 和固定 RBAC。
- trusted RequestContext、Workspace-scoped repository 和资源所有权检查。
- tenant-aware Plugin SDK protocol 与 Runtime installation binding。
- shared Plugin Runtime / Box Runtime 控制协议和 execution generation fence。
- stdio MCP 独立 gate。
- PostgreSQL shared schema、transaction-local scope、FORCE RLS 与 pgvector adapter。
- Cloud bootstrap 默认不可由普通配置激活,并对缺失安全能力失败关闭。
这些是代码能力边界,不等于完成闭源 SaaS 产品或生产部署验收。
### 15.2 尚未完成的激活门禁
以下事项完成并取得真实环境证据前,不得宣称 Cloud v2 production-ready
1. 闭源 Control Plane 的全局目录、注册、邀请、订阅、计费、entitlement 签发和签名 manifest bootstrap;横向扩展前 OAuth exchange 与目录投影还必须使用原子共享存储。
2. 普通业务写入贯穿 commit 的 generation-aware fence,以及与外部副作用同事务的 business outbox。
3. generation cutover 后稳定的 durable object identity 或原子对象引用迁移。
4. 所有 tenant-configurable outbound URL 的 SSRF 防护与 tenant-safe egressPlugin Runtime 还需在真实 Linux/nsjail/cgroup v2 环境验证 namespace、资源限制和文件隔离。
5. Plugin Runtime 已实现意外退出 worker 的 completion callback、有界 backoff 和自动恢复;Cloud 激活前增加全局重启风暴抑制并完成故障注入验证。
6. Plugin installation data 的 production hard disk quota provider,能够在写入边界原子拒绝超额,不能以目录扫描代替。
7. Box Runtime 的 production hard quota provider,包括 Workspace、Skill、root/tmp/home 的 byte 与 inode quota;真实部署还必须在启动和重连时通过共享卷 marker challenge。
8. PostgreSQL runtime credential 的专用 endpoint 或 HBA/proxy 跨 database 隔离证明、生产 migration/rollback 流程,以及 legacy pgvector migration 失败后精确恢复 RLS/FORCE 并可安全重试的集成证据。
9. 闭源目录事件、lease、snapshot、entitlement 和 usage/outbox 的重放、断流与灾难恢复验证。
10. 真实浏览器多 Account/RBAC/邀请/刷新场景已完成;仍需生产 Runtime 重启、worker crash、断流、异常回滚和闭源 Control Plane 的 fault-injection 验收。
### 15.3 有意暂缓的产品决策
- Workspace 创建后的休眠、释放、删除和保留策略。
- Workspace export 与单 Workspace restore。
- 非 Pro Workspace 的 BYOK E2B WebUI。
- 多副本 owner lease 的 store、TTL、fencing token 和转移顺序。
- PostgreSQL shard resolver、在线迁移和 dedicated shard 产品规则。
- artifact/cache 的签名来源、撤销、GC 和磁盘配额机制。
- custom roles、SSO、SCIM 和企业合规能力。
暂缓项不得被实现代码用隐式默认值提前固化。
## 16. 实施顺序
### Phase 0:契约和基线
- 固定术语、RequestContext、角色矩阵、edition policy 和错误语义。
- 建立升级备份、回滚和跨租户负向测试基线。
### Phase 1OSS tenancy kernel
- Account、Workspace、Membership、Invitation。
- singleton bootstrap、多用户邀请、RBAC 和前端权限。
### Phase 2:数据与入口隔离
- 为所有资源补充 Workspace scope。
- HTTP、API Key、Bot、Webhook、WebSocket、后台任务和 storage 统一上下文。
- SQLite migration recovery 与 PostgreSQL RLS 集成测试。
### Phase 3Runtime 与 SDK 隔离
- Plugin installation binding、nsjail、资源上限和 artifact replay。
- Box admission、session namespace、skill/attachment 文件边界。
- MCP gate、RAG/vector 与 long-running generation revalidation。
### Phase 4:闭源 SaaS 控制面
- signed manifest bootstrap。
- 全局目录、注册、邀请、Subscription、Entitlement 和 Usage ledger。
- projection、lease、outbox、reconciliation 和运维后台。
### Phase 5:生产部署激活
- 真实 Linux Plugin/Box hard isolation。
- PostgreSQL credential、migration、backup 和 rollback 验证。
- 完整浏览器/API/Runtime E2E 和故障注入。
- 所有激活门禁通过后才开启多 Workspace Cloud policy。
### Phase 6:同逻辑实例内部扩展
- 有容量证据后增加副本、owner lease 和 fencing。
- 有地域、合规或规模证据后增加 shared/dedicated shard。
- 保持外部身份、API 和 Workspace URL 不变。
## 17. 测试与验收
### 17.1 数据隔离
- 两个 Workspace 使用相同 resource UUID、name、vector ID 和 cache key,不发生冲突或越权。
- 故意遗漏应用层 Workspace filter 时,PostgreSQL RLS 仍阻止跨租户读写。
- 连接池复用、异常回滚、子任务、后台任务和 transaction pooling 不残留 tenant context。
- 跨 Workspace 猜测返回 404;同租户缺权限返回 403。
### 17.2 产品行为
- OSS 首个 Account 创建唯一 Workspace;第二个 Account 只能通过邀请加入;创建第二 Workspace 返回 edition error。
- 邀请覆盖有效、已使用、撤销、过期、邮箱不匹配和并发接受。
- owner/admin/developer/operator/viewer 的 API 和 WebUI 权限一致。
- SaaS 普通注册和邀请注册都创建个人 Workspace,但不创建专属部署或 Runtime。
### 17.3 Runtime
- 两个 Workspace 安装同一已验证 artifact 时只共享只读 code/env,进程、secret、home/tmp/data、日志和 Host API 完全隔离。
- cgroup、rlimit、namespace、egress 和 generation fence 在真实 Linux 环境生效。
- Runtime restart/cache loss 通过 durable desired state 与 binary storage 恢复。
- 两个 Workspace 的 Box session、files、process、skill、attachment 和 quota 完全隔离。
- stdio MCP gate 对 UI、API、bootstrap 和最终 execution 同时生效。
### 17.4 Control Plane 与故障
- DirectoryEvent 重复、乱序、缺口、snapshot + replay 和过期 lease 均安全。
- Entitlement 旧 revision、签名错误、过期和撤销均失败关闭。
- UsageEvent 重放不重复计费;业务事务回滚不发送副作用。
- Runtime、Core 或 Control Plane 重启不创建重复 Workspace、worker 或 sandbox。
- manifest、数据库安全校验或 hard quota 缺失时实例保持不可激活,而不是静默降级。
### 17.5 浏览器端到端
真实浏览器至少覆盖:
1. clean database 首位 owner 注册与 singleton Workspace bootstrap。
2. owner 创建邀请,第二个用户注册/登录并接受。
3. 角色在 viewer/operator/developer/admin 间变化时,导航、控制项和 API 结果同步变化。
4. Account/Workspace 切换清空前一 scope 状态,refresh 和新 tab 恢复正确 Workspace。
5. 第二 Workspace edition limit,以及 invitation used/revoked/expired/email mismatch 的可见错误。
6. 直接 API 越权、伪造 selector 和跨租户 UUID 猜测不能绕过 UI。
## 18. 最终结论
Cloud v2 的产品模型只有一个逻辑 LangBot 实例和实例内多个 Workspace。
当前选择单副本 MVP 是为了减少组件和新增租户成本,不是把单进程假设写进业务身份或协议。
未来需要容量或高可用时,在同一逻辑实例内部增加 Core/Runtime 副本和 PostgreSQL shard
Workspace 的 UUID、权限、数据边界和外部 API 均保持不变。
开源 Core 必须完整实现安全的 Workspace 隔离和 OSS 单 Workspace 多用户;闭源 Control Plane
管理 SaaS 的全局目录、订阅、权益、计费和生命周期。共享可信控制面、连接池、只读 artifact 和数据库组件,
同时让每个不可信插件进程、sandbox、secret、可写文件和 tenant transaction 保持独占边界,
才能在不增加每租户部署的前提下最大化降低新增用户成本。
在闭源控制面、事务 fence/outbox、真实 Runtime hard isolation、Box hard quota 和 PostgreSQL 生产隔离等门禁完成之前,
本架构仍处于隔离内核阶段,不应被描述为可上线的 SaaS 多租户部署。
+18 -19
View File
@@ -1,6 +1,6 @@
# Box 系统架构深度分析
> 更新日期: 2026-07-12
> 更新日期: 2026-06-02
> 状态更新: 自部署社区版已具备发布条件(box 可选、降级完善、无迁移欠债);工具调用循环上限、配额遍历异步化、`host_path` 挂载白名单等已落地。剩余多租户 / 安全硬化项见 [SaaS 阻塞项清单](./box-issues.md)。
> 分支: `feat/sandbox` (LangBot + langbot-plugin-sdk)
> 相关文档: [SaaS 阻塞项](./box-issues.md) | [Session 作用域](./box-session-scope.md) | [Runtime 对比](./box-vs-plugin-runtime.md) | [测试覆盖](./box-test-coverage.md) | [toB 分析](./box-tob-analysis.md)
@@ -13,9 +13,7 @@
┌──────────────────────────────────────────────────────────────────┐
│ LangBot 主进程 │
│ │
│ AgentRunner ──> SDK call_tool / scoped MCP bridge
│ │ │ │
│ └────────────────> ToolManager ──> NativeToolLoader │
LocalAgentRunner ──> ToolManager ──> NativeToolLoader
│ │ │ │ │
│ │ │ exec / read / write / edit │
│ │ │ glob / grep │
@@ -34,7 +32,7 @@
│ ├─ Host mount 校验 (allowed_mount_roots) │
│ ├─ Workspace quota 检查 │
│ ├─ 输出截断 (head+tail) │
│ ├─ Host scope 哈希 (resolve_box_session_id)
│ ├─ Session ID 模板解析 (resolve_box_session_id) │
│ ├─ 技能挂载组装 (build_skill_extra_mounts) │
│ ├─ 重连循环 (_reconnect_loop, 指数退避) │
│ └─ BoxRuntimeConnector │
@@ -87,8 +85,7 @@
**核心设计原则**:
- Box Runtime 作为独立进程运行,通过 Action RPC 与 LangBot 主进程通信,两者复用 SDK 的 IO 层(Handler → Connection → Controller
- 一个 session_id 对应一个容器/沙箱实例。同一 session 内可并存多条 mount 与多个 managed process
- AgentRunner 无权指定 session scope。SDK/Python `call_tool` 与 scoped MCP bridge 都发出同一个 `PluginToRuntimeAction.CALL_TOOL`,最终由 Host 的 ToolManager 执行,并使用当前 run 保存的同一个 execution Query
- Box 内托管的 stdio MCP server 使用独立的长期 `mcp-shared` session;它不是 AgentRunner 本次事件的 sandbox session(详见 [box-session-scope.md](./box-session-scope.md)
- Skill / 默认 exec / MCP Server 共享同一个 session 容器(详见 [box-session-scope.md](./box-session-scope.md)
---
@@ -96,7 +93,7 @@
### 2.1 BoxService (`pkg/box/service.py`, 722 行)
应用层门面,协调 Profile、安全校验、配额、连接、Skill 挂载与 Host scope 哈希
应用层门面,协调 Profile、安全校验、配额、连接、Skill 挂载与 Session 模板
主要公开方法(按定义顺序):
@@ -107,7 +104,7 @@ BoxService
├─ _reconnect_loop(connector) 指数退避重连
├─ available (property) 连接状态
├─ resolve_box_session_id(query) 哈希 Host 私有 scope,生成固定长度 session_id
├─ resolve_box_session_id(query) 从 pipeline 模板解析 session_id
├─ build_skill_extra_mounts(query) 组装 pipeline-bound skill 的挂载列表
├─ execute_tool(parameters, query) Agent 调用 exec 时的入口
@@ -140,8 +137,6 @@ BoxService
**输出截断**: 默认 4000 字符上限,保留前 60% + 后 40%,中间插入 `[...truncated...]`
**Session 所有权**: `resolve_box_session_id(query)` 只接受 Host 已确定的私有 scope 或 Query launcher/session identity,并输出 `lb-box-` + 64 位小写 SHA-256 十六进制摘要(固定 71 个 ASCII 字符)。哈希输入是 canonical JSON,包含 instance、workspace、bot、platform adapter、target type/id 与 thread;原始用户、群组、conversation 或 event id 不会出现在 Box session id 中。相同 Host scope 稳定复用,不同 target/thread/workspace/bot/adapter/instance 相互隔离;缺少可用 identity 时 fail closed。Pipeline、Agent 或 AgentRunner 配置都不能覆盖该规则。
**Skill 挂载合并**: `execute_tool()` 调用时,`build_skill_extra_mounts(query)` 会把当前 pipeline-bound 的所有 skill 的 `package_root` 作为 `extra_mounts` 加入 BoxSpec,挂在 `/workspace/.skills/<name>`。LLM 通过 `activate` 工具显式激活某个 skill 后,工具调用才允许引用这个 skill 的虚拟路径。
### 2.2 BoxRuntimeConnector (`pkg/box/connector.py`, 357 行)
@@ -419,9 +414,7 @@ ToolManager.initialize()
1. 验证 skill 已激活
2. 单次 exec 只能引用一个 skill 包
3. 若 skill 是 Python 项目(有 `requirements.txt``pyproject.toml`),命令会被 venv bootstrap 包裹(在 skill 挂载点内创建 `.venv`
4. 调用 `box_service.execute_tool()` → 走 Host 从当前事件生成的 session_id 与已组装好的 `extra_mounts`**不再为每 skill 起独立 session**
AgentRunner 可以直接通过 SDK/Python `AgentRunAPIProxy.call_tool` 调用这些工具,也可以让外部 harness 通过 SDK-owned scoped MCP bridge 回调。两条入口都发送 `PluginToRuntimeAction.CALL_TOOL`,共享同一个 run authorization、Host session 中保存的 execution Query、ToolManager 与 `resolve_box_session_id(query)` 规则;Runner 不能提交自定义 Box session id。Pipeline run 保存原 Query;纯 EBA run 由 Host 构造 `pipeline_config=None``pipeline_uuid=None` 的最小 Query。
4. 调用 `box_service.execute_tool()` → 走默认 session_id 与已组装好的 `extra_mounts`**不再为每 skill 起独立 session**
### 4.3 MCP-in-Box (`mcp_stdio.py`, 354 行)
@@ -429,7 +422,7 @@ AgentRunner 可以直接通过 SDK/Python `AgentRunAPIProxy.call_tool` 调用这
```
initialize()
1. 复用/创建共享 session (`session_id = mcp-shared`)
1. 复用/创建共享 session (session_id = _build_box_session_id())
- persistent=True,长期保持
2. workspace.execute_raw(install_cmd) 安装依赖 (可选)
3. 将每个 MCP server 文件 stage 到 /workspace/.mcp/<process_id>/
@@ -442,8 +435,6 @@ initialize()
每条 MCP server 是同一 session 中的一个 managed process,独立的 `process_id`、独立 attach URL,互不阻塞。
这里的 `mcp-shared` 只承载 LangBot 管理的 stdio MCP server 进程。AgentRunner 的 scoped MCP bridge 是回调 Host 工具的协议入口,不会把事件运行的 exec/read/write 改到 `mcp-shared`
---
## 5. 启动与生命周期
@@ -575,14 +566,22 @@ volumes:
| http/sse MCP server | 正常 | 正常(不依赖 Box) |
| Skill 列表/读取 (`list_skills`/`get_skill`/`read_skill_file`) | 走 Box runtime | 走 LangBot 本地 `data/skills/` 只读 fallback |
| Skill 创建/编辑/安装/写文件 | 走 Box runtime | **HTTP 400** + 明确错误信息(`_require_box_for_write`) |
| Pipeline AI 配置中 `box-session-id-template` | 正常生效 | **前端 banner** 提示字段无效 |
| Pipeline 扩展页 `enable_all_skills` / 绑定 skill | 可编辑 | **前端禁用** + banner |
| 仪表盘 Box 状态卡片 | 绿点 / "已连接" | 灰点 / "已禁用"(disabled) 或 红点 / "已断开"(failed) |
> 后端拒写的边界条件:如果 `ap.box_service` **完全没装**(老式 dev mode,没经过 BuildAppStage),`_require_box_for_write` 视作 no-op,保留 `data/skills/` 本地路径——以兼容历史测试与最小化设置。生产环境总会装 `ap.box_service`,因此该 fallback 不会被触发。
### Session scope
### Pipeline 配置 (templates/metadata/pipeline/ai.yaml)
Pipeline 与 AgentRunner 配置不再暴露 sandbox session 模板。Host 将当前平台会话/事件 scope 规范化后哈希成固定长度的 `lb-box-<sha256>`;相同 scope 稳定复用,不同 scope 隔离,缺少 identity 时拒绝执行。SDK/Python 与 scoped MCP bridge 的工具调用遵守同一规则。详见 [box-session-scope.md](./box-session-scope.md)。
`local-agent.config.box-session-id-template` 控制 session 作用域,预设:
- `{launcher_type}_{launcher_id}` — 每个会话 (推荐,默认)
- `{launcher_type}_{launcher_id}_{sender_id}` — 群聊每个用户
- `{launcher_type}_{launcher_id}_{conversation_id}` — 每个对话上下文
- `{query_id}` — 每条消息(完全隔离)
详见 [box-session-scope.md](./box-session-scope.md)。
### REST API
+365 -144
View File
@@ -1,181 +1,402 @@
# Box Session Scope
# Box Session Scope Design
> Last reviewed: 2026-07-12
> Status: implemented Host-owned, hashed execution scope; Runner/Pipeline session templates are removed.
> Date: 2026-04-18 (last reviewed 2026-06-02)
> Status (2026-06-02): the self-hosted community edition is release-ready (box optional, clean degradation, no migration debt). Tool-call loop cap, async quota scan, and the host_path mount allowlist have landed. Remaining multi-tenant / security hardening is tracked in [box-issues.md](./box-issues.md).
> Branch: `feat/sandbox` (LangBot + langbot-plugin-sdk)
> Related: [Box Architecture](./box-architecture.md) | [Box vs Plugin Runtime](./box-vs-plugin-runtime.md)
## 1. Decision
---
The LangBot Host owns the Box session used by an event run. A Pipeline, Agent,
or AgentRunner cannot choose a global, per-user, per-conversation, or per-query
sandbox mode.
## 0. Implementation Status (2026-05-19)
`BoxService.resolve_box_session_id(query)` always returns this shape:
This document was authored as a design proposal. The current `feat/sandbox` branch
has shipped the design largely as written:
```text
lb-box-<64 lowercase SHA-256 hex characters>
| Item | Status | Notes |
|------|--------|-------|
| `BoxMountSpec` + `BoxSpec.extra_mounts` | ✅ Shipped | SDK `box/models.py` |
| Docker / nsjail / E2B backends apply extra mounts | ✅ Shipped | Last gap closed by SDK commit `0fea9b1` (E2B) |
| `box-session-id-template` in `local-agent` pipeline config | ✅ Shipped | `templates/metadata/pipeline/ai.yaml`, default `{launcher_type}_{launcher_id}` |
| `BoxService.resolve_box_session_id(query)` | ✅ Shipped | `pkg/box/service.py:166` |
| `BoxService.build_skill_extra_mounts(query)` | ✅ Shipped | `pkg/box/service.py:189` |
| Skill exec uses unified container + extra mounts | ✅ Shipped | `pkg/provider/tools/loaders/native.py` skill branch |
| MCP-in-Box uses shared persistent session, multi-process | ✅ Shipped (earlier than originally scoped) | SDK commit `529088e`, LangBot `mcp_stdio.py:_build_box_session_id` |
| `BoxManagedProcessSpec.process_id` + multi-process per session | ✅ Shipped | `BoxRuntime` keeps `managed_processes: dict[pid, _ManagedProcess]` |
| Per-tenant / quota integration with templates | ❌ Not started | See [box-tob-analysis.md](./box-tob-analysis.md) |
The "Phase 2 deferred" note in §10 is **out of date** — MCP unification went in on
the same line. Pipeline-scoped (not user-scoped) MCP container is the realized
behavior: each pipeline's MCP servers share one `mcp-<pipeline>` session, and
user exec sessions use the template-derived id.
The remaining open work is multi-tenant overlays (tenant_id in session_id,
quota counters keyed by tenant), tracked in the toB analysis doc rather than here.
---
## 1. Problems
### 1.1 Default exec: per-message containers
Currently, `BoxService.execute_tool()` sets `session_id = str(query.query_id)` — an
auto-incrementing integer per incoming message. Every user message creates a new sandbox
container. Dependencies installed and in-container state are lost between messages.
### 1.2 Three isolated container pools
Default exec, skills, and MCP servers each manage their own containers with
independent session IDs:
| Path | Session ID | Container |
|--------------|-----------------------------------------------|-------------|
| Default exec | `str(query_id)` (per message) | Ephemeral |
| Skill exec | `skill-{launcher}_{id}-{skill_name}` | Per skill |
| MCP stdio | `mcp-{server_uuid}` | Per server |
This means a single logical user interaction can spawn 3+ containers that cannot
share state, see each other's files, or reuse installed dependencies.
### 1.3 Single bind mount limitation
`BoxSpec` currently supports only **one** `host_path``mount_path` bind mount.
This prevents mounting both a default workspace and skill directories into the
same container.
---
## 2. Concept Model
```
Platform Message
→ Query (query_id: int, auto-increment, per message)
→ Session (launcher_type + launcher_id, per chat window)
→ Conversation (uuid, per dialogue context within a Session)
```
The result is exactly 71 ASCII characters. Raw platform, user, group,
conversation, thread, and event identifiers never appear in the Box session
id. This avoids unsafe path characters, unbounded identifier length, and
identity leakage through runtime/container metadata.
| Concept | Key | Example | Scope |
|---------------|-------------------------------------|----------------------------|------------------------------|
| Query | `query_id` | `42` | Single message |
| Session | `launcher_type` + `launcher_id` | `group_123456` | Chat window (group or PM) |
| Conversation | `conversation_id` (UUID) | `a1b2c3d4-...` | Dialogue context within a Session |
| Sender | `sender_id` | `789` | Individual user |
This rule replaces all former concepts of:
Note: in a **group chat**, all users share the same Session (keyed by `group_id`). The
individual sender is tracked as `sender_id` but does not affect Session/Conversation routing.
- Pipeline or Runner `box-session-id-template` fields;
- a global forced session template;
- API fields that let a caller supply sandbox scope;
- LocalAgent-specific Host injection of Box availability, scope, or Pipeline id.
---
## 2. Canonical Host scope
## 3. Target Scenarios
Before hashing, the Host creates a canonical, sorted JSON scope with these
dimensions:
| # | Scenario | Box Granularity | Desired `session_id` |
|----|--------------------------------|------------------------------------------|---------------------------------------------------------|
| 1 | Personal assistant | 1 Box per user, long-lived | `{launcher_type}_{launcher_id}` |
| 2 | Customer service | 1 Box per customer, cross-pipeline | `{launcher_type}_{launcher_id}` |
| 3 | Internal employee tool | 1 Box per employee | `{launcher_type}_{launcher_id}` |
| 4 | Group chat shared assistant | 1 Box per group | `{launcher_type}_{launcher_id}` |
| 5 | Group chat isolated per user | 1 Box per user within a group | `{launcher_type}_{launcher_id}_{sender_id}` |
| 6 | Teaching (cross-channel) | 1 Box per student across groups/PMs | `{sender_id}` |
| 7 | One-off execution | 1 Box per message (current behavior) | `{query_id}` |
| 8 | Multi-project development | 1 Box per conversation context | `{launcher_type}_{launcher_id}_{conversation_id}` |
| Dimension | Purpose |
| --- | --- |
| `instance_id` | Isolate separate LangBot installations |
| `workspace_id` | Preserve workspace/tenant boundary when available |
| `bot_id` | Prevent two bots from sharing a sandbox accidentally |
| `platform_adapter` | Separate identical target ids from different adapters |
| `target_type` / `target_id` | Identify the platform session or event target |
| `thread_id` | Isolate threads within a target when available |
No single fixed granularity covers all scenarios. A template-based approach is needed.
The canonical JSON is domain-separated and hashed by the Host. Runner input,
runner config, and tool parameters are not trusted sources for this scope.
---
### 2.1 Target identity priority
## 4. Design Overview
The Host resolves `target_type` / `target_id` in this order:
Two key changes:
1. For a Pipeline-backed run, use the exact Query launcher tuple.
2. For a pure EBA run, use `delivery.reply_target.target_type/target_id`
(`launcher_type/launcher_id` aliases are accepted).
3. If there is no delivery target, use `conversation_id`.
4. For a non-message event without a conversation, use `event_id`, producing
an event-scoped sandbox.
1. **Unified container**: exec, skills, and MCP all share the same container per
session scope. No more separate container pools.
2. **Configurable session scope**: `session_id` is generated from a template with
pipeline variables, configurable per pipeline.
The adapter class or declared adapter capability supplies platform adapter
identity. The Host includes the active LangBot instance, workspace, bot, and
thread dimensions when they exist.
### 4.1 Unified Container with Multiple Mounts
### 2.2 Stability and isolation
A single container per session scope is created on first use. It has:
The same normalized scope always produces the same hash, so repeated runs in
the same platform conversation reuse the same Box workspace. A rotating
transcript/conversation id does not change the scope when an explicit platform
reply target remains the same.
- **Primary mount**: default workspace at `/workspace` (from `default_host_workspace`)
- **Skill mounts**: each pipeline-bound skill's `package_root` mounted at
`/workspace/.skills/{skill_name}/`
- **MCP servers**: run as managed processes inside the same container
A different target, thread, workspace, bot, platform adapter, or LangBot
instance changes the hash. If delivery target is unavailable and
`conversation_id` is the fallback, different conversations also produce
different hashes. Event-scoped fallback isolates unrelated non-message events.
### 2.3 Fail closed
If the private Host scope marker is present but empty or malformed, Box rejects
execution with `BoxValidationError`. A direct Query without either a valid
Host scope or launcher/session identity is also rejected. There is no
`unknown`, raw query id, global, or caller-selected fallback.
## 3. Host execution Query
AgentRunner callbacks need a Host-owned Query view because model/tool loaders
already consume that type. The Query is internal and is never exposed as a
Runner-controlled object.
- A Pipeline run stores the exact current Query in `AgentRunSession`.
- A pure EBA run builds a minimal Query with a valid Session and
`pipeline_config=None`, `pipeline_uuid=None`.
- The Host attaches canonical `_host_box_scope` and the authorized skill names
in `_pipeline_bound_skills`.
- `PluginToRuntimeAction.CALL_TOOL` restores this Query from the active
`run_id` before dispatching to `ToolManager`.
This gives Pipeline and pure EBA execution the same Host tool path without
inventing a fake Pipeline for an independent Agent.
## 4. AgentRunner callback paths
AgentRunner implementations may use either callback transport:
1. SDK/Python runners call `AgentRunAPIProxy.call_tool`.
2. External harnesses call the SDK-owned scoped MCP bridge.
Both transports emit the same `PluginToRuntimeAction.CALL_TOOL`. The Host then
validates the same run authorization, restores the same execution Query, and
dispatches to the same ToolManager and BoxService.
```text
AgentRunner
+-- AgentRunAPIProxy.call_tool --------+
| |
+-- SDK-owned scoped MCP bridge -------+--> PluginToRuntimeAction.CALL_TOOL
--> run authorization
--> execution Query
--> ToolManager
--> BoxService
--> lb-box-<sha256>
```
Container (session_id = "group_123456")
/workspace/ ← default workspace (bind mount, rw)
/workspace/.skills/web-search/ ← skill package (bind mount, rw)
/workspace/.skills/data-analysis/ ← skill package (bind mount, rw)
[managed process: mcp-server-a] ← MCP server running inside
[managed process: mcp-server-b] ← MCP server running inside
```
An AgentRunner is not required to use MCP. Local Python runners can use the SDK
directly; code-agent harnesses can use the bridge. The transports do not define
different authorization or sandbox semantics.
This requires extending `BoxSpec` to support multiple mounts (see §5).
## 5. Skills and mounts
### 4.2 Session ID Template
Native exec and skill-backed exec for one Host scope use the same hashed
session. `BoxService.build_skill_extra_mounts(query)` adds visible, authorized
skill packages under `/workspace/.skills/<name>` when the session is created.
A new field `box-session-id-template` in the `local-agent` pipeline runner config
controls the session scope:
Skill activation controls which skill-backed tools and paths are available. It
does not create a different session and does not grant the Runner authority to
change the session id.
```yaml
# templates/metadata/pipeline/ai.yaml (under local-agent.config)
- name: box-session-id-template
label:
en_US: Sandbox Scope
zh_Hans: 沙箱作用域
description:
en_US: >-
Determines how sandbox environments are shared. Use variables to
control isolation granularity.
zh_Hans: >-
决定沙箱环境的共享方式。使用变量控制隔离粒度。
type: select
required: false
default: "{launcher_type}_{launcher_id}"
options:
- value: "{launcher_type}_{launcher_id}"
label:
en_US: Per chat (Recommended)
zh_Hans: 每个会话(推荐)
- value: "{launcher_type}_{launcher_id}_{sender_id}"
label:
en_US: Per user in chat
zh_Hans: 会话中每个用户
- value: "{launcher_type}_{launcher_id}_{conversation_id}"
label:
en_US: Per conversation context
zh_Hans: 每个对话上下文
- value: "{query_id}"
label:
en_US: Per message (isolated)
zh_Hans: 每条消息(完全隔离)
```
## 6. `mcp-shared` is a different session
Available template variables (populated by PreProcessor in `query.variables`):
LangBot can host configured stdio MCP servers as managed processes inside Box.
Those long-lived infrastructure processes share the dedicated `mcp-shared`
session and are isolated from one another by `process_id`.
| Variable | Source | Example |
|---------------------|---------------------------------|----------------------|
| `{launcher_type}` | `query.session.launcher_type` | `person` / `group` |
| `{launcher_id}` | `query.session.launcher_id` | `123456` |
| `{sender_id}` | `query.sender_id` | `789` |
| `{conversation_id}` | `conversation.uuid` | `a1b2c3d4-...` |
| `{query_id}` | `query.query_id` | `42` |
This is separate from the scoped MCP bridge above:
Default `{launcher_type}_{launcher_id}` covers scenarios 14 out of the box.
| Path | Purpose | Session rule |
| --- | --- | --- |
| AgentRunner scoped MCP bridge | Call authorized Host tools for one active run | Host-owned `lb-box-<sha256>` from the run execution Query |
| MCP-in-Box stdio server | Keep configured MCP server processes running | Dedicated persistent `mcp-shared` session |
---
Calling a sandbox tool through the AgentRunner bridge never redirects the run
workspace into `mcp-shared`. Conversely, an MCP server's managed-process
lifecycle does not inherit the current event scope.
## 5. SDK Changes: Multi-Mount BoxSpec
## 7. Configuration and compatibility
### 5.1 Model Extension
There is no Box session scope field in Pipeline metadata, AgentRunner config,
or the public Pipeline/Runner API. Operators configure the Box subsystem itself
(`box.enabled`, backend/runtime settings, profiles, mount allowlists, quotas,
and workspace roots), not per-Runner session templates.
```python
# box/models.py
Old configuration containing `box-session-id-template` is unsupported in the
4.x contract. LangBot 4.x does not migrate LangBot 3.x configuration or
databases, so the removed field is not read as a compatibility fallback.
class BoxMountSpec(pydantic.BaseModel):
"""A single bind mount specification."""
host_path: str
mount_path: str
mode: BoxHostMountMode = BoxHostMountMode.READ_WRITE
## 8. Regression coverage
class BoxSpec(pydantic.BaseModel):
# ... existing fields ...
host_path: str | None = None # Primary mount (backward compat)
host_path_mode: BoxHostMountMode = BoxHostMountMode.READ_WRITE
mount_path: str = DEFAULT_BOX_MOUNT_PATH
extra_mounts: list[BoxMountSpec] = [] # NEW: additional mounts
```
Release tests should prove:
`extra_mounts` is additive — the existing `host_path` / `mount_path` pair remains
the primary mount for backward compatibility.
- every event-run session id matches `lb-box-[0-9a-f]{64}` and contains no raw
identity;
- the same canonical Host scope is stable while different targets,
conversations, threads, bots, adapters, workspaces, or instances are
isolated;
- Pipeline and pure EBA runs representing the same platform session produce
the same canonical scope;
- missing Host/Query identity fails closed;
- SDK/Python `call_tool` and the scoped MCP bridge both enter
`PluginToRuntimeAction.CALL_TOOL` and restore the run execution Query;
- Runner payload/config cannot override the session id;
- stdio MCP processes remain in `mcp-shared` and are isolated by process id;
- authorized skills are mounted into the hashed run session without creating
per-skill sessions.
### 5.2 Backend: Apply Extra Mounts
```python
# box/backend.py — CLISandboxBackend.start_session()
# Primary mount (unchanged)
if spec.host_path is not None and spec.host_path_mode != BoxHostMountMode.NONE:
args.extend(['-v', f'{spec.host_path}:{spec.mount_path}:{spec.host_path_mode.value}'])
# Extra mounts (NEW)
for mount in spec.extra_mounts:
if mount.mode != BoxHostMountMode.NONE:
args.extend(['-v', f'{mount.host_path}:{mount.mount_path}:{mount.mode.value}'])
```
Same pattern for nsjail backend.
---
## 6. LangBot Changes
### 6.1 Session ID Resolution
In `BoxService.execute_tool()`:
```python
# Before:
spec_payload.setdefault('session_id', str(query.query_id))
# After:
template = (query.pipeline_config or {}).get('ai', {}) \
.get('local-agent', {}).get('box-session-id-template',
'{launcher_type}_{launcher_id}')
variables = query.variables or {}
session_id = template.format_map(collections.defaultdict(
lambda: 'unknown', variables
))
spec_payload.setdefault('session_id', session_id)
```
### 6.2 Skill Exec: Use Same Container
Currently `native.py:_invoke_exec` creates a separate `BoxWorkspaceSession` per
skill with `host_path=package_root`. Instead:
1. Use the **same session_id** as default exec (from the template).
2. Pass the skill's `package_root` as an **extra mount** at
`/workspace/.skills/{skill_name}/` instead of replacing `/workspace`.
3. The container already has the default workspace at `/workspace`.
```python
# native.py — _invoke_exec, skill branch (REVISED)
# Same session_id as default exec
session_id = resolve_box_session_id(query)
spec_payload = {
'cmd': rewritten_command,
'workdir': rewritten_workdir,
'session_id': session_id,
'extra_mounts': [{
'host_path': package_root,
'mount_path': f'/workspace/.skills/{selected_skill_name}',
'mode': 'rw',
}],
}
result = await self.ap.box_service.execute_spec_payload(spec_payload, query)
```
The virtual path `/workspace/.skills/{name}` no longer needs rewriting at the
command level — it maps directly to the bind mount path inside the container.
### 6.3 MCP: Use Same Container
MCP servers should run inside the same container as exec and skills. Changes:
1. `BoxStdioSessionRuntime` uses the pipeline's session_id template instead of
`mcp-{server_uuid}`.
2. MCP server's working directory is a subdirectory (e.g. `/workspace/.mcp/{name}/`).
3. MCP server's dependencies are mounted or installed into that subdirectory.
4. The MCP server runs as a managed process inside the shared container.
Since MCP servers start at LangBot boot (not per-query), the session must be
created eagerly. The container will be kept alive by the managed process
exemption in TTL reaping (`runtime.py:259`).
**Note**: MCP sessions are pipeline-scoped (not per-launcher), so their session_id
should be a **fixed identifier per pipeline** rather than the user-facing template.
This means one shared MCP container per pipeline, with user exec sessions separate.
Alternatively, in a future iteration, MCP managed processes could be launched
lazily into the user's container on first MCP tool call. This is more complex
but maximizes sharing. For V1, keeping MCP containers at pipeline scope is
simpler and more predictable.
---
## 7. Mount Layout Summary
### Default exec (no skills activated)
```
Container (session_id from template)
/workspace/ ← default_host_workspace (rw)
```
### Exec with activated skills
```
Container (same session_id)
/workspace/ ← default_host_workspace (rw)
/workspace/.skills/web-search/ ← skill package_root (rw)
/workspace/.skills/data-analysis/ ← skill package_root (rw)
```
Extra mounts are **additive** — they are added when the container is first
created (or on the first exec that references a skill). Since Docker bind
mounts are specified at container creation time, skills must be known at
creation time.
**Resolution**: When creating a container, inject `extra_mounts` for **all
pipeline-bound skills** (from `extensions_preferences`), not just the
currently activated one. This way any skill can be activated later without
recreating the container.
### MCP servers (V1: pipeline-scoped)
```
Container (session_id = "mcp-pipeline-{pipeline_uuid}")
/workspace/ ← MCP shared workspace
/workspace/.mcp/server-a/ ← MCP server A files
/workspace/.mcp/server-b/ ← MCP server B files
[managed process: server-a]
[managed process: server-b]
```
---
## 8. Data Migration
Existing pipelines do not have `box-session-id-template`. The backend uses
`.get(..., default)` so missing keys fall back to `{launcher_type}_{launcher_id}`.
This changes behavior from per-message to per-launcher for existing pipelines.
Recommendation: **accept the behavior change** — per-launcher is the more
intuitive default, and the old per-message behavior was rarely desired.
---
## 9. Cloud Quota Implications
| Scope | Typical concurrent containers |
|-----------------------------------------------|-------------------------------|
| `{query_id}` (per message) | Many, short-lived |
| `{launcher_type}_{launcher_id}` (per chat) | = active chat count |
| `{sender_id}` (per user) | = active user count |
| `{conversation_id}` (per conversation) | Between per-chat and per-msg |
With the unified container model, each scope value maps to exactly **one**
container (instead of potentially 3+ per-message). This significantly reduces
resource usage.
Quota enforcement point: `BoxRuntime._get_or_create_session()` in the SDK.
---
## 10. Implementation Phases
### Phase 1: Session scope + skill unification (this PR)
1. **SDK**: Extend `BoxSpec` with `extra_mounts: list[BoxMountSpec]`.
2. **SDK**: Update Docker/nsjail backends to apply extra mounts.
3. **LangBot**: Add `box-session-id-template` to `local-agent` YAML metadata
and default pipeline config JSON.
4. **LangBot**: Update `BoxService.execute_tool()` to use template interpolation.
5. **LangBot**: Update `native.py:_invoke_exec` skill branch to use same
session_id + extra mounts instead of separate `BoxWorkspaceSession`.
6. **LangBot**: On container creation, inject extra mounts for all
pipeline-bound skills.
7. **Frontend**: No code change — `DynamicFormComponent` renders `select` fields.
8. **Tests**: Unit tests for template interpolation and multi-mount specs.
### Phase 2: MCP unification (future)
1. Refactor `BoxStdioSessionRuntime` to use pipeline-scoped shared container.
2. MCP servers become managed processes in the shared container.
3. Support multiple concurrent managed processes per container.
MCP unification is deferred because it requires changes to the managed process
model (currently 1 managed process per session) and has startup ordering
concerns (MCP servers start at boot, before any user query determines
a session_id).
+6 -7
View File
@@ -1,6 +1,6 @@
# Box 系统测试覆盖分析
> 更新日期: 2026-07-12
> 更新日期: 2026-06-02
> 状态更新: 自部署社区版已具备发布条件(box 可选、降级完善、无迁移欠债);工具调用循环上限、配额遍历异步化、`host_path` 挂载白名单等已落地。剩余多租户 / 安全硬化项见 [SaaS 阻塞项清单](./box-issues.md)。
> 分支: `feat/sandbox` (LangBot + langbot-plugin-sdk)
@@ -15,14 +15,13 @@
| `tests/unit_tests/box/test_box_connector.py` | 106 | 是 | Connector 传输决策、WS relay URL、dispose、心跳/重连 |
| `tests/unit_tests/box/test_box_service.py` | 1224 | 是 | Service 核心逻辑(最全面) |
| `tests/unit_tests/box/test_workspace.py` | 147 | 是 | WorkspaceSession 路径重写、payload 构建 |
| `tests/unit_tests/agent/test_execution_context.py` | 144 | 是 | Pipeline/纯 EBA execution Query、canonical Host scope、adapter/instance 隔离 |
| `tests/unit_tests/plugin/test_handler_actions.py` | 1071 | 是 | AgentRun CALL_TOOL 恢复 execution Query、纯 EBA native exec |
| `tests/unit_tests/provider/test_mcp_box_integration.py` | 707 | 是 | MCP Box 配置、路径重写、payload、shared-session/multi-process、runtime info |
| `tests/unit_tests/provider/test_localagent_sandbox_exec.py` | 444 | 是 | LocalAgent exec 流程、流式、Skill 激活 (Tool Call) |
| `tests/unit_tests/provider/test_tool_manager_native.py` | 249 | 是 | ToolManager 路由、native tool CRUD、路径穿越、6 工具暴露 |
| `tests/unit_tests/provider/test_skill_tools.py` | 582 | 是 | Skill 管理、Tool Call 激活、路径、authoring CRUD |
| `tests/unit_tests/test_skill_service.py` | 396 | 是 | HTTP serviceskill CRUD、zip/GitHub install、文件浏览 |
| `tests/unit_tests/test_paths.py` | 23 | 是 | paths 工具 |
| `tests/unit_tests/test_preproc.py` | 134 | 是 | PreProcessor 的模型、历史与 bound skill 解析 |
| `tests/unit_tests/test_preproc.py` | 134 | 是 | PreProcessor 注入 session 变量、bound skill 解析 |
| `tests/unit_tests/pipeline/test_chat_handler_logging.py` | 78 | 是 | Chat handler 日志相关回归 |
| `tests/integration_tests/box/test_box_integration.py` | 329 | **否** | 真实容器执行、超时、网络隔离 |
| `tests/integration_tests/box/test_box_mcp_integration.py` | 368 | **否** | Managed process、WS attach、shared-session 清理 |
@@ -36,7 +35,7 @@
| `tests/box/test_e2b_backend.py` | 482 | 是 | E2B SDK mock、session 生命周期、extra_mounts 同步 |
| `tests/box/test_skill_store.py` | 88 | 是 | zip preview/install、基础 file CRUD |
**说明**: 本表按当前主链列出 Box 相关测试;其中 2 个真实容器集成测试默认不在 CI 中运行。
**总计**: 17 个测试文件, ~6,500 行测试代码; 其中 2 个集成测试(约 700 行)在 CI 中运行。
> 较 2026-04-16 版增加:`test_skill_service.py`、`test_paths.py`、`test_preproc.py`、`test_chat_handler_logging.py` (LangBot)`test_backend_selection.py`、`test_e2b_backend.py`、`test_skill_store.py` (SDK)。`test_nsjail_backend.py` 增加 CLI 兼容性 case (commit `feed530`)。
@@ -52,7 +51,7 @@
| BoxService workspace quota | 优秀 | 前置/后置配额检查、超额清理 |
| BoxService 输出截断 | 优秀 | 短/精确边界/长输出、独立 stderr |
| BoxService 可观测性 | 优秀 | 状态报告、error ring buffer、buffer 上限 |
| BoxService Host-owned session | 优秀 | 覆盖 `lb-box-<sha256>` 固定格式、原始 identity 不泄露、同 scope 稳定、不同 conversation/scope 隔离、缺 identity fail closed |
| BoxService session 模板 | 良好 | `resolve_box_session_id` + `build_skill_extra_mounts` 在 service / native / mcp 三处都有覆盖 |
| RPC client/server 协议 | 优秀 | execute/get_sessions/delete/create/conflict error |
| BoxRuntimeConnector | 良好 | local/remote 模式、Docker 平台、relay URL、心跳与重连回调 |
| BoxWorkspaceSession | 良好 | payload 构建、managed process 路径重写、stage host file |
@@ -62,7 +61,7 @@
| Backend selection | 良好 | 显式 backend 优先级、local 探测顺序、配置变更触发 reselect |
| MCP Box 集成 | 良好 | config model、路径重写、payload、shared-session 多 process |
| Native tool loader | 良好 | 6 工具(exec/read/write/edit/glob/grep)、路径穿越拦截 |
| AgentRunner 工具入口 | 良好 | SDK proxy 与 MCP bridge 都映射到 `PluginToRuntimeAction.CALL_TOOL`Host action 测试覆盖 run-scoped execution Query 与纯 EBA native exec |
| LocalAgent exec 流程 | 良好 | 完整 tool call 循环、流式、system prompt 注入、Tool Call 激活 |
| Skill 系统 | 良好 | 加载、Tool Call 激活、marker、路径解析、authoring CRUD、HTTP service |
---
+4 -3
View File
@@ -1,6 +1,6 @@
[project]
name = "langbot"
version = "4.11.0"
version = "4.10.6"
description = "Production-grade platform for building agentic IM bots"
readme = "README.md"
license-files = ["LICENSE"]
@@ -39,6 +39,7 @@ dependencies = [
"quart>=0.20.0",
"quart-cors>=0.8.0",
"requests>=2.33.0",
"regex>=2026.1.15",
"slack-sdk>=3.35.0",
"alembic>=1.15.0",
"sqlalchemy[asyncio]>=2.0.40",
@@ -70,7 +71,7 @@ dependencies = [
"chromadb>=1.0.0,<2.0.0",
"qdrant-client (>=1.15.1,<2.0.0)",
"pyseekdb==1.1.0.post3",
"langbot-plugin==0.5.0a2",
"langbot-plugin @ git+https://github.com/langbot-app/langbot-plugin-sdk.git@1d65ed301a6afc52150a998043f73cd6032c8162",
"asyncpg>=0.30.0",
"line-bot-sdk>=3.19.0",
"matrix-nio>=0.25.2",
@@ -120,7 +121,7 @@ requires = ["setuptools>=61.0", "wheel"]
build-backend = "setuptools.build_meta"
[tool.setuptools]
package-data = { "langbot" = ["templates/**", "pkg/provider/modelmgr/requesters/*", "pkg/platform/sources/*", "pkg/platform/adapters/**", "web/dist/**", "pkg/persistence/alembic/**"] }
package-data = { "langbot" = ["templates/**", "pkg/provider/modelmgr/requesters/*", "pkg/platform/sources/*", "web/dist/**", "pkg/persistence/alembic/**"] }
[dependency-groups]
dev = [
+6
View File
@@ -13,6 +13,12 @@ testpaths = tests
# Asyncio configuration
asyncio_mode = auto
# Resource leaks are often reported during object finalization and wrapped by
# pytest. Keep both forms fatal so --disable-warnings cannot hide them.
filterwarnings =
error::ResourceWarning
error::pytest.PytestUnraisableExceptionWarning
# Output options
addopts =
-v
File diff suppressed because it is too large Load Diff
+466
View File
@@ -0,0 +1,466 @@
#!/usr/bin/env python3
"""Exercise long-lived Core registries and verify that they reach a plateau.
This probe is intentionally separate from the default test suite because the
audit profile creates tens of thousands of historical identities. It uses the
real admission, eviction, and cleanup code while replacing external platform
objects that are irrelevant to registry retention.
"""
from __future__ import annotations
import argparse
import asyncio
import gc
import json
import time
import tracemalloc
from dataclasses import asdict, dataclass
from types import SimpleNamespace
from unittest.mock import patch
import psutil
from langbot.pkg.api.http.context import ExecutionContext
# Import the Application graph before taskmgr. The production boot path has
# this same ordering; importing taskmgr first exposes its historical cycle
# through HTTP route annotations.
from langbot.pkg.core import app as _core_app # noqa: F401
from langbot.pkg.core.taskmgr import AsyncTaskManager
from langbot.pkg.pipeline.pool import QueryPool
from langbot.pkg.pipeline.ratelimit.algos.fixedwin import FixedWindowAlgo
from langbot.pkg.plugin.connector import PluginRuntimeConnector
from langbot.pkg.platform.sources.websocket_adapter import (
WebSocketMessage,
WebSocketSession,
)
from langbot.pkg.provider.modelmgr.modelmgr import ModelManager
from langbot.pkg.provider.session.sessionmgr import SessionManager
from langbot_plugin.api.entities.builtin.provider.session import LauncherTypes
@dataclass(frozen=True, slots=True)
class ProbeScale:
query_churn_per_phase: int
session_churn_per_phase: int
rate_limit_churn_per_phase: int
task_churn_per_phase: int
websocket_churn_per_phase: int
empty_workspace_churn_per_phase: int
SCALES = {
'quick': ProbeScale(
query_churn_per_phase=2_500,
session_churn_per_phase=500,
rate_limit_churn_per_phase=10_000,
task_churn_per_phase=1_000,
websocket_churn_per_phase=500,
empty_workspace_churn_per_phase=1_000,
),
'audit': ProbeScale(
query_churn_per_phase=25_000,
session_churn_per_phase=2_500,
rate_limit_churn_per_phase=10_000,
task_churn_per_phase=5_000,
websocket_churn_per_phase=2_500,
empty_workspace_churn_per_phase=10_000,
),
}
class _ProbeQuery:
"""Small weak-referenceable stand-in for SDK Query construction."""
def __init__(self, **values):
self.__dict__.update(values)
class _EmptyResult:
def all(self) -> list:
return []
class _EmptyPluginRuntimeHandler:
async def reconcile_plugin_installations(self, _states: tuple) -> dict:
return {
'applied': [],
'removed': [],
'missing_artifacts': [],
'failed_installations': [],
}
def unregister_installation_binding(self, _binding) -> None:
raise AssertionError('An empty Workspace exposed an installation binding')
@dataclass(frozen=True, slots=True)
class ProcessSample:
rss_bytes: int
traced_current_bytes: int
traced_peak_bytes: int
asyncio_tasks: int
threads: int
open_fds: int | None
def _sample_process() -> ProcessSample:
gc.collect()
process = psutil.Process()
try:
open_fds = process.num_fds()
except (AttributeError, psutil.Error):
open_fds = None
traced_current, traced_peak = tracemalloc.get_traced_memory()
return ProcessSample(
rss_bytes=process.memory_info().rss,
traced_current_bytes=traced_current,
traced_peak_bytes=traced_peak,
asyncio_tasks=len(asyncio.all_tasks()),
threads=process.num_threads(),
open_fds=open_fds,
)
def _execution_context(index: int, *, query_uuid: str | None = None) -> ExecutionContext:
return ExecutionContext(
instance_uuid='runtime-resource-probe',
workspace_uuid=f'workspace-{index}',
placement_generation=1,
bot_uuid='probe-bot',
pipeline_uuid='probe-pipeline',
query_uuid=query_uuid,
)
class CoreRuntimeProbe:
"""Own the same manager instances across two equal churn phases."""
def __init__(self) -> None:
self.query_pool = QueryPool(max_queries=100, max_queries_per_workspace=1)
app = SimpleNamespace(
event_loop=asyncio.get_running_loop(),
persistence_mgr=None,
instance_config=SimpleNamespace(
data={
'concurrency': {'session': 1},
'system': {
'session_retention': {
'idle_ttl_seconds': 86_400,
'max_entries': 200,
'max_entries_per_workspace': 200,
'max_conversations_per_session': 20,
'max_messages_per_conversation': 100,
},
'task_retention': {
'completed_limit': 200,
'max_log_chars': 4_096,
'max_active_user_tasks': 256,
'max_active_user_tasks_per_workspace': 8,
},
},
}
),
)
self.session_manager = SessionManager(app)
self.task_manager = AsyncTaskManager(app)
self.rate_limit = FixedWindowAlgo(SimpleNamespace())
self.websocket_session = WebSocketSession(
'resource-probe',
max_conversations=200,
max_messages=100,
)
logger = SimpleNamespace(
debug=lambda *_args, **_kwargs: None,
info=lambda *_args, **_kwargs: None,
warning=lambda *_args, **_kwargs: None,
error=lambda *_args, **_kwargs: None,
)
self.empty_model_queries = 0
async def execute_empty(_statement):
self.empty_model_queries += 1
return _EmptyResult()
model_app = SimpleNamespace(
logger=logger,
persistence_mgr=SimpleNamespace(execute_async=execute_empty),
)
self.empty_model_manager = ModelManager(model_app)
async def runtime_disconnect_callback(_connector) -> None:
return None
plugin_app = SimpleNamespace(
instance_config=SimpleNamespace(data={'plugin': {'enable': True}}),
deployment=SimpleNamespace(mode='cloud'),
logger=logger,
)
self.empty_plugin_connector = PluginRuntimeConnector(
plugin_app,
runtime_disconnect_callback,
)
self.empty_plugin_connector.handler = _EmptyPluginRuntimeHandler()
async def validate_context(context):
return context
async def load_desired_states(_context):
return []
self.empty_plugin_connector._validate_execution_context = validate_context
self.empty_plugin_connector._load_workspace_desired_states = load_desired_states
async def initialize(self) -> None:
await self.rate_limit.initialize()
async def run_phase(self, scale: ProbeScale, phase: int) -> None:
offsets = {
'query': (phase - 1) * scale.query_churn_per_phase,
'session': (phase - 1) * scale.session_churn_per_phase,
'rate': (phase - 1) * scale.rate_limit_churn_per_phase,
'task': (phase - 1) * scale.task_churn_per_phase,
'websocket': (phase - 1) * scale.websocket_churn_per_phase,
'empty_workspace': ((phase - 1) * scale.empty_workspace_churn_per_phase),
}
await self._churn_queries(offsets['query'], scale.query_churn_per_phase)
await self._churn_sessions(offsets['session'], scale.session_churn_per_phase)
await self._churn_rate_limits(offsets['rate'], scale.rate_limit_churn_per_phase)
await self._churn_tasks(offsets['task'], scale.task_churn_per_phase)
self._churn_websocket_history(
offsets['websocket'],
scale.websocket_churn_per_phase,
)
await self._churn_empty_workspaces(
offsets['empty_workspace'],
scale.empty_workspace_churn_per_phase,
)
await asyncio.sleep(0)
async def _churn_queries(self, start: int, count: int) -> None:
def make_query(**values):
return _ProbeQuery(**values)
with patch(
'langbot.pkg.pipeline.pool.pipeline_query.Query',
side_effect=make_query,
):
for index in range(start, start + count):
context = _execution_context(index)
query = await self.query_pool.add_query(
bot_uuid='probe-bot',
launcher_type=LauncherTypes.PERSON,
launcher_id=f'launcher-{index}',
sender_id=f'sender-{index}',
message_event=SimpleNamespace(),
message_chain=SimpleNamespace(),
adapter=None,
pipeline_uuid='probe-pipeline',
execution_context=context,
)
removed = await self.query_pool.remove_query(query)
if not removed:
raise AssertionError('Query cleanup failed')
async def _churn_sessions(self, start: int, count: int) -> None:
for index in range(start, start + count):
workspace_index = index % 100
context = _execution_context(
workspace_index,
query_uuid=f'session-query-{index}',
)
query = SimpleNamespace(
launcher_type=LauncherTypes.PERSON,
launcher_id=f'launcher-{index}',
sender_id=f'sender-{index}',
bot_uuid='probe-bot',
pipeline_uuid='probe-pipeline',
query_uuid=context.query_uuid,
_execution_context=context,
)
await self.session_manager.get_session(query)
async def _churn_rate_limits(self, start: int, count: int) -> None:
for index in range(start, start + count):
context = _execution_context(
index % 1_000,
query_uuid=f'rate-query-{index}',
)
query = SimpleNamespace(
bot_uuid='probe-bot',
pipeline_uuid='probe-pipeline',
_execution_context=context,
pipeline_config={
'safety': {
'rate-limit': {
'window-length': 60,
'limitation': 100_000,
'strategy': 'drop',
}
}
},
)
admitted = await self.rate_limit.require_access(
query,
LauncherTypes.PERSON,
f'rate-identity-{index}',
)
if not admitted:
raise AssertionError('Rate-limit registry rejected bounded churn')
async def _churn_tasks(self, start: int, count: int) -> None:
async def complete_immediately() -> None:
return None
for batch_start in range(start, start + count, 256):
batch_size = min(256, start + count - batch_start)
wrappers = [
self.task_manager.create_task(
complete_immediately(),
name=f'resource-probe-{batch_start + offset}',
)
for offset in range(batch_size)
]
await asyncio.gather(*(wrapper.task for wrapper in wrappers))
await asyncio.sleep(0)
def _churn_websocket_history(self, start: int, count: int) -> None:
for index in range(start, start + count):
conversation_key = f'conversation-{index}'
response_id = f'response-{index}'
indexes = self.websocket_session.get_stream_message_indexes(conversation_key)
indexes[response_id] = 0
self.websocket_session.append_message(
conversation_key,
WebSocketMessage(
id=self.websocket_session.next_message_id(conversation_key),
role='assistant',
content='probe',
message_chain=[],
timestamp='1970-01-01T00:00:00+00:00',
is_final=True,
),
)
async def _churn_empty_workspaces(self, start: int, count: int) -> None:
for index in range(start, start + count):
await self.empty_model_manager._load_workspace_models(_execution_context(index))
await self.empty_plugin_connector.reconcile_projected_workspaces(
_execution_context(index) for index in range(start, start + count)
)
def retained_state(self) -> dict[str, int]:
return {
'query_cached': len(self.query_pool.cached_queries),
'query_queued': len(self.query_pool.queries),
'query_active_workspaces': len(self.query_pool.active_query_count_by_workspace),
'query_scope_counters': len(self.query_pool.query_count_by_scope),
'sessions': len(self.session_manager.session_list),
'session_index': len(self.session_manager._session_index),
'rate_limit_containers': len(self.rate_limit.containers),
'task_records': len(self.task_manager.tasks),
'websocket_conversations': len(self.websocket_session.message_lists),
'websocket_stream_indexes': len(self.websocket_session.stream_message_indexes),
'empty_model_scopes': len(self.empty_model_manager._scope_generations),
'empty_model_providers': len(self.empty_model_manager.provider_dict),
'empty_model_llms': len(self.empty_model_manager.llm_model_dict),
'empty_plugin_workspace_sets': len(self.empty_plugin_connector._workspace_installations),
'empty_plugin_installations': len(self.empty_plugin_connector._known_desired_states),
}
def assert_bounded(self) -> None:
state = self.retained_state()
expected_maximums = {
'query_cached': 0,
'query_queued': 0,
'query_active_workspaces': 0,
'query_scope_counters': 100,
'sessions': 200,
'session_index': 200,
'rate_limit_containers': 10_000,
'task_records': 200,
'websocket_conversations': 200,
'websocket_stream_indexes': 200,
'empty_model_scopes': 0,
'empty_model_providers': 0,
'empty_model_llms': 0,
'empty_plugin_workspace_sets': 0,
'empty_plugin_installations': 0,
}
violations = {key: (state[key], maximum) for key, maximum in expected_maximums.items() if state[key] > maximum}
if violations:
raise AssertionError(f'Core retained-state limits failed: {violations}')
async def _run(args: argparse.Namespace) -> dict:
scale = SCALES[args.scale]
tracemalloc.start()
started_at = time.monotonic()
probe = CoreRuntimeProbe()
await probe.initialize()
baseline = _sample_process()
await probe.run_phase(scale, 1)
probe.assert_bounded()
phase_one = _sample_process()
state_one = probe.retained_state()
await probe.run_phase(scale, 2)
probe.assert_bounded()
phase_two = _sample_process()
state_two = probe.retained_state()
if state_two != state_one:
raise AssertionError(f'Core retained state did not plateau: phase_one={state_one}, phase_two={state_two}')
traced_growth = phase_two.traced_current_bytes - phase_one.traced_current_bytes
rss_growth = phase_two.rss_bytes - phase_one.rss_bytes
max_traced_growth = int(args.max_traced_growth_mib * 1024 * 1024)
max_rss_growth = int(args.max_rss_growth_mib * 1024 * 1024)
if traced_growth > max_traced_growth:
raise AssertionError(f'Second-phase traced memory grew by {traced_growth} bytes (limit {max_traced_growth})')
if rss_growth > max_rss_growth:
raise AssertionError(f'Second-phase RSS grew by {rss_growth} bytes (limit {max_rss_growth})')
return {
'component': 'langbot-core',
'scale': args.scale,
'work_per_phase': asdict(scale),
'elapsed_seconds': round(time.monotonic() - started_at, 3),
'samples': {
'baseline': asdict(baseline),
'phase_one': asdict(phase_one),
'phase_two': asdict(phase_two),
},
'second_phase_growth': {
'rss_bytes': rss_growth,
'traced_current_bytes': traced_growth,
},
'retained_state': {
'phase_one': state_one,
'phase_two': state_two,
},
'passed': True,
}
def _parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument('--scale', choices=tuple(SCALES), default='quick')
parser.add_argument('--max-traced-growth-mib', type=float, default=8.0)
parser.add_argument('--max-rss-growth-mib', type=float, default=64.0)
parser.add_argument('--json', action='store_true', help='Print compact JSON')
return parser.parse_args()
def main() -> None:
args = _parse_args()
result = asyncio.run(_run(args))
if args.json:
print(json.dumps(result, sort_keys=True))
else:
print(json.dumps(result, indent=2, sort_keys=True))
if __name__ == '__main__':
main()
+561
View File
@@ -0,0 +1,561 @@
#!/usr/bin/env python3
"""Measure populated Workspace runtime replacement cost and retention.
Unlike ``runtime_resource_probe.py``, which stresses historical request keys
and empty tenants, this probe keeps one representative Provider, LLM,
Embedding model, Rerank model, Pipeline, Bot, and Knowledge Base per Workspace.
It then advances every Workspace to a new placement generation and verifies
that old runtime objects are closed and collectible while active registry
cardinality remains constant.
"""
from __future__ import annotations
import argparse
import asyncio
import gc
import json
import time
import tracemalloc
import weakref
from dataclasses import asdict, dataclass
from types import SimpleNamespace
import psutil
from langbot.pkg.api.http.context import ExecutionContext
# Match the production import order; importing a leaf manager first exposes a
# historical annotation cycle that the application graph resolves.
from langbot.pkg.core import app as _core_app # noqa: F401
from langbot.pkg.entity.persistence import bot as persistence_bot
from langbot.pkg.entity.persistence import model as persistence_model
from langbot.pkg.entity.persistence import pipeline as persistence_pipeline
from langbot.pkg.entity.persistence import rag as persistence_rag
from langbot.pkg.pipeline.pipelinemgr import PipelineManager
from langbot.pkg.platform.botmgr import PlatformManager
from langbot.pkg.provider.modelmgr import requester
from langbot.pkg.provider.modelmgr.modelmgr import ModelManager
from langbot.pkg.provider.tools.loaders.mcp import MCPLoader
from langbot.pkg.rag.knowledge.kbmgr import RAGManager
from langbot.pkg.workspace.entities import WorkspaceExecutionBinding
@dataclass(frozen=True, slots=True)
class ProbeScale:
workspaces: int
SCALES = {
'quick': ProbeScale(workspaces=250),
'audit': ProbeScale(workspaces=5_000),
}
@dataclass(frozen=True, slots=True)
class ProcessSample:
rss_bytes: int
traced_current_bytes: int
traced_peak_bytes: int
asyncio_tasks: int
threads: int
open_fds: int | None
class _ProbeLogger:
def debug(self, *_args, **_kwargs) -> None:
return None
def info(self, *_args, **_kwargs) -> None:
return None
def warning(self, *_args, **_kwargs) -> None:
return None
def error(self, *_args, **_kwargs) -> None:
return None
class _ProbeWorkspaceService:
instance_uuid = 'runtime-capacity-probe'
def __init__(self) -> None:
self.generations: dict[str, int] = {}
self.binding_lookups = 0
async def get_execution_binding(
self,
workspace_uuid: str,
*,
expected_generation: int | None = None,
) -> WorkspaceExecutionBinding:
self.binding_lookups += 1
generation = self.generations[workspace_uuid]
if expected_generation is not None and expected_generation != generation:
raise AssertionError(f'stale probe generation {expected_generation} != {generation}')
return WorkspaceExecutionBinding(
instance_uuid=self.instance_uuid,
workspace_uuid=workspace_uuid,
placement_generation=generation,
write_fenced=False,
state='active',
)
class _ProbeRequester(requester.ProviderAPIRequester):
name = 'capacity-probe'
closed = 0
async def invoke_llm(
self,
query,
model,
messages,
funcs=None,
extra_args=None,
remove_think=False,
):
return None
async def aclose(self) -> None:
type(self).closed += 1
class _ProbeAdapter:
killed = 0
def __init__(self, _config, _logger) -> None:
self.listeners = []
def register_listener(self, event_type, listener) -> None:
self.listeners.append((event_type, listener))
async def kill(self) -> None:
type(self).killed += 1
class _ProbeMCPSession:
closed = 0
def __init__(self, server_name: str) -> None:
self.server_name = server_name
async def shutdown(self) -> None:
type(self).closed += 1
def _sample_process() -> ProcessSample:
gc.collect()
process = psutil.Process()
try:
open_fds = process.num_fds()
except (AttributeError, psutil.Error):
open_fds = None
traced_current, traced_peak = tracemalloc.get_traced_memory()
return ProcessSample(
rss_bytes=process.memory_info().rss,
traced_current_bytes=traced_current,
traced_peak_bytes=traced_peak,
asyncio_tasks=len(asyncio.all_tasks()),
threads=process.num_threads(),
open_fds=open_fds,
)
class PopulatedWorkspaceProbe:
def __init__(self) -> None:
_ProbeRequester.closed = 0
_ProbeAdapter.killed = 0
_ProbeMCPSession.closed = 0
self.workspace_service = _ProbeWorkspaceService()
self.logger = _ProbeLogger()
self.app = SimpleNamespace(
logger=self.logger,
workspace_service=self.workspace_service,
persistence_mgr=SimpleNamespace(
mode=SimpleNamespace(value='cloud_runtime'),
),
pipeline_config_meta_trigger={'name': 'trigger', 'stages': []},
pipeline_config_meta_safety={'name': 'safety', 'stages': []},
pipeline_config_meta_ai={'name': 'ai', 'stages': []},
pipeline_config_meta_output={'name': 'output', 'stages': []},
task_mgr=SimpleNamespace(
cancel_by_scope=lambda *_args, **_kwargs: None,
cancel_task=lambda *_args, **_kwargs: None,
),
)
self.model_manager = ModelManager(self.app)
self.model_manager.requester_dict = {
_ProbeRequester.name: _ProbeRequester,
}
self.pipeline_manager = PipelineManager(self.app)
self.pipeline_manager.stage_dict = {}
self.rag_manager = RAGManager(self.app)
self.mcp_loader = MCPLoader(self.app)
self.platform_manager = PlatformManager(self.app)
self.platform_manager.adapter_dict = {
'capacity-probe': _ProbeAdapter,
}
self.generation_refs: dict[
int,
list[weakref.ReferenceType],
] = {}
def _context(
self,
workspace_uuid: str,
generation: int,
*,
bot_uuid: str | None = None,
pipeline_uuid: str | None = None,
) -> ExecutionContext:
return ExecutionContext(
instance_uuid=self.workspace_service.instance_uuid,
workspace_uuid=workspace_uuid,
placement_generation=generation,
bot_uuid=bot_uuid,
pipeline_uuid=pipeline_uuid,
)
async def load_generation(self, workspaces: int, generation: int) -> None:
for index in range(workspaces):
workspace_uuid = f'workspace-{index}'
provider_uuid = f'provider-{index}'
llm_uuid = f'llm-{index}'
embedding_uuid = f'embedding-{index}'
rerank_uuid = f'rerank-{index}'
pipeline_uuid = f'pipeline-{index}'
bot_uuid = f'bot-{index}'
kb_uuid = f'knowledge-{index}'
mcp_server_name = f'mcp-{index}'
self.workspace_service.generations[workspace_uuid] = generation
context = self._context(workspace_uuid, generation)
runtime_provider = await self.model_manager.load_provider(
context,
persistence_model.ModelProvider(
uuid=provider_uuid,
workspace_uuid=workspace_uuid,
name='Capacity Provider',
requester=_ProbeRequester.name,
base_url='https://capacity.invalid',
api_keys=['probe'],
),
)
await self.model_manager.cache_provider(context, runtime_provider)
runtime_llm = await self.model_manager.load_llm_model_with_provider(
context,
persistence_model.LLMModel(
uuid=llm_uuid,
workspace_uuid=workspace_uuid,
name='Capacity LLM',
provider_uuid=provider_uuid,
abilities=['func_call'],
extra_args={'temperature': 0.1},
),
runtime_provider,
)
await self.model_manager.cache_llm_model(context, runtime_llm)
runtime_embedding = await self.model_manager.load_embedding_model_with_provider(
context,
persistence_model.EmbeddingModel(
uuid=embedding_uuid,
workspace_uuid=workspace_uuid,
name='Capacity Embedding',
provider_uuid=provider_uuid,
extra_args={'dimensions': 1_024},
),
runtime_provider,
)
await self.model_manager.cache_embedding_model(
context,
runtime_embedding,
)
runtime_rerank = await self.model_manager.load_rerank_model_with_provider(
context,
persistence_model.RerankModel(
uuid=rerank_uuid,
workspace_uuid=workspace_uuid,
name='Capacity Rerank',
provider_uuid=provider_uuid,
extra_args={},
),
runtime_provider,
)
await self.model_manager.cache_rerank_model(
context,
runtime_rerank,
)
pipeline_context = self._context(
workspace_uuid,
generation,
pipeline_uuid=pipeline_uuid,
)
await self.pipeline_manager.load_pipeline(
pipeline_context,
persistence_pipeline.LegacyPipeline(
uuid=pipeline_uuid,
workspace_uuid=workspace_uuid,
name='Capacity Pipeline',
description='',
for_version='probe',
is_default=True,
stages=[],
config={},
extensions_preferences={},
),
_binding_validated=True,
)
runtime_pipeline = self.pipeline_manager._pipelines_by_key[
(
self.workspace_service.instance_uuid,
workspace_uuid,
pipeline_uuid,
)
]
runtime_kb = await self.rag_manager.load_knowledge_base(
context,
persistence_rag.KnowledgeBase(
uuid=kb_uuid,
workspace_uuid=workspace_uuid,
name='Capacity Knowledge',
description='',
knowledge_engine_plugin_id=None,
collection_id=kb_uuid,
creation_settings={},
retrieval_settings={},
),
_binding_validated=True,
)
await self.mcp_loader._assert_execution_active(context)
runtime_mcp = _ProbeMCPSession(mcp_server_name)
self.mcp_loader._register_session(
context,
mcp_server_name,
runtime_mcp,
)
bot_context = self._context(
workspace_uuid,
generation,
bot_uuid=bot_uuid,
)
runtime_bot = await self.platform_manager.load_bot(
bot_context,
persistence_bot.Bot(
uuid=bot_uuid,
workspace_uuid=workspace_uuid,
name='Capacity Bot',
description='',
adapter='capacity-probe',
adapter_config={},
enable=True,
use_pipeline_uuid=pipeline_uuid,
pipeline_routing_rules=[],
),
_binding_validated=True,
)
self.generation_refs.setdefault(generation, []).extend(
(
weakref.ref(runtime_provider),
weakref.ref(runtime_llm),
weakref.ref(runtime_embedding),
weakref.ref(runtime_rerank),
weakref.ref(runtime_pipeline),
weakref.ref(runtime_kb),
weakref.ref(runtime_mcp),
weakref.ref(runtime_bot),
)
)
await asyncio.sleep(0)
def retained_state(self) -> dict[str, int]:
return {
'model_providers': len(self.model_manager.provider_dict),
'llm_models': len(self.model_manager.llm_model_dict),
'embedding_models': len(self.model_manager.embedding_model_dict),
'rerank_models': len(self.model_manager.rerank_model_dict),
'model_scopes': len(self.model_manager._scope_generations),
'pipelines': len(self.pipeline_manager._pipelines_by_key),
'pipeline_scopes': len(self.pipeline_manager._scope_generations),
'knowledge_bases': len(self.rag_manager.knowledge_bases),
'knowledge_scopes': len(self.rag_manager._scope_generations),
'mcp_sessions': len(self.mcp_loader.sessions),
'mcp_scopes': len(self.mcp_loader._scope_generations),
'bots': len(self.platform_manager._bots_by_key),
'bot_scopes': len(self.platform_manager._scope_generations),
'requesters_closed': _ProbeRequester.closed,
'adapters_killed': _ProbeAdapter.killed,
'mcp_sessions_closed': _ProbeMCPSession.closed,
'binding_lookups': self.workspace_service.binding_lookups,
}
def assert_generation_state(
self,
workspaces: int,
generation: int,
) -> None:
state = self.retained_state()
cardinality_keys = (
'model_providers',
'llm_models',
'embedding_models',
'rerank_models',
'model_scopes',
'pipelines',
'pipeline_scopes',
'knowledge_bases',
'knowledge_scopes',
'mcp_sessions',
'mcp_scopes',
'bots',
'bot_scopes',
)
invalid = {key: value for key in cardinality_keys if (value := state[key]) != workspaces}
if invalid:
raise AssertionError(f'populated Workspace cardinality mismatch: {invalid}')
expected_retired = (generation - 1) * workspaces
if state['requesters_closed'] != expected_retired:
raise AssertionError(f'retired requester count {state["requesters_closed"]} != {expected_retired}')
if state['adapters_killed'] != expected_retired:
raise AssertionError(f'retired adapter count {state["adapters_killed"]} != {expected_retired}')
if state['mcp_sessions_closed'] != expected_retired:
raise AssertionError(f'retired MCP session count {state["mcp_sessions_closed"]} != {expected_retired}')
def assert_generation_collected(self, generation: int) -> None:
gc.collect()
references = self.generation_refs.pop(generation)
retained = sum(reference() is not None for reference in references)
if retained:
raise AssertionError(f'{retained} generation-{generation} runtime objects remain reachable')
async def _run(args: argparse.Namespace) -> dict:
scale = SCALES[args.scale]
tracemalloc.start()
probe = PopulatedWorkspaceProbe()
baseline = _sample_process()
phase_one_started = time.monotonic()
await probe.load_generation(scale.workspaces, 1)
phase_one_seconds = time.monotonic() - phase_one_started
probe.assert_generation_state(scale.workspaces, 1)
phase_one = _sample_process()
phase_one_state = probe.retained_state()
phase_two_started = time.monotonic()
await probe.load_generation(scale.workspaces, 2)
phase_two_seconds = time.monotonic() - phase_two_started
probe.assert_generation_state(scale.workspaces, 2)
probe.assert_generation_collected(1)
phase_two = _sample_process()
phase_two_state = probe.retained_state()
phase_three_started = time.monotonic()
await probe.load_generation(scale.workspaces, 3)
phase_three_seconds = time.monotonic() - phase_three_started
probe.assert_generation_state(scale.workspaces, 3)
probe.assert_generation_collected(2)
phase_three = _sample_process()
phase_three_state = probe.retained_state()
cardinality_keys = (
'model_providers',
'llm_models',
'embedding_models',
'rerank_models',
'model_scopes',
'pipelines',
'pipeline_scopes',
'knowledge_bases',
'knowledge_scopes',
'mcp_sessions',
'mcp_scopes',
'bots',
'bot_scopes',
)
if any(
phase_two_state[key] != phase_one_state[key] or phase_three_state[key] != phase_one_state[key]
for key in cardinality_keys
):
raise AssertionError(
'populated Workspace registries did not plateau: '
f'phase_one={phase_one_state}, phase_two={phase_two_state}, '
f'phase_three={phase_three_state}'
)
traced_growth = phase_three.traced_current_bytes - phase_two.traced_current_bytes
rss_growth = phase_three.rss_bytes - phase_two.rss_bytes
max_traced_growth = int(args.max_traced_growth_mib * 1024 * 1024)
max_rss_growth = int(args.max_rss_growth_mib * 1024 * 1024)
if traced_growth > max_traced_growth:
raise AssertionError(f'replacement traced memory grew by {traced_growth} bytes (limit {max_traced_growth})')
if rss_growth > max_rss_growth:
raise AssertionError(f'replacement RSS grew by {rss_growth} bytes (limit {max_rss_growth})')
phase_ratio = max(
phase_two_seconds,
phase_three_seconds,
) / max(phase_one_seconds, 0.000_001)
if phase_ratio > args.max_replacement_time_ratio:
raise AssertionError(f'replacement phase ratio {phase_ratio:.3f} exceeds {args.max_replacement_time_ratio:.3f}')
return {
'component': 'langbot-populated-workspaces',
'scale': args.scale,
'workspaces': scale.workspaces,
'passed': True,
'phase_seconds': {
'initial': round(phase_one_seconds, 3),
'replacement_one': round(phase_two_seconds, 3),
'replacement_two': round(phase_three_seconds, 3),
'maximum_replacement_ratio': round(phase_ratio, 3),
},
'samples': {
'baseline': asdict(baseline),
'phase_one': asdict(phase_one),
'phase_two': asdict(phase_two),
'phase_three': asdict(phase_three),
},
'replacement_growth': {
'rss_bytes': rss_growth,
'traced_current_bytes': traced_growth,
},
'retained_state': {
'phase_one': phase_one_state,
'phase_two': phase_two_state,
'phase_three': phase_three_state,
},
}
def _parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument('--scale', choices=tuple(SCALES), default='quick')
parser.add_argument('--max-traced-growth-mib', type=float, default=16.0)
parser.add_argument('--max-rss-growth-mib', type=float, default=64.0)
parser.add_argument(
'--max-replacement-time-ratio',
type=float,
default=3.0,
)
parser.add_argument('--json', action='store_true')
return parser.parse_args()
def main() -> None:
args = _parse_args()
result = asyncio.run(_run(args))
if args.json:
print(json.dumps(result, sort_keys=True))
else:
print(json.dumps(result, indent=2, sort_keys=True))
if __name__ == '__main__':
main()
-17
View File
@@ -98,23 +98,6 @@ message truncation. Do not treat it as a long-term QA contract.
## 5. Run Gates
Cross-repository contract gate (no model calls):
```bash
bin/lbs suite plan langbot-workspace-contract-gate
bin/lbs suite run langbot-workspace-contract-gate --run-id langbot-workspace-contract-local
```
Top-down workspace release gate (browser workflows plus one complex Agent task):
```bash
bin/lbs suite plan langbot-workspace-release-gate
bin/lbs suite run langbot-workspace-release-gate --run-id langbot-workspace-release-local --include-manual-check
```
Run `agent-run-ledger-audit` immediately after the complex Agent case. Set
`LANGBOT_AGENT_RUN_ID` when the test instance has concurrent operator traffic.
Fast contract gate, no live service required:
```bash
+1 -1
View File
@@ -13,7 +13,7 @@
"pretest": "node scripts/bootstrap-lbs.mjs",
"precheck": "node scripts/bootstrap-lbs.mjs",
"lbs": "node src/lbs.ts",
"test": "node test/lbs-cli.test.ts && python3 -m unittest discover -s test -p 'test_*.py'",
"test": "node test/lbs-cli.test.ts",
"validate": "node src/lbs.ts validate",
"index": "node src/lbs.ts index",
"index:check": "node src/lbs.ts index --check",
-29
View File
@@ -257,22 +257,6 @@
"type": "string",
"enum": ["0", "1", "false", "true"]
},
"automation_debug_chat_load_require_success": {
"type": "string",
"enum": ["0", "1", "false", "true"]
},
"automation_debug_chat_load_provider_model_thresholds_json": {
"type": "string"
},
"automation_fake_provider_pipeline_name": {
"type": "string"
},
"automation_fake_provider_model_name": {
"type": "string"
},
"automation_fake_provider_fallback_model_names": {
"type": "string"
},
"automation_fake_provider_response_text": {
"type": "string"
},
@@ -291,9 +275,6 @@
"automation_fake_provider_fail_every_n": {
"type": "string"
},
"automation_fake_provider_fail_models": {
"type": "string"
},
"automation_fake_provider_fault_status": {
"type": "string"
},
@@ -301,16 +282,6 @@
"type": "string",
"enum": ["0", "1", "false", "true"]
},
"automation_fake_provider_fail_after_first_chunk_delay_ms": {
"type": "string"
},
"automation_fake_provider_fail_after_first_chunk_mode": {
"type": "string",
"enum": ["disconnect", "error_event"]
},
"automation_fake_provider_fail_after_first_chunk_models": {
"type": "string"
},
"automation_fake_provider_dynamic_response": {
"type": "string",
"enum": ["0", "1", "false", "true"]
@@ -1,86 +0,0 @@
#!/usr/bin/env node
import { spawn } from "node:child_process";
import { readFile } from "node:fs/promises";
import { join } from "node:path";
import { env } from "node:process";
import {
ensureEvidence,
evidencePaths,
exitCode,
loadEnvFiles,
localIsoWithOffset,
resolveLangBotRepo,
writeResult,
} from "./lib/langbot-e2e.mjs";
await loadEnvFiles();
const caseId = env.LBS_CASE_ID || "agent-run-ledger-audit";
const paths = evidencePaths(caseId);
await ensureEvidence(paths);
const startedAt = new Date();
const auditPath = join(paths.evidenceDir, "ledger-audit.json");
const result = {
source: "automation",
case_id: caseId,
run_id: paths.runId,
started_at: startedAt.toISOString(),
started_at_local: localIsoWithOffset(startedAt),
finished_at: "",
finished_at_local: "",
status: "fail",
reason: "",
audited_run: null,
metrics_summary: null,
failures: [],
warnings: [],
evidence: {
ledger_audit_json: auditPath,
automation_result_json: paths.automationResultJson,
result_json: paths.resultJson,
},
evidence_collected: ["filesystem", "metrics", "api_diagnostic"],
};
function run(command, args, cwd) {
return new Promise((resolvePromise) => {
const child = spawn(command, args, { cwd, env, stdio: ["ignore", "pipe", "pipe"] });
let stderr = "";
child.stderr.on("data", (chunk) => { stderr += chunk; });
child.on("error", (error) => resolvePromise({ status: null, stderr, error }));
child.on("close", (status) => resolvePromise({ status, stderr, error: null }));
});
}
try {
const repo = await resolveLangBotRepo();
const python = join(repo, ".venv", "bin", "python");
const args = ["scripts/e2e/agent-run-ledger-audit.py", "--repo", repo, "--output", auditPath];
if (env.LANGBOT_AGENT_RUN_ID) args.push("--run-id", env.LANGBOT_AGENT_RUN_ID);
if (env.LANGBOT_AGENT_TOOL_AUTHORIZATION_MODE) {
args.push("--tool-authorization-mode", env.LANGBOT_AGENT_TOOL_AUTHORIZATION_MODE);
}
const execution = await run(python, args, process.cwd());
const report = JSON.parse(await readFile(auditPath, "utf8"));
result.status = report.status;
result.reason = report.reason;
result.audited_run = report.run || null;
result.metrics_summary = report.metrics || null;
result.failures = report.failures || [];
result.warnings = report.warnings || [];
if (execution.error && result.status === "pass") {
result.status = "env_issue";
result.reason = execution.error.message;
}
} catch (error) {
result.status = /ENOENT|not found|No matching/i.test(error.message) ? "env_issue" : "fail";
result.reason = error.message;
} finally {
const finishedAt = new Date();
result.finished_at = finishedAt.toISOString();
result.finished_at_local = localIsoWithOffset(finishedAt);
await writeResult(paths, result);
console.log(JSON.stringify(result, null, 2));
}
process.exit(exitCode(result.status));
@@ -1,380 +0,0 @@
#!/usr/bin/env python3
"""Audit one persisted AgentRunner run without exposing authorization secrets."""
from __future__ import annotations
import argparse
import asyncio
import datetime
import json
import pathlib
import re
import sys
import urllib.parse
import sqlalchemy
import yaml
from sqlalchemy.ext.asyncio import create_async_engine
from agent_run_ledger_policy import (
classify_invalid_tool_argument_errors,
classify_tool_authorization,
invalid_tool_argument_error_signal,
load_ledger_json,
)
def database_url(repo: pathlib.Path) -> str:
config = yaml.safe_load((repo / "data/config.yaml").read_text(encoding="utf-8")) or {}
database = config.get("database", {})
kind = database.get("use", "sqlite")
if kind == "sqlite":
path = pathlib.Path(database.get("sqlite", {}).get("path", "data/langbot.db"))
if not path.is_absolute():
path = repo / path
return f"sqlite+aiosqlite:///{path}"
if kind in {"postgres", "postgresql"}:
values = database.get("postgresql", {})
user = urllib.parse.quote_plus(str(values.get("user", "postgres")))
password = urllib.parse.quote_plus(str(values.get("password", "postgres")))
host = values.get("host", "127.0.0.1")
port = values.get("port", 5432)
name = values.get("database", "postgres")
return f"postgresql+asyncpg://{user}:{password}@{host}:{port}/{name}"
raise RuntimeError(f"Unsupported database backend: {kind}")
def parse_created_after(value: str | None) -> datetime.datetime | None:
if not value:
return None
parsed = datetime.datetime.fromisoformat(value.replace("Z", "+00:00"))
if parsed.tzinfo is not None:
parsed = parsed.astimezone(datetime.timezone.utc).replace(tzinfo=None)
return parsed
def event_matches_tool_call(data_json: str | None, tool_name: str, parameters: dict | None) -> bool:
try:
data = json.loads(data_json or "{}")
except (TypeError, ValueError):
return False
if not isinstance(data, dict) or data.get("tool_name") != tool_name:
return False
return parameters is None or data.get("parameters") == parameters
def collect_result_texts(value: object) -> list[str]:
texts: list[str] = []
if isinstance(value, dict):
for key, item in value.items():
if key == "text" and isinstance(item, str):
texts.append(item)
else:
texts.extend(collect_result_texts(item))
elif isinstance(value, list):
for item in value:
texts.extend(collect_result_texts(item))
return texts
async def audit(
repo: pathlib.Path,
run_id: str | None,
*,
created_after: datetime.datetime | None = None,
expected_tool_name: str | None = None,
expected_parameters: dict | None = None,
expected_result_text: str | None = None,
tool_authorization_mode: str = "strict",
) -> dict:
engine = create_async_engine(database_url(repo))
failures: list[dict] = []
warnings: list[dict] = []
try:
async with engine.connect() as connection:
if run_id:
run_row = (await connection.execute(
sqlalchemy.text("SELECT * FROM agent_run WHERE run_id = :run_id"),
{"run_id": run_id},
)).mappings().first()
elif expected_tool_name:
query = "SELECT * FROM agent_run"
params = {}
if created_after is not None:
query += " WHERE created_at >= :created_after"
params["created_after"] = created_after
query += " ORDER BY id DESC LIMIT 100"
candidates = (await connection.execute(sqlalchemy.text(query), params)).mappings().all()
run_row = None
for candidate in candidates:
started_rows = (await connection.execute(
sqlalchemy.text(
"SELECT data_json FROM agent_run_event "
"WHERE run_id = :run_id AND type = 'tool.call.started' ORDER BY sequence"
),
{"run_id": str(candidate["run_id"])},
)).mappings().all()
if any(
event_matches_tool_call(row.get("data_json"), expected_tool_name, expected_parameters)
for row in started_rows
):
run_row = candidate
break
else:
run_row = (await connection.execute(
sqlalchemy.text("SELECT * FROM agent_run ORDER BY id DESC LIMIT 1")
)).mappings().first()
if run_row is None:
status = "fail" if expected_tool_name else "env_issue"
return {
"status": status,
"reason": "No AgentRunner run contains the expected tool call." if expected_tool_name else "No matching AgentRunner run exists.",
"failures": [{"kind": "expected_tool_call_missing"}] if expected_tool_name else [],
"warnings": [],
}
selected_run_id = str(run_row["run_id"])
event_rows = (await connection.execute(
sqlalchemy.text("SELECT sequence, type, data_json, metadata_json FROM agent_run_event WHERE run_id = :run_id ORDER BY sequence"),
{"run_id": selected_run_id},
)).mappings().all()
finally:
await engine.dispose()
authorization = load_ledger_json(
run_row.get("authorization_json"),
field="agent_run.authorization_json",
failures=failures,
)
tools = authorization.get("resources", {}).get("tools", []) if isinstance(authorization, dict) else []
allowed_tools: dict[str, dict] = {}
incomplete_tool_metadata: list[dict] = []
for tool in tools if isinstance(tools, list) else []:
if not isinstance(tool, dict):
incomplete_tool_metadata.append({"tool_name": "", "missing": ["tool object"]})
continue
name = str(tool.get("tool_name", ""))
missing = []
if not name:
missing.append("tool_name")
if not str(tool.get("description", "")).strip():
missing.append("description")
if not isinstance(tool.get("parameters"), dict):
missing.append("parameters")
if not (tool.get("source") or tool.get("tool_type") or tool.get("source_id")):
missing.append("owner")
if missing:
incomplete_tool_metadata.append({"tool_name": name, "missing": missing})
if name:
allowed_tools[name] = tool
if incomplete_tool_metadata:
failures.append({"kind": "incomplete_tool_metadata", "tools": incomplete_tool_metadata})
starts: dict[str, list[dict]] = {}
completions: dict[str, list[dict]] = {}
event_types: list[str] = []
invalid_event_json = 0
suspicious_errors: list[dict] = []
invalid_tool_argument_errors: list[dict] = []
successful_tool_completion_sequences: list[int] = []
forbidden_pattern = re.compile(
r"invalid json(?! arguments)|unauthori[sz]ed|permission denied|forbidden|timed?\s*out|timeout",
re.I,
)
def error_surface(value: object) -> list[str]:
"""Collect diagnostic fields without treating normal tool parameters as errors."""
collected: list[str] = []
if not isinstance(value, dict):
return collected
for key, item in value.items():
normalized = str(key).lower()
if normalized in {"error", "code", "status", "reason", "error_message"} and item is not None and item != "":
collected.append(str(item))
if isinstance(item, dict):
collected.extend(error_surface(item))
return collected
for row in event_rows:
event_type = str(row["type"])
event_types.append(event_type)
before = len(failures)
data = load_ledger_json(
row.get("data_json"),
field=f"agent_run_event[{row['sequence']}].data_json",
failures=failures,
)
invalid_event_json += int(len(failures) > before)
if not isinstance(data, dict):
failures.append({"kind": "invalid_event_payload", "sequence": row["sequence"], "type": event_type})
continue
if event_type in {"tool.call.started", "tool.call.completed"}:
call_id = str(data.get("tool_call_id", ""))
item = {"sequence": row["sequence"], "tool_name": str(data.get("tool_name", "")), "data": data}
if not call_id:
failures.append({"kind": "missing_tool_call_id", "sequence": row["sequence"], "type": event_type})
elif event_type == "tool.call.started":
starts.setdefault(call_id, []).append(item)
else:
completions.setdefault(call_id, []).append(item)
if not data.get("error") and data.get("result") is not None:
successful_tool_completion_sequences.append(row["sequence"])
diagnostic_text = "\n".join(error_surface(data))
if event_type == "run.failed":
diagnostic_text += "\n" + json.dumps(data, ensure_ascii=True)
match = forbidden_pattern.search(diagnostic_text)
if match:
suspicious_errors.append({"sequence": row["sequence"], "type": event_type, "signal": match.group(0)})
elif event_type == "tool.call.completed":
signal = invalid_tool_argument_error_signal(diagnostic_text)
if signal:
invalid_tool_argument_errors.append(
{"sequence": row["sequence"], "type": event_type, "signal": signal}
)
if run_row["status"] != "completed":
failures.append({"kind": "run_status", "actual": run_row["status"], "expected": "completed"})
if "run.completed" not in event_types:
failures.append({"kind": "missing_run_completed_event"})
if "run.failed" in event_types:
failures.append({"kind": "run_failed_event"})
all_call_ids = sorted(set(starts) | set(completions))
unauthorized_calls = []
for call_id in all_call_ids:
started = starts.get(call_id, [])
completed = completions.get(call_id, [])
if len(started) != 1 or len(completed) != 1:
failures.append({"kind": "tool_call_pairing", "tool_call_id": call_id, "started": len(started), "completed": len(completed)})
continue
if started[0]["tool_name"] != completed[0]["tool_name"]:
failures.append({"kind": "tool_name_mismatch", "tool_call_id": call_id})
if started[0]["sequence"] >= completed[0]["sequence"]:
failures.append({"kind": "tool_call_order", "tool_call_id": call_id})
if started[0]["tool_name"] not in allowed_tools:
unauthorized_calls.append({"tool_call_id": call_id, "tool_name": started[0]["tool_name"]})
authorization_failures, authorization_warnings = classify_tool_authorization(
unauthorized_calls,
authorization_mode=tool_authorization_mode,
)
failures.extend(authorization_failures)
warnings.extend(authorization_warnings)
unrecovered_argument_errors, recovered_argument_warnings = classify_invalid_tool_argument_errors(
invalid_tool_argument_errors,
successful_tool_completion_sequences=successful_tool_completion_sequences,
run_completed=(
run_row["status"] == "completed"
and "run.completed" in event_types
and "run.failed" not in event_types
),
)
if unrecovered_argument_errors:
suspicious_errors.extend(unrecovered_argument_errors)
warnings.extend(recovered_argument_warnings)
if suspicious_errors:
failures.append({"kind": "forbidden_error_signals", "events": suspicious_errors})
if not event_rows:
failures.append({"kind": "missing_run_events"})
if not tools:
warnings.append({"kind": "no_authorized_tools", "reason": "The run authorization snapshot exposes no tools."})
expected_call_summary = None
if expected_tool_name:
matching_starts = [
item
for items in starts.values()
for item in items
if item["tool_name"] == expected_tool_name
and (expected_parameters is None or item["data"].get("parameters") == expected_parameters)
]
if len(matching_starts) != 1:
failures.append({"kind": "expected_tool_call_count", "actual": len(matching_starts), "expected": 1})
matching_completions = []
for started in matching_starts:
call_id = str(started["data"].get("tool_call_id", ""))
matching_completions.extend(completions.get(call_id, []))
result_text_match = expected_result_text is None or any(
expected_result_text in collect_result_texts(completed["data"].get("result"))
for completed in matching_completions
)
if expected_result_text is not None and not result_text_match:
failures.append({"kind": "expected_tool_result_text_missing"})
expected_call_summary = {
"tool_name": expected_tool_name,
"parameters_match_required": expected_parameters is not None,
"matched_started_count": len(matching_starts),
"matched_completed_count": len(matching_completions),
"result_text_match_required": expected_result_text is not None,
"result_text_match": result_text_match,
}
metrics = {
"event_count": len(event_rows),
"tool_call_started": sum(len(items) for items in starts.values()),
"tool_call_completed": sum(len(items) for items in completions.values()),
"tool_call_ids": len(all_call_ids),
"authorized_tool_count": len(allowed_tools),
"tool_authorization_mode": tool_authorization_mode,
"runner_native_tool_call_count": len(unauthorized_calls) if tool_authorization_mode == "runner-native" else 0,
"invalid_event_json": invalid_event_json,
"suspicious_error_count": len(suspicious_errors),
"recovered_tool_argument_error_count": len(recovered_argument_warnings),
}
return {
"status": "pass" if not failures else "fail",
"reason": "Agent run ledger audit passed." if not failures else f"Agent run ledger audit found {len(failures)} invariant failure(s).",
"run": {
"run_id": selected_run_id,
"runner_id": run_row["runner_id"],
"status": run_row["status"],
"created_at": str(run_row["created_at"]),
"finished_at": str(run_row["finished_at"]),
},
"metrics": metrics,
"expected_tool_call": expected_call_summary,
"failures": failures,
"warnings": warnings,
}
def main() -> int:
parser = argparse.ArgumentParser()
parser.add_argument("--repo", required=True)
parser.add_argument("--run-id")
parser.add_argument("--created-after")
parser.add_argument("--expected-tool-name")
parser.add_argument("--expected-parameters-json")
parser.add_argument("--expected-result-text")
parser.add_argument(
"--tool-authorization-mode",
choices=("strict", "runner-native"),
default="strict",
)
parser.add_argument("--output", required=True)
args = parser.parse_args()
try:
expected_parameters = None
if args.expected_parameters_json:
expected_parameters = json.loads(args.expected_parameters_json)
if not isinstance(expected_parameters, dict):
raise ValueError("--expected-parameters-json must decode to an object")
if (expected_parameters is not None or args.expected_result_text) and not args.expected_tool_name:
raise ValueError("--expected-tool-name is required with expected parameters or result text")
report = asyncio.run(audit(
pathlib.Path(args.repo).resolve(),
args.run_id,
created_after=parse_created_after(args.created_after),
expected_tool_name=args.expected_tool_name,
expected_parameters=expected_parameters,
expected_result_text=args.expected_result_text,
tool_authorization_mode=args.tool_authorization_mode,
))
except Exception as exc: # noqa: BLE001 - probe must classify environment failures
report = {"status": "env_issue", "reason": str(exc), "failures": [], "warnings": []}
pathlib.Path(args.output).write_text(json.dumps(report, indent=2) + "\n", encoding="utf-8")
print(json.dumps(report))
return 0 if report["status"] == "pass" else 2 if report["status"] == "env_issue" else 1
if __name__ == "__main__":
sys.exit(main())
@@ -1,265 +0,0 @@
#!/usr/bin/env node
import {
apiJson,
createBrowser,
ensureAuthenticatedBrowser,
ensureEvidence,
evidencePaths,
exitCode,
loadEnvFiles,
localIsoWithOffset,
safeScreenshot,
scanBrowserDiagnostics,
writeResult,
} from "./lib/langbot-e2e.mjs";
const caseId = "agent-runner-health-visibility";
await loadEnvFiles();
const paths = evidencePaths(caseId);
await ensureEvidence(paths);
const readyScreenshot = paths.screenshot.replace(/\.png$/, "-ready.png");
const unavailableScreenshot = paths.screenshot.replace(
/\.png$/,
"-unavailable.png",
);
const startedAt = new Date();
const frontendUrl = process.env.LANGBOT_FRONTEND_URL || "";
const backendUrl = process.env.LANGBOT_BACKEND_URL || "";
const missingRunnerId = "plugin:qa/missing-runner/default";
let browser;
let token = "";
let agentId = "";
const result = {
source: "automation",
case_id: caseId,
run_id: paths.runId,
started_at: startedAt.toISOString(),
started_at_local: localIsoWithOffset(startedAt),
finished_at: "",
finished_at_local: "",
status: "fail",
reason: "",
url: "",
visible_signals: [],
api: {},
diagnostics: null,
cleanup: null,
evidence: {
console_log: paths.consoleLog,
network_log: paths.networkLog,
screenshot: paths.screenshot,
ready_screenshot: readyScreenshot,
unavailable_screenshot: unavailableScreenshot,
automation_result_json: paths.automationResultJson,
result_json: paths.resultJson,
},
evidence_collected: ["ui", "screenshot", "console", "api_diagnostic"],
};
function schemaDefaults(items = []) {
return Object.fromEntries(
items
.filter((item) => item.name && Object.hasOwn(item, "default"))
.map((item) => [item.name, item.default]),
);
}
async function openRunnerSettings(page, agentUrl) {
await page.goto(agentUrl, { waitUntil: "domcontentloaded" });
await page
.getByRole("button", { name: /Runner|运行器|ランナー/ })
.first()
.waitFor({ timeout: 15_000 });
await page
.getByRole("button", { name: /Runner|运行器|ランナー/ })
.first()
.click();
}
try {
if (!frontendUrl) throw new Error("LANGBOT_FRONTEND_URL is not configured.");
if (!backendUrl) throw new Error("LANGBOT_BACKEND_URL is not configured.");
browser = await createBrowser(paths);
const { page } = browser;
await page.goto(frontendUrl, { waitUntil: "domcontentloaded" });
const auth = await ensureAuthenticatedBrowser(page, {
frontendUrl,
backendUrl,
});
if (auth.status !== "pass") {
result.status = auth.status;
throw new Error(auth.reason);
}
token = await page.evaluate(() => localStorage.getItem("token") || "");
if (!token) {
result.status = "blocked";
throw new Error("Authenticated browser has no reusable local token.");
}
const pluginStatus = await apiJson(
backendUrl,
"/api/v1/system/status/plugin-system",
{ token },
);
const pluginData = pluginStatus.json.data || {};
result.api.plugin_status = {
http_status: pluginStatus.status,
code: pluginStatus.json.code ?? null,
is_enable: pluginData.is_enable ?? null,
is_connected: pluginData.is_connected ?? null,
};
if (pluginStatus.status >= 400 || pluginStatus.json.code !== 0) {
result.status = "env_issue";
throw new Error(pluginStatus.json.msg || "Plugin status request failed.");
}
if (!pluginData.is_enable || !pluginData.is_connected) {
result.status = "env_issue";
throw new Error("The plugin runtime is not enabled and connected.");
}
const metadata = await apiJson(backendUrl, "/api/v1/agents/_/metadata", {
token,
});
const runnerTab = metadata.json.data?.runner_config;
const runnerStage = runnerTab?.stages?.find(
(stage) => stage.name === "runner",
);
const runnerOptions =
runnerStage?.config?.find((item) => item.name === "id")?.options || [];
const runner = runnerOptions[0];
result.api.agent_metadata = {
http_status: metadata.status,
code: metadata.json.code ?? null,
runner_count: runnerOptions.length,
selected_runner: runner?.name || null,
};
if (metadata.status >= 400 || metadata.json.code !== 0) {
throw new Error(metadata.json.msg || "Agent metadata request failed.");
}
if (!runner?.name) {
result.status = "blocked";
throw new Error("No registered AgentRunner is available for the UI check.");
}
const runnerConfigStage = runnerTab.stages.find(
(stage) => stage.name === runner.name,
);
const create = await apiJson(backendUrl, "/api/v1/agents", {
method: "POST",
token,
body: {
kind: "agent",
name: `Runner Health ${paths.runId.slice(-40)}`,
description: "Temporary AgentRunner health visibility fixture",
emoji: "H",
component_ref: runner.name,
config: {
runner: { id: runner.name, "expire-time": 0 },
runner_config: {
[runner.name]: schemaDefaults(runnerConfigStage?.config),
},
},
enabled: true,
supported_event_patterns: ["message.*"],
},
});
agentId = create.json.data?.uuid || "";
result.api.create_agent = {
http_status: create.status,
code: create.json.code ?? null,
};
if (create.status >= 400 || create.json.code !== 0 || !agentId) {
throw new Error(create.json.msg || "Failed to create the temporary Agent.");
}
const agentUrl = `${frontendUrl.replace(/\/$/, "")}/home/agents?id=${encodeURIComponent(agentId)}`;
result.url = agentUrl;
await openRunnerSettings(page, agentUrl);
await page
.getByText(/Runner ready|运行器已就绪|Runner の準備完了/, { exact: true })
.waitFor({ timeout: 15_000 });
result.visible_signals.push("registered-runner-ready");
await safeScreenshot(page, readyScreenshot);
const staleUpdate = await apiJson(
backendUrl,
`/api/v1/agents/${encodeURIComponent(agentId)}`,
{
method: "PUT",
token,
body: {
component_ref: missingRunnerId,
config: {
runner: { id: missingRunnerId, "expire-time": 0 },
runner_config: { [missingRunnerId]: {} },
},
},
},
);
result.api.set_stale_runner = {
http_status: staleUpdate.status,
code: staleUpdate.json.code ?? null,
};
if (staleUpdate.status >= 400 || staleUpdate.json.code !== 0) {
throw new Error(
staleUpdate.json.msg || "Failed to set the stale runner fixture.",
);
}
await openRunnerSettings(page, agentUrl);
await page
.getByText(
/Selected runner is unavailable|所选运行器不可用|選択した Runner は利用できません/,
{ exact: true },
)
.waitFor({ timeout: 15_000 });
await page.getByRole("link", { name: /Extensions|扩展|拡張機能/ }).waitFor();
result.visible_signals.push("stale-runner-unavailable", "recovery-action");
await safeScreenshot(page, unavailableScreenshot);
await safeScreenshot(page, paths.screenshot);
result.diagnostics = await scanBrowserDiagnostics(paths);
if (result.diagnostics.status !== "pass") {
throw new Error(result.diagnostics.reason);
}
result.status = "pass";
result.reason =
"Agent Runner settings visibly distinguished a registered runner from a stale binding.";
} catch (error) {
if (!["blocked", "env_issue"].includes(result.status)) result.status = "fail";
result.reason = result.reason || error.message;
if (browser?.page) await safeScreenshot(browser.page, paths.screenshot);
} finally {
const cleanup = {};
if (agentId && token && backendUrl) {
const deletedAgent = await apiJson(
backendUrl,
`/api/v1/agents/${encodeURIComponent(agentId)}`,
{ method: "DELETE", token },
).catch((error) => ({
status: 0,
json: { code: null, msg: error.message },
}));
cleanup.agent_deleted =
deletedAgent.status < 400 && deletedAgent.json.code === 0;
cleanup.agent_http_status = deletedAgent.status;
}
result.cleanup = cleanup;
if (agentId && !cleanup.agent_deleted && result.status === "pass") {
result.status = "fail";
result.reason = "The temporary runner health Agent was not deleted.";
}
if (browser) await browser.close().catch(() => {});
const finishedAt = new Date();
result.finished_at = finishedAt.toISOString();
result.finished_at_local = localIsoWithOffset(finishedAt);
await writeResult(paths, result);
console.log(JSON.stringify(result, null, 2));
}
process.exit(exitCode(result.status));
@@ -70,7 +70,7 @@ const startedAt = new Date();
const targets = [
{
id: "local-agent",
expected_runner_id: "plugin:langbot-team/LocalAgent/default",
expected_runner_id: "plugin:langbot/local-agent/default",
pipeline_url: firstEnv("LANGBOT_LOCAL_AGENT_PIPELINE_URL"),
pipeline_name: firstEnv("LANGBOT_LOCAL_AGENT_PIPELINE_NAME"),
require_func_call_model: true,
@@ -79,7 +79,7 @@ const targets = [
},
{
id: "acp-agent-runner",
expected_runner_id: "plugin:langbot-team/ACPAgentRunner/default",
expected_runner_id: "plugin:langbot/acp-agent-runner/default",
pipeline_url: firstEnv("LANGBOT_ACP_AGENT_RUNNER_PIPELINE_URL", "LANGBOT_AGENT_RUNNER_PIPELINE_URL"),
pipeline_name: firstEnv("LANGBOT_ACP_AGENT_RUNNER_PIPELINE_NAME", "LANGBOT_AGENT_RUNNER_PIPELINE_NAME"),
require_func_call_model: false,
@@ -219,7 +219,7 @@ async function run() {
return metadata.author && metadata.name ? `${metadata.author}/${metadata.name}` : "";
})
.filter(Boolean);
const requiredPlugins = ["langbot-team/LocalAgent", "langbot-team/ACPAgentRunner", "qa/plugin-smoke"];
const requiredPlugins = ["langbot/local-agent", "langbot/acp-agent-runner", "qa/plugin-smoke"];
const pluginPresence = Object.fromEntries(requiredPlugins.map((id) => [id, installedPluginIds.includes(id)]));
for (const [id, present] of Object.entries(pluginPresence)) {
addCheck(`plugin:${id}`, present ? "pass" : "blocked", { plugin_id: id, reason: present ? "" : "Required plugin is not listed by /api/v1/plugins." });
@@ -309,7 +309,7 @@ async function run() {
const config = pipeline.config || {};
const aiConfig = config.ai && typeof config.ai === "object" ? config.ai : {};
const runner = aiConfig.runner && typeof aiConfig.runner === "object" ? aiConfig.runner : {};
const runnerId = runner.id || "";
const runnerId = runner.id || runner.runner || "";
const runnerConfigs = aiConfig.runner_config && typeof aiConfig.runner_config === "object" ? aiConfig.runner_config : {};
const runnerConfig = runnerConfigs[runnerId] && typeof runnerConfigs[runnerId] === "object" ? runnerConfigs[runnerId] : {};
const pipelineSummary = {
@@ -1,78 +0,0 @@
"""Policy helpers for classifying AgentRunner ledger error signals."""
from __future__ import annotations
import json
import re
_INVALID_TOOL_ARGUMENT_PATTERN = re.compile(
r"invalid json arguments|\b\d+\s+validation errors?\s+for\s+[A-Za-z_][A-Za-z0-9_]*Args\b",
re.IGNORECASE,
)
def load_ledger_json(value: str | None, *, field: str, failures: list[dict]) -> object:
"""Decode persisted ledger JSON and retain corruption as an invariant failure."""
if not value:
return {}
try:
return json.loads(value)
except (TypeError, ValueError) as exc:
failures.append({"kind": "invalid_json", "field": field, "reason": str(exc)})
return {}
def invalid_tool_argument_error_signal(value: str) -> str:
"""Return the persisted signal for malformed model-supplied tool arguments."""
match = _INVALID_TOOL_ARGUMENT_PATTERN.search(value)
return match.group(0) if match else ""
def classify_invalid_tool_argument_errors(
events: list[dict],
*,
successful_tool_completion_sequences: list[int],
run_completed: bool,
) -> tuple[list[dict], list[dict]]:
"""Split malformed tool arguments into recovered warnings and hard failures."""
failures: list[dict] = []
warnings: list[dict] = []
for event in events:
recovered = run_completed and any(
sequence > event["sequence"]
for sequence in successful_tool_completion_sequences
)
if recovered:
warnings.append(
{
"kind": "recovered_tool_argument_error",
"event": event,
"reason": "The model continued with a later successful tool call and the run completed.",
}
)
else:
failures.append(event)
return failures, warnings
def classify_tool_authorization(
calls: list[dict],
*,
authorization_mode: str,
) -> tuple[list[dict], list[dict]]:
"""Classify tool names absent from the Host authorization snapshot."""
if not calls:
return [], []
if authorization_mode == "runner-native":
return [], [
{
"kind": "runner_native_tool_calls",
"calls": calls,
"reason": (
"External runner tool telemetry is not a LangBot Host tool call; "
"the runner's own permission system governs it."
),
}
]
return [{"kind": "unauthorized_tool_calls", "calls": calls}], []
@@ -1,345 +0,0 @@
#!/usr/bin/env node
import {
apiJson,
bodyText,
createBrowser,
ensureAuthenticatedBrowser,
ensureEvidence,
evidencePaths,
exitCode,
loadEnvFiles,
localIsoWithOffset,
safeScreenshot,
scanBrowserDiagnostics,
writeResult,
} from "./lib/langbot-e2e.mjs";
const caseId = "bot-event-routing-product-flow";
await loadEnvFiles();
const paths = evidencePaths(caseId);
await ensureEvidence(paths);
const mobileScreenshot = paths.screenshot.replace(/\.png$/, "-mobile.png");
const scenarioMenuScreenshot = paths.screenshot.replace(
/\.png$/,
"-scenario-menu.png",
);
const startedAt = new Date();
const frontendUrl = process.env.LANGBOT_FRONTEND_URL || "";
const backendUrl = process.env.LANGBOT_BACKEND_URL || "";
const fixtureName = `EBA Product Flow ${paths.runId}`;
let browser;
let token = "";
let botId = "";
const result = {
source: "automation",
case_id: caseId,
run_id: paths.runId,
started_at: startedAt.toISOString(),
started_at_local: localIsoWithOffset(startedAt),
finished_at: "",
finished_at_local: "",
status: "fail",
reason: "",
url: "",
bot_id: "",
visible_signals: [],
api: {},
diagnostics: null,
cleanup: null,
evidence: {
console_log: paths.consoleLog,
network_log: paths.networkLog,
screenshot: paths.screenshot,
mobile_screenshot: mobileScreenshot,
scenario_menu_screenshot: scenarioMenuScreenshot,
automation_result_json: paths.automationResultJson,
result_json: paths.resultJson,
},
evidence_collected: ["ui", "screenshot", "console", "api_diagnostic"],
};
try {
if (!frontendUrl) throw new Error("LANGBOT_FRONTEND_URL is not configured.");
if (!backendUrl) throw new Error("LANGBOT_BACKEND_URL is not configured.");
browser = await createBrowser(paths);
const { page } = browser;
await page.goto(frontendUrl, { waitUntil: "domcontentloaded" });
const auth = await ensureAuthenticatedBrowser(page, {
frontendUrl,
backendUrl,
});
if (auth.status !== "pass") {
result.status = auth.status;
throw new Error(auth.reason);
}
token = await page.evaluate(() => localStorage.getItem("token") || "");
if (!token) {
result.status = "blocked";
throw new Error("Authenticated browser has no reusable local token.");
}
await page.goto(`${frontendUrl.replace(/\/$/, "")}/home/bots?id=new`, {
waitUntil: "domcontentloaded",
});
await page
.getByText(/Create Bot|创建机器人|ボットを作成/)
.first()
.waitFor({ timeout: 15_000 });
await page.getByRole("combobox").first().click();
await page
.getByRole("option", { name: /HTTP Bot|HTTP 通用接入|HTTP ボット/ })
.click();
await page
.getByText(/Event Routing|事件路由|イベントルーティング/)
.first()
.waitFor();
const addBehavior = page.getByRole("button", {
name: /Add behavior|添加行为|動作を追加/,
});
await addBehavior.waitFor();
await addBehavior.click();
const messageBehavior = page.getByRole("menuitem", {
name: /Reply to messages|回复收到的消息|受信メッセージに返信/,
});
const noEventRoutes = page.getByText(
/No event routes|暂无事件路由|イベントルートはありません/,
{ exact: true },
);
await noEventRoutes.waitFor();
try {
await messageBehavior.waitFor({ timeout: 5_000 });
} catch {
// Async adapter fields can rerender once after the first menu click.
await addBehavior.click();
await messageBehavior.waitFor();
}
await page.waitForTimeout(250);
await safeScreenshot(page, scenarioMenuScreenshot);
await messageBehavior.click();
await noEventRoutes.waitFor({ state: "hidden" });
result.visible_signals.push("create-mode-routing", "scenario-route-added");
const create = await apiJson(backendUrl, "/api/v1/platform/bots", {
method: "POST",
token,
body: {
name: fixtureName,
description: "Temporary EBA product-flow fixture",
adapter: "http_bot",
adapter_config: {
inbound_secret: "eba-product-flow-local-only",
callback_url: "http://127.0.0.1:9/langbot-e2e-unused",
outbound_secret: "",
default_session_type: "person",
signature_required: false,
callback_timeout: 1,
callback_max_retries: 0,
},
enable: true,
event_bindings: [
{
id: "qa-message-discard",
event_pattern: "message.received",
target_type: "discard",
target_uuid: "",
filters: [],
priority: 0,
enabled: true,
description: "Discard the deterministic QA event",
order: 0,
},
{
id: "qa-message-discard-shadowed",
event_pattern: "message.received",
target_type: "discard",
target_uuid: "",
filters: [],
priority: 0,
enabled: true,
description: "Verify visible route conflict guidance",
order: 1,
},
],
},
});
botId = create.json.data?.uuid || "";
result.api.create_bot = {
http_status: create.status,
code: create.json.code ?? null,
};
result.bot_id = botId;
if (create.status >= 400 || create.json.code !== 0 || !botId) {
throw new Error(create.json.msg || "Failed to create the temporary Bot.");
}
const botUrl = `${frontendUrl.replace(/\/$/, "")}/home/bots?id=${encodeURIComponent(botId)}`;
await page.goto(botUrl, { waitUntil: "domcontentloaded" });
await page
.waitForLoadState("networkidle", { timeout: 10_000 })
.catch(() => {});
result.url = page.url();
await page
.getByText(/Event Routing|事件路由|イベントルーティング/)
.first()
.waitFor({ timeout: 15_000 });
await page
.getByText(
/Events this adapter can receive|此适配器可接收的事件|このアダプターが受信できるイベント/,
)
.waitFor();
await page
.getByText(/Message received|收到消息|メッセージを受信/)
.first()
.waitFor();
await page
.getByText(
/Some routes overlap|部分路由存在覆盖冲突|一部のルートが重複しています/,
)
.waitFor();
await page
.getByText(
/Events that match no route are ignored|未命中任何路由的事件会被忽略|どのルートにも一致しないイベントは無視されます/,
)
.waitFor();
result.visible_signals.push(
"event-routing",
"adapter-capabilities",
"friendly-event-name",
"route-conflict-guidance",
"fallback-guidance",
);
await page
.getByRole("button", { name: /Test route|测试路由|ルートをテスト/ })
.click();
await page.getByRole("dialog").waitFor();
await page
.getByRole("button", { name: /Preview route|预览路由|ルートをプレビュー/ })
.click();
await page
.getByText(/Route matched|已命中路由|ルートに一致しました/)
.waitFor({ timeout: 15_000 });
await page
.getByText(/Discard|丢弃|破棄/)
.first()
.waitFor();
result.visible_signals.push("dry-run-matched", "discard-target");
await page
.getByRole("button", {
name: /Run saved route|运行已保存路由|保存済みルートを実行/,
})
.click();
await page
.getByText(
/saved route ran successfully|已保存路由运行成功|保存済みルートを実行しました/,
)
.waitFor({ timeout: 20_000 });
result.visible_signals.push("test-event-dispatched");
await page
.getByRole("button", { name: /Close|关闭|閉じる/ })
.first()
.click();
await page.getByRole("dialog").waitFor({ state: "hidden" });
await page
.getByText(/Discarded|已丢弃|破棄済み/)
.first()
.waitFor({ timeout: 10_000 });
result.visible_signals.push("route-status-discarded");
const text = await bodyText(page);
if (/\bEBA event\b/.test(text)) {
throw new Error(
"The primary route surface exposes an internal EBA runtime message.",
);
}
const status = await apiJson(
backendUrl,
`/api/v1/platform/bots/${encodeURIComponent(botId)}/event-routes/status`,
{ token },
);
const latest = status.json.data?.routes?.find(
(route) => route.binding_id === "qa-message-discard",
);
result.api.route_status = {
http_status: status.status,
code: status.json.code ?? null,
last_status: latest?.last_status || null,
};
if (
status.status >= 400 ||
status.json.code !== 0 ||
latest?.last_status !== "discarded"
) {
throw new Error(
"Route status API did not confirm the discarded synthetic event.",
);
}
await safeScreenshot(page, paths.screenshot);
await page.setViewportSize({ width: 390, height: 844 });
await page.waitForTimeout(250);
const horizontalOverflow = await page.evaluate(
() => document.documentElement.scrollWidth - window.innerWidth,
);
if (horizontalOverflow > 1) {
throw new Error(
`The mobile route editor overflows horizontally by ${horizontalOverflow}px.`,
);
}
await page
.getByText(
/Some routes overlap|部分路由存在覆盖冲突|一部のルートが重複しています/,
)
.waitFor();
await safeScreenshot(page, mobileScreenshot);
result.visible_signals.push("mobile-layout");
result.diagnostics = await scanBrowserDiagnostics(paths);
if (result.diagnostics.status !== "pass") {
throw new Error(result.diagnostics.reason);
}
result.status = "pass";
result.reason =
"Bot event routing, dry-run, synthetic dispatch, and visible route status passed in the WebUI.";
} catch (error) {
if (!["blocked", "env_issue"].includes(result.status)) result.status = "fail";
result.reason = result.reason || error.message;
if (browser?.page) await safeScreenshot(browser.page, paths.screenshot);
} finally {
if (botId && token && backendUrl) {
const cleanup = await apiJson(
backendUrl,
`/api/v1/platform/bots/${encodeURIComponent(botId)}`,
{ method: "DELETE", token },
).catch((error) => ({
status: 0,
json: { code: null, msg: error.message },
}));
result.cleanup = {
http_status: cleanup.status,
code: cleanup.json.code ?? null,
deleted: cleanup.status < 400 && cleanup.json.code === 0,
};
if (!result.cleanup.deleted && result.status === "pass") {
result.status = "fail";
result.reason =
cleanup.json.msg || "The temporary Bot could not be deleted.";
}
}
if (browser) await browser.close().catch(() => {});
const finishedAt = new Date();
result.finished_at = finishedAt.toISOString();
result.finished_at_local = localIsoWithOffset(finishedAt);
await writeResult(paths, result);
console.log(JSON.stringify(result, null, 2));
}
process.exit(exitCode(result.status));
@@ -12,7 +12,7 @@ import {
writeResult,
} from "./lib/langbot-e2e.mjs";
const RUNNER_ID = "plugin:langbot-team/ACPAgentRunner/default";
const RUNNER_ID = "plugin:langbot/acp-agent-runner/default";
const DEFAULT_PIPELINE_NAME = "Agent QA ACP Claude Debug Chat";
const DEFAULT_LOCAL_PASSWORD = "LangBotE2ELocalPass!2026";
const caseId = "ensure-acp-agent-runner-pipeline";
@@ -102,7 +102,7 @@ try {
});
Object.assign(result, prepared);
if (result.pipeline_id) {
result.pipeline_url = `${frontendUrl.replace(/\/$/, "")}/home/agents?id=${encodeURIComponent(result.pipeline_id)}`;
result.pipeline_url = `${frontendUrl.replace(/\/$/, "")}/home/pipelines?id=${encodeURIComponent(result.pipeline_id)}`;
}
if (writeEnv && result.pipeline_id) {
@@ -14,7 +14,7 @@ import {
writeResult,
} from "./lib/langbot-e2e.mjs";
const RUNNER_ID = "plugin:langbot-team/LocalAgent/default";
const RUNNER_ID = "local-agent";
const DEFAULT_LOCAL_PASSWORD = "LangBotE2ELocalPass!2026";
const DEFAULT_PIPELINE_NAME = "LangBot QA Fake Provider Debug Chat";
const DEFAULT_PROVIDER_NAME = "LangBot QA Fake OpenAI Provider";
@@ -41,8 +41,6 @@ const pipelineName = env.LANGBOT_FAKE_PROVIDER_PIPELINE_NAME || DEFAULT_PIPELINE
const providerName = env.LANGBOT_FAKE_PROVIDER_NAME || DEFAULT_PROVIDER_NAME;
const requester = env.LANGBOT_FAKE_PROVIDER_REQUESTER || DEFAULT_REQUESTER;
const modelName = env.LANGBOT_FAKE_PROVIDER_MODEL_NAME || DEFAULT_MODEL_NAME;
const fallbackModelNames = textList(env.LANGBOT_FAKE_PROVIDER_FALLBACK_MODEL_NAMES)
.filter((name) => name !== modelName);
const result = {
source: "automation",
@@ -77,7 +75,6 @@ const result = {
test_status: "not_run",
test_reason: "",
},
fallback_models: [],
pipeline_id: "",
pipeline_name: pipelineName,
pipeline_url: "",
@@ -144,26 +141,14 @@ try {
});
result.model = model;
const fallbackModels = [];
for (const fallbackModelName of fallbackModelNames) {
fallbackModels.push(await ensureModel({
backendUrl,
token: auth.token,
providerUuid: provider.uuid,
name: fallbackModelName,
}));
}
result.fallback_models = fallbackModels;
const pipeline = await ensurePipeline({
backendUrl,
token: auth.token,
name: pipelineName,
modelUuid: model.uuid,
fallbackModelUuids: fallbackModels.map((item) => item.uuid),
});
Object.assign(result, pipeline);
result.pipeline_url = `${frontendUrl.replace(/\/$/, "")}/home/agents?id=${encodeURIComponent(pipeline.pipeline_id)}`;
result.pipeline_url = `${frontendUrl.replace(/\/$/, "")}/home/pipelines?id=${encodeURIComponent(pipeline.pipeline_id)}`;
const runConfig = await configureFakeProvider(fakeProvider.url, targetFakeProviderConfig(), true);
result.fake_provider.config = runConfig.config || targetFakeProviderConfig();
@@ -176,7 +161,6 @@ try {
LANGBOT_FAKE_PROVIDER_PID: fakeProvider.pid ? String(fakeProvider.pid) : "",
LANGBOT_FAKE_PROVIDER_PROVIDER_UUID: provider.uuid,
LANGBOT_FAKE_PROVIDER_MODEL_UUID: model.uuid,
LANGBOT_FAKE_PROVIDER_FALLBACK_MODEL_UUIDS: fallbackModels.map((item) => item.uuid).join(","),
LANGBOT_FAKE_PROVIDER_PIPELINE_URL: result.pipeline_url,
LANGBOT_FAKE_PROVIDER_PIPELINE_NAME: pipelineName,
});
@@ -184,7 +168,7 @@ try {
}
result.status = "pass";
result.reason = `Fake provider pipeline is configured with ${requester}/${modelName} and ${fallbackModels.length} fallback(s).`;
result.reason = `Fake provider pipeline is configured with ${requester}/${modelName}.`;
} catch (error) {
result.status = result.status === "env_issue" ? "env_issue" : "fail";
result.reason = result.reason || safeReason(error.message);
@@ -344,11 +328,7 @@ function healthyFakeProviderConfig() {
fault_status: 500,
fail_first_n: 0,
fail_every_n: 0,
fail_models: [],
fail_after_first_chunk: false,
fail_after_first_chunk_delay_ms: 0,
fail_after_first_chunk_mode: "disconnect",
fail_after_first_chunk_models: [],
dynamic_response: true,
};
}
@@ -362,14 +342,7 @@ function targetFakeProviderConfig() {
fault_status: httpFaultStatus(env.LANGBOT_FAKE_PROVIDER_FAULT_STATUS, 500),
fail_first_n: nonNegativeInteger(env.LANGBOT_FAKE_PROVIDER_FAIL_FIRST_N, 0),
fail_every_n: nonNegativeInteger(env.LANGBOT_FAKE_PROVIDER_FAIL_EVERY_N, 0),
fail_models: textList(env.LANGBOT_FAKE_PROVIDER_FAIL_MODELS),
fail_after_first_chunk: envBool(env.LANGBOT_FAKE_PROVIDER_FAIL_AFTER_FIRST_CHUNK, false),
fail_after_first_chunk_delay_ms: nonNegativeInteger(
env.LANGBOT_FAKE_PROVIDER_FAIL_AFTER_FIRST_CHUNK_DELAY_MS,
0,
),
fail_after_first_chunk_mode: streamFaultMode(env.LANGBOT_FAKE_PROVIDER_FAIL_AFTER_FIRST_CHUNK_MODE),
fail_after_first_chunk_models: textList(env.LANGBOT_FAKE_PROVIDER_FAIL_AFTER_FIRST_CHUNK_MODELS),
dynamic_response: envBool(env.LANGBOT_FAKE_PROVIDER_DYNAMIC_RESPONSE, true),
};
}
@@ -452,7 +425,7 @@ async function ensureModel({ backendUrl, token, providerUuid, name }) {
const body = {
name,
provider_uuid: providerUuid,
abilities: ["func_call", "vision"],
abilities: [],
context_length: positiveInteger(env.LANGBOT_FAKE_PROVIDER_CONTEXT_LENGTH, 8192),
extra_args: {},
prefered_ranking: 0,
@@ -503,7 +476,7 @@ async function ensureModel({ backendUrl, token, providerUuid, name }) {
};
}
async function ensurePipeline({ backendUrl, token, name, modelUuid, fallbackModelUuids = [] }) {
async function ensurePipeline({ backendUrl, token, name, modelUuid }) {
const list = await apiJson(backendUrl, "/api/v1/pipelines", { token });
if (isApiFailure(list)) {
throw new Error(list.json.msg || "Failed to list pipelines.");
@@ -538,17 +511,15 @@ async function ensurePipeline({ backendUrl, token, name, modelUuid, fallbackMode
const config = pipeline.config && typeof pipeline.config === "object" ? pipeline.config : {};
const ai = config.ai && typeof config.ai === "object" ? config.ai : {};
const runnerConfigs = ai.runner_config && typeof ai.runner_config === "object"
? ai.runner_config
: {};
const existingLocalAgentConfig = runnerConfigs[RUNNER_ID] && typeof runnerConfigs[RUNNER_ID] === "object"
? runnerConfigs[RUNNER_ID]
const existingLocalAgentConfig = ai["local-agent"] && typeof ai["local-agent"] === "object"
? ai["local-agent"]
: {};
const localAgentConfig = {
timeout: 60,
prompt: [{ role: "system", content: "You are a deterministic QA assistant. Reply exactly as instructed." }],
"remove-think": false,
"knowledge-bases": [],
"box-session-id-template": "{launcher_type}_{launcher_id}",
"retrieval-top-k": 5,
"rerank-model": "",
"rerank-top-k": 5,
@@ -565,7 +536,7 @@ async function ensurePipeline({ backendUrl, token, name, modelUuid, fallbackMode
"max-round": positiveInteger(existingLocalAgentConfig["max-round"], 10),
model: {
primary: modelUuid,
fallbacks: fallbackModelUuids,
fallbacks: [],
},
};
const updatedConfig = {
@@ -575,12 +546,10 @@ async function ensurePipeline({ backendUrl, token, name, modelUuid, fallbackMode
runner: {
...(ai.runner && typeof ai.runner === "object" ? ai.runner : {}),
id: RUNNER_ID,
runner: RUNNER_ID,
"expire-time": 0,
},
runner_config: {
...runnerConfigs,
[RUNNER_ID]: localAgentConfig,
},
"local-agent": localAgentConfig,
},
};
@@ -632,17 +601,6 @@ function envBool(value, fallback) {
return fallback;
}
function textList(value) {
return String(value || "")
.split(/\r?\n|,/)
.map((item) => item.trim())
.filter(Boolean);
}
function streamFaultMode(value) {
return String(value || "").trim().toLowerCase() === "error_event" ? "error_event" : "disconnect";
}
function sleep(ms) {
return new Promise((resolve) => setTimeout(resolve, ms));
}
@@ -21,7 +21,7 @@ await ensureEvidence(paths);
const backendUrl = env.LANGBOT_BACKEND_URL || "";
const user = env.LANGBOT_E2E_LOGIN_USER || "";
const password = env.LANGBOT_E2E_LOGIN_PASSWORD || "LangBotE2ELocalPass!2026";
const expectedText = env.LANGBOT_E2E_RAG_EXPECTED_TEXT || "azalea-cobalt-7421";
const expectedText = env.LANGBOT_E2E_EXPECTED_TEXT || "azalea-cobalt-7421";
const query = env.LANGBOT_E2E_RETRIEVE_QUERY || "What is the local agent runner retrieval sentinel?";
const writeEnv = process.argv.includes("--write-env");
const checkOnly = process.argv.includes("--check-only");
@@ -18,14 +18,12 @@ import {
writeResult,
} from "./lib/langbot-e2e.mjs";
const RUNNER_ID = "plugin:langbot-team/LocalAgent/default";
const RUNNER_ID = "local-agent";
const SPACE_PROVIDER_UUID = "00000000-0000-0000-0000-000000000000";
const DEFAULT_PIPELINE_NAME = "Agent QA Local Agent Debug Chat";
const DEFAULT_LOCAL_PASSWORD = "LangBotE2ELocalPass!2026";
const DEFAULT_MODEL_TEST_LIMIT = 8;
const DEFAULT_MODEL_FALLBACK_COUNT = 3;
const DEFAULT_FAKE_MODEL_UUID = "langbot-e2e-fake-local-agent-model";
const DEFAULT_FAKE_PROVIDER_NAME = "LangBot E2E Fake OpenAI Provider";
const caseId = "ensure-local-agent-pipeline";
await loadEnvFiles();
@@ -33,10 +31,7 @@ const paths = evidencePaths(caseId);
await ensureEvidence(paths);
const writeEnv = process.argv.includes("--write-env");
const pipelineName =
env.LANGBOT_E2E_CREATE_PIPELINE_NAME ||
env.LANGBOT_LOCAL_AGENT_PIPELINE_NAME ||
DEFAULT_PIPELINE_NAME;
const pipelineName = env.LANGBOT_E2E_CREATE_PIPELINE_NAME || env.LANGBOT_LOCAL_AGENT_PIPELINE_NAME || DEFAULT_PIPELINE_NAME;
const frontendUrl = env.LANGBOT_FRONTEND_URL || "";
const backendUrl = env.LANGBOT_BACKEND_URL || "";
const envLocalPath = resolve("skills/.env.local");
@@ -56,13 +51,11 @@ const result = {
selected_model_id: "",
selected_model_name: "",
fallback_model_ids: [],
fake_provider: null,
model_count: 0,
space_model_count: 0,
scanned_space_model_count: 0,
tested_model_count: 0,
model_tests: [],
plugin_setup: null,
created: false,
updated: false,
wrote_env: false,
@@ -90,9 +83,7 @@ try {
const password = env.LANGBOT_E2E_LOGIN_PASSWORD || DEFAULT_LOCAL_PASSWORD;
if (!user) {
result.status = "env_issue";
throw new Error(
"LANGBOT_E2E_LOGIN_USER is required so this setup can create/update the pipeline via backend API.",
);
throw new Error("LANGBOT_E2E_LOGIN_USER is required so this setup can create/update the pipeline via backend API.");
}
const auth = await resetAndAuthLocalUser({ backendUrl, user, password });
@@ -102,23 +93,11 @@ try {
backend_token_check: auth.check,
};
const pluginSetup = await ensureLocalAgentRunner({
backendUrl,
token: auth.token,
});
result.plugin_setup = pluginSetup;
if (pluginSetup.status !== "pass") {
result.status = pluginSetup.status === "env_issue" ? "env_issue" : "fail";
throw new Error(pluginSetup.reason || "Failed to prepare the LocalAgent runner plugin.");
}
const wizard = await skipWizard({ backendUrl, token: auth.token });
result.wizard = wizard;
if (wizard.status !== "pass") {
result.status = "fail";
throw new Error(
wizard.reason || "Failed to mark the local QA wizard as skipped.",
);
throw new Error(wizard.reason || "Failed to mark the local QA wizard as skipped.");
}
const prepared = await ensureLocalAgentPipeline({
@@ -129,7 +108,7 @@ try {
});
Object.assign(result, prepared);
if (result.pipeline_id) {
result.pipeline_url = `${frontendUrl.replace(/\/$/, "")}/home/agents?id=${encodeURIComponent(result.pipeline_id)}`;
result.pipeline_url = `${frontendUrl.replace(/\/$/, "")}/home/pipelines?id=${encodeURIComponent(result.pipeline_id)}`;
}
if (writeEnv && result.pipeline_id) {
@@ -139,12 +118,10 @@ try {
LANGBOT_PIPELINE_NAME: result.pipeline_name || pipelineName,
LANGBOT_LOCAL_AGENT_PIPELINE_URL: result.pipeline_url,
LANGBOT_LOCAL_AGENT_PIPELINE_NAME: result.pipeline_name || pipelineName,
...(result.selected_model_id
? {
LANGBOT_LOCAL_AGENT_MODEL_UUID: result.selected_model_id,
LANGBOT_E2E_MODEL_UUID: result.selected_model_id,
}
: {}),
...(result.selected_model_id ? {
LANGBOT_LOCAL_AGENT_MODEL_UUID: result.selected_model_id,
LANGBOT_E2E_MODEL_UUID: result.selected_model_id,
} : {}),
});
result.wrote_env = true;
}
@@ -155,21 +132,12 @@ try {
const browserCheck = await verifyBrowserToken(page, backendUrl);
result.browser_token_check = browserCheck;
if (!browserCheck.authenticated) {
throw new Error(
browserCheck.reason || "Browser token check failed after setup.",
);
throw new Error(browserCheck.reason || "Browser token check failed after setup.");
}
await page.goto(result.pipeline_url || frontendUrl, {
waitUntil: "domcontentloaded",
});
await page
.waitForLoadState("networkidle", { timeout: 10_000 })
.catch(() => {});
await page.goto(result.pipeline_url || frontendUrl, { waitUntil: "domcontentloaded" });
await page.waitForLoadState("networkidle", { timeout: 10_000 }).catch(() => {});
const text = await bodyText(page);
result.page_signal =
["Pipelines", "流水线", pipelineName].find((signal) =>
text.includes(signal),
) || "";
result.page_signal = ["Pipelines", "流水线", pipelineName].find((signal) => text.includes(signal)) || "";
} catch (error) {
result.status = result.status === "env_issue" ? "env_issue" : "fail";
result.reason = result.reason || error.message;
@@ -180,281 +148,24 @@ try {
console.log(JSON.stringify(result, null, 2));
}
process.exit(
result.status === "pass" ? 0 : result.status === "env_issue" ? 2 : 1,
);
process.exit(result.status === "pass" ? 0 : result.status === "env_issue" ? 2 : 1);
async function skipWizard({ backendUrl, token }) {
const response = await apiJson(
backendUrl,
"/api/v1/system/wizard/completed",
{
method: "POST",
token,
body: { status: "skipped" },
},
);
const response = await apiJson(backendUrl, "/api/v1/system/wizard/completed", {
method: "POST",
token,
body: { status: "skipped" },
});
const ok = response.status < 400 && response.json.code === 0;
return {
status: ok ? "pass" : "fail",
http_status: response.status,
code: response.json.code ?? null,
reason: ok
? "Wizard marked skipped for local QA."
: response.json.msg || "Wizard status update failed.",
reason: ok ? "Wizard marked skipped for local QA." : response.json.msg || "Wizard status update failed.",
};
}
async function ensureLocalAgentRunner({ backendUrl, token }) {
const [author, name] = RUNNER_ID.replace(/^plugin:/, "").split("/");
const existingRunnerIds = await listRunnerIds(backendUrl, token);
if (existingRunnerIds.includes(RUNNER_ID)) {
return {
status: "pass",
reason: "LocalAgent runner is already registered.",
plugin_id: `${author}/${name}`,
runner_id: RUNNER_ID,
installed: false,
};
}
const pluginsResponse = await apiJson(backendUrl, "/api/v1/plugins", {
token,
});
if (isApiFailure(pluginsResponse)) {
return {
status: "env_issue",
reason:
pluginsResponse.json.msg ||
"Failed to inspect installed plugins before preparing LocalAgent.",
};
}
const pluginId = `${author}/${name}`;
const pluginPresent = (pluginsResponse.json.data?.plugins || []).some(
(plugin) => {
const metadata =
plugin.manifest?.manifest?.metadata ||
plugin.manifest?.metadata ||
plugin.metadata ||
{};
return `${metadata.author}/${metadata.name}` === pluginId;
},
);
if (pluginPresent) {
const registered = await waitForRunnerRegistration({
backendUrl,
token,
runnerId: RUNNER_ID,
timeoutMs: 30_000,
});
return registered
? {
status: "pass",
reason: "Installed LocalAgent plugin finished Runner registration.",
plugin_id: pluginId,
runner_id: RUNNER_ID,
installed: false,
}
: {
status: "fail",
reason: `${pluginId} is installed but ${RUNNER_ID} did not register.`,
plugin_id: pluginId,
runner_id: RUNNER_ID,
installed: false,
};
}
const spaceUrl = String(env.LANGBOT_SPACE_URL || "https://space.langbot.app").replace(
/\/$/,
"",
);
let detailResponse;
try {
detailResponse = await fetch(
`${spaceUrl}/api/v1/marketplace/plugins/${encodeURIComponent(author)}/${encodeURIComponent(name)}`,
);
} catch (error) {
return {
status: "env_issue",
reason: `Could not reach LangBot Space for ${pluginId}: ${error.message}`,
};
}
const detail = await detailResponse.json().catch(() => ({}));
const version = detail?.data?.plugin?.latest_version || "";
if (detailResponse.status >= 400 || detail.code !== 0 || !version) {
return {
status: "env_issue",
reason:
detail.msg ||
`LangBot Space did not return an installable version for ${pluginId}.`,
};
}
const install = await apiJson(
backendUrl,
"/api/v1/plugins/install/marketplace",
{
method: "POST",
token,
body: {
plugin_author: author,
plugin_name: name,
plugin_version: version,
},
},
);
const taskId = install.json.data?.task_id ?? null;
if (isApiFailure(install) || !taskId) {
return {
status: "fail",
reason:
install.json.msg ||
`Marketplace install did not create a task for ${pluginId}.`,
plugin_id: pluginId,
version,
};
}
const task = await waitForTask({ backendUrl, token, taskId });
if (!taskComplete(task)) {
return {
status: taskFailed(task) ? "fail" : "env_issue",
reason:
task?.runtime?.exception ||
task?.error ||
`Marketplace install task for ${pluginId} did not complete.`,
plugin_id: pluginId,
version,
task_id: taskId,
};
}
const registered = await waitForRunnerRegistration({
backendUrl,
token,
runnerId: RUNNER_ID,
timeoutMs: 60_000,
});
return registered
? {
status: "pass",
reason: `Installed ${pluginId} ${version} and registered ${RUNNER_ID}.`,
plugin_id: pluginId,
runner_id: RUNNER_ID,
version,
task_id: taskId,
installed: true,
}
: {
status: "fail",
reason: `Installed ${pluginId} ${version}, but ${RUNNER_ID} did not register.`,
plugin_id: pluginId,
runner_id: RUNNER_ID,
version,
task_id: taskId,
installed: true,
};
}
async function listRunnerIds(backendUrl, token) {
const response = await apiJson(backendUrl, "/api/v1/pipelines/_/metadata", {
token,
});
if (isApiFailure(response)) return [];
return (response.json.data?.configs || [])
.flatMap((section) => section.stages || [])
.flatMap((stage) => stage.config || [])
.filter((item) => item.name === "id")
.flatMap((item) => item.options || [])
.map((option) => option.name || option.value || option.id || "")
.filter(Boolean);
}
async function waitForRunnerRegistration({
backendUrl,
token,
runnerId,
timeoutMs,
}) {
const deadline = Date.now() + timeoutMs;
while (Date.now() < deadline) {
if ((await listRunnerIds(backendUrl, token)).includes(runnerId)) return true;
await sleep(1000);
}
return false;
}
async function waitForTask({ backendUrl, token, taskId }) {
const deadline =
Date.now() + Number(env.LANGBOT_PLUGIN_INSTALL_TIMEOUT_MS || 120_000);
let task = null;
while (Date.now() < deadline) {
const response = await apiJson(
backendUrl,
`/api/v1/system/tasks/${encodeURIComponent(taskId)}`,
{ token },
);
task = response.json.data || response.json;
if (taskComplete(task) || taskFailed(task)) return task;
await sleep(1000);
}
return task;
}
function taskComplete(task) {
const status = String(task?.status || task?.state || "").toLowerCase();
const runtimeStatus = String(
task?.runtime?.status || task?.runtime?.state || "",
).toLowerCase();
return (
["done", "completed", "success", "succeeded", "finished"].includes(
status,
) ||
["done", "completed", "success", "succeeded", "finished"].includes(
runtimeStatus,
) ||
task?.done === true ||
task?.completed === true ||
(task?.runtime?.done === true && !task?.runtime?.exception)
);
}
function taskFailed(task) {
const status = String(task?.status || task?.state || "").toLowerCase();
const runtimeStatus = String(
task?.runtime?.status || task?.runtime?.state || "",
).toLowerCase();
return (
["failed", "error", "cancelled", "canceled"].includes(status) ||
["failed", "error", "cancelled", "canceled"].includes(runtimeStatus) ||
task?.failed === true ||
Boolean(task?.error) ||
Boolean(task?.runtime?.exception)
);
}
function sleep(ms) {
return new Promise((resolvePromise) => setTimeout(resolvePromise, ms));
}
async function ensureLocalAgentPipeline({
backendUrl,
token,
pipelineName,
runnerId,
}) {
// Fake-provider use must be explicit for this pipeline. The generic fake
// provider variables persist after load tests and must not silently replace
// a real LocalAgent model in later QA runs.
const fakeProviderBaseUrl = env.LANGBOT_E2E_FAKE_PROVIDER_BASE_URL || "";
let fakeModel = null;
if (fakeProviderBaseUrl) {
fakeModel = await ensureFakeProviderModel({
backendUrl,
token,
baseUrl: fakeProviderBaseUrl,
});
if (fakeModel.status !== "pass") return fakeModel;
}
async function ensureLocalAgentPipeline({ backendUrl, token, pipelineName, runnerId }) {
const [pipelineList, modelList] = await Promise.all([
apiJson(backendUrl, "/api/v1/pipelines", { token }),
apiJson(backendUrl, "/api/v1/provider/models/llm", { token }),
@@ -488,9 +199,7 @@ async function ensureLocalAgentPipeline({
.map((item) => item.trim())
.filter(Boolean),
);
const spaceModels = models.filter(
(model) => isSpaceModel(model) && !skippedModelIds.has(model.uuid),
);
const spaceModels = models.filter((model) => isSpaceModel(model) && !skippedModelIds.has(model.uuid));
const pipelines = pipelineList.json.data?.pipelines || [];
let pipeline = pipelines.find((item) => item.name === pipelineName) || null;
let created = false;
@@ -501,8 +210,7 @@ async function ensureLocalAgentPipeline({
token,
body: {
name: pipelineName,
description:
"Local QA pipeline for AgentRunner Debug Chat smoke tests.",
description: "Local QA pipeline for AgentRunner Debug Chat smoke tests.",
emoji: "QA",
},
});
@@ -516,11 +224,7 @@ async function ensureLocalAgentPipeline({
};
}
const pipelineId = createdResponse.json.data?.uuid || "";
const loaded = await apiJson(
backendUrl,
`/api/v1/pipelines/${encodeURIComponent(pipelineId)}`,
{ token },
);
const loaded = await apiJson(backendUrl, `/api/v1/pipelines/${encodeURIComponent(pipelineId)}`, { token });
pipeline = loaded.json.data?.pipeline || null;
created = true;
}
@@ -534,11 +238,7 @@ async function ensureLocalAgentPipeline({
};
}
const loaded = await apiJson(
backendUrl,
`/api/v1/pipelines/${encodeURIComponent(pipeline.uuid)}`,
{ token },
);
const loaded = await apiJson(backendUrl, `/api/v1/pipelines/${encodeURIComponent(pipeline.uuid)}`, { token });
if (isApiFailure(loaded) || !loaded.json.data?.pipeline) {
return {
status: "fail",
@@ -551,53 +251,32 @@ async function ensureLocalAgentPipeline({
}
pipeline = loaded.json.data.pipeline;
const config =
pipeline.config && typeof pipeline.config === "object"
? pipeline.config
: {};
const config = pipeline.config && typeof pipeline.config === "object" ? pipeline.config : {};
const ai = config.ai && typeof config.ai === "object" ? config.ai : {};
const runnerConfigs =
ai.runner_config && typeof ai.runner_config === "object"
? ai.runner_config
: {};
const rawExistingLocalAgentConfig =
runnerConfigs[runnerId] && typeof runnerConfigs[runnerId] === "object"
? runnerConfigs[runnerId]
: {};
const rawExistingLocalAgentConfig = ai["local-agent"] && typeof ai["local-agent"] === "object"
? ai["local-agent"]
: {};
const existingLocalAgentConfig = rawExistingLocalAgentConfig;
const existingModel =
existingLocalAgentConfig.model &&
typeof existingLocalAgentConfig.model === "object"
? existingLocalAgentConfig.model
: {};
const requestedModelId =
env.LANGBOT_LOCAL_AGENT_MODEL_UUID || env.LANGBOT_E2E_MODEL_UUID || "";
const selected = fakeModel
? {
status: "pass",
reason: "",
selected_model_id: fakeModel.model_uuid,
selected_model_name: fakeModel.model_name,
fallback_model_ids: [],
scanned_space_model_count: 0,
tested_model_count: 0,
model_tests: [],
}
: await selectWorkingSpaceModel({
backendUrl,
token,
models,
skippedModelIds,
skippedModelNames,
requestedModelId,
existingModelId: existingModel.primary || "",
});
const existingModel = existingLocalAgentConfig.model && typeof existingLocalAgentConfig.model === "object"
? existingLocalAgentConfig.model
: {};
const requestedModelId = env.LANGBOT_LOCAL_AGENT_MODEL_UUID || env.LANGBOT_E2E_MODEL_UUID || "";
const selected = await selectWorkingSpaceModel({
backendUrl,
token,
models,
skippedModelIds,
skippedModelNames,
requestedModelId,
existingModelId: existingModel.primary || "",
});
const selectedModelId = selected.selected_model_id || "";
const localAgentConfig = {
timeout: 300,
prompt: [{ role: "system", content: "You are a helpful assistant." }],
"remove-think": false,
"knowledge-bases": [],
"box-session-id-template": "{launcher_type}_{launcher_id}",
"retrieval-top-k": 5,
"rerank-model": "",
"rerank-top-k": 5,
@@ -609,6 +288,7 @@ async function ensureLocalAgentPipeline({
"context-reserve-tokens": 16384,
"context-keep-recent-tokens": 20000,
"context-summary-tokens": 8000,
...existingLocalAgentConfig,
// Current backend truncation still reads this field directly.
"max-round": positiveInteger(existingLocalAgentConfig["max-round"], 10),
model: {
@@ -623,30 +303,23 @@ async function ensureLocalAgentPipeline({
runner: {
...(ai.runner && typeof ai.runner === "object" ? ai.runner : {}),
id: runnerId,
runner: runnerId,
"expire-time": 0,
},
runner_config: {
...runnerConfigs,
[runnerId]: localAgentConfig,
},
"local-agent": localAgentConfig,
},
};
const updateResponse = await apiJson(
backendUrl,
`/api/v1/pipelines/${encodeURIComponent(pipeline.uuid)}`,
{
method: "PUT",
token,
body: {
name: pipelineName,
description:
"Local QA pipeline for AgentRunner Debug Chat smoke tests.",
emoji: "QA",
config: updatedConfig,
},
const updateResponse = await apiJson(backendUrl, `/api/v1/pipelines/${encodeURIComponent(pipeline.uuid)}`, {
method: "PUT",
token,
body: {
name: pipelineName,
description: "Local QA pipeline for AgentRunner Debug Chat smoke tests.",
emoji: "QA",
config: updatedConfig,
},
);
});
if (isApiFailure(updateResponse)) {
return {
status: "fail",
@@ -668,8 +341,7 @@ async function ensureLocalAgentPipeline({
status: selectedModelId ? "pass" : "env_issue",
reason: selectedModelId
? `Local-agent pipeline is configured for Debug Chat with Space model ${selected.selected_model_name || selectedModelId} and ${selected.fallback_model_ids.length} fallback(s).`
: selected.reason ||
"No working Space LLM model is configured in this LangBot instance.",
: selected.reason || "No working Space LLM model is configured in this LangBot instance.",
pipeline_id: pipeline.uuid,
pipeline_name: pipelineName,
model_count: models.length,
@@ -680,214 +352,21 @@ async function ensureLocalAgentPipeline({
selected_model_id: selectedModelId,
selected_model_name: selected.selected_model_name,
fallback_model_ids: selected.fallback_model_ids,
fake_provider: fakeModel,
created,
updated: true,
};
}
async function ensureFakeProviderModel({ backendUrl, token, baseUrl }) {
const modelUuid = env.LANGBOT_E2E_FAKE_MODEL_UUID || DEFAULT_FAKE_MODEL_UUID;
const modelName =
env.LANGBOT_E2E_FAKE_MODEL_NAME ||
env.LANGBOT_FAKE_PROVIDER_MODEL ||
env.LANGBOT_FAKE_PROVIDER_MODEL_NAME ||
"langbot-e2e-fake-model";
const providerName =
env.LANGBOT_E2E_FAKE_PROVIDER_NAME || DEFAULT_FAKE_PROVIDER_NAME;
const providerRequester =
env.LANGBOT_E2E_FAKE_PROVIDER_REQUESTER || "openai-chat-completions";
const apiKey =
env.LANGBOT_E2E_FAKE_PROVIDER_API_KEY ||
env.LANGBOT_FAKE_PROVIDER_API_KEY ||
"fake-key";
const providersResponse = await apiJson(
backendUrl,
"/api/v1/provider/providers",
{ token },
);
if (isApiFailure(providersResponse)) {
return {
status: "fail",
reason:
providersResponse.json.msg ||
"Failed to list providers before creating fake provider.",
provider_status: providersResponse.status,
};
}
const normalizedBaseUrl = String(baseUrl || "").replace(/\/$/, "");
const providers = providersResponse.json.data?.providers || [];
let provider = providers.find(
(item) =>
item.name === providerName ||
(item.requester === providerRequester &&
String(item.base_url || "").replace(/\/$/, "") === normalizedBaseUrl),
);
const providerBody = {
name: providerName,
requester: providerRequester,
base_url: normalizedBaseUrl,
api_keys: [apiKey],
};
if (provider?.uuid) {
const updateResponse = await apiJson(
backendUrl,
`/api/v1/provider/providers/${encodeURIComponent(provider.uuid)}`,
{
method: "PUT",
token,
body: providerBody,
},
);
if (isApiFailure(updateResponse)) {
return {
status: "fail",
reason: updateResponse.json.msg || "Failed to update fake provider.",
provider_status: updateResponse.status,
};
}
} else {
const createProviderResponse = await apiJson(
backendUrl,
"/api/v1/provider/providers",
{
method: "POST",
token,
body: providerBody,
},
);
if (isApiFailure(createProviderResponse)) {
return {
status: "fail",
reason:
createProviderResponse.json.msg || "Failed to create fake provider.",
provider_status: createProviderResponse.status,
};
}
provider = { uuid: createProviderResponse.json.data?.uuid || "" };
}
if (!provider?.uuid) {
return {
status: "fail",
reason: "Fake provider did not return a provider uuid.",
};
}
let resolvedModelUuid = modelUuid;
let modelResponse = await apiJson(
backendUrl,
`/api/v1/provider/models/llm/${encodeURIComponent(modelUuid)}`,
{
token,
},
);
if (modelResponse.status === 404) {
const providerModelsResponse = await apiJson(
backendUrl,
`/api/v1/provider/models/llm?provider_uuid=${encodeURIComponent(provider.uuid)}`,
{ token },
);
if (!isApiFailure(providerModelsResponse)) {
const existingModel = (
providerModelsResponse.json.data?.models || []
).find((item) => item.name === modelName);
if (existingModel?.uuid) {
resolvedModelUuid = existingModel.uuid;
modelResponse = await apiJson(
backendUrl,
`/api/v1/provider/models/llm/${encodeURIComponent(resolvedModelUuid)}`,
{
token,
},
);
}
}
}
const modelBody = {
name: modelName,
provider_uuid: provider.uuid,
abilities: ["func_call", "vision"],
context_length: 8192,
extra_args: {},
prefered_ranking: 0,
};
if (modelResponse.status === 404) {
const createModelResponse = await apiJson(
backendUrl,
"/api/v1/provider/models/llm",
{
method: "POST",
token,
body: {
uuid: modelUuid,
...modelBody,
},
},
);
if (isApiFailure(createModelResponse)) {
return {
status: "fail",
reason: createModelResponse.json.msg || "Failed to create fake model.",
model_status: createModelResponse.status,
};
}
resolvedModelUuid = createModelResponse.json.data?.uuid || modelUuid;
} else if (isApiFailure(modelResponse)) {
return {
status: "fail",
reason: modelResponse.json.msg || "Failed to load fake model.",
model_status: modelResponse.status,
};
} else {
const updateModelResponse = await apiJson(
backendUrl,
`/api/v1/provider/models/llm/${encodeURIComponent(resolvedModelUuid)}`,
{
method: "PUT",
token,
body: modelBody,
},
);
if (isApiFailure(updateModelResponse)) {
return {
status: "fail",
reason: updateModelResponse.json.msg || "Failed to update fake model.",
model_status: updateModelResponse.status,
};
}
}
return {
status: "pass",
provider_uuid: provider.uuid,
provider_requester: providerRequester,
base_url: normalizedBaseUrl,
model_uuid: resolvedModelUuid,
model_name: modelName,
};
}
function isApiFailure(response) {
return (
response.status >= 400 ||
(response.json.code !== undefined && response.json.code !== 0)
);
return response.status >= 400 || (response.json.code !== undefined && response.json.code !== 0);
}
function isSpaceModel(model) {
const provider =
model?.provider && typeof model.provider === "object" ? model.provider : {};
return (
model?.provider_uuid === SPACE_PROVIDER_UUID ||
provider.uuid === SPACE_PROVIDER_UUID ||
provider.requester === "space-chat-completions" ||
provider.name === "LangBot Models"
);
const provider = model?.provider && typeof model.provider === "object" ? model.provider : {};
return model?.provider_uuid === SPACE_PROVIDER_UUID
|| provider.uuid === SPACE_PROVIDER_UUID
|| provider.requester === "space-chat-completions"
|| provider.name === "LangBot Models";
}
async function selectWorkingSpaceModel({
@@ -900,24 +379,15 @@ async function selectWorkingSpaceModel({
existingModelId,
}) {
const modelTests = [];
const testLimit = positiveInteger(
env.LANGBOT_E2E_MODEL_TEST_LIMIT,
DEFAULT_MODEL_TEST_LIMIT,
);
const fallbackCount = positiveInteger(
env.LANGBOT_E2E_MODEL_FALLBACK_COUNT,
DEFAULT_MODEL_FALLBACK_COUNT,
);
const testLimit = positiveInteger(env.LANGBOT_E2E_MODEL_TEST_LIMIT, DEFAULT_MODEL_TEST_LIMIT);
const fallbackCount = positiveInteger(env.LANGBOT_E2E_MODEL_FALLBACK_COUNT, DEFAULT_MODEL_FALLBACK_COUNT);
const workingModels = [];
const spaceModels = rankModels(
models.filter(
(model) =>
model.uuid &&
isSpaceModel(model) &&
!skippedModelIds.has(model.uuid) &&
!skippedModelNames.has(model.name),
),
);
const spaceModels = rankModels(models.filter((model) => (
model.uuid
&& isSpaceModel(model)
&& !skippedModelIds.has(model.uuid)
&& !skippedModelNames.has(model.name)
)));
const requestedModel = requestedModelId
? spaceModels.find((model) => model.uuid === requestedModelId) || null
: null;
@@ -926,9 +396,7 @@ async function selectWorkingSpaceModel({
: null;
const candidates = uniqueCandidates([
...(requestedModel ? [existingCandidate(requestedModel, "requested")] : []),
...(existingModel
? [existingCandidate(existingModel, "existing-pipeline")]
: []),
...(existingModel ? [existingCandidate(existingModel, "existing-pipeline")] : []),
...spaceModels.map((model) => existingCandidate(model, "configured-space")),
]);
@@ -937,16 +405,9 @@ async function selectWorkingSpaceModel({
scanResult = await scanSpaceModels({ backendUrl, token });
if (scanResult.status === "pass") {
const knownNames = new Set(spaceModels.map((model) => model.name));
candidates.push(
...scanResult.models
.filter(
(model) =>
model.name &&
!knownNames.has(model.name) &&
!skippedModelNames.has(model.name),
)
.map((model) => scannedCandidate(model)),
);
candidates.push(...scanResult.models
.filter((model) => model.name && !knownNames.has(model.name) && !skippedModelNames.has(model.name))
.map((model) => scannedCandidate(model)));
}
}
@@ -974,16 +435,14 @@ async function selectWorkingSpaceModel({
};
}
const baseReason =
unique.length === 0
? scanResult.reason || "No Space LLM model candidates are available."
: `No working Space LLM model found after testing ${modelTests.length} candidate(s).`;
const baseReason = unique.length === 0
? scanResult.reason || "No Space LLM model candidates are available."
: `No working Space LLM model found after testing ${modelTests.length} candidate(s).`;
return {
status: "env_issue",
reason:
requestedModelId && !requestedModel
? `Requested Space LLM model ${requestedModelId} is missing or skipped; ${baseReason}`
: baseReason,
reason: requestedModelId && !requestedModel
? `Requested Space LLM model ${requestedModelId} is missing or skipped; ${baseReason}`
: baseReason,
selected_model_id: "",
selected_model_name: "",
fallback_model_ids: [],
@@ -1003,11 +462,7 @@ async function scanSpaceModels({ backendUrl, token }) {
return {
status: "env_issue",
models: [],
reason: safeReason(
response.json.msg ||
response.json.message ||
"Failed to scan Space LLM models.",
),
reason: safeReason(response.json.msg || response.json.message || "Failed to scan Space LLM models."),
};
}
return {
@@ -1037,42 +492,28 @@ async function ensureAndTestModel({ backendUrl, token, candidate }) {
if (isApiFailure(create) || !modelUuid) {
return modelTestResult(candidate, {
status: "fail",
reason: safeReason(
create.json.msg || "Failed to create scanned Space model.",
),
reason: safeReason(create.json.msg || "Failed to create scanned Space model."),
http_status: create.status,
});
}
created = true;
}
const test = await apiJson(
backendUrl,
`/api/v1/provider/models/llm/${encodeURIComponent(modelUuid)}/test`,
{
method: "POST",
token,
body: { extra_args: {} },
},
);
const test = await apiJson(backendUrl, `/api/v1/provider/models/llm/${encodeURIComponent(modelUuid)}/test`, {
method: "POST",
token,
body: { extra_args: {} },
});
const passed = !isApiFailure(test);
if (!passed && created) {
await apiJson(
backendUrl,
`/api/v1/provider/models/llm/${encodeURIComponent(modelUuid)}`,
{
method: "DELETE",
token,
},
).catch(() => {});
await apiJson(backendUrl, `/api/v1/provider/models/llm/${encodeURIComponent(modelUuid)}`, {
method: "DELETE",
token,
}).catch(() => {});
}
return modelTestResult(candidate, {
status: passed ? "pass" : "fail",
reason: passed
? ""
: safeReason(
test.json.msg || test.json.message || "Space model test failed.",
),
reason: passed ? "" : safeReason(test.json.msg || test.json.message || "Space model test failed."),
http_status: test.status,
model_uuid: modelUuid,
created,
@@ -1117,9 +558,7 @@ function uniqueCandidates(candidates) {
const seen = new Set();
const result = [];
for (const candidate of candidates) {
const key = candidate.uuid
? `uuid:${candidate.uuid}`
: `name:${candidate.name}`;
const key = candidate.uuid ? `uuid:${candidate.uuid}` : `name:${candidate.name}`;
if (!candidate.name || seen.has(key)) continue;
seen.add(key);
result.push(candidate);
@@ -1129,12 +568,8 @@ function uniqueCandidates(candidates) {
function rankModels(models) {
return [...models].sort((left, right) => {
const leftRank = Number.isFinite(Number(left.prefered_ranking))
? Number(left.prefered_ranking)
: 9999;
const rightRank = Number.isFinite(Number(right.prefered_ranking))
? Number(right.prefered_ranking)
: 9999;
const leftRank = Number.isFinite(Number(left.prefered_ranking)) ? Number(left.prefered_ranking) : 9999;
const rightRank = Number.isFinite(Number(right.prefered_ranking)) ? Number(right.prefered_ranking) : 9999;
if (leftRank !== rightRank) return leftRank - rightRank;
return String(left.name || "").localeCompare(String(right.name || ""));
});
@@ -1170,9 +605,5 @@ async function upsertEnvLocal(path, updates) {
for (const [key, value] of Object.entries(updates)) {
if (!seen.has(key)) next.push(`${key}=${value}`);
}
await writeFile(
path,
`${next.filter((line, index) => line !== "" || index < next.length - 1).join("\n")}\n`,
"utf8",
);
await writeFile(path, `${next.filter((line, index) => line !== "" || index < next.length - 1).join("\n")}\n`, "utf8");
}
@@ -74,7 +74,7 @@ try {
});
Object.assign(result, prepared);
if (result.pipeline_id) {
result.pipeline_url = `${frontendUrl.replace(/\/$/, "")}/home/agents?id=${encodeURIComponent(result.pipeline_id)}`;
result.pipeline_url = `${frontendUrl.replace(/\/$/, "")}/home/pipelines?id=${encodeURIComponent(result.pipeline_id)}`;
}
if (writeEnv && result.pipeline_id) {
+17 -409
View File
@@ -9,43 +9,7 @@ const args = parseArgs(process.argv.slice(2));
const host = args.host || env.LANGBOT_FAKE_PROVIDER_HOST || "127.0.0.1";
const port = integer(args.port ?? env.LANGBOT_FAKE_PROVIDER_PORT, 0);
const stateFile = args["state-file"] || env.LANGBOT_FAKE_PROVIDER_STATE_FILE || "";
const modelName = args.model || env.LANGBOT_FAKE_PROVIDER_MODEL_NAME || env.LANGBOT_FAKE_PROVIDER_MODEL || "gpt-4o-mini";
const COMPACTION_SENTINEL = "qa_compaction_sentinel_7391";
const COMBO_SENTINEL = "qa_combo_compaction_sentinel_2406";
const RAG_SENTINEL = "azalea-cobalt-7421";
const COMBO_TOOL_INPUT = "combo-tool-ok-local-agent";
const COMBO_TOOL_RESULT = `qa-plugin-smoke:${COMBO_TOOL_INPUT}`;
const COMBO_FINAL = `COMBO_FINAL ${COMBO_SENTINEL} ${RAG_SENTINEL} ${COMBO_TOOL_RESULT}`;
const MULTITOOL_SENTINEL = "qa_multitool_compaction_sentinel_6718";
const MULTITOOL_TOOL_A_INPUT = "multi-tool-a-local-agent";
const MULTITOOL_TOOL_B_INPUT = "multi-tool-b-local-agent";
const MULTITOOL_TOOL_A_RESULT = `qa-plugin-smoke:${MULTITOOL_TOOL_A_INPUT}`;
const MULTITOOL_TOOL_B_RESULT = `qa-plugin-smoke:${MULTITOOL_TOOL_B_INPUT}`;
const MULTITOOL_FINAL = [
"MULTITOOL_COMBO_FINAL",
MULTITOOL_SENTINEL,
RAG_SENTINEL,
MULTITOOL_TOOL_A_RESULT,
MULTITOOL_TOOL_B_RESULT,
].join(" ");
const PARALLEL_SENTINEL = "qa_parallel_compaction_sentinel_8142";
const PARALLEL_TOOL_A_INPUT = "parallel-tool-a-local-agent";
const PARALLEL_TOOL_B_INPUT = "parallel-tool-b-local-agent";
const PARALLEL_TOOL_A_RESULT = `qa-plugin-smoke:${PARALLEL_TOOL_A_INPUT}`;
const PARALLEL_TOOL_B_RESULT = `qa-plugin-smoke:${PARALLEL_TOOL_B_INPUT}`;
const PARALLEL_FINAL = [
"PARALLEL_COMBO_FINAL",
PARALLEL_SENTINEL,
RAG_SENTINEL,
PARALLEL_TOOL_A_RESULT,
PARALLEL_TOOL_B_RESULT,
].join(" ");
const LOOP_LIMIT_INPUT = "loop-limit-repeat-local-agent";
const TOOL_ERROR_INPUT = "tool-error-recovery-local-agent";
const TOOL_ERROR_FINAL = "TOOL_ERROR_RECOVERY_FINAL qa-plugin-smoke forced failure observed";
const STEERING_FOLLOWUP_SENTINEL = "qa_steering_sentinel_6194";
const STEERING_SLEEP_INPUT = "steering-e2e-anchor";
const STEERING_SLEEP_RESULT = `qa-plugin-smoke:sleep:8:${STEERING_SLEEP_INPUT}`;
const modelName = env.LANGBOT_FAKE_PROVIDER_MODEL_NAME || "gpt-4o-mini";
const config = {
response_text: env.LANGBOT_FAKE_PROVIDER_RESPONSE_TEXT || "OK",
first_token_delay_ms: integer(env.LANGBOT_FAKE_PROVIDER_FIRST_TOKEN_DELAY_MS, 25),
@@ -54,11 +18,7 @@ const config = {
fault_status: integer(env.LANGBOT_FAKE_PROVIDER_FAULT_STATUS, 500),
fail_first_n: integer(env.LANGBOT_FAKE_PROVIDER_FAIL_FIRST_N, 0),
fail_every_n: integer(env.LANGBOT_FAKE_PROVIDER_FAIL_EVERY_N, 0),
fail_models: textList(env.LANGBOT_FAKE_PROVIDER_FAIL_MODELS),
fail_after_first_chunk: bool(env.LANGBOT_FAKE_PROVIDER_FAIL_AFTER_FIRST_CHUNK, false),
fail_after_first_chunk_delay_ms: integer(env.LANGBOT_FAKE_PROVIDER_FAIL_AFTER_FIRST_CHUNK_DELAY_MS, 0),
fail_after_first_chunk_mode: faultMode(env.LANGBOT_FAKE_PROVIDER_FAIL_AFTER_FIRST_CHUNK_MODE),
fail_after_first_chunk_models: textList(env.LANGBOT_FAKE_PROVIDER_FAIL_AFTER_FIRST_CHUNK_MODELS),
dynamic_response: !/^(0|false|no|off)$/i.test(env.LANGBOT_FAKE_PROVIDER_DYNAMIC_RESPONSE || ""),
request_log_limit: integer(env.LANGBOT_FAKE_PROVIDER_REQUEST_LOG_LIMIT, 500),
};
@@ -128,7 +88,6 @@ const server = createServer(async (request, response) => {
created: 1,
owned_by: "langbot-qa",
type: "llm",
context_length: 8192,
},
],
});
@@ -139,20 +98,15 @@ const server = createServer(async (request, response) => {
requestCount += 1;
const body = await readJson(request);
const requestId = `chatcmpl-langbot-fake-${requestCount}`;
const requestModel = String(body.model || modelName);
const shouldFail = requestCount <= config.fail_first_n
|| (config.fail_every_n > 0 && requestCount % config.fail_every_n === 0)
|| config.fail_models.includes(requestModel);
const failAfterFirstChunk = config.fail_after_first_chunk
|| config.fail_after_first_chunk_models.includes(requestModel);
const replyMessage = buildResponse(body);
const replyText = replyMessage.content || "";
|| (config.fail_every_n > 0 && requestCount % config.fail_every_n === 0);
const replyText = responseTextForBody(body);
requestRecord = recordRequest({
id: requestId,
request_number: requestCount,
path: url.pathname,
stream: Boolean(body.stream),
model: requestModel,
model: body.model || "",
message_count: Array.isArray(body.messages) ? body.messages.length : 0,
should_fail: shouldFail,
status: "running",
@@ -186,8 +140,8 @@ const server = createServer(async (request, response) => {
await streamCompletion(response, {
requestId,
model: body.model || modelName,
message: replyMessage,
failAfterFirstChunk,
content: replyText,
failAfterFirstChunk: config.fail_after_first_chunk,
requestRecord,
startedPerf,
});
@@ -196,7 +150,7 @@ const server = createServer(async (request, response) => {
sendJson(response, 200, completionPayload({
requestId,
model: body.model || modelName,
message: replyMessage,
content: replyText,
}));
markRequestTiming(requestRecord, "first_chunk", startedPerf);
markRequestTiming(requestRecord, "first_content_chunk", startedPerf);
@@ -316,8 +270,8 @@ function sendJson(response, status, payload) {
response.end(text);
}
function completionPayload({ requestId, model, message }) {
const completionTokens = tokenEstimate(message.content || JSON.stringify(message.tool_calls || []));
function completionPayload({ requestId, model, content }) {
const completionTokens = tokenEstimate(content);
return {
id: requestId,
object: "chat.completion",
@@ -326,8 +280,11 @@ function completionPayload({ requestId, model, message }) {
choices: [
{
index: 0,
message,
finish_reason: message.tool_calls?.length ? "tool_calls" : "stop",
message: {
role: "assistant",
content,
},
finish_reason: "stop",
},
],
usage: {
@@ -341,7 +298,7 @@ function completionPayload({ requestId, model, message }) {
async function streamCompletion(response, {
requestId,
model,
message,
content,
failAfterFirstChunk: failMidStream,
requestRecord,
startedPerf,
@@ -362,56 +319,7 @@ async function streamCompletion(response, {
choices: [{ index: 0, delta: { role: "assistant" }, finish_reason: null }],
});
const chunks = message.tool_calls?.length ? [] : splitContent(message.content || "");
if (message.tool_calls?.length) {
await sleep(config.chunk_delay_ms);
markRequestTiming(requestRecord, "first_content_chunk", startedPerf);
requestRecord.content_chunk_count = (requestRecord.content_chunk_count || 0) + 1;
writeSse(response, {
id: requestId,
object: "chat.completion.chunk",
created: Math.floor(Date.now() / 1000),
model,
choices: [
{
index: 0,
delta: {
tool_calls: message.tool_calls.map((call, index) => ({
index,
id: call.id,
type: call.type,
function: {
name: call.function.name,
arguments: call.function.arguments,
},
})),
},
finish_reason: null,
},
],
});
await sleep(config.chunk_delay_ms);
writeSse(response, {
id: requestId,
object: "chat.completion.chunk",
created: Math.floor(Date.now() / 1000),
model,
choices: [{ index: 0, delta: {}, finish_reason: "tool_calls" }],
usage: {
prompt_tokens: 8,
completion_tokens: tokenEstimate(JSON.stringify(message.tool_calls)),
total_tokens: 8 + tokenEstimate(JSON.stringify(message.tool_calls)),
},
});
response.write("data: [DONE]\n\n");
response.end();
finishRequestRecord(requestRecord, startedPerf, {
status: "ok",
http_status: 200,
});
return;
}
const chunks = splitContent(content);
for (let index = 0; index < chunks.length; index += 1) {
await sleep(config.chunk_delay_ms);
if (index === 0) markRequestTiming(requestRecord, "first_content_chunk", startedPerf);
@@ -424,22 +332,6 @@ async function streamCompletion(response, {
choices: [{ index: 0, delta: { content: chunks[index] }, finish_reason: null }],
});
if (failMidStream && index === 0) {
await sleep(config.fail_after_first_chunk_delay_ms);
if (config.fail_after_first_chunk_mode === "error_event") {
response.write(`data: ${JSON.stringify({
error: {
message: "LangBot fake provider injected a mid-stream error event",
type: "fake_provider_stream_fault",
code: "fake_provider_stream_fault",
},
})}\n\n`);
response.end();
finishRequestRecord(requestRecord, startedPerf, {
status: "mid_stream_error_event",
http_status: 200,
});
return;
}
finishRequestRecord(requestRecord, startedPerf, {
status: "mid_stream_disconnect",
http_status: 200,
@@ -450,7 +342,7 @@ async function streamCompletion(response, {
}
await sleep(config.chunk_delay_ms);
const completionTokens = tokenEstimate(message.content || "");
const completionTokens = tokenEstimate(content);
writeSse(response, {
id: requestId,
object: "chat.completion.chunk",
@@ -507,266 +399,6 @@ function responseTextForBody(body) {
return config.response_text;
}
function messageText(messages = []) {
return messages
.map((message) => {
const parts = [contentText(message?.content)];
if (Array.isArray(message?.tool_calls)) {
parts.push(
message.tool_calls
.map((call) => `${call?.function?.name || call?.name || ""} ${call?.function?.arguments || ""}`)
.join("\n"),
);
}
return parts.filter(Boolean).join("\n");
})
.join("\n");
}
function contentText(content) {
if (typeof content === "string") return content;
if (Array.isArray(content)) return content.map((part) => contentText(part)).join("");
if (content && typeof content === "object") {
for (const key of ["text", "content", "message", "error", "value"]) {
if (content[key] !== undefined && content[key] !== null) {
return contentText(content[key]);
}
}
return JSON.stringify(content);
}
return "";
}
function latestUserText(messages = []) {
const latest = [...messages].reverse().find((message) => message?.role === "user");
return contentText(latest?.content);
}
function toolNames(tools = []) {
return tools
.map((tool) => tool?.function?.name || tool?.name || "")
.filter(Boolean);
}
function firstToolName(tools = [], candidates = []) {
const names = toolNames(tools);
return candidates.find((candidate) => names.includes(candidate)) || "";
}
function isSummarizationRequest(messages = []) {
const text = messageText(messages);
return /context summarization assistant|conversation to summarize|<previous-summary>/i.test(text);
}
function buildSummaryResponse(text) {
const references = [];
if (text.includes(COMPACTION_SENTINEL)) references.push(COMPACTION_SENTINEL);
if (text.includes(COMBO_SENTINEL)) references.push(COMBO_SENTINEL);
if (text.includes(MULTITOOL_SENTINEL)) references.push(MULTITOOL_SENTINEL);
if (text.includes(PARALLEL_SENTINEL)) references.push(PARALLEL_SENTINEL);
if (/RAG_TOOL_COMBO_GOAL/i.test(text)) references.push("RAG_TOOL_COMBO_GOAL");
if (/MULTITOOL_RAG_GOAL/i.test(text)) references.push("MULTITOOL_RAG_GOAL");
if (/PARALLEL_RAG_GOAL/i.test(text)) references.push("PARALLEL_RAG_GOAL");
const critical = references.length
? references.map((reference) => `- ${reference}`).join("\n")
: "- (none)";
return {
role: "assistant",
content: [
"## Goal",
"Preserve deterministic LangBot E2E context across compaction.",
"",
"## Constraints & Preferences",
"- Keep exact sentinel values verbatim.",
"",
"## Progress",
"### Done",
"- [x] Earlier conversation was compacted by the fake provider.",
"",
"### In Progress",
"- [ ] Continue the current Debug Chat run.",
"",
"### Blocked",
"- (none)",
"",
"## Key Decisions",
"- **Deterministic QA**: Return only known sentinel values from compacted input.",
"",
"## Next Steps",
"1. Answer the current user request using preserved context.",
"",
"## Critical Context",
critical,
].join("\n"),
};
}
function buildResponse(payload) {
const text = messageText(payload.messages || []);
const current = latestUserText(payload.messages || []);
const tools = payload.tools || [];
if (isSummarizationRequest(payload.messages || [])) {
return buildSummaryResponse(text);
}
if (/qa-effective-prompt/i.test(current) && /PROMPT_PREPROCESS_OK/.test(text)) {
return { role: "assistant", content: "PROMPT_PREPROCESS_OK" };
}
const pluginTool = firstToolName(tools, ["qa_plugin_echo"]);
const pluginFailTool = firstToolName(tools, ["qa_plugin_fail"]);
const pluginSleepTool = firstToolName(tools, ["qa_plugin_sleep"]);
const mcpTool = firstToolName(tools, ["qa_mcp_echo"]);
const requestedMcpEchoText = current.match(
/qa_mcp_echo[\s\S]*?exactly this text:\s*([A-Za-z0-9_:-]+)/i,
)?.[1] || "mcp-ok-local-agent";
const expectedMcpEchoResult = `qa_mcp_echo:${requestedMcpEchoText}`;
if (/STEERING_NO_FOLLOWUP|qa_plugin_sleep|steering-e2e-anchor|qa_steering_sentinel_6194/i.test(current || text)) {
if (text.includes(STEERING_FOLLOWUP_SENTINEL) && text.includes(STEERING_SLEEP_RESULT)) {
return { role: "assistant", content: STEERING_FOLLOWUP_SENTINEL };
}
if (text.includes(STEERING_SLEEP_RESULT) && !text.includes(STEERING_FOLLOWUP_SENTINEL)) {
return { role: "assistant", content: "STEERING_NO_FOLLOWUP" };
}
if (pluginSleepTool && !text.includes(STEERING_SLEEP_RESULT)) {
return toolCall("call_qa_plugin_sleep_steering", pluginSleepTool, { seconds: 8, text: STEERING_SLEEP_INPUT });
}
}
if (/测试暗号是什么|original sentinel|first.*sentinel/i.test(current)) {
return {
role: "assistant",
content: text.includes(COMPACTION_SENTINEL) ? COMPACTION_SENTINEL : "COMPACTION_SENTINEL_MISSING",
};
}
if ((current.includes(COMPACTION_SENTINEL)
|| current.includes(COMBO_SENTINEL)
|| current.includes(MULTITOOL_SENTINEL)
|| current.includes(PARALLEL_SENTINEL))
&& /请只回复 MEMORY_SET|only reply MEMORY_SET/i.test(current)) {
return { role: "assistant", content: "MEMORY_SET" };
}
if (/PARALLEL_CONTEXT_PRESSURE_READY/.test(current)) return { role: "assistant", content: "PARALLEL_CONTEXT_PRESSURE_READY" };
if (/MULTITOOL_CONTEXT_PRESSURE_READY/.test(current)) return { role: "assistant", content: "MULTITOOL_CONTEXT_PRESSURE_READY" };
if (/COMBO_CONTEXT_PRESSURE_READY/.test(current)) return { role: "assistant", content: "COMBO_CONTEXT_PRESSURE_READY" };
if (/CONTEXT_PRESSURE_READY/.test(current)) return { role: "assistant", content: "CONTEXT_PRESSURE_READY" };
if (text.includes(expectedMcpEchoResult)) return { role: "assistant", content: expectedMcpEchoResult };
if (/qa_mcp_echo|mcp-ok-local-agent/i.test(current || text) && mcpTool && !text.includes(expectedMcpEchoResult)) {
return toolCall("call_qa_mcp_echo", mcpTool, { text: requestedMcpEchoText });
}
if (/LOOP_LIMIT|loop-limit-repeat-local-agent|iteration limit/i.test(current || text) && pluginTool) {
return toolCall(`call_qa_plugin_echo_loop_limit_${Date.now()}`, pluginTool, { text: LOOP_LIMIT_INPUT });
}
if (/TOOL_ERROR_RECOVERY|tool-error-recovery-local-agent|qa_plugin_fail/i.test(current || text)) {
if (/(?:Error:|Tool execution failed:|ActionCallError:|RuntimeError:)/i.test(text)
&& /(?:qa-plugin-smoke forced failure|qa_plugin_fail|tool-error-recovery-local-agent)/i.test(text)) {
return { role: "assistant", content: TOOL_ERROR_FINAL };
}
if (pluginFailTool) {
return toolCall("call_qa_plugin_fail_recovery", pluginFailTool, { text: TOOL_ERROR_INPUT });
}
}
const isParallelComboRequest = /PARALLEL_COMBO|parallel-tool-a-local-agent|parallel-tool-b-local-agent|PARALLEL_RAG_GOAL/i.test(current || text);
if (isParallelComboRequest && pluginTool && !text.includes(PARALLEL_TOOL_A_RESULT) && !text.includes(PARALLEL_TOOL_B_RESULT)) {
return toolCalls([
["call_qa_plugin_echo_parallel_a", pluginTool, { text: PARALLEL_TOOL_A_INPUT }],
["call_qa_plugin_echo_parallel_b", pluginTool, { text: PARALLEL_TOOL_B_INPUT }],
]);
}
if (isParallelComboRequest && (text.includes(PARALLEL_TOOL_A_RESULT) || text.includes(PARALLEL_TOOL_B_RESULT))) {
return missingOrFinal(text, [
[PARALLEL_SENTINEL, "parallel-memory"],
[RAG_SENTINEL, "rag"],
[PARALLEL_TOOL_A_RESULT, "tool-a"],
[PARALLEL_TOOL_B_RESULT, "tool-b"],
], "PARALLEL_COMBO_MISSING", PARALLEL_FINAL);
}
const isMultiToolComboRequest = /MULTITOOL_COMBO|multi-tool-a-local-agent|multi-tool-b-local-agent|MULTITOOL_RAG_GOAL/i.test(current || text);
if (isMultiToolComboRequest && pluginTool && !text.includes(MULTITOOL_TOOL_A_RESULT)) {
return toolCall("call_qa_plugin_echo_multi_a", pluginTool, { text: MULTITOOL_TOOL_A_INPUT });
}
if (isMultiToolComboRequest && pluginTool && text.includes(MULTITOOL_TOOL_A_RESULT) && !text.includes(MULTITOOL_TOOL_B_RESULT)) {
return toolCall("call_qa_plugin_echo_multi_b", pluginTool, { text: MULTITOOL_TOOL_B_INPUT });
}
if (isMultiToolComboRequest && text.includes(MULTITOOL_TOOL_A_RESULT) && text.includes(MULTITOOL_TOOL_B_RESULT)) {
return missingOrFinal(text, [
[MULTITOOL_SENTINEL, "multi-memory"],
[RAG_SENTINEL, "rag"],
[MULTITOOL_TOOL_A_RESULT, "tool-a"],
[MULTITOOL_TOOL_B_RESULT, "tool-b"],
], "MULTITOOL_COMBO_MISSING", MULTITOOL_FINAL);
}
const isComboRequest = /qa_combo|\bCOMBO_FINAL\b|combo-tool-ok-local-agent|RAG_TOOL_COMBO_GOAL/i.test(current || text);
if (isComboRequest && pluginTool && !text.includes(COMBO_TOOL_RESULT)) {
return toolCall("call_qa_plugin_echo_combo", pluginTool, { text: COMBO_TOOL_INPUT });
}
if (isComboRequest && text.includes(COMBO_TOOL_RESULT)) {
return missingOrFinal(text, [
[COMBO_SENTINEL, "combo-memory"],
[RAG_SENTINEL, "rag"],
[COMBO_TOOL_RESULT, "tool-result"],
], "COMBO_MISSING", COMBO_FINAL);
}
if (/qa-plugin-smoke:plugin-tool-ok-local-agent/.test(text)) {
return { role: "assistant", content: "qa-plugin-smoke:plugin-tool-ok-local-agent" };
}
if (/qa_plugin_echo|plugin-tool-ok-local-agent/i.test(current || text) && pluginTool && !/qa-plugin-smoke:plugin-tool-ok-local-agent/.test(text)) {
return toolCall("call_qa_plugin_echo", pluginTool, { text: "plugin-tool-ok-local-agent" });
}
const e2eTool = firstToolName(tools, ["e2e_lookup"]);
if (/Use the e2e lookup tool|e2e lookup tool|tool loop/i.test(current || text) && e2eTool && !/tool-result:alpha/.test(text)) {
return toolCall("call_e2e_lookup", e2eTool, { query: "alpha" });
}
if (/tool-result:alpha/.test(text)) return { role: "assistant", content: "Tool loop final answer after tool-result:alpha" };
if (/RAG_SENTINEL/.test(text)) return { role: "assistant", content: "RAG final answer with RAG_SENTINEL" };
if (/NONSTREAM_OK/i.test(current || text)) return { role: "assistant", content: "NONSTREAM_OK" };
if (/IMAGE_OK/i.test(current || text)
&& /(?:\[Image\]|langbot_input_attachments|data:image\/|\"type\":\s*\"image\")/i.test(current || text)) {
return { role: "assistant", content: "IMAGE_OK" };
}
if (/azalea-cobalt-7421/.test(text)) return { role: "assistant", content: "azalea-cobalt-7421" };
return { role: "assistant", content: responseTextForBody(payload) };
}
function toolCall(id, name, args) {
return toolCalls([[id, name, args]]);
}
function toolCalls(calls) {
return {
role: "assistant",
content: "",
tool_calls: calls.map(([id, name, args]) => ({
id,
type: "function",
function: {
name,
arguments: JSON.stringify(args),
},
})),
};
}
function missingOrFinal(text, requirements, prefix, finalText) {
const missing = requirements
.filter(([needle]) => !text.includes(needle))
.map(([, label]) => label);
return {
role: "assistant",
content: missing.length ? `${prefix}_${missing.join("_")}` : finalText,
};
}
function flattenContent(content) {
if (typeof content === "string") return content;
if (Array.isArray(content)) {
@@ -839,18 +471,12 @@ function applyConfig(updates) {
assignNonNegativeInteger(updates, "chunk_count");
assignNonNegativeInteger(updates, "fail_first_n");
assignNonNegativeInteger(updates, "fail_every_n");
assignTextList(updates, "fail_models");
assignNonNegativeInteger(updates, "request_log_limit");
if (updates.fault_status !== undefined) {
const parsed = Number.parseInt(String(updates.fault_status), 10);
if (Number.isInteger(parsed) && parsed >= 400 && parsed <= 599) config.fault_status = parsed;
}
assignBoolean(updates, "fail_after_first_chunk");
assignNonNegativeInteger(updates, "fail_after_first_chunk_delay_ms");
if (updates.fail_after_first_chunk_mode !== undefined) {
config.fail_after_first_chunk_mode = faultMode(updates.fail_after_first_chunk_mode);
}
assignTextList(updates, "fail_after_first_chunk_models");
assignBoolean(updates, "dynamic_response");
}
@@ -868,21 +494,3 @@ function assignBoolean(updates, key) {
if (updates[key] === undefined) return;
config[key] = bool(updates[key], config[key]);
}
function assignTextList(updates, key) {
if (updates[key] === undefined) return;
config[key] = Array.isArray(updates[key])
? updates[key].map(String).map((item) => item.trim()).filter(Boolean)
: textList(updates[key]);
}
function textList(value) {
return String(value || "")
.split(/\r?\n|,/)
.map((item) => item.trim())
.filter(Boolean);
}
function faultMode(value) {
return String(value || "").trim().toLowerCase() === "error_event" ? "error_event" : "disconnect";
}
+18 -65
View File
@@ -26,14 +26,8 @@ const packagePath = resolve(
|| "skills/langbot-testing/fixtures/plugins/qa-plugin-smoke/dist/qa-plugin-smoke-0.1.0.lbpkg",
);
const expectedPluginId = env.LANGBOT_E2E_EXPECTED_PLUGIN_ID || "qa/plugin-smoke";
const expectedTools = (env.LANGBOT_E2E_EXPECTED_TOOLS || env.LANGBOT_E2E_EXPECTED_TOOL || (
expectedPluginId === "qa/plugin-smoke" ? "qa_plugin_echo" : ""
))
.split(",")
.map((item) => item.trim())
.filter(Boolean);
const expectedTool = env.LANGBOT_E2E_EXPECTED_TOOL || (expectedPluginId === "qa/plugin-smoke" ? "qa_plugin_echo" : "");
const expectedRunnerId = env.LANGBOT_E2E_EXPECTED_RUNNER_ID || "";
const forceReinstall = /^(?:1|true|yes|on)$/i.test(env.LANGBOT_E2E_FORCE_REINSTALL || "");
const result = {
source: "automation",
@@ -46,7 +40,6 @@ const result = {
package_preview: null,
task_id: null,
task: null,
reinstall_reason: "",
plugin_present_before: false,
plugin_present_after: false,
tool_names: [],
@@ -71,25 +64,21 @@ try {
}
result.plugin_present_before = await hasPlugin(backendUrl, auth.token);
if (expectedTools.length > 0) {
result.tool_names = await listToolNames(backendUrl, auth.token);
}
const missingToolsBefore = expectedTools.filter((tool) => !result.tool_names.includes(tool));
if (result.plugin_present_before && (forceReinstall || missingToolsBefore.length > 0)) {
result.reinstall_reason = forceReinstall
? "Explicit reinstall requested by LANGBOT_E2E_FORCE_REINSTALL."
: `Installed plugin is missing expected tools: ${missingToolsBefore.join(", ")}`;
const removeTask = await removePlugin(backendUrl, auth.token);
if (!isTaskComplete(removeTask)) {
throw new Error(`Plugin reinstall cleanup did not complete successfully: ${JSON.stringify(removeTask)}`);
}
result.plugin_present_before = false;
await sleep(1000);
}
if (!result.plugin_present_before) {
result.task = await installPlugin(backendUrl, auth.token, bytes, packagePath);
result.task_id = result.task?.id || result.task_id;
const form = new FormData();
form.set("file", new Blob([bytes]), packagePath.split("/").pop());
const response = await fetch(`${backendUrl.replace(/\/$/, "")}/api/v1/plugins/install/local`, {
method: "POST",
headers: { Authorization: `Bearer ${auth.token}` },
body: form,
});
const json = await response.json().catch(() => ({}));
if (response.status >= 400 || json.code !== 0) {
throw new Error(json.msg || `Plugin install request failed with HTTP ${response.status}.`);
}
result.task_id = json.data?.task_id ?? null;
if (!result.task_id) throw new Error("Plugin install response did not include task_id.");
result.task = await waitForTask(backendUrl, auth.token, result.task_id);
if (!isTaskComplete(result.task)) {
throw new Error(`Plugin install task did not complete successfully: ${JSON.stringify(result.task)}`);
}
@@ -98,11 +87,10 @@ try {
await sleep(1000);
result.plugin_present_after = await hasPlugin(backendUrl, auth.token);
if (!result.plugin_present_after) throw new Error(`${expectedPluginId} is not listed by /api/v1/plugins after install.`);
if (expectedTools.length > 0) {
if (expectedTool) {
result.tool_names = await listToolNames(backendUrl, auth.token);
const missingTools = expectedTools.filter((tool) => !result.tool_names.includes(tool));
if (missingTools.length > 0) {
throw new Error(`${missingTools.join(", ")} is not listed by /api/v1/tools after install.`);
if (!result.tool_names.includes(expectedTool)) {
throw new Error(`${expectedTool} is not listed by /api/v1/tools after install.`);
}
}
if (expectedRunnerId) {
@@ -133,41 +121,6 @@ async function hasPlugin(backendUrl, token) {
});
}
async function installPlugin(backendUrl, token, bytes, packagePath) {
const form = new FormData();
form.set("file", new Blob([bytes]), packagePath.split("/").pop());
const response = await fetch(`${backendUrl.replace(/\/$/, "")}/api/v1/plugins/install/local`, {
method: "POST",
headers: { Authorization: `Bearer ${token}` },
body: form,
});
const json = await response.json().catch(() => ({}));
if (response.status >= 400 || json.code !== 0) {
throw new Error(json.msg || `Plugin install request failed with HTTP ${response.status}.`);
}
const taskId = json.data?.task_id ?? null;
if (!taskId) throw new Error("Plugin install response did not include task_id.");
const task = await waitForTask(backendUrl, token, taskId);
return { id: taskId, ...task };
}
async function removePlugin(backendUrl, token) {
const [author, pluginName] = expectedPluginId.split("/");
if (!author || !pluginName) throw new Error(`Invalid expected plugin id: ${expectedPluginId}`);
const response = await apiJson(
backendUrl,
`/api/v1/plugins/${encodeURIComponent(author)}/${encodeURIComponent(pluginName)}?delete_data=false`,
{ method: "DELETE", token },
);
if (response.status >= 400 || response.json.code !== 0) {
throw new Error(response.json.msg || `Plugin delete request failed with HTTP ${response.status}.`);
}
const taskId = response.json.data?.task_id ?? null;
if (!taskId) throw new Error("Plugin delete response did not include task_id.");
const task = await waitForTask(backendUrl, token, taskId);
return { id: taskId, ...task };
}
async function previewPackage(backendUrl, token, bytes, packagePath) {
const form = new FormData();
form.set("file", new Blob([bytes]), packagePath.split("/").pop());
+29 -229
View File
@@ -1,7 +1,6 @@
import {
bodyText,
clickFirstVisible,
clickFirstVisibleLocator,
countOccurrences,
gotoFrontend,
isLoginUrl,
@@ -29,11 +28,6 @@ export function findNewFailureSignal(beforeText, afterText, failureSignals = DEB
return failureSignals.find((signal) => countOccurrences(afterText, signal) > countOccurrences(beforeText, signal)) || "";
}
export function hasDebugChatOutcome(text, expectedText, minExpectedCount, failureBaselines = []) {
if (countOccurrences(text, expectedText) >= minExpectedCount) return true;
return failureBaselines.some(({ signal, count }) => countOccurrences(text, signal) > count);
}
function findFailureSignalInText(text, failureSignals = DEBUG_CHAT_FAILURE_SIGNALS) {
return failureSignals.find((signal) => String(text || "").includes(signal)) || "";
}
@@ -51,24 +45,22 @@ function debugChatInput(page) {
}
async function clickDebugChatTab(page) {
const label = /^(?:Debug Chat|调试聊天|调试对话|对话调试)$/i;
const configuredTimeout = Number.parseInt(
process.env.LANGBOT_E2E_UI_READY_TIMEOUT_MS
|| process.env.LANGBOT_E2E_NAVIGATION_TIMEOUT_MS
|| "30000",
10,
);
const timeout = Number.isFinite(configuredTimeout) && configuredTimeout > 0
? configuredTimeout
: 30_000;
return await clickFirstVisibleLocator(page, [
page.getByRole("tab", { name: label }),
page.locator('[data-slot="tabs-trigger"]').filter({ hasText: label }),
page.getByText(label, { exact: true }),
], timeout);
const tabByRole = page.getByRole("tab", { name: /Debug Chat|调试聊天|调试对话|Debug|调试/i }).first();
if (await tabByRole.isVisible({ timeout: 3_000 }).catch(() => false)) {
await tabByRole.click();
return true;
}
const tabBySelector = page.locator('[role="tab"]').filter({ hasText: /Debug Chat|调试聊天|调试对话|Debug|调试/i }).first();
if (await tabBySelector.isVisible({ timeout: 2_000 }).catch(() => false)) {
await tabBySelector.click();
return true;
}
return Boolean(await clickFirstVisible(page, ["Debug Chat", "调试聊天", "调试对话"], 2_000));
}
export async function waitForDebugChatReady(page, timeout = 20_000) {
async function waitForDebugChatReady(page, timeout = 20_000) {
const input = debugChatInput(page);
const visible = await input.isVisible({ timeout }).catch(() => false);
if (!visible) {
@@ -78,13 +70,7 @@ export async function waitForDebugChatReady(page, timeout = 20_000) {
};
}
const deadline = Date.now() + timeout;
let enabled = false;
while (Date.now() < deadline) {
enabled = await input.isEnabled().catch(() => false);
if (enabled) break;
await page.waitForTimeout(Math.min(250, Math.max(1, deadline - Date.now())));
}
const enabled = await input.isEnabled({ timeout }).catch(() => false);
if (!enabled) {
return {
ready: false,
@@ -99,22 +85,14 @@ export function classifyDebugChatResult({
beforeText,
afterText,
expectedText,
expectedTexts = null,
prompt,
latestExpectedLeaf,
latestFailureLeaf,
beforeMessages = null,
afterMessages = null,
latestAssistantText = "",
latestAssistantIsFinal = null,
maxNewAssistantMessages = null,
failureSignals = DEBUG_CHAT_FAILURE_SIGNALS,
}) {
const requiredExpectedTexts = [...new Set(
(Array.isArray(expectedTexts) && expectedTexts.length > 0 ? expectedTexts : [expectedText])
.map(String)
.filter(Boolean),
)];
const minExpectedCount = minExpectedOccurrences(beforeText, expectedText, prompt);
const finalCount = countOccurrences(afterText, expectedText);
const failureText = findNewFailureSignal(beforeText, afterText, failureSignals);
@@ -126,25 +104,11 @@ export function classifyDebugChatResult({
const afterAssistantExpectedCount = hasMessageEvidence
? countExpectedInMessages(afterMessages, expectedText)
: null;
const beforeAssistantMessageCount = hasMessageEvidence
? beforeMessages.filter((message) => message.role === "assistant").length
: null;
const afterAssistantMessageCount = hasMessageEvidence
? afterMessages.filter((message) => message.role === "assistant").length
: null;
const newAssistantMessageCount = hasMessageEvidence
? afterAssistantMessageCount - beforeAssistantMessageCount
: null;
const assistantMessageEvidence = {
before_assistant_message_count: beforeAssistantMessageCount,
after_assistant_message_count: afterAssistantMessageCount,
new_assistant_message_count: newAssistantMessageCount,
};
const assistantExpectedIncreased = hasMessageEvidence
? afterAssistantExpectedCount > beforeAssistantExpectedCount
: false;
if (hasMessageEvidence) {
const missingExpectedTexts = requiredExpectedTexts.filter(
(text) => !String(latestAssistantText).includes(text),
);
const latestAssistantFailure = findFailureSignalInText(latestAssistantText, failureSignals);
if (latestAssistantFailure) {
return {
@@ -155,44 +119,16 @@ export function classifyDebugChatResult({
failure_signal: latestAssistantFailure,
before_assistant_expected_count: beforeAssistantExpectedCount,
after_assistant_expected_count: afterAssistantExpectedCount,
...assistantMessageEvidence,
};
}
if (latestAssistantIsFinal === false) {
return {
status: "fail",
reason: "The latest assistant message contained the expected text but was not final.",
min_expected_count: minExpectedCount,
final_count: finalCount,
before_assistant_expected_count: beforeAssistantExpectedCount,
after_assistant_expected_count: afterAssistantExpectedCount,
...assistantMessageEvidence,
latest_assistant_is_final: false,
};
}
if (maxNewAssistantMessages !== null && newAssistantMessageCount > maxNewAssistantMessages) {
return {
status: "fail",
reason: `Debug Chat created ${newAssistantMessageCount} assistant messages; expected at most ${maxNewAssistantMessages}.`,
min_expected_count: minExpectedCount,
final_count: finalCount,
before_assistant_expected_count: beforeAssistantExpectedCount,
after_assistant_expected_count: afterAssistantExpectedCount,
...assistantMessageEvidence,
};
}
if (newAssistantMessageCount > 0 && missingExpectedTexts.length === 0) {
if (assistantExpectedIncreased && String(latestAssistantText).includes(expectedText)) {
return {
status: "pass",
reason: requiredExpectedTexts.length === 1
? `Expected text appeared in a new assistant message: ${expectedText}`
: `All ${requiredExpectedTexts.length} expected text fragments appeared in a new assistant message.`,
reason: `Expected text appeared in a new assistant message: ${expectedText}`,
min_expected_count: minExpectedCount,
final_count: finalCount,
before_assistant_expected_count: beforeAssistantExpectedCount,
after_assistant_expected_count: afterAssistantExpectedCount,
...assistantMessageEvidence,
missing_expected_texts: [],
};
}
if (failureText) {
@@ -204,20 +140,15 @@ export function classifyDebugChatResult({
failure_signal: failureText,
before_assistant_expected_count: beforeAssistantExpectedCount,
after_assistant_expected_count: afterAssistantExpectedCount,
...assistantMessageEvidence,
};
}
return {
status: "fail",
reason: missingExpectedTexts.length > 0
? `A new assistant message was missing expected text: ${missingExpectedTexts.join(", ")}`
: `Expected text did not appear in a new assistant message. Expected assistant occurrences to increase above ${beforeAssistantExpectedCount}, saw ${afterAssistantExpectedCount}.`,
reason: `Expected text did not appear in a new assistant message. Expected assistant occurrences to increase above ${beforeAssistantExpectedCount}, saw ${afterAssistantExpectedCount}.`,
min_expected_count: minExpectedCount,
final_count: finalCount,
before_assistant_expected_count: beforeAssistantExpectedCount,
after_assistant_expected_count: afterAssistantExpectedCount,
...assistantMessageEvidence,
missing_expected_texts: missingExpectedTexts,
};
}
if (failureText) {
@@ -257,16 +188,8 @@ export function classifyDebugChatResult({
export async function openPipelineDebugChat(page, { pipelineUrl, pipelineName, envHint = "LANGBOT_PIPELINE_URL or LANGBOT_PIPELINE_NAME" }) {
if (pipelineUrl) {
let alreadyAtPipeline = false;
try {
alreadyAtPipeline = new URL(page.url()).href === new URL(pipelineUrl).href;
} catch {
// Invalid URLs are handled by page.goto below.
}
if (!alreadyAtPipeline) {
await page.goto(pipelineUrl, { waitUntil: "commit" });
await page.waitForLoadState("networkidle", { timeout: 10_000 }).catch(() => {});
}
await page.goto(pipelineUrl, { waitUntil: "domcontentloaded" });
await page.waitForLoadState("networkidle", { timeout: 10_000 }).catch(() => {});
} else {
if (!pipelineName) {
return {
@@ -367,37 +290,12 @@ export async function visibleDebugChatMessages(page) {
});
}
export async function waitForExpectedDebugChatText(page, {
expectedText,
expectedTexts = null,
minExpectedCount,
minExpectedCounts = null,
timeoutMs,
beforeText = "",
failureSignals = DEBUG_CHAT_FAILURE_SIGNALS,
}) {
const requiredExpectedTexts = [...new Set(
(Array.isArray(expectedTexts) && expectedTexts.length > 0 ? expectedTexts : [expectedText])
.map(String)
.filter(Boolean),
)];
const expectedRequirements = requiredExpectedTexts.map((text, index) => ({
text,
min: Array.isArray(minExpectedCounts) && Number.isFinite(minExpectedCounts[index])
? minExpectedCounts[index]
: (text === expectedText ? minExpectedCount : minExpectedOccurrences(beforeText, text, "")),
}));
const failureBaselines = failureSignals.map((signal) => ({
signal,
count: countOccurrences(beforeText, signal),
}));
export async function waitForExpectedDebugChatText(page, { expectedText, minExpectedCount, timeoutMs }) {
await page.waitForFunction(
({ requirements, failures }) => {
const text = document.body.innerText;
if (requirements.every((item) => text.split(item.text).length - 1 >= item.min)) return true;
return failures.some(({ signal, count }) => text.split(signal).length - 1 > count);
({ expected, min }) => {
return document.body.innerText.split(expected).length - 1 >= min;
},
{ requirements: expectedRequirements, failures: failureBaselines },
{ expected: expectedText, min: minExpectedCount },
{ timeout: timeoutMs },
).catch(() => {});
}
@@ -418,62 +316,6 @@ export async function waitForDebugChatTextStable(page, { timeoutMs = 5_000, quie
}
}
async function fetchDebugChatHistory(page, { backendUrl, pipelineId, sessionType }) {
if (!backendUrl || !pipelineId || !sessionType) {
return { status: "not_required", messages: [] };
}
return await page.evaluate(async ({ backendUrl, pipelineId, sessionType }) => {
const token = localStorage.getItem("token") || "";
const response = await fetch(
`${backendUrl.replace(/\/$/, "")}/api/v1/pipelines/${encodeURIComponent(pipelineId)}/ws/messages/${encodeURIComponent(sessionType)}`,
{ headers: token ? { Authorization: `Bearer ${token}` } : {} },
);
const json = await response.json().catch(() => ({}));
return {
status: response.ok && json.code === 0 ? "ready" : "fail",
http_status: response.status,
code: json.code ?? null,
messages: json.data?.messages || [],
reason: response.ok && json.code === 0 ? "" : json.msg || `Debug Chat history returned HTTP ${response.status}.`,
};
}, { backendUrl, pipelineId, sessionType });
}
async function waitForFinalDebugChatAssistant(page, {
backendUrl,
pipelineId,
sessionType,
beforeAssistantCount,
timeoutMs,
}) {
if (!backendUrl || !pipelineId || !sessionType) {
return { status: "not_required", latest_assistant_is_final: null };
}
const deadline = Date.now() + Math.max(1, timeoutMs);
let lastHistory = null;
while (Date.now() < deadline) {
lastHistory = await fetchDebugChatHistory(page, { backendUrl, pipelineId, sessionType });
if (lastHistory.status === "fail") return lastHistory;
const assistants = lastHistory.messages.filter((message) => message.role === "assistant");
const latest = assistants.at(-1);
if (assistants.length > beforeAssistantCount && latest?.is_final === true) {
return {
status: "pass",
latest_assistant_is_final: true,
assistant_message_count: assistants.length,
};
}
await page.waitForTimeout(Math.min(250, Math.max(1, deadline - Date.now())));
}
const assistants = (lastHistory?.messages || []).filter((message) => message.role === "assistant");
return {
status: "fail",
reason: "Timed out waiting for the new assistant message to become final.",
latest_assistant_is_final: assistants.at(-1)?.is_final === true,
assistant_message_count: assistants.length,
};
}
export async function attachDebugChatImage(page, imagePath) {
if (!imagePath) return { status: "not_required", reason: "" };
const input = page.locator('input[type="file"][accept*="image"], input[type="file"]').first();
@@ -503,54 +345,21 @@ export async function sendDebugChatPrompt(page, prompt, imagePath = "") {
return true;
}
export async function runDebugChatPrompt(page, {
prompt,
expectedText,
expectedTexts = null,
responseTimeoutMs,
imagePath = "",
backendUrl = "",
pipelineId = "",
sessionType = "person",
maxNewAssistantMessages = null,
failureSignals = DEBUG_CHAT_FAILURE_SIGNALS,
}) {
export async function runDebugChatPrompt(page, { prompt, expectedText, responseTimeoutMs, imagePath = "", failureSignals = DEBUG_CHAT_FAILURE_SIGNALS }) {
const beforeText = await bodyText(page);
const beforeMessages = await visibleDebugChatMessages(page);
const beforeHistory = await fetchDebugChatHistory(page, { backendUrl, pipelineId, sessionType });
const beforeHistoryAssistantCount = beforeHistory.messages.filter((message) => message.role === "assistant").length;
const requiredExpectedTexts = [...new Set(
(Array.isArray(expectedTexts) && expectedTexts.length > 0 ? expectedTexts : [expectedText])
.map(String)
.filter(Boolean),
)];
const minExpectedCount = minExpectedOccurrences(beforeText, expectedText, prompt);
const minExpectedCounts = requiredExpectedTexts.map(
(text) => minExpectedOccurrences(beforeText, text, prompt),
);
const sent = await sendDebugChatPrompt(page, prompt, imagePath);
if (sent !== true) {
if (sent && typeof sent === "object" && typeof sent.reason === "string") return sent;
return { status: "fail", reason: "Could not find a Debug Chat text input." };
}
const responseStartedAt = Date.now();
await waitForExpectedDebugChatText(page, {
expectedText,
expectedTexts: requiredExpectedTexts,
minExpectedCount,
minExpectedCounts,
prompt,
timeoutMs: responseTimeoutMs,
beforeText,
failureSignals,
});
const finalAssistant = await waitForFinalDebugChatAssistant(page, {
backendUrl,
pipelineId,
sessionType,
beforeAssistantCount: beforeHistoryAssistantCount,
timeoutMs: Math.max(1, responseTimeoutMs - (Date.now() - responseStartedAt)),
});
await waitForDebugChatTextStable(page);
@@ -561,27 +370,18 @@ export async function runDebugChatPrompt(page, {
const failureText = findNewFailureSignal(beforeText, afterText, failureSignals);
const latestFailureLeaf = failureText ? await latestVisibleLeafText(page, [failureText]) : "";
const classified = classifyDebugChatResult({
return classifyDebugChatResult({
beforeText,
afterText,
expectedText,
expectedTexts: requiredExpectedTexts,
prompt,
latestExpectedLeaf,
latestFailureLeaf,
beforeMessages,
afterMessages,
latestAssistantText,
latestAssistantIsFinal: finalAssistant.latest_assistant_is_final,
maxNewAssistantMessages,
failureSignals,
});
return {
...classified,
latest_assistant_is_final: finalAssistant.latest_assistant_is_final,
final_assistant_wait_status: finalAssistant.status,
final_assistant_wait_reason: finalAssistant.reason || "",
};
}
export async function setDebugChatStreamOutput(page, desired) {
@@ -1,200 +0,0 @@
import { execFile } from "node:child_process";
import { closeSync, openSync } from "node:fs";
import { mkdtemp, readFile, rm } from "node:fs/promises";
import { createServer } from "node:net";
import { tmpdir } from "node:os";
import { dirname, join, resolve } from "node:path";
import { fileURLToPath } from "node:url";
import { promisify } from "node:util";
import { spawn } from "node:child_process";
import { pathExists, resolveLangBotRepo } from "./langbot-e2e.mjs";
const execFileAsync = promisify(execFile);
const proxyKeys = [
"ALL_PROXY",
"all_proxy",
"HTTP_PROXY",
"http_proxy",
"HTTPS_PROXY",
"https_proxy",
];
function withoutProxy(source = process.env) {
const next = { ...source };
for (const key of proxyKeys) delete next[key];
next.NO_PROXY = "127.0.0.1,localhost";
next.no_proxy = "127.0.0.1,localhost";
return next;
}
async function getFreePort() {
const server = createServer();
await new Promise((resolvePromise, reject) => {
server.once("error", reject);
server.listen(0, "127.0.0.1", resolvePromise);
});
const address = server.address();
const port = typeof address === "object" && address ? address.port : 0;
await new Promise((resolvePromise) => server.close(resolvePromise));
if (!port) throw new Error("Could not allocate an isolated LangBot port.");
return port;
}
async function waitForHttp(url, child, timeoutMs, label) {
const deadline = Date.now() + timeoutMs;
while (Date.now() < deadline) {
if (child.exitCode !== null) {
throw new Error(`${label} exited before becoming ready (code ${child.exitCode}).`);
}
try {
const response = await fetch(url, { redirect: "manual" });
if (response.status < 500) return response;
} catch {
// The service is still starting.
}
await new Promise((resolvePromise) => setTimeout(resolvePromise, 500));
}
throw new Error(`${label} did not become ready within ${timeoutMs}ms.`);
}
function spawnLogged(command, args, { cwd, env, logPath }) {
const logFd = openSync(logPath, "a");
try {
return spawn(command, args, {
cwd,
env,
detached: true,
stdio: ["ignore", logFd, logFd],
});
} finally {
closeSync(logFd);
}
}
async function stopProcess(child) {
if (!child || child.exitCode !== null) return;
const closed = new Promise((resolvePromise) => child.once("close", resolvePromise));
try {
process.kill(-child.pid, "SIGTERM");
} catch {
child.kill("SIGTERM");
}
const graceful = await Promise.race([
closed.then(() => true),
new Promise((resolvePromise) => setTimeout(() => resolvePromise(false), 5_000)),
]);
if (graceful) return;
try {
process.kill(-child.pid, "SIGKILL");
} catch {
child.kill("SIGKILL");
}
await closed.catch(() => {});
}
async function initializeUser(backendUrl, user, password) {
const response = await fetch(`${backendUrl}/api/v1/user/init`, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ user, password }),
});
const payload = await response.json().catch(() => ({}));
if (response.status >= 400 || ![0, 1].includes(payload.code)) {
throw new Error(payload.msg || `Could not initialize isolated user (HTTP ${response.status}).`);
}
}
export async function startIsolatedLangBotInstance({ evidenceDir }) {
const repo = await resolveLangBotRepo();
const webRepo = process.env.LANGBOT_WEB_REPO || join(repo, "web");
const python = join(repo, ".venv", "bin", "python");
const configScript = resolve(
dirname(fileURLToPath(import.meta.url)),
"../prepare-isolated-langbot-config.py",
);
if (!(await pathExists(python))) throw new Error(`LangBot virtualenv Python is missing: ${python}`);
if (!(await pathExists(join(webRepo, "node_modules")))) {
throw new Error(`LangBot frontend dependencies are missing: ${join(webRepo, "node_modules")}`);
}
const [backendPort, frontendPort, pluginDebugPort] = await Promise.all([
getFreePort(),
getFreePort(),
getFreePort(),
]);
const instanceRoot = await mkdtemp(join(tmpdir(), "langbot-clean-catalog-"));
const backendUrl = `http://127.0.0.1:${backendPort}`;
const frontendUrl = `http://127.0.0.1:${frontendPort}`;
const backendLog = join(evidenceDir, "isolated-backend.log");
const frontendLog = join(evidenceDir, "isolated-frontend.log");
const user = "langbot-e2e@example.invalid";
const password = "LangBotIsolatedE2E!2026";
let backend;
let frontend;
const stop = async () => {
await stopProcess(frontend);
await stopProcess(backend);
await rm(instanceRoot, { recursive: true, force: true });
};
try {
await execFileAsync(
python,
[
configScript,
"--instance-root",
instanceRoot,
"--port",
String(backendPort),
"--plugin-debug-port",
String(pluginDebugPort),
],
{ cwd: repo, env: withoutProxy(), timeout: 30_000 },
);
backend = spawnLogged(python, ["-m", "langbot"], {
cwd: instanceRoot,
env: {
...withoutProxy(),
PYTHONPATH: join(repo, "src"),
LANGBOT_DATA_ROOT: join(instanceRoot, "data"),
API__PORT: String(backendPort),
API__WEBHOOK_PREFIX: backendUrl,
SPACE__DISABLE_TELEMETRY: "true",
SPACE__DISABLE_MODELS_SERVICE: "true",
},
logPath: backendLog,
});
await waitForHttp(`${backendUrl}/api/v1/system/info`, backend, 180_000, "Isolated LangBot backend");
const configText = await readFile(join(instanceRoot, "data", "config.yaml"), "utf8");
const recoveryKey = configText.match(/^\s*recovery_key:\s*['"]?([^'"\s#]+)['"]?\s*$/m)?.[1] || "";
if (!recoveryKey) throw new Error("Isolated LangBot did not generate a recovery key.");
await initializeUser(backendUrl, user, password);
frontend = spawnLogged(
"pnpm",
["exec", "vite", "--host", "127.0.0.1", "--port", String(frontendPort), "--strictPort"],
{
cwd: webRepo,
env: { ...withoutProxy(), VITE_API_BASE_URL: backendUrl },
logPath: frontendLog,
},
);
await waitForHttp(frontendUrl, frontend, 60_000, "Isolated LangBot frontend");
return {
backendUrl,
frontendUrl,
recoveryKey,
user,
password,
logs: { backend: backendLog, frontend: frontendLog },
stop,
};
} catch (error) {
await stop();
throw error;
}
}
+23 -228
View File
@@ -1,5 +1,5 @@
import { appendFile, mkdir, readFile, stat, writeFile } from "node:fs/promises";
import { basename, join, resolve } from "node:path";
import { join, resolve } from "node:path";
import { env } from "node:process";
const secretRe = /(?:authorization|bearer|token|secret|password|api[_-]?key|jwt|oauth)\s*[:=]\s*["']?[^"',\s]+/gi;
@@ -50,38 +50,6 @@ export async function ensureEvidence(paths) {
await appendFile(paths.networkLog, "", "utf8");
}
export async function beginBackendLogCapture(evidenceDir, sourcePath = env.LANGBOT_BACKEND_LOG || "") {
if (!sourcePath) return null;
const source = resolve(sourcePath);
try {
const info = await stat(source);
return {
source,
start_offset: info.size,
target: resolve(evidenceDir, "backend.log"),
};
} catch {
return null;
}
}
export async function finishBackendLogCapture(capture) {
if (!capture) return null;
try {
const content = await readFile(capture.source);
const start = content.length >= capture.start_offset ? capture.start_offset : 0;
const window = content.subarray(start);
if (window.length === 0) return null;
await writeFile(capture.target, window);
return {
path: capture.target,
bytes: window.length,
};
} catch {
return null;
}
}
export async function pathExists(path) {
try {
await stat(path);
@@ -103,86 +71,6 @@ export async function writeResult(paths, result) {
}
}
function browserDiagnosticFindings(source, text) {
const findings = [];
const lines = String(text || "").split(/\r?\n/);
for (let index = 0; index < lines.length; index += 1) {
const line = lines[index];
if (!line) continue;
const lineNumber = index + 1;
if (source === "console") {
const checks = [
["pageerror", /\[pageerror\]/i],
["frontend_uncaught_error", /\[error\].*(?:\bUncaught\b|Unhandled(?: promise rejection|Rejection)|TypeError|ReferenceError)/i],
["http_5xx", /Failed to load resource: the server responded with a status of 5\d\d/i],
["api_server_error", /\[error\].*Server error:/i],
["plugin_runtime_timeout", /\[error\].*Action [A-Za-z0-9_]+ call timed out/i],
];
for (const [kind, regex] of checks) {
if (!regex.test(line)) continue;
findings.push({
source,
severity: "fail",
kind,
line: lineNumber,
excerpt: redact(line.trim()),
});
break;
}
continue;
}
if (source === "network") {
if (/\[response\]\s+5\d\d\b/i.test(line)) {
findings.push({
source,
severity: "fail",
kind: "http_5xx",
line: lineNumber,
excerpt: redact(line.trim()),
});
continue;
}
if (/\[requestfailed\]/i.test(line) && !/net::ERR_ABORTED/i.test(line)) {
findings.push({
source,
severity: "warning",
kind: "request_failed",
line: lineNumber,
excerpt: redact(line.trim()),
});
}
}
}
return findings;
}
export async function scanBrowserDiagnostics(paths) {
const sources = [
["console", paths.consoleLog],
["network", paths.networkLog],
];
const findings = [];
for (const [source, path] of sources) {
let text = "";
try {
text = await readFile(path, "utf8");
} catch {
continue;
}
findings.push(...browserDiagnosticFindings(source, text));
}
const hasFailure = findings.some((finding) => finding.severity === "fail");
return {
status: hasFailure ? "fail" : "pass",
findings,
reason: hasFailure
? `Browser diagnostics found ${findings.filter((finding) => finding.severity === "fail").length} failing signal(s).`
: "No failing browser diagnostics found.",
};
}
export async function loadEnvFiles(paths = ["skills/.env", "skills/.env.local"]) {
const processEnvKeys = new Set(Object.keys(env));
for (const path of paths) {
@@ -204,28 +92,8 @@ export async function loadEnvFiles(paths = ["skills/.env", "skills/.env.local"])
}
}
export async function resolveLangBotRepo(repo = env.LANGBOT_REPO || "", cwd = process.cwd()) {
if (repo) return resolve(repo);
const candidates = [
resolve(cwd),
basename(cwd) === "skills" ? resolve(cwd, "..") : "",
resolve(cwd, "../LangBot"),
resolve(cwd, "LangBot"),
].filter(Boolean);
const seen = new Set();
for (const candidate of candidates) {
if (seen.has(candidate)) continue;
seen.add(candidate);
if (await pathExists(resolve(candidate, "data/config.yaml"))) return candidate;
}
return resolve(cwd, "../LangBot");
}
export async function readRecoveryKey(repo = env.LANGBOT_REPO || "") {
const configPath = resolve(await resolveLangBotRepo(repo), "data/config.yaml");
export async function readRecoveryKey(repo = env.LANGBOT_REPO || "../LangBot") {
const configPath = resolve(repo, "data/config.yaml");
const config = await readFile(configPath, "utf8");
const match = config.match(/^\s*recovery_key:\s*['"]?([^'"\s#]+)['"]?\s*$/m);
return match?.[1] || "";
@@ -334,62 +202,6 @@ export async function verifyBrowserToken(page, backendUrl) {
}, backendUrl);
}
export async function ensureAuthenticatedBrowser(page, {
frontendUrl = env.LANGBOT_FRONTEND_URL || "",
backendUrl = env.LANGBOT_BACKEND_URL || "",
user = env.LANGBOT_E2E_LOGIN_USER || "",
password = env.LANGBOT_E2E_LOGIN_PASSWORD || "LangBotE2ELocalPass!2026",
recoveryKey = "",
} = {}) {
if (!frontendUrl) return { status: "env_issue", reason: "LANGBOT_FRONTEND_URL is not configured." };
if (!backendUrl) return { status: "env_issue", reason: "LANGBOT_BACKEND_URL is not configured." };
const current = await verifyBrowserToken(page, backendUrl).catch((error) => ({
authenticated: false,
reason: error.message,
}));
if (current.authenticated) {
return {
status: "pass",
reason: "Existing browser token is valid.",
backend_token_check: null,
browser_token_check: current,
injected: false,
};
}
if (!user) {
return {
status: "blocked",
reason: "Browser profile is not authenticated for LANGBOT_FRONTEND_URL, and LANGBOT_E2E_LOGIN_USER is not configured for automatic local login.",
backend_token_check: null,
browser_token_check: current,
injected: false,
};
}
const auth = await resetAndAuthLocalUser({ backendUrl, user, password, recoveryKey });
await setBrowserToken(page, frontendUrl, auth.token);
const browserCheck = await verifyBrowserToken(page, backendUrl);
if (!browserCheck.authenticated) {
return {
status: "blocked",
reason: browserCheck.reason || "Browser token check failed after automatic local login.",
backend_token_check: auth.check,
browser_token_check: browserCheck,
injected: true,
};
}
return {
status: "pass",
reason: "Browser token injected from local recovery login.",
backend_token_check: auth.check,
browser_token_check: browserCheck,
injected: true,
};
}
export function exitCode(status) {
if (status === "pass") return 0;
if (status === "blocked" || status === "env_issue") return 2;
@@ -428,10 +240,6 @@ export async function createBrowser(paths) {
context = await browser.newContext({ viewport: { width: 1440, height: 960 } });
}
const page = context.pages()[0] || await context.newPage();
const navigationTimeoutMs = Number.parseInt(env.LANGBOT_E2E_NAVIGATION_TIMEOUT_MS || "30000", 10);
if (Number.isFinite(navigationTimeoutMs) && navigationTimeoutMs > 0) {
page.setDefaultNavigationTimeout(navigationTimeoutMs);
}
page.on("console", (message) => {
appendLine(paths.consoleLog, `[${message.type()}] ${message.text()}`).catch(() => {});
@@ -487,40 +295,27 @@ export function countOccurrences(haystack, needle) {
return String(haystack).split(needle).length - 1;
}
async function clickVisibleCandidate(page, candidates, timeout) {
const deadline = Date.now() + Math.max(1, timeout);
do {
for (const candidate of candidates) {
const count = await candidate.locator.count().catch(() => 0);
for (let index = 0; index < count; index += 1) {
const element = candidate.locator.nth(index);
if (!await element.isVisible().catch(() => false)) continue;
const remaining = Math.max(1, deadline - Date.now());
const clicked = await element.click({ timeout: Math.min(1_000, remaining) })
.then(() => true)
.catch(() => false);
if (clicked) return candidate.value;
}
}
const remaining = deadline - Date.now();
if (remaining <= 0) break;
await page.waitForTimeout(Math.min(100, remaining));
} while (Date.now() < deadline);
return null;
}
export async function clickFirstVisibleLocator(page, locators, timeout = 2_000) {
const candidates = locators.map((locator) => ({ locator, value: true }));
return Boolean(await clickVisibleCandidate(page, candidates, timeout));
}
export async function clickFirstVisible(page, labels, timeout = 2_000) {
const candidates = labels.flatMap((label) => [
{ locator: page.getByRole("button", { name: label }), value: label },
{ locator: page.getByRole("link", { name: label }), value: label },
{ locator: page.getByText(label, { exact: false }), value: label },
]);
return await clickVisibleCandidate(page, candidates, timeout);
for (const label of labels) {
const roleButton = page.getByRole("button", { name: label }).first();
if (await roleButton.isVisible({ timeout }).catch(() => false)) {
await roleButton.click();
return label;
}
const roleLink = page.getByRole("link", { name: label }).first();
if (await roleLink.isVisible({ timeout }).catch(() => false)) {
await roleLink.click();
return label;
}
const text = page.getByText(label, { exact: false }).first();
if (await text.isVisible({ timeout }).catch(() => false)) {
await text.click();
return label;
}
}
return null;
}
export async function fillFirstTextInput(page, value) {
@@ -10,13 +10,10 @@ import {
waitForDebugChatTextStable,
} from "./lib/debug-chat.mjs";
import {
beginBackendLogCapture,
createBrowser,
ensureAuthenticatedBrowser,
ensureEvidence,
evidencePaths,
exitCode,
finishBackendLogCapture,
localIsoWithOffset,
loadEnvFiles,
pathExists,
@@ -29,12 +26,11 @@ await loadEnvFiles();
const caseId = env.LBS_CASE_ID || "local-agent-steering-debug-chat";
const paths = evidencePaths(caseId);
await ensureEvidence(paths);
const backendLogCapture = await beginBackendLogCapture(paths.evidenceDir);
const backendUrl = (env.LANGBOT_BACKEND_URL || "").replace(/\/$/, "");
const pipelineUrl = env.LANGBOT_E2E_PIPELINE_URL || env.LANGBOT_LOCAL_AGENT_PIPELINE_URL || env.LANGBOT_PIPELINE_URL || "";
const pipelineName = env.LANGBOT_E2E_PIPELINE_NAME || env.LANGBOT_LOCAL_AGENT_PIPELINE_NAME || env.LANGBOT_PIPELINE_NAME || "";
const expectedRunnerId = env.LANGBOT_E2E_EXPECTED_RUNNER_ID || "plugin:langbot-team/LocalAgent/default";
const expectedRunnerId = env.LANGBOT_E2E_EXPECTED_RUNNER_ID || "plugin:langbot/local-agent/default";
const expectedText = env.LANGBOT_E2E_EXPECTED_TEXT || "qa_steering_sentinel_6194";
const responseTimeoutMs = positiveInt(env.LANGBOT_E2E_RESPONSE_TIMEOUT_MS, 240000);
const followupDelayMs = 1000;
@@ -78,7 +74,6 @@ const result = {
pipeline_config: null,
debug_chat_reset: null,
tool_diagnostic: null,
browser_auth: null,
steering: null,
evidence: {
console_log: paths.consoleLog,
@@ -100,18 +95,6 @@ try {
browser = await createBrowser(paths);
const { page } = browser;
const authDiagnostic = await ensureAuthenticatedBrowser(page, {
frontendUrl: env.LANGBOT_FRONTEND_URL || "",
backendUrl,
});
result.browser_auth = authDiagnostic;
if (!result.evidence_collected.includes("api_diagnostic")) result.evidence_collected.push("api_diagnostic");
if (authDiagnostic.status === "env_issue" || authDiagnostic.status === "blocked" || authDiagnostic.status === "fail") {
result.status = authDiagnostic.status;
result.reason = authDiagnostic.reason || "Browser authentication failed.";
throw new Error(result.reason);
}
const openResult = await openPipelineDebugChat(page, {
pipelineUrl,
pipelineName,
@@ -193,12 +176,6 @@ try {
} finally {
if (browser?.page) await safeScreenshot(browser.page, paths.screenshot);
if (browser) await browser.close().catch(() => {});
const backendLog = await finishBackendLogCapture(backendLogCapture);
if (backendLog) {
result.evidence.backend_log = backendLog.path;
result.backend_log = backendLog;
if (!result.evidence_collected.includes("backend_log")) result.evidence_collected.push("backend_log");
}
const finishedAt = new Date();
result.finished_at = finishedAt.toISOString();
result.finished_at_local = localIsoWithOffset(finishedAt);
@@ -455,7 +432,7 @@ async function inspectPipeline(page, { backendUrl, pipelineUrl, pipelineName, ex
}
const config = pipeline.config || {};
const runner = config.ai?.runner || {};
const runnerId = runner.id || "";
const runnerId = runner.id || runner.runner || "";
if (!runnerId) {
return {
status: "blocked",
@@ -464,7 +441,7 @@ async function inspectPipeline(page, { backendUrl, pipelineUrl, pipelineName, ex
pipeline_id: pipelineId,
pipeline_name: pipeline.name,
matched_by: matchedBy,
reason: "Pipeline has no ai.runner.id.",
reason: "Pipeline has no ai.runner.id or legacy ai.runner.runner.",
};
}
if (expectedRunnerId && runnerId !== expectedRunnerId) {
+3 -45
View File
@@ -1,7 +1,7 @@
#!/usr/bin/env node
import { existsSync, readFileSync } from "node:fs";
import { readFile, writeFile } from "node:fs/promises";
import { writeFile } from "node:fs/promises";
import { resolve } from "node:path";
import { env } from "node:process";
import {
@@ -46,8 +46,6 @@ const startupTimeoutSec = Number(env.LANGBOT_MCP_STARTUP_TIMEOUT_SEC || "300");
const readyTimeoutMs = Number(env.LANGBOT_MCP_READY_TIMEOUT_MS || "360000");
const backendUrl = env.LANGBOT_BACKEND_URL || "";
const apiDiagnosticPath = resolve(paths.evidenceDir, "api-diagnostic.json");
const envLocalPath = resolve("skills/.env.local");
const serverUuidEnvKey = env.LANGBOT_MCP_SERVER_UUID_ENV_KEY || "LANGBOT_MCP_QA_STDIO_SERVER_UUID";
let browser;
const result = {
@@ -71,8 +69,7 @@ const result = {
automation_result_json: paths.automationResultJson,
result_json: paths.resultJson,
},
evidence_collected: ["screenshot", "console", "network", "api_diagnostic"],
wrote_env: false,
evidence_collected: ["api_diagnostic"],
};
async function run() {
@@ -162,7 +159,6 @@ async function run() {
const deadline = Date.now() + readyTimeoutMs;
let lastTools = [];
let lastRuntime = null;
let lastServer = null;
while (Date.now() < deadline) {
await new Promise((resolveReady) => setTimeout(resolveReady, 500));
const tools = await getJson("/api/v1/tools");
@@ -171,8 +167,7 @@ async function run() {
.map((tool) => tool.name || tool.tool_name || tool.function?.name || "")
.filter(Boolean)
.sort();
lastServer = server.json.data?.server || null;
lastRuntime = lastServer?.runtime_info || null;
lastRuntime = server.json.data?.server?.runtime_info || null;
if (lastTools.includes(expectedTool)) break;
}
@@ -182,7 +177,6 @@ async function run() {
save_status: save.status,
save_code: save.json.code ?? null,
save_msg: save.json.msg || "",
server_uuid: lastServer?.uuid || save.json.data?.uuid || "",
tool_names: lastTools,
has_expected_tool: lastTools.includes(expectedTool),
runtime_status: lastRuntime?.status || null,
@@ -218,47 +212,11 @@ async function run() {
result.reason = `MCP server ${serverName} did not expose ${expectedTool}. See ${apiDiagnosticPath}.`;
return;
}
if (!diagnostic.server_uuid) {
result.status = "fail";
result.reason = `MCP server ${serverName} exposed ${expectedTool}, but the server UUID was not returned. See ${apiDiagnosticPath}.`;
return;
}
await upsertEnvLocal(envLocalPath, {
[serverUuidEnvKey]: diagnostic.server_uuid,
});
result.wrote_env = true;
result.server_uuid = diagnostic.server_uuid;
result.server_uuid_env_key = serverUuidEnvKey;
result.status = "pass";
result.reason = `MCP server ${serverName} is connected and exposes ${expectedTool} through LangBot /api/v1/tools.`;
}
async function upsertEnvLocal(path, updates) {
let text = "";
try {
text = await readFile(path, "utf8");
} catch {
text = "";
}
const lines = text ? text.split(/\r?\n/) : [];
const seen = new Set();
const updated = lines.map((line) => {
const match = line.match(/^([A-Z][A-Z0-9_]*)=/);
if (!match || !Object.prototype.hasOwnProperty.call(updates, match[1])) return line;
seen.add(match[1]);
return `${match[1]}=${updates[match[1]]}`;
});
for (const [key, value] of Object.entries(updates)) {
if (!seen.has(key)) updated.push(`${key}=${value}`);
}
await writeFile(path, `${updated.filter((line, index) => line || index < updated.length - 1).join("\n")}\n`, "utf8");
}
try {
await run();
} catch (error) {
@@ -1,93 +0,0 @@
#!/usr/bin/env python3
"""Drive one deterministic event through an enabled OneBot reverse WebSocket."""
from __future__ import annotations
import argparse
import asyncio
import json
import time
import websockets
async def run(port: int) -> None:
uri = f'ws://127.0.0.1:{port}/ws'
headers = {
'X-Self-ID': '900001',
'X-Client-Role': 'Universal',
'User-Agent': 'LangBot-E2E-OneBot/1.0',
}
connect_deadline = time.monotonic() + 15
connection = None
while time.monotonic() < connect_deadline:
try:
connection = await websockets.connect(uri, additional_headers=headers)
break
except OSError:
await asyncio.sleep(0.25)
if connection is None:
raise RuntimeError(f'OneBot reverse WebSocket did not open on port {port}')
actions: list[str] = []
event_deadline = time.monotonic() + 20
async with connection as ws:
await ws.send(
json.dumps(
{
'post_type': 'notice',
'notice_type': 'group_increase',
'sub_type': 'invite',
'time': int(time.time()),
'self_id': 900001,
'group_id': 20001,
'operator_id': 10002,
'user_id': 10003,
}
)
)
delivered = False
while time.monotonic() < event_deadline and not delivered:
try:
raw = await asyncio.wait_for(ws.recv(), timeout=2)
except asyncio.TimeoutError:
continue
request = json.loads(raw)
action = request.get('action', '')
actions.append(action)
if action == 'get_group_info':
data = {'group_id': 20001, 'group_name': 'LangBot Runtime QA'}
elif action == 'get_group_member_info':
data = {
'group_id': 20001,
'user_id': 10003,
'nickname': 'Runtime QA Member',
'card': 'Runtime QA Member',
}
elif action == 'send_group_msg':
data = {'message_id': 70001}
delivered = True
else:
data = {}
await ws.send(
json.dumps(
{
'status': 'ok',
'retcode': 0,
'data': data,
'echo': request.get('echo'),
}
)
)
if not delivered:
raise RuntimeError(f'Agent output was not delivered; actions={actions}')
print(json.dumps({'connected': True, 'actions': actions, 'delivered': True}))
if __name__ == '__main__':
parser = argparse.ArgumentParser()
parser.add_argument('--port', type=int, required=True)
args = parser.parse_args()
asyncio.run(run(args.port))
+23 -373
View File
@@ -10,24 +10,19 @@ import {
setDebugChatStreamOutput,
} from "./lib/debug-chat.mjs";
import {
beginBackendLogCapture,
createBrowser,
ensureAuthenticatedBrowser,
ensureEvidence,
evidencePaths,
exitCode,
finishBackendLogCapture,
localIsoWithOffset,
pathExists,
safeScreenshot,
scanBrowserDiagnostics,
writeResult,
} from "./lib/langbot-e2e.mjs";
const caseId = env.LBS_CASE_ID || "pipeline-debug-chat";
const paths = evidencePaths(caseId);
await ensureEvidence(paths);
const backendLogCapture = await beginBackendLogCapture(paths.evidenceDir);
const expectedText = env.LANGBOT_E2E_EXPECTED_TEXT || "OK";
const prompt = env.LANGBOT_E2E_PROMPT || `请只回复 ${expectedText},用于前端调试测试。`;
@@ -55,20 +50,15 @@ const pipelineName = pipelineRequired
const expectedRunnerId = env.LANGBOT_E2E_EXPECTED_RUNNER_ID || "";
const resetDebugChat = boolFromEnv(env.LANGBOT_E2E_RESET_DEBUG_CHAT, false);
const restoreRunnerConfig = boolFromEnv(env.LANGBOT_E2E_RESTORE_RUNNER_CONFIG, true);
const restoreExtensions = boolFromEnv(env.LANGBOT_E2E_RESTORE_EXTENSIONS, true);
const debugChatSessionType = env.LANGBOT_E2E_DEBUG_CHAT_SESSION_TYPE || "person";
const maxNewAssistantMessages = Number.parseInt(env.LANGBOT_E2E_MAX_NEW_ASSISTANT_MESSAGES || "1", 10);
const pipelineConfigDiagnosticPath = resolve(paths.evidenceDir, "pipeline-config-diagnostic.json");
const pipelineExtensionsDiagnosticPath = resolve(paths.evidenceDir, "pipeline-extensions-diagnostic.json");
const debugChatResetDiagnosticPath = resolve(paths.evidenceDir, "debug-chat-reset-diagnostic.json");
const pipelineConfigRestoreDiagnosticPath = resolve(paths.evidenceDir, "pipeline-config-restore-diagnostic.json");
const pipelineExtensionsRestoreDiagnosticPath = resolve(paths.evidenceDir, "pipeline-extensions-restore-diagnostic.json");
const metricsPath = resolve(paths.evidenceDir, "metrics.json");
const startedAt = new Date();
let browser;
let restoreConfigPlan = null;
let restoreExtensionsPlan = null;
let restorePlan = null;
let result = {
source: "automation",
case_id: caseId,
@@ -116,9 +106,7 @@ function parseJsonEnv(key, fallback) {
}
function positiveNumberEnv(key, fallback) {
const raw = env[key];
if (raw === undefined || raw === "") return fallback;
const value = Number(raw);
const value = Number(env[key] || "");
return Number.isFinite(value) && value >= 0 ? value : fallback;
}
@@ -143,28 +131,22 @@ function stats(values) {
function promptStepsFromEnv() {
const rawSteps = parseJsonEnv("LANGBOT_E2E_PROMPTS_JSON", null);
if (rawSteps === null) {
return [{ prompt, expectedText, expectedTexts: [expectedText], responseTimeoutMs: safeResponseTimeoutMs }];
return [{ prompt, expectedText, responseTimeoutMs: safeResponseTimeoutMs }];
}
if (!Array.isArray(rawSteps) || rawSteps.length === 0) {
throw new Error("LANGBOT_E2E_PROMPTS_JSON must be a non-empty JSON array.");
}
return rawSteps.map((item, index) => {
if (typeof item === "string") {
return { prompt: item, expectedText, expectedTexts: [expectedText], responseTimeoutMs: safeResponseTimeoutMs };
return { prompt: item, expectedText, responseTimeoutMs: safeResponseTimeoutMs };
}
if (!item || typeof item !== "object" || typeof item.prompt !== "string" || !item.prompt) {
throw new Error(`LANGBOT_E2E_PROMPTS_JSON[${index}] must be a string or an object with a prompt string.`);
}
const stepTimeout = Number.parseInt(String(item.response_timeout_ms || item.responseTimeoutMs || safeResponseTimeoutMs), 10);
const stepExpectedText = String(item.expected_text || item.expectedText || expectedText);
const additionalExpectedTexts = item.expected_texts || item.expectedTexts || [];
if (!Array.isArray(additionalExpectedTexts)) {
throw new Error(`LANGBOT_E2E_PROMPTS_JSON[${index}].expected_texts must be an array.`);
}
return {
prompt: item.prompt,
expectedText: stepExpectedText,
expectedTexts: [...new Set([stepExpectedText, ...additionalExpectedTexts.map(String)].filter(Boolean))],
expectedText: String(item.expected_text || item.expectedText || expectedText),
responseTimeoutMs: Number.isFinite(stepTimeout) && stepTimeout > 0 ? stepTimeout : safeResponseTimeoutMs,
};
});
@@ -249,7 +231,6 @@ async function runFilesystemChecks(checks) {
}
const contains = textList(check.contains);
const notContains = textList(check.not_contains || check.notContains);
const expectedJson = check.json_equals ?? check.jsonEquals;
const expectedExitCode = Number.isInteger(check.exit_code)
? check.exit_code
: Number.isInteger(check.expected_exit_code)
@@ -268,31 +249,18 @@ async function runFilesystemChecks(checks) {
}
const missing = contains.filter((needle) => !text.includes(needle));
const forbidden = notContains.filter((needle) => text.includes(needle));
let jsonError = "";
let jsonMatches = null;
if (expectedJson !== undefined) {
try {
jsonMatches = JSON.stringify(sortJson(JSON.parse(text))) === JSON.stringify(sortJson(expectedJson));
if (!jsonMatches) jsonError = "JSON content does not match json_equals.";
} catch (error) {
jsonMatches = false;
jsonError = `Invalid JSON: ${error.message}`;
}
}
const failed = missing.length > 0 || forbidden.length > 0 || jsonMatches === false;
results.push({
index,
status: failed ? "fail" : "pass",
status: missing.length || forbidden.length ? "fail" : "pass",
type: "file",
path,
missing,
forbidden,
json_matches: jsonMatches,
reason: missing.length
? `Missing expected text: ${missing.join(", ")}`
: forbidden.length
? `Found forbidden text: ${forbidden.join(", ")}`
: jsonError,
: "",
});
continue;
}
@@ -334,16 +302,6 @@ async function runFilesystemChecks(checks) {
};
}
function sortJson(value) {
if (Array.isArray(value)) return value.map(sortJson);
if (!value || typeof value !== "object") return value;
return Object.fromEntries(
Object.keys(value)
.sort()
.map((key) => [key, sortJson(value[key])]),
);
}
function pipelineIdFromUrl(url) {
if (!url) return "";
try {
@@ -359,11 +317,6 @@ function sanitizePipelineDiagnostic(diagnostic) {
return safe;
}
function sanitizePipelineExtensionsDiagnostic(diagnostic) {
const { restore_extensions: _restoreExtensions, ...safe } = diagnostic || {};
return safe;
}
async function prepareImageFixture(paths) {
if (imagePathEnv) return resolve(imagePathEnv);
if (!imageBase64Path) return "";
@@ -465,7 +418,7 @@ async function inspectAndPatchPipelineConfig(page, {
const config = JSON.parse(JSON.stringify(pipeline.config || {}));
const aiConfig = config.ai && typeof config.ai === "object" ? config.ai : {};
const runner = aiConfig.runner && typeof aiConfig.runner === "object" ? aiConfig.runner : {};
const runnerId = runner.id || "";
const runnerId = runner.id || runner.runner || "";
if (!runnerId) {
return {
status: "blocked",
@@ -474,7 +427,7 @@ async function inspectAndPatchPipelineConfig(page, {
pipeline_id: pipelineId,
pipeline_name: pipeline.name,
matched_by: matchedBy,
reason: "Pipeline has no ai.runner.id.",
reason: "Pipeline has no ai.runner.id or legacy ai.runner.runner.",
};
}
if (expectedRunnerId && runnerId !== expectedRunnerId) {
@@ -498,31 +451,6 @@ async function inspectAndPatchPipelineConfig(page, {
? runnerConfigs[runnerId]
: {};
const patchKeys = Object.keys(runnerConfigPatch || {});
const sensitiveConfigKeyRe = /(?:api[_-]?key|authorization|bearer|credential|jwt|oauth|password|secret)/i;
const sanitizeConfigValue = (key, value, depth = 0) => {
if (sensitiveConfigKeyRe.test(String(key))) return "[redacted]";
if (typeof value === "string") {
return value.length > 1200 ? `${value.slice(0, 1200)}...[truncated ${value.length - 1200} chars]` : value;
}
if (Array.isArray(value)) {
const items = value.slice(0, 50).map((item) => sanitizeConfigValue(key, item, depth + 1));
if (value.length > 50) items.push(`[truncated ${value.length - 50} items]`);
return items;
}
if (value && typeof value === "object") {
if (depth >= 3) return "[object truncated]";
return Object.fromEntries(
Object.entries(value).map(([childKey, childValue]) => [
childKey,
sanitizeConfigValue(childKey, childValue, depth + 1),
]),
);
}
return value;
};
const pickPatchValues = (config) => Object.fromEntries(
patchKeys.map((key) => [key, sanitizeConfigValue(key, config?.[key])]),
);
const baseDiagnostic = {
status: "ready",
authenticated: true,
@@ -534,7 +462,6 @@ async function inspectAndPatchPipelineConfig(page, {
expected_runner_id: expectedRunnerId || "",
patch_keys: patchKeys,
runner_config_before_keys: Object.keys(currentRunnerConfig),
runner_config_patch_before: pickPatchValues(currentRunnerConfig),
patched: patchKeys.length > 0,
};
@@ -579,7 +506,6 @@ async function inspectAndPatchPipelineConfig(page, {
put_status: update.status,
put_code: update.json.code ?? null,
runner_config_after_keys: Object.keys(updatedRunnerConfig),
runner_config_patch_after: pickPatchValues(updatedRunnerConfig),
restore_config: config,
};
}, {
@@ -622,201 +548,6 @@ async function restorePipelineConfig(page, { backendUrl, pipelineId, config }) {
}, { backendUrl, pipelineId, config });
}
async function inspectAndPatchPipelineExtensions(page, {
backendUrl,
pipelineId,
extensionsPatch,
}) {
return await page.evaluate(async ({
backendUrl,
pipelineId,
extensionsPatch,
}) => {
const token = localStorage.getItem("token");
if (!token) {
return {
status: "blocked",
authenticated: false,
pipeline_id: pipelineId,
reason: "Browser profile has no localStorage token.",
};
}
const headers = {
Authorization: `Bearer ${token}`,
"Content-Type": "application/json",
};
const getJson = async (path) => {
const response = await fetch(`${backendUrl}${path}`, { headers });
return {
status: response.status,
json: await response.json().catch(() => ({})),
};
};
const putJson = async (path, body) => {
const response = await fetch(`${backendUrl}${path}`, {
method: "PUT",
headers,
body: JSON.stringify(body),
});
return {
status: response.status,
json: await response.json().catch(() => ({})),
};
};
if (!pipelineId) {
return {
status: "blocked",
authenticated: true,
pipeline_resolved: false,
reason: "Pipeline id is required before patching extensions.",
};
}
const before = await getJson(`/api/v1/pipelines/${encodeURIComponent(pipelineId)}/extensions`);
let extensions = before.json.data || {};
let getStatus = before.status;
let getCode = before.json.code ?? null;
let fallbackReason = "";
if (before.status >= 400 || before.json.code !== 0) {
const pipelineResponse = await getJson(`/api/v1/pipelines/${encodeURIComponent(pipelineId)}`);
const pipeline = pipelineResponse.json.data?.pipeline || {};
const prefs = pipeline.extensions_preferences || {};
if (pipelineResponse.status >= 400 || pipelineResponse.json.code !== 0 || !pipeline.uuid) {
return {
status: "fail",
authenticated: true,
pipeline_id: pipelineId,
get_status: before.status,
get_code: before.json.code ?? null,
fallback_pipeline_status: pipelineResponse.status,
fallback_pipeline_code: pipelineResponse.json.code ?? null,
reason: before.json.msg || "Could not load pipeline extensions.",
};
}
fallbackReason = before.json.msg || "Could not load pipeline extensions; restored from pipeline preferences.";
extensions = {
enable_all_plugins: prefs.enable_all_plugins ?? true,
enable_all_mcp_servers: prefs.enable_all_mcp_servers ?? true,
enable_all_skills: prefs.enable_all_skills ?? true,
bound_plugins: prefs.plugins || [],
bound_mcp_servers: prefs.mcp_servers || [],
bound_skills: prefs.skills || [],
};
}
const patchKeys = Object.keys(extensionsPatch || {});
const restoreExtensions = {
enable_all_plugins: extensions.enable_all_plugins ?? true,
enable_all_mcp_servers: extensions.enable_all_mcp_servers ?? true,
enable_all_skills: extensions.enable_all_skills ?? true,
bound_plugins: extensions.bound_plugins || [],
bound_mcp_servers: extensions.bound_mcp_servers || [],
bound_skills: extensions.bound_skills || [],
};
const baseDiagnostic = {
status: "ready",
authenticated: true,
pipeline_id: pipelineId,
patch_keys: patchKeys,
patched: patchKeys.length > 0,
get_status: getStatus,
get_code: getCode,
fallback_reason: fallbackReason,
extensions_before: {
enable_all_plugins: restoreExtensions.enable_all_plugins,
enable_all_mcp_servers: restoreExtensions.enable_all_mcp_servers,
enable_all_skills: restoreExtensions.enable_all_skills,
bound_plugins: restoreExtensions.bound_plugins,
bound_mcp_servers: restoreExtensions.bound_mcp_servers,
bound_skills: restoreExtensions.bound_skills,
},
};
if (patchKeys.length === 0) {
return baseDiagnostic;
}
const updateBody = {
enable_all_plugins: Object.prototype.hasOwnProperty.call(extensionsPatch, "enable_all_plugins")
? Boolean(extensionsPatch.enable_all_plugins)
: restoreExtensions.enable_all_plugins,
enable_all_mcp_servers: Object.prototype.hasOwnProperty.call(extensionsPatch, "enable_all_mcp_servers")
? Boolean(extensionsPatch.enable_all_mcp_servers)
: restoreExtensions.enable_all_mcp_servers,
enable_all_skills: Object.prototype.hasOwnProperty.call(extensionsPatch, "enable_all_skills")
? Boolean(extensionsPatch.enable_all_skills)
: restoreExtensions.enable_all_skills,
bound_plugins: Object.prototype.hasOwnProperty.call(extensionsPatch, "bound_plugins")
? extensionsPatch.bound_plugins
: restoreExtensions.bound_plugins,
bound_mcp_servers: Object.prototype.hasOwnProperty.call(extensionsPatch, "bound_mcp_servers")
? extensionsPatch.bound_mcp_servers
: restoreExtensions.bound_mcp_servers,
bound_skills: Object.prototype.hasOwnProperty.call(extensionsPatch, "bound_skills")
? extensionsPatch.bound_skills
: restoreExtensions.bound_skills,
};
const update = await putJson(`/api/v1/pipelines/${encodeURIComponent(pipelineId)}/extensions`, updateBody);
if (update.status >= 400 || update.json.code !== 0) {
return {
...baseDiagnostic,
status: "fail",
put_status: update.status,
put_code: update.json.code ?? null,
reason: update.json.msg || "Pipeline extensions update failed.",
};
}
return {
...baseDiagnostic,
put_status: update.status,
put_code: update.json.code ?? null,
extensions_after: updateBody,
restore_extensions: restoreExtensions,
};
}, {
backendUrl,
pipelineId,
extensionsPatch,
});
}
async function restorePipelineExtensions(page, { backendUrl, pipelineId, extensions }) {
return await page.evaluate(async ({ backendUrl, pipelineId, extensions }) => {
const token = localStorage.getItem("token");
if (!token) {
return {
status: "blocked",
authenticated: false,
pipeline_id: pipelineId,
reason: "Browser profile has no localStorage token.",
};
}
const response = await fetch(`${backendUrl}/api/v1/pipelines/${encodeURIComponent(pipelineId)}/extensions`, {
method: "PUT",
headers: {
Authorization: `Bearer ${token}`,
"Content-Type": "application/json",
},
body: JSON.stringify(extensions),
});
const json = await response.json().catch(() => ({}));
return {
status: response.status >= 400 || json.code !== 0 ? "fail" : "ready",
authenticated: true,
pipeline_id: pipelineId,
put_status: response.status,
put_code: json.code ?? null,
reason: response.status >= 400 || json.code !== 0
? json.msg || "Pipeline extensions restore failed."
: "Pipeline extensions restored.",
};
}, { backendUrl, pipelineId, extensions });
}
async function resetPipelineDebugChat(page, { backendUrl, pipelineId, sessionType }) {
return await page.evaluate(async ({ backendUrl, pipelineId, sessionType }) => {
const token = localStorage.getItem("token");
@@ -859,13 +590,11 @@ try {
const promptSteps = promptStepsFromEnv();
const filesystemChecks = parseJsonEnv("LANGBOT_E2E_FILESYSTEM_CHECKS_JSON", []);
const runnerConfigPatch = parseJsonEnv("LANGBOT_E2E_RUNNER_CONFIG_PATCH_JSON", {});
const extensionsPatch = parseJsonEnv("LANGBOT_E2E_EXTENSIONS_PATCH_JSON", {});
const runnerPatchKeys = Object.keys(runnerConfigPatch);
const extensionsPatchKeys = Object.keys(extensionsPatch);
if (runnerPatchKeys.length > 0 || extensionsPatchKeys.length > 0 || resetDebugChat || expectedRunnerId) {
if (runnerPatchKeys.length > 0 || resetDebugChat || expectedRunnerId) {
if (!backendUrl) {
result.status = "env_issue";
result.reason = "LANGBOT_BACKEND_URL is required for runner config patch, extensions patch, runner assertion, or Debug Chat reset.";
result.reason = "LANGBOT_BACKEND_URL is required for runner config patch, runner assertion, or Debug Chat reset.";
throw new Error(result.reason);
}
}
@@ -873,30 +602,13 @@ try {
result.prompt = promptSteps.length === 1 ? promptSteps[0].prompt : `${promptSteps.length} prompts`;
result.expected_text = promptSteps.at(-1)?.expectedText || expectedText;
const authDiagnostic = await ensureAuthenticatedBrowser(page, {
frontendUrl: pipelineUrl || env.LANGBOT_FRONTEND_URL || "",
backendUrl,
const openResult = await openPipelineDebugChat(page, {
pipelineUrl,
pipelineName,
envHint: pipelineRequired
? "case-specific pipeline env mapped to LANGBOT_E2E_PIPELINE_URL or LANGBOT_E2E_PIPELINE_NAME"
: "LANGBOT_PIPELINE_URL or LANGBOT_PIPELINE_NAME",
});
result.browser_auth = authDiagnostic;
if (!result.evidence_collected.includes("api_diagnostic")) result.evidence_collected.push("api_diagnostic");
const authFailed = authDiagnostic.status === "env_issue" || authDiagnostic.status === "blocked" || authDiagnostic.status === "fail";
if (authFailed) {
result.status = authDiagnostic.status;
result.reason = authDiagnostic.reason || "Browser authentication failed.";
} else {
result.status = "running";
result.reason = "";
}
const openResult = authFailed
? { opened: false, status: result.status, reason: result.reason }
: await openPipelineDebugChat(page, {
pipelineUrl,
pipelineName,
envHint: pipelineRequired
? "case-specific pipeline env mapped to LANGBOT_E2E_PIPELINE_URL or LANGBOT_E2E_PIPELINE_NAME"
: "LANGBOT_PIPELINE_URL or LANGBOT_PIPELINE_NAME",
});
result.url = page.url();
if (!openResult.opened) {
@@ -905,7 +617,7 @@ try {
} else {
result.status = "running";
result.reason = "";
if (runnerPatchKeys.length > 0 || extensionsPatchKeys.length > 0 || resetDebugChat || expectedRunnerId) {
if (runnerPatchKeys.length > 0 || resetDebugChat || expectedRunnerId) {
const pipelineDiagnostic = await inspectAndPatchPipelineConfig(page, {
backendUrl,
pipelineUrl,
@@ -924,35 +636,13 @@ try {
result.reason = pipelineDiagnostic.reason || "Pipeline config preparation failed.";
} else {
if (pipelineDiagnostic.restore_config && restoreRunnerConfig) {
restoreConfigPlan = {
restorePlan = {
backendUrl,
pipelineId: pipelineDiagnostic.pipeline_id,
config: pipelineDiagnostic.restore_config,
};
}
if (extensionsPatchKeys.length > 0) {
const extensionsDiagnostic = await inspectAndPatchPipelineExtensions(page, {
backendUrl,
pipelineId: pipelineDiagnostic.pipeline_id,
extensionsPatch,
});
const safeExtensionsDiagnostic = sanitizePipelineExtensionsDiagnostic(extensionsDiagnostic);
await writeFile(pipelineExtensionsDiagnosticPath, `${JSON.stringify(safeExtensionsDiagnostic, null, 2)}\n`, "utf8");
result.evidence.pipeline_extensions_diagnostic_json = pipelineExtensionsDiagnosticPath;
result.pipeline_extensions = safeExtensionsDiagnostic;
if (!result.evidence_collected.includes("api_diagnostic")) result.evidence_collected.push("api_diagnostic");
if (extensionsDiagnostic.status === "fail" || extensionsDiagnostic.status === "blocked") {
result.status = extensionsDiagnostic.status;
result.reason = extensionsDiagnostic.reason || "Pipeline extensions preparation failed.";
} else if (extensionsDiagnostic.restore_extensions && restoreExtensions) {
restoreExtensionsPlan = {
backendUrl,
pipelineId: extensionsDiagnostic.pipeline_id,
extensions: extensionsDiagnostic.restore_extensions,
};
}
}
if (!["fail", "blocked", "env_issue"].includes(result.status) && resetDebugChat) {
if (resetDebugChat) {
const resetDiagnostic = await resetPipelineDebugChat(page, {
backendUrl,
pipelineId: pipelineDiagnostic.pipeline_id,
@@ -997,22 +687,14 @@ try {
const chatResult = await runDebugChatPrompt(page, {
prompt: step.prompt,
expectedText: step.expectedText,
expectedTexts: step.expectedTexts,
responseTimeoutMs: step.responseTimeoutMs,
imagePath: index === 0 ? imagePath : "",
backendUrl,
pipelineId: result.pipeline_config?.pipeline_id || pipelineIdFromUrl(pipelineUrl),
sessionType: debugChatSessionType,
maxNewAssistantMessages: Number.isFinite(maxNewAssistantMessages) && maxNewAssistantMessages >= 0
? maxNewAssistantMessages
: 1,
failureSignals: failureSignals.length > 0 ? failureSignals : undefined,
});
const promptDurationMs = Date.now() - promptStartedAt;
result.chat_results.push({
index,
expected_text: step.expectedText,
expected_texts: step.expectedTexts,
status: chatResult.status,
reason: chatResult.reason,
response_duration_ms: promptDurationMs,
@@ -1020,14 +702,7 @@ try {
final_count: chatResult.final_count,
before_assistant_expected_count: chatResult.before_assistant_expected_count,
after_assistant_expected_count: chatResult.after_assistant_expected_count,
before_assistant_message_count: chatResult.before_assistant_message_count,
after_assistant_message_count: chatResult.after_assistant_message_count,
new_assistant_message_count: chatResult.new_assistant_message_count,
latest_assistant_is_final: chatResult.latest_assistant_is_final,
final_assistant_wait_status: chatResult.final_assistant_wait_status,
final_assistant_wait_reason: chatResult.final_assistant_wait_reason,
failure_signal: chatResult.failure_signal || "",
missing_expected_texts: chatResult.missing_expected_texts || [],
});
result.status = chatResult.status;
result.reason = `Prompt ${index + 1}/${promptSteps.length}: ${chatResult.reason}`;
@@ -1045,15 +720,6 @@ try {
result.reason = filesystemResult.reason || "Filesystem checks failed.";
}
}
if (result.status === "pass" && !boolFromEnv(env.LANGBOT_E2E_ALLOW_BROWSER_ERRORS, false)) {
const browserDiagnostics = await scanBrowserDiagnostics(paths);
result.browser_diagnostics = browserDiagnostics;
if (browserDiagnostics.status === "fail") {
result.status = "fail";
result.reason = browserDiagnostics.reason;
}
}
}
} catch (error) {
if (!["env_issue", "blocked", "fail", "pass"].includes(result.status) || !result.reason) {
@@ -1062,20 +728,10 @@ try {
result.reason = result.reason || error.message;
} finally {
if (browser?.page) await safeScreenshot(browser.page, paths.screenshot);
if (browser?.page && restoreExtensionsPlan) {
const restoreDiagnostic = await restorePipelineExtensions(browser.page, restoreExtensionsPlan).catch((error) => ({
if (browser?.page && restorePlan) {
const restoreDiagnostic = await restorePipelineConfig(browser.page, restorePlan).catch((error) => ({
status: "fail",
pipeline_id: restoreExtensionsPlan.pipelineId,
reason: error.message,
}));
await writeFile(pipelineExtensionsRestoreDiagnosticPath, `${JSON.stringify(restoreDiagnostic, null, 2)}\n`, "utf8");
result.evidence.pipeline_extensions_restore_diagnostic_json = pipelineExtensionsRestoreDiagnosticPath;
result.pipeline_extensions_restore = restoreDiagnostic;
}
if (browser?.page && restoreConfigPlan) {
const restoreDiagnostic = await restorePipelineConfig(browser.page, restoreConfigPlan).catch((error) => ({
status: "fail",
pipeline_id: restoreConfigPlan.pipelineId,
pipeline_id: restorePlan.pipelineId,
reason: error.message,
}));
await writeFile(pipelineConfigRestoreDiagnosticPath, `${JSON.stringify(restoreDiagnostic, null, 2)}\n`, "utf8");
@@ -1083,12 +739,6 @@ try {
result.pipeline_config_restore = restoreDiagnostic;
}
if (browser) await browser.close().catch(() => {});
const backendLog = await finishBackendLogCapture(backendLogCapture);
if (backendLog) {
result.evidence.backend_log = backendLog.path;
result.backend_log = backendLog;
if (!result.evidence_collected.includes("backend_log")) result.evidence_collected.push("backend_log");
}
const finishedAt = new Date();
result.finished_at = finishedAt.toISOString();
result.finished_at_local = localIsoWithOffset(finishedAt);
@@ -1,45 +0,0 @@
#!/usr/bin/env python3
from __future__ import annotations
import argparse
import sys
from pathlib import Path
import yaml
PROJECT_ROOT = Path(__file__).resolve().parents[3]
sys.path.insert(0, str(PROJECT_ROOT))
from tests.e2e.utils.config_factory import ( # noqa: E402
create_minimal_config,
create_test_directories,
)
def main() -> None:
parser = argparse.ArgumentParser()
parser.add_argument('--instance-root', required=True, type=Path)
parser.add_argument('--port', required=True, type=int)
parser.add_argument('--plugin-debug-port', required=True, type=int)
args = parser.parse_args()
config_path = create_minimal_config(args.instance_root, port=args.port)
create_test_directories(args.instance_root)
with config_path.open('r', encoding='utf-8') as file:
config = yaml.safe_load(file)
config['plugin']['enable'] = True
config['plugin']['enable_marketplace'] = True
config['plugin']['display_plugin_debug_url'] = (
f'ws://127.0.0.1:{args.plugin_debug_port}/plugin/debug/ws'
)
with config_path.open('w', encoding='utf-8') as file:
yaml.safe_dump(config, file, allow_unicode=True, sort_keys=False)
if __name__ == '__main__':
main()
@@ -1,62 +0,0 @@
#!/usr/bin/env node
import { cp, mkdir, rm } from "node:fs/promises";
import { dirname, resolve } from "node:path";
import { fileURLToPath } from "node:url";
import { env } from "node:process";
import { spawnSync } from "node:child_process";
import {
ensureEvidence,
evidencePaths,
loadEnvFiles,
writeResult,
} from "./lib/langbot-e2e.mjs";
await loadEnvFiles();
const paths = evidencePaths("reset-complex-agent-task");
await ensureEvidence(paths);
const scriptDir = dirname(fileURLToPath(import.meta.url));
const source = resolve(
scriptDir,
"../../skills/langbot-testing/fixtures/complex-agent-task/workspace",
);
const repo = env.LANGBOT_REPO || "";
const target = repo ? resolve(repo, "data/box/default/order-orchestrator") : "";
const result = {
source: "setup_automation",
case_id: "reset-complex-agent-task",
run_id: paths.runId,
status: "fail",
reason: "",
target,
initial_test_exit_code: null,
evidence_collected: ["filesystem"],
};
try {
if (!repo) throw new Error("LANGBOT_REPO is required.");
await rm(target, { recursive: true, force: true });
await mkdir(dirname(target), { recursive: true });
await cp(source, target, { recursive: true });
const baseline = spawnSync(
"python3",
["-m", "unittest", "discover", "-s", "tests", "-v"],
{ cwd: target, encoding: "utf8", timeout: 60_000 },
);
result.initial_test_exit_code = baseline.status;
result.initial_failure_preview = `${baseline.stdout || ""}\n${baseline.stderr || ""}`.slice(0, 4000);
if (baseline.error) throw baseline.error;
if (baseline.status === 0) throw new Error("Complex task baseline unexpectedly passes; the fixture must start failing.");
result.status = "pass";
result.reason = "Complex task workspace reset and failing baseline confirmed.";
} catch (error) {
result.status = /required|ENOENT/.test(error.message) ? "env_issue" : "fail";
result.reason = error.message;
}
await writeResult(paths, result);
console.log(JSON.stringify(result, null, 2));
process.exit(result.status === "pass" ? 0 : result.status === "env_issue" ? 2 : 1);
@@ -1,76 +0,0 @@
#!/usr/bin/env node
import { env } from "node:process";
import {
apiJson,
ensureEvidence,
evidencePaths,
loadEnvFiles,
resetAndAuthLocalUser,
writeResult,
} from "./lib/langbot-e2e.mjs";
const DEFAULT_LOCAL_PASSWORD = "LangBotE2ELocalPass!2026";
const caseId = "reset-sandbox-skill";
await loadEnvFiles();
const paths = evidencePaths(caseId);
await ensureEvidence(paths);
const nameArgument = process.argv.find((argument) => argument.startsWith("--name="));
const skillName = nameArgument?.slice("--name=".length).trim() || "";
const backendUrl = env.LANGBOT_BACKEND_URL || "";
const result = {
source: "setup_automation",
case_id: caseId,
run_id: paths.runId,
status: "fail",
reason: "",
skill_name: skillName,
existed: false,
deleted: false,
evidence_collected: ["api_diagnostic"],
};
try {
if (!backendUrl) throw new Error("LANGBOT_BACKEND_URL is not configured.");
if (!skillName || !/^[A-Za-z0-9_-]+$/.test(skillName)) {
throw new Error("--name must contain only letters, numbers, hyphens, or underscores.");
}
const user = env.LANGBOT_E2E_LOGIN_USER || "";
if (!user) throw new Error("LANGBOT_E2E_LOGIN_USER is required.");
const password = env.LANGBOT_E2E_LOGIN_PASSWORD || DEFAULT_LOCAL_PASSWORD;
const auth = await resetAndAuthLocalUser({ backendUrl, user, password });
const listed = await apiJson(backendUrl, "/api/v1/skills", { token: auth.token });
if (listed.status >= 400 || listed.json.code !== 0) {
throw new Error(listed.json.msg || `Skill listing failed with HTTP ${listed.status}.`);
}
result.existed = (listed.json.data?.skills || []).some((skill) => skill?.name === skillName);
if (result.existed) {
const deleted = await apiJson(
backendUrl,
`/api/v1/skills/${encodeURIComponent(skillName)}`,
{ method: "DELETE", token: auth.token },
);
if (deleted.status >= 400 || deleted.json.code !== 0) {
throw new Error(deleted.json.msg || `Skill deletion failed with HTTP ${deleted.status}.`);
}
result.deleted = true;
}
result.status = "pass";
result.reason = result.deleted
? `Removed stale sandbox skill ${skillName}.`
: `Sandbox skill ${skillName} was already absent.`;
} catch (error) {
result.status = /ECONNREFUSED|fetch failed|not configured|required/.test(error.message)
? "env_issue"
: "fail";
result.reason = error.message;
}
await writeResult(paths, result);
console.log(JSON.stringify(result, null, 2));
process.exit(result.status === "pass" ? 0 : result.status === "env_issue" ? 2 : 1);
@@ -1,355 +0,0 @@
#!/usr/bin/env node
import { spawn } from "node:child_process";
import { createServer } from "node:net";
import { dirname, join } from "node:path";
import { fileURLToPath } from "node:url";
import {
createBrowser,
ensureAuthenticatedBrowser,
ensureEvidence,
evidencePaths,
exitCode,
loadEnvFiles,
localIsoWithOffset,
pathExists,
resolveLangBotRepo,
safeScreenshot,
scanBrowserDiagnostics,
writeResult,
} from "./lib/langbot-e2e.mjs";
const caseId = "wizard-onebot-agent-runtime";
await loadEnvFiles();
const paths = evidencePaths(caseId);
await ensureEvidence(paths);
const startedAt = new Date();
const frontendUrl = process.env.LANGBOT_FRONTEND_URL || "";
const backendUrl = process.env.LANGBOT_BACKEND_URL || "";
let browser;
let botId = "";
let agentId = "";
let token = "";
const result = {
source: "automation",
case_id: caseId,
run_id: paths.runId,
started_at: startedAt.toISOString(),
started_at_local: localIsoWithOffset(startedAt),
finished_at: "",
finished_at_local: "",
status: "fail",
reason: "",
url: "",
bot_id: "",
agent_id: "",
visible_signals: [],
api: {},
runtime: {},
cleanup: {},
diagnostics: null,
evidence: {
screenshot: paths.screenshot,
console_log: paths.consoleLog,
network_log: paths.networkLog,
automation_result_json: paths.automationResultJson,
result_json: paths.resultJson,
},
evidence_collected: ["ui", "screenshot", "console", "api_diagnostic"],
};
async function getFreePort() {
const server = createServer();
await new Promise((resolve, reject) => {
server.once("error", reject);
server.listen(0, "127.0.0.1", resolve);
});
const address = server.address();
const port = typeof address === "object" && address ? address.port : 0;
await new Promise((resolve) => server.close(resolve));
if (!port) throw new Error("Could not allocate a temporary OneBot port.");
return port;
}
async function api(page, path, options = {}) {
return await page.evaluate(
async ({ baseUrl, path, options, token }) => {
const response = await fetch(`${baseUrl.replace(/\/$/, "")}${path}`, {
method: options.method || "GET",
headers: {
Authorization: `Bearer ${token}`,
"Content-Type": "application/json",
},
body:
options.body === undefined ? undefined : JSON.stringify(options.body),
});
return {
status: response.status,
json: await response.json().catch(() => ({})),
};
},
{ baseUrl: backendUrl, path, options, token },
);
}
async function runProbe(port) {
const repo = await resolveLangBotRepo();
const python = join(repo, ".venv", "bin", "python");
const script = join(
dirname(fileURLToPath(import.meta.url)),
"onebot-runtime-probe.py",
);
if (!(await pathExists(python))) {
throw new Error(`LangBot virtualenv Python is missing: ${python}`);
}
return await new Promise((resolve, reject) => {
const child = spawn(python, [script, "--port", String(port)], {
cwd: repo,
});
let stdout = "";
let stderr = "";
child.stdout.on("data", (chunk) => (stdout += chunk));
child.stderr.on("data", (chunk) => (stderr += chunk));
child.on("error", reject);
child.on("close", (code) => {
if (code === 0) {
resolve(JSON.parse(stdout.trim().split("\n").at(-1)));
} else {
reject(
new Error(
stderr.trim() || stdout.trim() || `OneBot probe exited ${code}`,
),
);
}
});
});
}
try {
if (!frontendUrl) throw new Error("LANGBOT_FRONTEND_URL is not configured.");
if (!backendUrl) throw new Error("LANGBOT_BACKEND_URL is not configured.");
const port = await getFreePort();
browser = await createBrowser(paths);
const { page } = browser;
await page.goto(frontendUrl, { waitUntil: "domcontentloaded" });
const auth = await ensureAuthenticatedBrowser(page, {
frontendUrl,
backendUrl,
});
if (auth.status !== "pass") {
result.status = auth.status;
throw new Error(auth.reason);
}
token = await page.evaluate(() => localStorage.getItem("token") || "");
if (!token) {
result.status = "blocked";
throw new Error("Authenticated browser token is unavailable.");
}
// Keep this case isolated from the instance's real onboarding state.
await page.route("**/api/v1/system/wizard/progress", async (route) => {
await route.fulfill({
status: 200,
contentType: "application/json",
body: JSON.stringify({ code: 0, msg: "ok", data: {} }),
});
});
await page.goto(`${frontendUrl.replace(/\/$/, "")}/wizard`, {
waitUntil: "domcontentloaded",
});
await page
.getByRole("button", {
name: /Welcome new members|欢迎新成员|新しいメンバーを歓迎/,
})
.click();
await page.getByText("OneBot v11", { exact: true }).last().click();
const botResponse = page.waitForResponse(
(response) =>
response.request().method() === "POST" &&
/\/api\/v1\/platform\/bots$/.test(new URL(response.url()).pathname),
);
await page
.getByRole("button", {
name: /Confirm, Create Bot|确定,创建机器人|確定、ボットを作成/,
})
.click();
const botPayload = await (await botResponse).json();
botId = botPayload.data?.uuid || "";
if (!botId) throw new Error("Wizard did not create a Bot.");
result.bot_id = botId;
result.api.create_bot = { code: botPayload.code ?? null };
await page
.getByText(/Configure Your Bot|配置机器人|ボットを設定/)
.first()
.waitFor({ timeout: 15_000 });
const portInput = page.getByRole("spinbutton").first();
await portInput.fill(String(port));
await portInput.press("Tab");
await page
.getByRole("button", {
name: /Save & Enable Bot|保存并启用|保存して有効化/,
})
.click();
await page
.getByText(
/Bot configuration saved and enabled|机器人配置已保存并启用|ボット設定が保存され、有効になりました/,
)
.waitFor({ timeout: 15_000 });
result.visible_signals.push("bot-created", "adapter-enabled");
await page.getByRole("button", { name: /Next|下一步|次へ/ }).click();
const localAgentTitle = page
.getByText(/^(Local Agent|本地 Agent)$/)
.first();
const localAgentCard = localAgentTitle.locator(
'xpath=ancestor::*[@data-slot="card"][1]',
);
await localAgentCard
.getByRole("button", {
name: /Use This Runner|使用此运行器|この Runner を使用/,
})
.click();
await page.waitForFunction(
() =>
Array.from(document.querySelectorAll("textarea")).some((element) => {
const value = element.value || "";
return [
"Welcome new group members",
"欢迎新群成员",
"新しいグループメンバー",
].some((expected) => value.includes(expected));
}),
null,
{ timeout: 15_000 },
);
result.visible_signals.push("scenario-prompt-visible");
await page
.getByText(/QA Deterministic Runner|QA 确定性 Runner/)
.first()
.click();
const agentResponse = page.waitForResponse(
(response) =>
response.request().method() === "POST" &&
/\/api\/v1\/agents$/.test(new URL(response.url()).pathname),
);
await page
.getByRole("button", {
name: /Create & Deploy|创建并部署|作成&デプロイ/,
})
.click();
const agentPayload = await (await agentResponse).json();
agentId = agentPayload.data?.uuid || "";
if (!agentId) throw new Error("Wizard did not create an Agent.");
result.agent_id = agentId;
result.api.create_agent = { code: agentPayload.code ?? null };
await page.getByText(/All Set!|一切就绪!/).waitFor({ timeout: 15_000 });
result.visible_signals.push("agent-created", "wizard-complete");
const savedBot = await api(page, `/api/v1/platform/bots/${botId}`);
const route = savedBot.json.data?.bot?.event_bindings?.[0];
result.api.saved_route = {
http_status: savedBot.status,
event_pattern: route?.event_pattern || null,
target_type: route?.target_type || null,
target_matches: route?.target_uuid === agentId,
};
if (
route?.event_pattern !== "group.member_joined" ||
route?.target_type !== "agent" ||
route?.target_uuid !== agentId
) {
throw new Error("Wizard did not persist the expected Agent route.");
}
result.visible_signals.push("route-persisted");
result.runtime.onebot = await runProbe(port);
let routeStatus;
for (let attempt = 0; attempt < 30; attempt += 1) {
routeStatus = await api(
page,
`/api/v1/platform/bots/${botId}/event-routes/status`,
);
if (routeStatus.json.data?.routes?.[0]?.last_status === "delivered") break;
await page.waitForTimeout(250);
}
const latest = routeStatus?.json?.data?.routes?.[0];
result.runtime.route_status = latest || null;
if (latest?.last_status !== "delivered") {
throw new Error(
`Runtime route did not reach delivered status: ${latest?.last_status || "none"}`,
);
}
result.visible_signals.push(
"platform-event-converted",
"agent-ran",
"reply-delivered",
);
const botUrl = `${frontendUrl.replace(/\/$/, "")}/home/bots?id=${encodeURIComponent(botId)}`;
await page.goto(botUrl, { waitUntil: "domcontentloaded" });
result.url = page.url();
const deliveredStatus = page
.getByText(/Delivered|已投递|配信済み/, { exact: true })
.first();
await deliveredStatus.waitFor({ timeout: 15_000 });
await deliveredStatus.scrollIntoViewIfNeeded();
await page.waitForTimeout(250);
result.visible_signals.push("delivered-status-visible");
await safeScreenshot(page, paths.screenshot);
result.diagnostics = await scanBrowserDiagnostics(paths);
if (result.diagnostics.status !== "pass") {
throw new Error(result.diagnostics.reason);
}
result.status = "pass";
result.reason =
"Quick Start created an enabled OneBot scenario that converted a platform event, ran an Agent, delivered its reply, and showed the final status.";
} catch (error) {
if (!["blocked", "env_issue"].includes(result.status)) result.status = "fail";
result.reason = result.reason || error.message;
if (browser?.page) await safeScreenshot(browser.page, paths.screenshot);
} finally {
if (browser?.page && token) {
if (botId) {
const deleted = await api(
browser.page,
`/api/v1/platform/bots/${encodeURIComponent(botId)}`,
{ method: "DELETE" },
).catch(() => ({ status: 0, json: {} }));
result.cleanup.bot_deleted =
deleted.status < 400 && deleted.json.code === 0;
}
if (agentId) {
const deleted = await api(
browser.page,
`/api/v1/agents/${encodeURIComponent(agentId)}`,
{ method: "DELETE" },
).catch(() => ({ status: 0, json: {} }));
result.cleanup.agent_deleted =
deleted.status < 400 && deleted.json.code === 0;
}
}
const cleanupFailed =
(botId && !result.cleanup.bot_deleted) ||
(agentId && !result.cleanup.agent_deleted);
if (cleanupFailed && result.status === "pass") {
result.status = "fail";
result.reason = "The temporary OneBot wizard resources were not deleted.";
}
if (browser) await browser.close().catch(() => {});
const finishedAt = new Date();
result.finished_at = finishedAt.toISOString();
result.finished_at_local = localIsoWithOffset(finishedAt);
await writeResult(paths, result);
console.log(JSON.stringify(result, null, 2));
}
process.exit(exitCode(result.status));
@@ -1,404 +0,0 @@
#!/usr/bin/env node
import {
apiJson,
createBrowser,
ensureAuthenticatedBrowser,
ensureEvidence,
evidencePaths,
exitCode,
loadEnvFiles,
localIsoWithOffset,
safeScreenshot,
scanBrowserDiagnostics,
writeResult,
} from "./lib/langbot-e2e.mjs";
import { startIsolatedLangBotInstance } from "./lib/isolated-langbot-instance.mjs";
const caseId = "wizard-runner-marketplace-catalog";
await loadEnvFiles();
const paths = evidencePaths(caseId);
await ensureEvidence(paths);
const mobileScreenshot = paths.screenshot.replace(/\.png$/, "-mobile.png");
const installedScreenshot = paths.screenshot.replace(/\.png$/, "-installed.png");
const startedAt = new Date();
let frontendUrl = "";
let backendUrl = "";
let browser;
let isolated;
let token = "";
let botId = "";
const result = {
source: "automation",
case_id: caseId,
run_id: paths.runId,
started_at: startedAt.toISOString(),
started_at_local: localIsoWithOffset(startedAt),
finished_at: "",
finished_at_local: "",
status: "fail",
reason: "",
url: "",
visible_signals: [],
api: {},
marketplace_request: null,
marketplace_response: null,
diagnostics: null,
cleanup: {},
evidence: {
screenshot: paths.screenshot,
mobile_screenshot: mobileScreenshot,
installed_screenshot: installedScreenshot,
console_log: paths.consoleLog,
network_log: paths.networkLog,
isolated_backend_log: `${paths.evidenceDir}/isolated-backend.log`,
isolated_frontend_log: `${paths.evidenceDir}/isolated-frontend.log`,
automation_result_json: paths.automationResultJson,
result_json: paths.resultJson,
},
evidence_collected: ["ui", "screenshot", "console", "api_diagnostic"],
};
try {
isolated = await startIsolatedLangBotInstance({
evidenceDir: paths.evidenceDir,
});
frontendUrl = isolated.frontendUrl;
backendUrl = isolated.backendUrl;
result.api.isolated_instance = {
backend_url: backendUrl,
frontend_url: frontendUrl,
};
browser = await createBrowser(paths);
const { page } = browser;
page.on("request", (request) => {
if (
!/\/api\/v1\/marketplace\/(extensions|plugins)\/search$/.test(
new URL(request.url()).pathname,
)
) {
return;
}
try {
const payload = request.postDataJSON();
if (payload?.component_filter === "AgentRunner") {
result.marketplace_request = {
endpoint: new URL(request.url()).pathname,
component_filter: payload.component_filter,
type_filter: payload.type_filter || null,
page_size: payload.page_size || null,
};
}
} catch {
// The assertion below reports a missing or malformed catalog request.
}
});
page.on("response", async (response) => {
const pathname = new URL(response.url()).pathname;
if (!/\/api\/v1\/marketplace\/(extensions|plugins)\/search$/.test(pathname)) {
return;
}
try {
const payload = await response.json();
const entries = payload?.data?.extensions || payload?.data?.plugins || [];
const localAgent = entries.find(
(entry) => `${entry.author}/${entry.name}` === "langbot-team/LocalAgent",
);
result.marketplace_response = {
endpoint: pathname,
total: payload?.data?.total ?? entries.length,
local_agent_present: Boolean(localAgent),
local_agent_version: localAgent?.latest_version || null,
};
} catch {
// The UI assertions below surface malformed Marketplace responses.
}
});
await page.goto(frontendUrl, { waitUntil: "domcontentloaded" });
const auth = await ensureAuthenticatedBrowser(page, {
frontendUrl,
backendUrl,
user: isolated.user,
password: isolated.password,
recoveryKey: isolated.recoveryKey,
});
if (auth.status !== "pass") {
result.status = auth.status;
throw new Error(auth.reason);
}
token = await page.evaluate(() => localStorage.getItem("token") || "");
if (!token) {
result.status = "blocked";
throw new Error("Authenticated browser token is unavailable.");
}
const [plugins, metadata, system] = await Promise.all([
apiJson(backendUrl, "/api/v1/plugins", { token }),
apiJson(backendUrl, "/api/v1/pipelines/_/metadata", { token }),
apiJson(backendUrl, "/api/v1/system/info", { token }),
]);
const runnerStage = metadata.json.data?.configs
?.find((config) => config.name === "ai")
?.stages?.find((stage) => stage.name === "runner");
const runnerOptions =
runnerStage?.config?.find((item) => item.name === "id")?.options || [];
const installedPlugins = plugins.json.data?.plugins || [];
result.api.clean_state = {
plugin_count: installedPlugins.length,
runner_count: runnerOptions.length,
wizard_status: system.json.data?.wizard_status || null,
};
if (installedPlugins.length !== 0 || runnerOptions.length !== 0) {
result.status = "blocked";
throw new Error(
`This case requires zero plugins and zero runners; found ${installedPlugins.length} plugins and ${runnerOptions.length} runners.`,
);
}
if (system.json.data?.wizard_status !== "none") {
result.status = "blocked";
throw new Error(
`This case requires a first-run instance; wizard status is ${system.json.data?.wizard_status}.`,
);
}
const suffix = paths.runId.slice(-40);
const bot = await apiJson(backendUrl, "/api/v1/platform/bots", {
method: "POST",
token,
body: {
name: `Runner Catalog Wizard ${suffix}`,
description: "Temporary clean-runner catalog fixture",
adapter: "aiocqhttp",
adapter_config: {
host: "127.0.0.1",
port: 2280,
"access-token": "",
},
enable: false,
event_bindings: [],
},
});
botId = bot.json.data?.uuid || "";
if (bot.status >= 400 || bot.json.code !== 0 || !botId) {
throw new Error(bot.json.msg || "Failed to create the temporary Bot.");
}
const progress = await apiJson(backendUrl, "/api/v1/system/wizard/progress", {
method: "PUT",
token,
body: {
step: 2,
selected_scenario: "message_reply",
selected_adapter: "aiocqhttp",
created_bot_uuid: botId,
bot_saved: true,
selected_runner: null,
},
});
if (progress.status >= 400 || progress.json.code !== 0) {
throw new Error(progress.json.msg || "Failed to prepare Wizard progress.");
}
await page.goto(`${frontendUrl.replace(/\/$/, "")}/wizard`, {
waitUntil: "domcontentloaded",
});
result.url = page.url();
await page
.getByText(/Select an AI Engine|选择 AI 引擎|AIエンジンを選択/, {
exact: true,
})
.waitFor({ timeout: 15_000 });
const localAgentIdentity = page.getByText("langbot-team/LocalAgent", {
exact: true,
});
await localAgentIdentity.waitFor({ timeout: 30_000 });
const localAgentCard = localAgentIdentity.locator(
"xpath=ancestor::*[@data-slot='card'][1]",
);
const installButton = localAgentCard.getByRole("button", {
name: /Install & Continue|安装并继续|インストールして続行/,
});
await installButton.waitFor();
const browseLink = page.getByRole("link", {
name: /Browse Runner Extensions|浏览运行器扩展|Runner 拡張機能を見る/,
});
await browseLink.waitFor();
const href = await browseLink.getAttribute("href");
if (href !== "/home/extensions?type=plugin&component=AgentRunner") {
throw new Error(`Unexpected Runner marketplace URL: ${href}`);
}
const nextButton = page.getByRole("button", {
name: /Create & Deploy|创建并部署|作成&デプロイ/,
});
if (!(await nextButton.isDisabled())) {
throw new Error("Wizard allowed continuing without an installed Runner.");
}
if (!result.marketplace_request) {
throw new Error(
"Wizard did not request the AgentRunner Marketplace catalog.",
);
}
if (
!result.marketplace_response?.local_agent_present ||
!result.marketplace_response?.local_agent_version
) {
throw new Error(
"The Marketplace catalog did not return an installable langbot-team/LocalAgent version.",
);
}
result.visible_signals.push(
"first-run-ai-engine-step",
"published-local-agent-runner",
"install-and-continue-action",
"runner-marketplace-link",
"next-disabled-without-runner",
);
await safeScreenshot(page, paths.screenshot);
await page.setViewportSize({ width: 390, height: 844 });
await page.waitForTimeout(250);
const overflow = await page.evaluate(
() => document.documentElement.scrollWidth - window.innerWidth,
);
if (overflow > 1) {
throw new Error(`Runner catalog Wizard overflows mobile by ${overflow}px.`);
}
await browseLink.scrollIntoViewIfNeeded();
await safeScreenshot(page, mobileScreenshot);
result.visible_signals.push("mobile-layout");
await page.setViewportSize({ width: 1440, height: 1000 });
await installButton.click();
const installingButton = localAgentCard.getByRole("button", {
name: /Installing\.\.\.|正在安装\.\.\.|インストール中\.\.\./,
});
await installingButton.waitFor({ timeout: 15_000 });
const installOutcome = await Promise.race([
nextButton
.click({ trial: true, timeout: 180_000 })
.then(() => "registered"),
installButton
.waitFor({ state: "visible", timeout: 180_000 })
.then(() => "failed"),
]);
if (installOutcome === "failed") {
throw new Error("LocalAgent installation failed before Runner registration.");
}
if (await nextButton.isDisabled()) {
throw new Error("Create & Deploy remained disabled after LocalAgent installation.");
}
const [installedPluginsResponse, installedMetadataResponse] =
await Promise.all([
apiJson(backendUrl, "/api/v1/plugins", { token }),
apiJson(backendUrl, "/api/v1/pipelines/_/metadata", { token }),
]);
const postInstallPlugins =
installedPluginsResponse.json.data?.plugins || [];
const installedRunnerStage = installedMetadataResponse.json.data?.configs
?.find((config) => config.name === "ai")
?.stages?.find((stage) => stage.name === "runner");
const installedRunnerOptions =
installedRunnerStage?.config?.find((item) => item.name === "id")?.options ||
[];
const localAgentInstalled = postInstallPlugins.some((plugin) => {
const metadata = plugin.manifest?.manifest?.metadata || {};
return `${metadata.author}/${metadata.name}` === "langbot-team/LocalAgent";
});
const localAgentRegistered = installedRunnerOptions.some(
(option) => option.name === "plugin:langbot-team/LocalAgent/default",
);
result.api.installed_state = {
plugin_count: postInstallPlugins.length,
runner_count: installedRunnerOptions.length,
local_agent_installed: localAgentInstalled,
local_agent_registered: localAgentRegistered,
};
if (!localAgentInstalled || !localAgentRegistered) {
throw new Error(
"LocalAgent installation completed in the UI but the plugin or Runner registration is missing.",
);
}
await safeScreenshot(page, installedScreenshot);
result.visible_signals.push(
"local-agent-installed",
"local-agent-runner-registered",
"local-agent-selected",
"create-and-deploy-enabled",
);
result.diagnostics = await scanBrowserDiagnostics(paths);
if (result.diagnostics.status !== "pass") {
throw new Error(result.diagnostics.reason);
}
result.status = "pass";
result.reason =
"A clean first-run instance discovered LocalAgent in the AgentRunner catalog, installed and registered it, selected it, and enabled Create & Deploy.";
} catch (error) {
if (!["blocked", "env_issue"].includes(result.status)) result.status = "fail";
result.reason = result.reason || error.message;
if (browser?.page) await safeScreenshot(browser.page, paths.screenshot);
} finally {
const cleanup = {};
if (token && backendUrl) {
const resetProgress = await apiJson(
backendUrl,
"/api/v1/system/wizard/progress",
{
method: "PUT",
token,
body: {
step: 0,
selected_scenario: null,
selected_adapter: null,
created_bot_uuid: null,
bot_saved: false,
selected_runner: null,
},
},
).catch(() => ({ status: 0, json: {} }));
cleanup.progress_reset =
resetProgress.status < 400 && resetProgress.json.code === 0;
}
if (botId && token && backendUrl) {
const deletedBot = await apiJson(
backendUrl,
`/api/v1/platform/bots/${encodeURIComponent(botId)}`,
{ method: "DELETE", token },
).catch(() => ({ status: 0, json: {} }));
cleanup.bot_deleted = deletedBot.status < 400 && deletedBot.json.code === 0;
}
result.cleanup = cleanup;
if (
result.status === "pass" &&
(!cleanup.progress_reset || (botId && !cleanup.bot_deleted))
) {
result.status = "fail";
result.reason = "The clean Wizard fixtures were not fully reset.";
}
if (browser) await browser.close().catch(() => {});
if (isolated) {
try {
await isolated.stop();
cleanup.isolated_instance_stopped = true;
} catch (error) {
cleanup.isolated_instance_stopped = false;
if (result.status === "pass") {
result.status = "fail";
result.reason = `Could not stop the isolated LangBot instance: ${error.message}`;
}
}
}
const finishedAt = new Date();
result.finished_at = finishedAt.toISOString();
result.finished_at_local = localIsoWithOffset(finishedAt);
await writeResult(paths, result);
console.log(JSON.stringify(result, null, 2));
}
process.exit(exitCode(result.status));
@@ -1,349 +0,0 @@
#!/usr/bin/env node
import {
apiJson,
createBrowser,
ensureAuthenticatedBrowser,
ensureEvidence,
evidencePaths,
exitCode,
loadEnvFiles,
localIsoWithOffset,
safeScreenshot,
scanBrowserDiagnostics,
writeResult,
} from "./lib/langbot-e2e.mjs";
const caseId = "wizard-scenario-routing";
await loadEnvFiles();
const paths = evidencePaths(caseId);
await ensureEvidence(paths);
const messageScreenshot = paths.screenshot.replace(
/\.png$/,
"-message-scenario.png",
);
const mobileScreenshot = paths.screenshot.replace(/\.png$/, "-mobile.png");
const startedAt = new Date();
const frontendUrl = process.env.LANGBOT_FRONTEND_URL || "";
const backendUrl = process.env.LANGBOT_BACKEND_URL || "";
let browser;
let token = "";
let botId = "";
let agentId = "";
const result = {
source: "automation",
case_id: caseId,
run_id: paths.runId,
started_at: startedAt.toISOString(),
started_at_local: localIsoWithOffset(startedAt),
finished_at: "",
finished_at_local: "",
status: "fail",
reason: "",
url: "",
visible_signals: [],
api: {},
diagnostics: null,
cleanup: null,
evidence: {
console_log: paths.consoleLog,
network_log: paths.networkLog,
screenshot: paths.screenshot,
message_scenario_screenshot: messageScreenshot,
mobile_screenshot: mobileScreenshot,
automation_result_json: paths.automationResultJson,
result_json: paths.resultJson,
},
evidence_collected: ["ui", "screenshot", "console", "api_diagnostic"],
};
try {
if (!frontendUrl) throw new Error("LANGBOT_FRONTEND_URL is not configured.");
if (!backendUrl) throw new Error("LANGBOT_BACKEND_URL is not configured.");
browser = await createBrowser(paths);
const { page } = browser;
await page.goto(frontendUrl, { waitUntil: "domcontentloaded" });
const auth = await ensureAuthenticatedBrowser(page, {
frontendUrl,
backendUrl,
});
if (auth.status !== "pass") {
result.status = auth.status;
throw new Error(auth.reason);
}
token = await page.evaluate(() => localStorage.getItem("token") || "");
if (!token) {
result.status = "blocked";
throw new Error("Authenticated browser has no reusable local token.");
}
const adapters = await apiJson(backendUrl, "/api/v1/platform/adapters", {
token,
});
const catalog = adapters.json.data?.adapters || [];
const httpBot = catalog.find((adapter) => adapter.name === "http_bot");
const welcomeAdapters = catalog.filter((adapter) =>
(adapter.spec?.supported_events || []).includes("group.member_joined"),
);
result.api.adapters = {
http_status: adapters.status,
code: adapters.json.code ?? null,
http_bot_supports_message:
!httpBot?.spec?.supported_events?.length ||
httpBot.spec.supported_events.includes("message.received"),
http_bot_supports_member_joined:
httpBot?.spec?.supported_events?.includes("group.member_joined") || false,
welcome_adapter_count: welcomeAdapters.length,
};
if (adapters.status >= 400 || adapters.json.code !== 0) {
throw new Error(adapters.json.msg || "Adapter discovery failed.");
}
if (!httpBot || welcomeAdapters.length === 0) {
throw new Error(
"The adapter catalog does not contain the fixtures required by this case.",
);
}
if (result.api.adapters.http_bot_supports_member_joined) {
throw new Error(
"HTTP Bot unexpectedly declares group.member_joined support; update the case expectation.",
);
}
const progressUrl = `${backendUrl.replace(/\/$/, "")}/api/v1/system/wizard/progress`;
await page.route(progressUrl, async (route) => {
await route.fulfill({
status: 200,
contentType: "application/json",
body: JSON.stringify({ code: 0, msg: "ok", data: {} }),
});
});
await page.goto(`${frontendUrl.replace(/\/$/, "")}/wizard`, {
waitUntil: "domcontentloaded",
});
result.url = page.url();
await page
.getByText(
/What should this bot do\?|这个机器人要完成什么?|このボットで何を実現しますか?/,
)
.waitFor({ timeout: 15_000 });
await page
.getByRole("button", {
name: /Reply to messages|回复收到的消息|受信メッセージに返信/,
})
.click();
await page
.getByText(/HTTP Bot|HTTP 通用接入|HTTP ボット/, { exact: true })
.waitFor();
result.visible_signals.push("scenario-first", "message-http-compatible");
await safeScreenshot(page, messageScreenshot);
await page
.getByRole("button", {
name: /Welcome new members|欢迎新成员|新しいメンバーを歓迎/,
})
.click();
await page
.getByText(/Discord/, { exact: true })
.first()
.waitFor();
const httpBotCount = await page
.getByText(/HTTP Bot|HTTP 通用接入|HTTP ボット/, { exact: true })
.count();
if (httpBotCount !== 0) {
throw new Error("HTTP Bot remained visible for group.member_joined.");
}
const selectedScenario = page.getByRole("button", {
name: /Welcome new members|欢迎新成员|新しいメンバーを歓迎/,
});
if (
!(await selectedScenario.getByText("Agent", { exact: true }).isVisible())
) {
throw new Error(
"The non-message scenario is not visibly labeled as Agent.",
);
}
result.visible_signals.push(
"welcome-compatible-channel",
"http-filtered",
"agent-behavior-label",
);
await safeScreenshot(page, paths.screenshot);
await page.setViewportSize({ width: 390, height: 844 });
await page.waitForTimeout(250);
const horizontalOverflow = await page.evaluate(
() => document.documentElement.scrollWidth - window.innerWidth,
);
if (horizontalOverflow > 1) {
throw new Error(
`The mobile wizard overflows horizontally by ${horizontalOverflow}px.`,
);
}
await selectedScenario.scrollIntoViewIfNeeded();
await page
.getByText(/Discord/, { exact: true })
.first()
.waitFor();
await safeScreenshot(page, mobileScreenshot);
result.visible_signals.push("mobile-layout");
const fixtureSuffix = paths.runId.slice(-48);
const agent = await apiJson(backendUrl, "/api/v1/agents", {
method: "POST",
token,
body: {
kind: "agent",
name: `Wizard Scenario Agent ${fixtureSuffix}`,
description: "Temporary wizard scenario routing fixture",
emoji: "W",
enabled: true,
supported_event_patterns: ["group.member_joined"],
},
});
agentId = agent.json.data?.uuid || "";
result.api.create_agent = {
http_status: agent.status,
code: agent.json.code ?? null,
};
if (agent.status >= 400 || agent.json.code !== 0 || !agentId) {
throw new Error(agent.json.msg || "Failed to create the temporary Agent.");
}
const bot = await apiJson(backendUrl, "/api/v1/platform/bots", {
method: "POST",
token,
body: {
name: `Wizard Scenario Bot ${fixtureSuffix}`,
description: "Temporary disabled wizard scenario routing fixture",
adapter: "aiocqhttp",
adapter_config: {
host: "127.0.0.1",
port: 2280,
"access-token": "",
},
enable: false,
event_bindings: [],
},
});
botId = bot.json.data?.uuid || "";
result.api.create_bot = {
http_status: bot.status,
code: bot.json.code ?? null,
};
if (bot.status >= 400 || bot.json.code !== 0 || !botId) {
throw new Error(bot.json.msg || "Failed to create the temporary Bot.");
}
const update = await apiJson(
backendUrl,
`/api/v1/platform/bots/${encodeURIComponent(botId)}`,
{
method: "PUT",
token,
body: {
event_bindings: [
{
event_pattern: "group.member_joined",
target_type: "agent",
target_uuid: agentId,
filters: [],
priority: 0,
enabled: true,
description: "Welcome new members",
},
],
},
},
);
result.api.bind_route = {
http_status: update.status,
code: update.json.code ?? null,
};
if (update.status >= 400 || update.json.code !== 0) {
throw new Error(update.json.msg || "Failed to bind the temporary route.");
}
const savedBot = await apiJson(
backendUrl,
`/api/v1/platform/bots/${encodeURIComponent(botId)}`,
{ token },
);
const savedRoute = savedBot.json.data?.bot?.event_bindings?.[0];
result.api.read_route = {
http_status: savedBot.status,
code: savedBot.json.code ?? null,
event_pattern: savedRoute?.event_pattern || null,
target_type: savedRoute?.target_type || null,
target_matches: savedRoute?.target_uuid === agentId,
};
if (
savedBot.status >= 400 ||
savedBot.json.code !== 0 ||
savedRoute?.event_pattern !== "group.member_joined" ||
savedRoute?.target_type !== "agent" ||
savedRoute?.target_uuid !== agentId
) {
throw new Error("The saved Bot did not return the expected Agent route.");
}
result.visible_signals.push("agent-route-persisted");
result.diagnostics = await scanBrowserDiagnostics(paths);
if (result.diagnostics.status !== "pass") {
throw new Error(result.diagnostics.reason);
}
result.status = "pass";
result.reason =
"Quick Start visibly filtered channels by scenario at desktop and mobile widths.";
} catch (error) {
if (!["blocked", "env_issue"].includes(result.status)) result.status = "fail";
result.reason = result.reason || error.message;
if (browser?.page) await safeScreenshot(browser.page, paths.screenshot);
} finally {
const cleanup = {};
if (botId && token && backendUrl) {
const deletedBot = await apiJson(
backendUrl,
`/api/v1/platform/bots/${encodeURIComponent(botId)}`,
{ method: "DELETE", token },
).catch((error) => ({
status: 0,
json: { code: null, msg: error.message },
}));
cleanup.bot_deleted = deletedBot.status < 400 && deletedBot.json.code === 0;
cleanup.bot_http_status = deletedBot.status;
}
if (agentId && token && backendUrl) {
const deletedAgent = await apiJson(
backendUrl,
`/api/v1/agents/${encodeURIComponent(agentId)}`,
{ method: "DELETE", token },
).catch((error) => ({
status: 0,
json: { code: null, msg: error.message },
}));
cleanup.agent_deleted =
deletedAgent.status < 400 && deletedAgent.json.code === 0;
cleanup.agent_http_status = deletedAgent.status;
}
result.cleanup = cleanup;
const cleanupFailed =
(botId && !cleanup.bot_deleted) || (agentId && !cleanup.agent_deleted);
if (cleanupFailed && result.status === "pass") {
result.status = "fail";
result.reason = "The temporary wizard routing fixtures were not deleted.";
}
if (browser) await browser.close().catch(() => {});
const finishedAt = new Date();
result.finished_at = finishedAt.toISOString();
result.finished_at_local = localIsoWithOffset(finishedAt);
await writeResult(paths, result);
console.log(JSON.stringify(result, null, 2));
}
process.exit(exitCode(result.status));
@@ -1,163 +0,0 @@
#!/usr/bin/env node
import { spawn } from "node:child_process";
import { access, readFile, writeFile } from "node:fs/promises";
import { delimiter, dirname, join, resolve } from "node:path";
import { env } from "node:process";
import {
ensureEvidence,
evidencePaths,
exitCode,
loadEnvFiles,
localIsoWithOffset,
resolveLangBotRepo,
writeResult,
} from "./lib/langbot-e2e.mjs";
await loadEnvFiles();
const caseId = env.LBS_CASE_ID || "workspace-compatibility-preflight";
const paths = evidencePaths(caseId);
await ensureEvidence(paths);
const startedAt = new Date();
const detailsPath = join(paths.evidenceDir, "workspace-preflight.json");
const result = {
source: "automation",
case_id: caseId,
run_id: paths.runId,
started_at: startedAt.toISOString(),
started_at_local: localIsoWithOffset(startedAt),
finished_at: "",
finished_at_local: "",
status: "fail",
reason: "",
checks: [],
warnings: [],
evidence: {
workspace_preflight_json: detailsPath,
automation_result_json: paths.automationResultJson,
result_json: paths.resultJson,
},
evidence_collected: ["filesystem", "api_diagnostic"],
};
const repositories = [
{ id: "langbot", directory: "LangBot", envKey: "LANGBOT_REPO", manifest: false },
{ id: "plugin-sdk", directory: "langbot-plugin-sdk", envKey: "LANGBOT_PLUGIN_SDK_REPO", manifest: false },
{ id: "agent-runner", directory: "langbot-agent-runner", envKey: "LANGBOT_AGENT_RUNNER_REPO", manifest: false },
{ id: "local-agent", directory: "langbot-local-agent", envKey: "LANGBOT_LOCAL_AGENT_REPO", identity: "langbot-team/LocalAgent" },
{ id: "control-plane", directory: "langbot-agent-control-plane", envKey: "LANGBOT_AGENT_CONTROL_PLANE_REPO", identity: "langbot/agent-control-plane" },
{ id: "longterm-memory", directory: "langbot-longterm-memory", envKey: "LANGBOT_LONGTERM_MEMORY_REPO", identity: "langbot-team/LongTermMemory" },
{ id: "parser", directory: "langbot-parser", envKey: "LANGBOT_PARSER_PLUGIN_REPO", identity: "langbot-team/GeneralParsers" },
{ id: "rag", directory: "langbot-rag", envKey: "LANGBOT_RAG_PLUGIN_REPO", identity: "langbot-team/LangRAG" },
{ id: "skill-authoring", directory: "langbot-skill-authoring", envKey: "LANGBOT_SKILL_AUTHORING_REPO", identity: "huanghuoguoguo/skill-authoring" },
];
function run(command, args, options = {}) {
return new Promise((resolvePromise) => {
const child = spawn(command, args, { ...options, encoding: "utf8", stdio: ["ignore", "pipe", "pipe"] });
let stdout = "";
let stderr = "";
child.stdout.on("data", (chunk) => { stdout += chunk; });
child.stderr.on("data", (chunk) => { stderr += chunk; });
child.on("error", (error) => resolvePromise({ status: null, stdout, stderr, error }));
child.on("close", (status) => resolvePromise({ status, stdout, stderr, error: null }));
});
}
function addCheck(name, status, detail = {}) {
result.checks.push({ name, status, ...detail });
}
try {
const langbotRepo = await resolveLangBotRepo();
const workspaceRoot = resolve(env.LANGBOT_WORKSPACE_ROOT || dirname(langbotRepo));
const resolved = {};
for (const repository of repositories) {
const path = resolve(env[repository.envKey] || (repository.id === "langbot" ? langbotRepo : join(workspaceRoot, repository.directory)));
resolved[repository.id] = path;
try {
await access(join(path, ".git"));
} catch {
addCheck(`repo:${repository.id}`, "fail", { path, reason: "Git checkout is missing." });
continue;
}
const branch = await run("git", ["branch", "--show-current"], { cwd: path });
const branchName = branch.stdout.trim();
const compatible = branch.status === 0 && /^(?:main|dev\/4\.11\.x)$/.test(branchName);
addCheck(`repo:${repository.id}`, compatible ? "pass" : "fail", {
path,
branch: branchName,
reason: compatible ? "" : "Expected main or dev/4.11.x compatibility branch.",
});
const dirty = await run("git", ["status", "--short"], { cwd: path });
if (dirty.stdout.trim()) {
result.warnings.push({ name: `dirty:${repository.id}`, path, entries: dirty.stdout.trim().split(/\r?\n/).length });
}
if (repository.identity) {
const manifestPath = join(path, "manifest.yaml");
try {
const manifest = await readFile(manifestPath, "utf8");
const author = manifest.match(/^\s{2}author:\s*([^\s#]+)/m)?.[1] || "";
const name = manifest.match(/^\s{2}name:\s*([^\s#]+)/m)?.[1] || "";
const identity = `${author}/${name}`;
addCheck(`manifest:${repository.id}`, identity === repository.identity ? "pass" : "fail", {
path: manifestPath,
identity,
expected_identity: repository.identity,
});
} catch (error) {
addCheck(`manifest:${repository.id}`, "fail", { path: manifestPath, reason: error.message });
}
}
}
const python = join(resolved.langbot, ".venv", "bin", "python");
try {
await access(python);
addCheck("langbot-venv", "pass", { python });
} catch {
addCheck("langbot-venv", "fail", { python, reason: "LangBot virtualenv Python is missing." });
}
const sdkSrc = join(resolved["plugin-sdk"], "src");
const importProbe = await run(python, ["-c", [
"import json, pathlib, langbot_plugin",
"from langbot_plugin.api.entities.builtin.agent_runner.input import AgentInput",
"from langbot_plugin.api.entities.builtin.agent_runner.result import AgentRunResult",
"print(json.dumps({'path': str(pathlib.Path(langbot_plugin.__file__).resolve()), 'entities': [AgentInput.__name__, AgentRunResult.__name__]}))",
].join("; ")], {
cwd: resolved.langbot,
env: { ...env, PYTHONPATH: [sdkSrc, env.PYTHONPATH].filter(Boolean).join(delimiter) },
});
let importDetail = {};
try { importDetail = JSON.parse(importProbe.stdout.trim()); } catch { importDetail = { stderr: importProbe.stderr.trim() }; }
const localSdkLoaded = importProbe.status === 0 && resolve(importDetail.path || "").startsWith(resolve(sdkSrc));
addCheck("local-sdk-import", localSdkLoaded ? "pass" : "fail", {
expected_root: resolve(sdkSrc),
...importDetail,
reason: localSdkLoaded ? "" : "langbot_plugin did not load from the workspace SDK source tree.",
});
const failures = result.checks.filter((check) => check.status === "fail");
result.status = failures.length === 0 ? "pass" : "fail";
result.reason = failures.length === 0
? `Workspace compatibility preflight passed with ${result.warnings.length} non-blocking dirty-worktree warning(s).`
: `Workspace compatibility preflight found ${failures.length} blocking check(s).`;
} catch (error) {
result.status = /missing|ENOENT|not found/i.test(error.message) ? "env_issue" : "fail";
result.reason = error.message;
} finally {
const finishedAt = new Date();
result.finished_at = finishedAt.toISOString();
result.finished_at_local = localIsoWithOffset(finishedAt);
await writeFile(detailsPath, `${JSON.stringify({ checks: result.checks, warnings: result.warnings }, null, 2)}\n`, "utf8");
await writeResult(paths, result);
console.log(JSON.stringify(result, null, 2));
}
process.exit(exitCode(result.status));
@@ -1,186 +0,0 @@
#!/usr/bin/env node
import { join } from "node:path";
import { env } from "node:process";
import {
apiJson,
createBrowser,
ensureAuthenticatedBrowser,
ensureEvidence,
evidencePaths,
exitCode,
loadEnvFiles,
localIsoWithOffset,
safeScreenshot,
scanBrowserDiagnostics,
writeResult,
} from "./lib/langbot-e2e.mjs";
await loadEnvFiles();
const caseId = env.LBS_CASE_ID || "workspace-plugin-pages-smoke";
const paths = evidencePaths(caseId);
await ensureEvidence(paths);
const startedAt = new Date();
const frontendUrl = env.LANGBOT_FRONTEND_URL || "";
const backendUrl = env.LANGBOT_BACKEND_URL || "";
const targets = [
{
id: "control-plane",
author: "langbot",
plugin: "agent-control-plane",
page: "control-plane",
signal: "Agent Control Plane",
endpoint: "/health",
method: "GET",
},
{
id: "memory-console",
author: "langbot-team",
plugin: "LongTermMemory",
page: "memory_console",
signal: "Memory Console",
endpoint: "/summary",
method: "GET",
},
{
id: "langrag-observability",
author: "langbot-team",
plugin: "LangRAG",
page: "observability",
signal: "LangRAG Observability",
endpoint: "/snapshot",
method: "GET",
},
{
id: "skill-authoring",
author: "huanghuoguoguo",
plugin: "skill-authoring",
page: "authoring",
signal: "Skill Authoring",
endpoint: "/health",
method: "GET",
},
];
const result = {
source: "automation",
case_id: caseId,
run_id: paths.runId,
started_at: startedAt.toISOString(),
started_at_local: localIsoWithOffset(startedAt),
finished_at: "",
finished_at_local: "",
status: "fail",
reason: "",
frontend_url: frontendUrl,
backend_url: backendUrl,
pages: [],
browser_diagnostics: null,
evidence: {
console_log: paths.consoleLog,
network_log: paths.networkLog,
screenshot: paths.screenshot,
automation_result_json: paths.automationResultJson,
result_json: paths.resultJson,
},
evidence_collected: ["ui", "screenshot", "console", "network", "api_diagnostic"],
};
let browser;
try {
if (!frontendUrl || !backendUrl) throw new Error("LANGBOT_FRONTEND_URL and LANGBOT_BACKEND_URL must be configured.");
browser = await createBrowser(paths);
const { page } = browser;
await page.goto(frontendUrl, { waitUntil: "domcontentloaded" });
const auth = await ensureAuthenticatedBrowser(page, { frontendUrl, backendUrl });
if (auth.status !== "pass") {
result.status = auth.status;
throw new Error(auth.reason);
}
const token = await page.evaluate(() => localStorage.getItem("token") || "");
const pluginsResponse = await apiJson(backendUrl, "/api/v1/plugins", { token });
if (pluginsResponse.status >= 400 || pluginsResponse.json.code !== 0) {
throw new Error(pluginsResponse.json.msg || "Failed to list installed plugins.");
}
const installed = new Map((pluginsResponse.json.data?.plugins || []).map((item) => {
const metadata = item.manifest?.manifest?.metadata || item.manifest?.metadata || {};
return [`${metadata.author || ""}/${metadata.name || ""}`, item];
}));
const missing = targets.filter((target) => {
const plugin = installed.get(`${target.author}/${target.plugin}`);
if (!plugin) return true;
const pages = plugin.manifest?.manifest?.spec?.pages || plugin.manifest?.spec?.pages || [];
return !pages.some((pageSpec) => pageSpec.id === target.page);
});
if (missing.length) {
result.status = "blocked";
throw new Error(`Required plugin page fixture(s) are not installed or do not register the expected Page ID: ${missing.map((item) => item.id).join(", ")}`);
}
for (const target of targets) {
const routeId = `${target.author}/${target.plugin}/${target.page}`;
const route = `${frontendUrl.replace(/\/$/, "")}/home/plugin-pages?id=${encodeURIComponent(routeId)}`;
const pageResult = {
id: target.id,
route,
visible_signal: target.signal,
ui_status: "fail",
api_status: "fail",
screenshot: join(paths.evidenceDir, `${target.id}.png`),
page_api: null,
};
result.pages.push(pageResult);
await page.goto(route, { waitUntil: "domcontentloaded" });
await page.waitForLoadState("networkidle", { timeout: 15_000 }).catch(() => {});
const frame = page.frameLocator("iframe").first();
await frame.getByText(target.signal, { exact: true }).first().waitFor({ state: "visible", timeout: 30_000 });
pageResult.ui_status = "pass";
await safeScreenshot(page, pageResult.screenshot);
const apiResponse = await apiJson(
backendUrl,
`/api/v1/plugins/${encodeURIComponent(target.author)}/${encodeURIComponent(target.plugin)}/page-api`,
{
method: "POST",
token,
body: { page_id: target.page, endpoint: target.endpoint, method: target.method, body: null },
},
);
pageResult.page_api = {
endpoint: target.endpoint,
method: target.method,
http_status: apiResponse.status,
code: apiResponse.json.code ?? null,
has_data: apiResponse.json.data !== undefined && apiResponse.json.data !== null,
};
if (apiResponse.status >= 400 || apiResponse.json.code !== 0 || apiResponse.json.data == null) {
throw new Error(`${target.id} Page API ${target.method} ${target.endpoint} failed.`);
}
pageResult.api_status = "pass";
}
await safeScreenshot(page, paths.screenshot);
result.browser_diagnostics = await scanBrowserDiagnostics(paths);
if (result.browser_diagnostics.status !== "pass") {
throw new Error(result.browser_diagnostics.reason);
}
result.status = "pass";
result.reason = "All workspace plugin pages rendered in the LangBot WebUI and their read-only Page APIs passed.";
} catch (error) {
if (result.status === "fail" && /configured|Playwright is not installed|ECONNREFUSED|ERR_CONNECTION/i.test(error.message)) {
result.status = "env_issue";
}
result.reason = error.message;
} finally {
if (browser?.page) await safeScreenshot(browser.page, paths.screenshot);
if (browser) await browser.close().catch(() => {});
const finishedAt = new Date();
result.finished_at = finishedAt.toISOString();
result.finished_at_local = localIsoWithOffset(finishedAt);
await writeResult(paths, result);
console.log(JSON.stringify(result, null, 2));
}
process.exit(exitCode(result.status));
@@ -1,136 +0,0 @@
#!/usr/bin/env node
import { spawn } from "node:child_process";
import { access, mkdir, writeFile } from "node:fs/promises";
import { basename, delimiter, dirname, join, resolve } from "node:path";
import { env } from "node:process";
import {
ensureEvidence,
evidencePaths,
exitCode,
loadEnvFiles,
localIsoWithOffset,
resolveLangBotRepo,
writeResult,
} from "./lib/langbot-e2e.mjs";
await loadEnvFiles();
const caseId = env.LBS_CASE_ID || "workspace-repository-contracts";
const paths = evidencePaths(caseId);
await ensureEvidence(paths);
const startedAt = new Date();
const logsDir = join(paths.evidenceDir, "repository-contracts");
await mkdir(logsDir, { recursive: true });
const result = {
source: "automation",
case_id: caseId,
run_id: paths.runId,
started_at: startedAt.toISOString(),
started_at_local: localIsoWithOffset(startedAt),
finished_at: "",
finished_at_local: "",
status: "fail",
reason: "",
repositories: [],
evidence: {
contracts_dir: logsDir,
automation_result_json: paths.automationResultJson,
result_json: paths.resultJson,
},
evidence_collected: ["filesystem", "metrics"],
};
function run(command, args, { cwd, processEnv, timeoutMs = 900_000 }) {
return new Promise((resolvePromise) => {
const started = Date.now();
const child = spawn(command, args, { cwd, env: processEnv, stdio: ["ignore", "pipe", "pipe"] });
let stdout = "";
let stderr = "";
let timedOut = false;
const timer = setTimeout(() => {
timedOut = true;
child.kill("SIGTERM");
}, timeoutMs);
child.stdout.on("data", (chunk) => { stdout += chunk; });
child.stderr.on("data", (chunk) => { stderr += chunk; });
child.on("error", (error) => {
clearTimeout(timer);
resolvePromise({ status: null, stdout, stderr, error, timedOut, durationMs: Date.now() - started });
});
child.on("close", (status) => {
clearTimeout(timer);
resolvePromise({ status, stdout, stderr, error: null, timedOut, durationMs: Date.now() - started });
});
});
}
try {
const langbotRepo = await resolveLangBotRepo();
const workspaceRoot = resolve(env.LANGBOT_WORKSPACE_ROOT || dirname(langbotRepo));
const sdkRepo = resolve(env.LANGBOT_PLUGIN_SDK_REPO || join(workspaceRoot, "langbot-plugin-sdk"));
const python = join(langbotRepo, ".venv", "bin", "python");
await access(python);
await access(join(sdkRepo, "src", "langbot_plugin"));
const processEnv = {
...env,
PYTHONPATH: [join(sdkRepo, "src"), env.PYTHONPATH].filter(Boolean).join(delimiter),
};
const specs = [
["agent-runner", "langbot-agent-runner", ["-m", "pytest", "tests", "-q"]],
["control-plane", "langbot-agent-control-plane", ["-m", "pytest", "tests", "-q"]],
["longterm-memory", "langbot-longterm-memory", ["-m", "pytest", "tests", "-q"]],
["parser", "langbot-parser", ["-m", "pytest", "tests", "-q"]],
["rag", "langbot-rag", ["-m", "pytest", "tests", "-q"]],
["skill-authoring", "langbot-skill-authoring", ["-m", "pytest", "tests", "-q"]],
["local-agent", "langbot-local-agent", ["-m", "pytest", "tests", "-q"]],
["langbot-agent-provider", "LangBot", ["-m", "pytest", "tests/unit_tests/agent", "tests/unit_tests/provider", "-q"]],
["skills-cli", "LangBot/skills", ["--test", "test/lbs-cli.test.ts"]],
["plugin-sdk-runtime", "langbot-plugin-sdk", ["-m", "pytest", "tests", "-q", "--ignore=tests/packaging/test_installed_cli_blackbox.py"]],
["plugin-sdk-packaging", "langbot-plugin-sdk", ["-m", "pytest", "tests/packaging/test_installed_cli_blackbox.py", "-q"]],
];
for (const [id, directory, args] of specs) {
const cwd = resolve(workspaceRoot, directory);
const command = id === "skills-cli" ? env.LANGBOT_NODE || "node" : python;
const execution = await run(command, args, { cwd, processEnv });
const stdoutPath = join(logsDir, `${id}.stdout.log`);
const stderrPath = join(logsDir, `${id}.stderr.log`);
await writeFile(stdoutPath, execution.stdout, "utf8");
await writeFile(stderrPath, execution.stderr, "utf8");
const combined = `${execution.stdout}\n${execution.stderr}`;
const packagingNetworkIssue = id === "plugin-sdk-packaging"
&& execution.status !== 0
&& /hatchling|PyPI|TLS|SSL|network|timed out|Temporary failure|connection|EOF/i.test(combined);
const status = execution.status === 0 ? "pass" : packagingNetworkIssue ? "env_issue" : "fail";
result.repositories.push({
id,
repository: basename(cwd),
status,
exit_status: execution.status,
timed_out: execution.timedOut,
duration_ms: execution.durationMs,
stdout_log: stdoutPath,
stderr_log: stderrPath,
reason: execution.error?.message || (execution.timedOut ? "Test command timed out." : packagingNetworkIssue ? "Packaging blackbox could not fetch its isolated build dependency." : ""),
});
}
const productFailures = result.repositories.filter((item) => item.status === "fail");
const envIssues = result.repositories.filter((item) => item.status === "env_issue");
result.status = productFailures.length === 0 ? "pass" : "fail";
result.reason = productFailures.length === 0
? `All repository product contracts passed; ${envIssues.length} packaging environment issue(s) were classified separately.`
: `${productFailures.length} repository product contract group(s) failed.`;
} catch (error) {
result.status = "env_issue";
result.reason = error.message;
} finally {
const finishedAt = new Date();
result.finished_at = finishedAt.toISOString();
result.finished_at_local = localIsoWithOffset(finishedAt);
await writeResult(paths, result);
console.log(JSON.stringify(result, null, 2));
}
process.exit(exitCode(result.status));
+16 -725
View File
@@ -64,7 +64,7 @@
{
"directory": "langbot-mcp-ops",
"name": "langbot-mcp-ops",
"description": "Operate a LangBot instance through its built-in MCP (Model Context Protocol) server. Use when an AI agent needs to manage LangBot — list/create/update/delete bots, agents, pipelines, models, knowledge bases, MCP servers, and skills — over MCP instead of raw HTTP. Covers the /mcp endpoint, API-key auth (web-UI lbk_ keys and the config.yaml global key), the tool surface, and client configuration. Triggers on \"langbot mcp\", \"manage langbot via mcp\", \"langbot /mcp\", \"langbot mcp server\".",
"description": "Operate a LangBot instance through its built-in MCP (Model Context Protocol) server. Use when an AI agent needs to manage LangBot — list/create/update/delete bots, pipelines, models, knowledge bases, MCP servers, and skills — over MCP instead of raw HTTP. Covers the /mcp endpoint, API-key auth (web-UI lbk_ keys and the config.yaml global key), the tool surface, and client configuration. Triggers on \"langbot mcp\", \"manage langbot via mcp\", \"langbot /mcp\", \"langbot mcp server\".",
"references": [],
"cases": [],
"case_summaries": [],
@@ -134,18 +134,14 @@
"references/pipeline-debug-chat.md",
"references/plugin-e2e-smoke.md",
"references/sandbox-skill-authoring.md",
"references/skill-all-tool-acceptance.md",
"references/troubleshooting.md",
"references/web-ui-testing.md",
"references/workspace-release-testing.md"
"references/web-ui-testing.md"
],
"cases": [
"acp-agent-runner-debug-chat",
"agent-run-ledger-audit",
"agent-runner-async-db-readiness",
"agent-runner-behavior-matrix",
"agent-runner-fixture-contract",
"agent-runner-health-visibility",
"agent-runner-ledger-concurrency",
"agent-runner-ledger-contention",
"agent-runner-ledger-invariants",
@@ -154,8 +150,6 @@
"agent-runner-qa-debug-chat",
"agent-runner-release-preflight",
"agent-runner-runtime-chaos",
"bot-event-routing-product-flow",
"box-mcp-heartbeat-recovery",
"dify-agent-debug-chat",
"langbot-fake-provider-debug-chat-cross-pipeline-isolation",
"langbot-fake-provider-debug-chat-fault-recovery",
@@ -171,22 +165,14 @@
"langrag-parser-golden-e2e",
"langrag-sentinel-kb-discover",
"local-agent-basic-debug-chat",
"local-agent-combo-rag-compaction-tool-debug-chat",
"local-agent-complex-coding-task-debug-chat",
"local-agent-context-compaction-debug-chat",
"local-agent-effective-prompt-debug-chat",
"local-agent-model-fallback-before-first-chunk-debug-chat",
"local-agent-multimodal-debug-chat",
"local-agent-multitool-rag-compaction-debug-chat",
"local-agent-nonstreaming-debug-chat",
"local-agent-parallel-tools-rag-compaction-debug-chat",
"local-agent-plugin-tool-call-debug-chat",
"local-agent-rag-debug-chat",
"local-agent-rag-multimodal-debug-chat",
"local-agent-steering-debug-chat",
"local-agent-streaming-post-commit-failure-debug-chat",
"local-agent-tool-error-recovery-debug-chat",
"local-agent-tool-loop-limit-debug-chat",
"mcp-stdio-register",
"mcp-stdio-tool-call",
"pipeline-debug-chat",
@@ -196,14 +182,7 @@
"qa-plugin-smoke-live-install",
"sandbox-skill-authoring-e2e",
"sandbox-skill-authoring-edit-existing-e2e",
"skill-discovery-via-mcp-gateway",
"webui-login-state",
"wizard-onebot-agent-runtime",
"wizard-runner-marketplace-catalog",
"wizard-scenario-routing",
"workspace-compatibility-preflight",
"workspace-plugin-pages-smoke",
"workspace-repository-contracts"
"webui-login-state"
],
"case_summaries": [
{
@@ -236,30 +215,6 @@
"backend_log"
]
},
{
"id": "agent-run-ledger-audit",
"title": "Persisted AgentRunner run ledger passes end-to-end invariants",
"mode": "probe",
"area": "agent",
"type": "regression",
"priority": "p0",
"risk": "high",
"ci_eligible": false,
"tags": [
"agent-runner",
"ledger",
"audit",
"tools"
],
"automation": "scripts/e2e/agent-run-ledger-audit.mjs",
"setup_automation": [],
"setup_provides_env": [],
"evidence_required": [
"filesystem",
"metrics",
"api_diagnostic"
]
},
{
"id": "agent-runner-async-db-readiness",
"title": "AgentRunner async DB readiness probe",
@@ -326,31 +281,6 @@
"filesystem"
]
},
{
"id": "agent-runner-health-visibility",
"title": "Agent configuration shows whether its selected runner is usable",
"mode": "agent-browser",
"area": "agent",
"type": "feature",
"priority": "p0",
"risk": "medium",
"ci_eligible": false,
"tags": [
"agent-runner",
"agent",
"productization",
"release-gate"
],
"automation": "scripts/e2e/agent-runner-health-visibility.mjs",
"setup_automation": [],
"setup_provides_env": [],
"evidence_required": [
"ui",
"screenshot",
"console",
"api_diagnostic"
]
},
{
"id": "agent-runner-ledger-concurrency",
"title": "AgentRunner run ledger concurrency and auth pytest probe",
@@ -545,62 +475,6 @@
"filesystem"
]
},
{
"id": "bot-event-routing-product-flow",
"title": "Bot event routing can be configured and tested from the WebUI",
"mode": "agent-browser",
"area": "bot",
"type": "feature",
"priority": "p0",
"risk": "medium",
"ci_eligible": false,
"tags": [
"bot",
"eba",
"event-routing",
"productization"
],
"automation": "scripts/e2e/bot-event-routing-product-flow.mjs",
"setup_automation": [],
"setup_provides_env": [],
"evidence_required": [
"ui",
"screenshot",
"console",
"api_diagnostic"
]
},
{
"id": "box-mcp-heartbeat-recovery",
"title": "Box heartbeat restores an existing stdio MCP tool session",
"mode": "probe",
"area": "reliability",
"type": "chaos",
"priority": "p1",
"risk": "high",
"ci_eligible": false,
"tags": [
"box",
"mcp",
"stdio",
"reliability",
"chaos",
"fault-injection"
],
"automation": "skills/langbot-testing/probes/box-mcp-heartbeat-recovery.mjs",
"setup_automation": [],
"setup_provides_env": [],
"evidence_required": [
"ui",
"screenshot",
"console",
"network",
"metrics",
"api_diagnostic",
"resource_log",
"filesystem"
]
},
{
"id": "dify-agent-debug-chat",
"title": "Dify AgentRunner returns a response through Pipeline Debug Chat",
@@ -955,8 +829,7 @@
"ui",
"screenshot",
"console",
"network",
"api_diagnostic"
"backend_log"
]
},
{
@@ -1036,79 +909,6 @@
"backend_log"
]
},
{
"id": "local-agent-combo-rag-compaction-tool-debug-chat",
"title": "Local Agent preserves RAG, compacted history, and plugin tool result in one Debug Chat run",
"mode": "agent-browser",
"area": "pipeline",
"type": "regression",
"priority": "p1",
"risk": "high",
"ci_eligible": false,
"tags": [
"local-agent",
"langrag",
"context",
"compaction",
"plugin",
"tools",
"pipeline"
],
"automation": "scripts/e2e/pipeline-debug-chat.mjs",
"setup_automation": [
"node:scripts/e2e/ensure-local-agent-pipeline.mjs --write-env",
"node:scripts/e2e/ensure-langrag-sentinel-kb.mjs --write-env",
"case:qa-plugin-smoke-live-install"
],
"setup_provides_env": [
"LANGBOT_LOCAL_AGENT_PIPELINE_URL",
"LANGBOT_LOCAL_AGENT_PIPELINE_NAME",
"LANGBOT_LOCAL_AGENT_RAG_KB_UUID"
],
"evidence_required": [
"ui",
"screenshot",
"console",
"backend_log",
"api_diagnostic"
]
},
{
"id": "local-agent-complex-coding-task-debug-chat",
"title": "Local Agent completes a multi-file coding task with iterative verification",
"mode": "agent-browser",
"area": "pipeline",
"type": "regression",
"priority": "p1",
"risk": "high",
"ci_eligible": false,
"tags": [
"local-agent",
"agentic",
"coding",
"sandbox",
"tools",
"long-running"
],
"automation": "scripts/e2e/pipeline-debug-chat.mjs",
"setup_automation": [
"node:scripts/e2e/ensure-local-agent-pipeline.mjs --write-env",
"node:scripts/e2e/reset-complex-agent-task.mjs"
],
"setup_provides_env": [
"LANGBOT_LOCAL_AGENT_PIPELINE_URL",
"LANGBOT_LOCAL_AGENT_PIPELINE_NAME"
],
"evidence_required": [
"ui",
"screenshot",
"console",
"network",
"api_diagnostic",
"filesystem",
"metrics"
]
},
{
"id": "local-agent-context-compaction-debug-chat",
"title": "Local Agent compacts long Debug Chat history and preserves older facts",
@@ -1170,38 +970,6 @@
"backend_log"
]
},
{
"id": "local-agent-model-fallback-before-first-chunk-debug-chat",
"title": "Local Agent falls back when the primary model fails before streaming starts",
"mode": "probe",
"area": "pipeline",
"type": "chaos",
"priority": "p1",
"risk": "high",
"ci_eligible": false,
"tags": [
"local-agent",
"fallback",
"streaming",
"fake-provider",
"fault-injection"
],
"automation": "skills/langbot-testing/probes/langbot-debug-chat-concurrency.mjs",
"setup_automation": [
"node:scripts/e2e/ensure-fake-provider-pipeline.mjs --write-env"
],
"setup_provides_env": [
"LANGBOT_FAKE_PROVIDER_URL",
"LANGBOT_FAKE_PROVIDER_PIPELINE_URL",
"LANGBOT_FAKE_PROVIDER_PIPELINE_NAME"
],
"evidence_required": [
"metrics",
"network",
"api_diagnostic",
"filesystem"
]
},
{
"id": "local-agent-multimodal-debug-chat",
"title": "Local Agent Debug Chat preserves uploaded image input",
@@ -1232,47 +1000,9 @@
"backend_log"
]
},
{
"id": "local-agent-multitool-rag-compaction-debug-chat",
"title": "Local Agent preserves RAG, compacted history, and two-step plugin tool loop",
"mode": "agent-browser",
"area": "pipeline",
"type": "regression",
"priority": "p1",
"risk": "high",
"ci_eligible": false,
"tags": [
"local-agent",
"langrag",
"context",
"compaction",
"plugin",
"tools",
"tool-loop",
"pipeline"
],
"automation": "scripts/e2e/pipeline-debug-chat.mjs",
"setup_automation": [
"node:scripts/e2e/ensure-local-agent-pipeline.mjs --write-env",
"node:scripts/e2e/ensure-langrag-sentinel-kb.mjs --write-env",
"case:qa-plugin-smoke-live-install"
],
"setup_provides_env": [
"LANGBOT_LOCAL_AGENT_PIPELINE_URL",
"LANGBOT_LOCAL_AGENT_PIPELINE_NAME",
"LANGBOT_LOCAL_AGENT_RAG_KB_UUID"
],
"evidence_required": [
"ui",
"screenshot",
"console",
"backend_log",
"api_diagnostic"
]
},
{
"id": "local-agent-nonstreaming-debug-chat",
"title": "Local Agent Debug Chat returns a deterministic response with UI streaming disabled",
"title": "Local Agent Debug Chat returns a deterministic non-streaming response",
"mode": "agent-browser",
"area": "pipeline",
"type": "regression",
@@ -1298,44 +1028,6 @@
"backend_log"
]
},
{
"id": "local-agent-parallel-tools-rag-compaction-debug-chat",
"title": "Local Agent preserves RAG, compacted history, and parallel plugin tools",
"mode": "agent-browser",
"area": "pipeline",
"type": "regression",
"priority": "p1",
"risk": "high",
"ci_eligible": false,
"tags": [
"local-agent",
"langrag",
"context",
"compaction",
"plugin",
"tools",
"parallel-tools",
"pipeline"
],
"automation": "scripts/e2e/pipeline-debug-chat.mjs",
"setup_automation": [
"node:scripts/e2e/ensure-local-agent-pipeline.mjs --write-env",
"node:scripts/e2e/ensure-langrag-sentinel-kb.mjs --write-env",
"case:qa-plugin-smoke-live-install"
],
"setup_provides_env": [
"LANGBOT_LOCAL_AGENT_PIPELINE_URL",
"LANGBOT_LOCAL_AGENT_PIPELINE_NAME",
"LANGBOT_LOCAL_AGENT_RAG_KB_UUID"
],
"evidence_required": [
"ui",
"screenshot",
"console",
"backend_log",
"api_diagnostic"
]
},
{
"id": "local-agent-plugin-tool-call-debug-chat",
"title": "Local Agent can call a plugin-provided tool",
@@ -1463,105 +1155,6 @@
"api_diagnostic"
]
},
{
"id": "local-agent-streaming-post-commit-failure-debug-chat",
"title": "Local Agent does not fall back after a committed stream fails",
"mode": "probe",
"area": "pipeline",
"type": "chaos",
"priority": "p1",
"risk": "high",
"ci_eligible": false,
"tags": [
"local-agent",
"fallback",
"streaming",
"fake-provider",
"fault-injection"
],
"automation": "skills/langbot-testing/probes/langbot-debug-chat-concurrency.mjs",
"setup_automation": [
"node:scripts/e2e/ensure-fake-provider-pipeline.mjs --write-env"
],
"setup_provides_env": [
"LANGBOT_FAKE_PROVIDER_URL",
"LANGBOT_FAKE_PROVIDER_PIPELINE_URL",
"LANGBOT_FAKE_PROVIDER_PIPELINE_NAME"
],
"evidence_required": [
"metrics",
"network",
"api_diagnostic",
"filesystem"
]
},
{
"id": "local-agent-tool-error-recovery-debug-chat",
"title": "Local Agent feeds plugin tool errors back to the model",
"mode": "agent-browser",
"area": "pipeline",
"type": "regression",
"priority": "p1",
"risk": "high",
"ci_eligible": false,
"tags": [
"local-agent",
"plugin",
"tools",
"tool-error",
"pipeline"
],
"automation": "scripts/e2e/pipeline-debug-chat.mjs",
"setup_automation": [
"node:scripts/e2e/ensure-local-agent-pipeline.mjs --write-env",
"case:qa-plugin-smoke-live-install"
],
"setup_provides_env": [
"LANGBOT_LOCAL_AGENT_PIPELINE_URL",
"LANGBOT_LOCAL_AGENT_PIPELINE_NAME"
],
"evidence_required": [
"ui",
"screenshot",
"console",
"backend_log",
"api_diagnostic"
]
},
{
"id": "local-agent-tool-loop-limit-debug-chat",
"title": "Local Agent stops repeated tool calls at max-tool-iterations",
"mode": "agent-browser",
"area": "pipeline",
"type": "regression",
"priority": "p1",
"risk": "high",
"ci_eligible": false,
"tags": [
"local-agent",
"plugin",
"tools",
"tool-loop",
"limit",
"pipeline"
],
"automation": "scripts/e2e/pipeline-debug-chat.mjs",
"setup_automation": [
"node:scripts/e2e/ensure-local-agent-pipeline.mjs --write-env",
"case:qa-plugin-smoke-live-install"
],
"setup_provides_env": [
"LANGBOT_LOCAL_AGENT_PIPELINE_URL",
"LANGBOT_LOCAL_AGENT_PIPELINE_NAME"
],
"evidence_required": [
"ui",
"screenshot",
"console",
"backend_log",
"api_diagnostic"
]
},
{
"id": "mcp-stdio-register",
"title": "MCP stdio fixture is registered and exposes qa_mcp_echo",
@@ -1603,21 +1196,18 @@
],
"automation": "scripts/e2e/pipeline-debug-chat.mjs",
"setup_automation": [
"node:scripts/e2e/ensure-fake-provider-pipeline.mjs --write-env",
"node:scripts/e2e/ensure-local-agent-pipeline.mjs --write-env",
"case:mcp-stdio-register"
],
"setup_provides_env": [
"LANGBOT_FAKE_PROVIDER_PIPELINE_URL",
"LANGBOT_FAKE_PROVIDER_PIPELINE_NAME",
"LANGBOT_MCP_QA_STDIO_SERVER_UUID"
"LANGBOT_LOCAL_AGENT_PIPELINE_URL",
"LANGBOT_LOCAL_AGENT_PIPELINE_NAME"
],
"evidence_required": [
"ui",
"screenshot",
"console",
"network",
"api_diagnostic",
"metrics"
"backend_log",
"api_diagnostic"
]
},
{
@@ -1788,14 +1378,8 @@
"edit"
],
"automation": "scripts/e2e/pipeline-debug-chat.mjs",
"setup_automation": [
"node:scripts/e2e/ensure-local-agent-pipeline.mjs --write-env",
"node:scripts/e2e/reset-sandbox-skill.mjs --name=lb-skill-edit-regression"
],
"setup_provides_env": [
"LANGBOT_LOCAL_AGENT_PIPELINE_URL",
"LANGBOT_LOCAL_AGENT_PIPELINE_NAME"
],
"setup_automation": [],
"setup_provides_env": [],
"evidence_required": [
"ui",
"backend_log",
@@ -1803,31 +1387,6 @@
"filesystem"
]
},
{
"id": "skill-discovery-via-mcp-gateway",
"title": "External harness discovers LangBot skills via langbot_list_assets (all-tool model)",
"mode": "agent-browser",
"area": "sandbox",
"type": "regression",
"priority": "p2",
"risk": "medium",
"ci_eligible": false,
"tags": [
"skills",
"mcp-gateway",
"acp-agent-runner",
"all-tool-model",
"tools"
],
"automation": "scripts/e2e/pipeline-debug-chat.mjs",
"setup_automation": [],
"setup_provides_env": [],
"evidence_required": [
"ui",
"screenshot",
"backend_log"
]
},
{
"id": "webui-login-state",
"title": "Configured frontend opens with authenticated LangBot WebUI state",
@@ -1849,169 +1408,17 @@
"screenshot",
"console"
]
},
{
"id": "wizard-onebot-agent-runtime",
"title": "Quick Start runs a real OneBot member event through an Agent",
"mode": "agent-browser",
"area": "wizard",
"type": "feature",
"priority": "p0",
"risk": "high",
"ci_eligible": false,
"tags": [
"wizard",
"eba",
"onebot",
"agent-runner",
"release-gate"
],
"automation": "scripts/e2e/wizard-onebot-agent-runtime.mjs",
"setup_automation": [],
"setup_provides_env": [],
"evidence_required": [
"ui",
"screenshot",
"console",
"api_diagnostic"
]
},
{
"id": "wizard-runner-marketplace-catalog",
"title": "Quick Start installs a published AgentRunner on a clean instance",
"mode": "agent-browser",
"area": "wizard",
"type": "feature",
"priority": "p0",
"risk": "high",
"ci_eligible": false,
"tags": [
"wizard",
"agent-runner",
"marketplace",
"clean-install",
"productization"
],
"automation": "scripts/e2e/wizard-runner-marketplace-catalog.mjs",
"setup_automation": [],
"setup_provides_env": [],
"evidence_required": [
"ui",
"screenshot",
"console",
"api_diagnostic"
]
},
{
"id": "wizard-scenario-routing",
"title": "Quick Start filters channels by the selected bot scenario",
"mode": "agent-browser",
"area": "wizard",
"type": "feature",
"priority": "p0",
"risk": "medium",
"ci_eligible": false,
"tags": [
"wizard",
"eba",
"event-routing",
"productization"
],
"automation": "scripts/e2e/wizard-scenario-routing.mjs",
"setup_automation": [],
"setup_provides_env": [],
"evidence_required": [
"ui",
"screenshot",
"console",
"api_diagnostic"
]
},
{
"id": "workspace-compatibility-preflight",
"title": "LangBot workspace repositories and local SDK are mutually compatible",
"mode": "probe",
"area": "release",
"type": "smoke",
"priority": "p0",
"risk": "high",
"ci_eligible": true,
"tags": [
"workspace",
"compatibility",
"preflight",
"sdk"
],
"automation": "scripts/e2e/workspace-compatibility-preflight.mjs",
"setup_automation": [],
"setup_provides_env": [],
"evidence_required": [
"filesystem",
"api_diagnostic"
]
},
{
"id": "workspace-plugin-pages-smoke",
"title": "Major workspace plugin pages render and answer read-only Page APIs",
"mode": "agent-browser",
"area": "plugin",
"type": "smoke",
"priority": "p0",
"risk": "high",
"ci_eligible": false,
"tags": [
"workspace",
"plugin",
"page",
"browser"
],
"automation": "scripts/e2e/workspace-plugin-pages-smoke.mjs",
"setup_automation": [],
"setup_provides_env": [],
"evidence_required": [
"ui",
"screenshot",
"console",
"network",
"api_diagnostic"
]
},
{
"id": "workspace-repository-contracts",
"title": "Changed LangBot workspace repositories pass their contract suites",
"mode": "probe",
"area": "release",
"type": "regression",
"priority": "p0",
"risk": "high",
"ci_eligible": true,
"tags": [
"workspace",
"contracts",
"pytest",
"sdk"
],
"automation": "scripts/e2e/workspace-repository-contracts.mjs",
"setup_automation": [],
"setup_provides_env": [],
"evidence_required": [
"filesystem",
"metrics"
]
}
],
"suites": [
"agent-runner-release-gate",
"core-smoke",
"eba-release-gate",
"langbot-debug-chat-isolation-gate",
"langbot-debug-chat-load-gate",
"langbot-live-backend-gate",
"langbot-performance-contract-gate",
"langbot-performance-reliability-gate",
"langbot-user-path-performance-gate",
"langbot-workspace-contract-gate",
"langbot-workspace-release-gate",
"local-agent-gate"
],
"suite_summaries": [
@@ -2049,11 +1456,6 @@
"local-agent-context-compaction-debug-chat",
"local-agent-rag-debug-chat",
"local-agent-plugin-tool-call-debug-chat",
"local-agent-tool-error-recovery-debug-chat",
"local-agent-tool-loop-limit-debug-chat",
"local-agent-combo-rag-compaction-tool-debug-chat",
"local-agent-multitool-rag-compaction-debug-chat",
"local-agent-parallel-tools-rag-compaction-debug-chat",
"mcp-stdio-register",
"mcp-stdio-tool-call",
"local-agent-nonstreaming-debug-chat",
@@ -2079,26 +1481,6 @@
"local-agent-basic-debug-chat"
]
},
{
"id": "eba-release-gate",
"title": "Event-based bot product release gate",
"description": "Browser and runtime gate for scenario onboarding, route diagnostics, and real platform-to-Agent execution.",
"type": "release_gate",
"priority": "p0",
"tags": [
"eba",
"event-routing",
"wizard",
"release-gate"
],
"cases": [
"wizard-scenario-routing",
"wizard-runner-marketplace-catalog",
"agent-runner-health-visibility",
"bot-event-routing-product-flow",
"wizard-onebot-agent-runtime"
]
},
{
"id": "langbot-debug-chat-isolation-gate",
"title": "LangBot Debug Chat isolation gate",
@@ -2206,57 +1588,6 @@
"pipeline-debug-chat-performance"
]
},
{
"id": "langbot-workspace-contract-gate",
"title": "LangBot workspace contract gate",
"description": "Low-cost cross-repository gate for branch/SDK compatibility and deterministic product contracts before live workflow testing.",
"type": "regression",
"priority": "p0",
"tags": [
"workspace",
"contracts",
"pr-gate"
],
"cases": [
"workspace-compatibility-preflight",
"agent-runner-fixture-contract",
"agent-runner-behavior-matrix",
"agent-runner-ledger-invariants",
"langbot-fault-taxonomy-contract",
"langbot-overhead-accounting-contract",
"workspace-repository-contracts"
]
},
{
"id": "langbot-workspace-release-gate",
"title": "LangBot workspace top-down release gate",
"description": "Broad release gate combining deterministic repository contracts with representative browser workflows, plugin pages, RAG/parser, EBA, external AgentRunner, and one complex LocalAgent task.",
"type": "release_gate",
"priority": "p0",
"tags": [
"workspace",
"release-gate",
"browser",
"agent"
],
"cases": [
"workspace-compatibility-preflight",
"workspace-repository-contracts",
"agent-runner-release-preflight",
"webui-login-state",
"wizard-scenario-routing",
"bot-event-routing-product-flow",
"agent-runner-health-visibility",
"agent-runner-qa-debug-chat",
"local-agent-complex-coding-task-debug-chat",
"agent-run-ledger-audit",
"langrag-kb-retrieve",
"workspace-plugin-pages-smoke",
"mcp-stdio-register",
"mcp-stdio-tool-call",
"wizard-onebot-agent-runtime"
]
},
{
"id": "local-agent-gate",
"title": "Local Agent runner regression gate",
@@ -2270,18 +1601,11 @@
],
"cases": [
"local-agent-basic-debug-chat",
"local-agent-model-fallback-before-first-chunk-debug-chat",
"local-agent-streaming-post-commit-failure-debug-chat",
"qa-plugin-smoke-live-install",
"local-agent-effective-prompt-debug-chat",
"local-agent-context-compaction-debug-chat",
"local-agent-rag-debug-chat",
"local-agent-plugin-tool-call-debug-chat",
"local-agent-tool-error-recovery-debug-chat",
"local-agent-tool-loop-limit-debug-chat",
"local-agent-combo-rag-compaction-tool-debug-chat",
"local-agent-multitool-rag-compaction-debug-chat",
"local-agent-parallel-tools-rag-compaction-debug-chat",
"local-agent-steering-debug-chat",
"mcp-stdio-tool-call",
"local-agent-nonstreaming-debug-chat",
@@ -2331,8 +1655,7 @@
"kind": "python",
"path": "fixtures/mcp/qa_mcp_echo_server.py",
"related_cases": [
"mcp-stdio-tool-call",
"box-mcp-heartbeat-recovery"
"mcp-stdio-tool-call"
]
},
{
@@ -2342,10 +1665,7 @@
"path": "fixtures/rag/sentinel-doc.txt",
"related_cases": [
"langrag-kb-retrieve",
"local-agent-rag-debug-chat",
"local-agent-combo-rag-compaction-tool-debug-chat",
"local-agent-multitool-rag-compaction-debug-chat",
"local-agent-parallel-tools-rag-compaction-debug-chat"
"local-agent-rag-debug-chat"
]
},
{
@@ -2377,11 +1697,6 @@
"plugin-e2e-smoke",
"local-agent-effective-prompt-debug-chat",
"local-agent-plugin-tool-call-debug-chat",
"local-agent-tool-error-recovery-debug-chat",
"local-agent-tool-loop-limit-debug-chat",
"local-agent-combo-rag-compaction-tool-debug-chat",
"local-agent-multitool-rag-compaction-debug-chat",
"local-agent-parallel-tools-rag-compaction-debug-chat",
"local-agent-steering-debug-chat"
]
},
@@ -2392,8 +1707,7 @@
"path": "fixtures/plugins/qa-plugin-smoke/dist/qa-plugin-smoke-0.1.0.lbpkg",
"related_cases": [
"qa-plugin-smoke-live-install",
"plugin-e2e-smoke",
"local-agent-tool-error-recovery-debug-chat"
"plugin-e2e-smoke"
]
}
],
@@ -2402,12 +1716,10 @@
"aiosqlite-connect-hangs",
"ambiguous-runner-default-label",
"backend-not-listening",
"box-runtime-silent-action-timeout",
"box-session-conflict-logical-metadata",
"debug-chat-history-contaminates-automation",
"dynamic-form-missing-config-id",
"e2b-extra-mount-sync-missing",
"local-agent-mcp-text-result-empty-follow-up",
"local-agent-model-route-unavailable",
"marketplace-network-flaky",
"mcp-stdio-args-not-applied",
@@ -2462,16 +1774,6 @@
"local-agent-basic-debug-chat"
]
},
{
"id": "box-runtime-silent-action-timeout",
"title": "Box runtime process stays alive while all action RPC calls time out",
"category": "product",
"related_cases": [
"local-agent-complex-coding-task-debug-chat",
"agent-run-ledger-audit",
"mcp-stdio-tool-call"
]
},
{
"id": "box-session-conflict-logical-metadata",
"title": "BoxSessionConflictError after a successful first exec",
@@ -2507,16 +1809,6 @@
"sandbox-skill-authoring-e2e"
]
},
{
"id": "local-agent-mcp-text-result-empty-follow-up",
"title": "MCP tool succeeds but Local Agent completes without a visible reply",
"category": "product",
"related_cases": [
"mcp-stdio-register",
"mcp-stdio-tool-call",
"agent-run-ledger-audit"
]
},
{
"id": "local-agent-model-route-unavailable",
"title": "Local Agent model route is unavailable for the requested run shape",
@@ -2526,8 +1818,7 @@
"local-agent-plugin-tool-call-debug-chat",
"mcp-stdio-tool-call",
"local-agent-multimodal-debug-chat",
"local-agent-nonstreaming-debug-chat",
"local-agent-complex-coding-task-debug-chat"
"local-agent-nonstreaming-debug-chat"
]
},
{
+11
View File
@@ -27,6 +27,17 @@ The `all` / `box` profile starts three services:
- `langbot_box` — Box sandbox runtime (`:5410`). Uses the host Docker socket to
spawn sandbox containers, so the **Box root host path and in-container path
must be identical** (`BOX__LOCAL__HOST_ROOT=${LANGBOT_BOX_ROOT:-${PWD}/data/box}`).
Its RPC and managed-process relay require a shared
`LANGBOT_BOX_CONTROL_TOKEN` (at least 32 non-whitespace characters) in both
the LangBot and Box containers. Generate it once with `openssl rand -hex 32`;
never put it in `box.runtime.endpoint` or commit it to config.
Every Compose deployment also needs one
`LANGBOT_PLUGIN_RUNTIME_CONTROL_TOKEN` shared by `langbot` and
`langbot_plugin_runtime`. Generate it with `openssl rand -hex 32` and export it
before `docker compose up`; the external Plugin Runtime fails closed when the
token is empty or weak. Kubernetes uses the `langbot-plugin-runtime-control`
Secret shown in `docker/kubernetes.yaml`.
With Box off, the dashboard/skills list stays visible (read-only) but sandbox
tools, skill add/edit, and stdio MCP are disabled. Set `box.enabled: false`
+17 -4
View File
@@ -65,10 +65,23 @@ Route auth is declared per-route via `AuthType` in
- `API_KEY``X-API-Key` or `Authorization: Bearer <key>`.
- `USER_TOKEN_OR_API_KEY` — either.
API keys are verified by `apikey_service.verify_api_key()`, which accepts:
1. the **global key** from `config.yaml` `api.global_api_key` (no DB, no login,
no `lbk_` prefix required), then
2. **web-UI keys** (DB-stored, `lbk_` prefix).
Authenticated routes receive an immutable `RequestContext` containing the
principal, authorized Workspace membership, fixed-role permissions, instance,
request id, and placement generation. A browser's `X-Workspace-Id` is only a
selector and is always checked against the Account membership. Tenant services
must accept this context (or an explicit trusted execution context) and fail
closed when it is absent.
API-key authentication accepts:
1. the **global key** from `config.yaml` `api.global_api_key` only for a
community instance with exactly one local Workspace, then
2. **web-UI keys** whose one-time `lbk_` secret is stored only as a hash and is
bound to one Workspace, explicit scopes, status, and optional expiry.
An API key derives its Workspace from the key record and ignores a caller's
Workspace selector. Public Bot/Webhook routes similarly derive Workspace from
the opaque owning resource rather than a header.
Route groups self-register via `@group.group_class(name, path)` and are
discovered by `importutil.import_modules_in_pkg`.
@@ -24,7 +24,7 @@ Healthy startup includes:
```text
Running on http://0.0.0.0:<backend-port>
Connected to plugin runtime.
Plugin langbot-team/LocalAgent initialized
Plugin langbot/local-agent initialized
```
Quick check:
+17 -16
View File
@@ -1,6 +1,6 @@
---
name: langbot-mcp-ops
description: Operate a LangBot instance through its built-in MCP (Model Context Protocol) server. Use when an AI agent needs to manage LangBot — list/create/update/delete bots, agents, pipelines, models, knowledge bases, MCP servers, and skills — over MCP instead of raw HTTP. Covers the /mcp endpoint, API-key auth (web-UI lbk_ keys and the config.yaml global key), the tool surface, and client configuration. Triggers on "langbot mcp", "manage langbot via mcp", "langbot /mcp", "langbot mcp server".
description: Operate a LangBot instance through its built-in MCP (Model Context Protocol) server. Use when an AI agent needs to manage LangBot — list/create/update/delete bots, pipelines, models, knowledge bases, MCP servers, and skills — over MCP instead of raw HTTP. Covers the /mcp endpoint, API-key auth (web-UI lbk_ keys and the config.yaml global key), the tool surface, and client configuration. Triggers on "langbot mcp", "manage langbot via mcp", "langbot /mcp", "langbot mcp server".
---
# LangBot MCP Operations
@@ -29,13 +29,19 @@ Authorization: Bearer <api-key>
Two kinds of key are accepted:
1. **Web-UI key** — created in the web UI (sidebar → API Keys), prefixed `lbk_`,
stored in the database.
1. **Web-UI key** — created in the web UI (sidebar → API Keys), prefixed `lbk_`.
The secret is shown once; only its SHA-256 hash is stored. Each key is bound
to one Workspace and has explicit scopes, status, optional expiry, and
last-used metadata. The key determines the Workspace; callers cannot switch
it with `X-Workspace-Id`.
2. **Global API key** — set in `data/config.yaml` under `api.global_api_key`.
Requires no login session and no DB record; does not need the `lbk_` prefix.
Leave empty to disable. See the `langbot-deploy` skill for config details.
It is accepted only by a community instance with exactly one local
Workspace and is disabled for SaaS multi-Workspace operation. Leave empty to
disable. See the `langbot-deploy` skill for config details.
Requests without a valid key get `401 Unauthorized`.
Invalid, revoked, or expired keys get `401 Unauthorized`. A valid key whose
scopes do not authorize a tool gets `403 Forbidden`.
## Client configuration
@@ -58,8 +64,6 @@ The tools wrap the LangBot service layer. Current tools (v1):
| --- | --- |
| `get_system_info` | Version, edition, instance id |
| `list_bots` / `get_bot` / `create_bot` / `update_bot` / `delete_bot` | Manage messaging-platform bots (secrets redacted on read) |
| `list_bot_event_route_statuses` / `test_bot_event_route` | Inspect bot event-route runtime status and dispatch a synthetic test event through saved routes without sending real outbound platform messages |
| `list_processors` / `get_processor` / `create_processor` / `update_processor` / `delete_processor` | Manage the peer Agent and Pipeline processor types |
| `list_pipelines` / `get_pipeline` / `create_pipeline` / `update_pipeline` / `delete_pipeline` | Manage pipelines |
| `list_llm_models` / `get_llm_model` / `list_embedding_models` / `list_model_providers` | Inspect models & providers |
| `list_knowledge_bases` / `get_knowledge_base` / `retrieve_knowledge_base` | RAG knowledge bases (incl. semantic search) |
@@ -68,13 +72,9 @@ The tools wrap the LangBot service layer. Current tools (v1):
Mutating tools (`create_*`, `update_*`) take a JSON object matching the same
shape as the corresponding HTTP API request body. Discover resources with the
`list_*` / `get_*` tools before mutating; identifiers are UUIDs.
`test_bot_event_route` uses the bot's saved runtime route table, injects a
synthetic event such as `message.received`, and suppresses platform delivery.
It still executes the selected processor, so tools and external services may
have side effects. Use `payload` for sample event fields, for example
`{"message_text": "hello", "chat_type": "private", "chat_id": "u1"}`.
`list_*` / `get_*` tools before mutating; identifiers are UUIDs. Reads require
`resource.view`; mutations require `resource.manage`. All service calls inherit
the immutable Workspace context authenticated at the MCP transport boundary.
## How to use
@@ -101,7 +101,8 @@ have side effects. Use `payload` for sample event fields, for example
- `/mcp` is the **server** LangBot exposes. The `/api/v1/mcp` routes are the
**client** side (managing external MCP servers LangBot connects to). Don't
confuse them.
- A `401` means the key is wrong, missing, or (for the global key)
`api.global_api_key` is empty in config.yaml.
- A `401` means the key is wrong, missing, revoked, expired, or (for the global
key) `api.global_api_key` is empty or the instance is not an OSS singleton.
- A `403` means the key is valid but lacks the permission required by the tool.
- The global key is plaintext in config.yaml — only enable it on trusted/internal
deployments and serve over HTTPS.
+1 -1
View File
@@ -422,7 +422,7 @@ When a plugin doesn't work:
3. **Verify plugin loaded**: `GET /api/v1/plugins` — should list your plugin
4. **Test person mode first**: `session_type=person` always triggers pipeline, isolating trigger rule issues
5. **Check trigger rules**: Group mode requires @bot, prefix match, or random% to enter pipeline
6. **Verify model configured**: Pipeline's `config.ai.runner_config[config.ai.runner.id].model.primary` must point to a valid model UUID with working API keys
6. **Verify model configured**: Pipeline's `config.ai.local-agent.model.primary` must point to a valid model UUID with working API keys
## Publishing Plugins
-1
View File
@@ -22,7 +22,6 @@ Use this skill when an agent needs to verify LangBot behavior through the WebUI
- **LangRAG knowledge bases**: read `references/langrag-knowledge-base.md`.
- **MCP stdio tool testing**: read `references/mcp-stdio-testing.md`.
- **Performance, reliability, or chaos probes**: read `references/performance-reliability-testing.md`.
- **Cross-repository workspace and release gates**: read `references/workspace-release-testing.md`.
- **Drive a live instance over MCP (not raw HTTP)**: use the `langbot-mcp-ops` skill — the instance exposes an MCP server at `http://<host>:5300/mcp` (reuses API keys). Useful for setting up bots/pipelines/models as test fixtures programmatically.
- **Known failures and fixes**: read `references/troubleshooting.md`.
- **Reusable test groups**: run `bin/lbs suite list` and `bin/lbs suite plan <suite-id>` before manually assembling a case set.

Some files were not shown because too many files have changed in this diff Show More