Files
LangBot/docs/multi-tenant/cloud-runtime-soak-gate.md
T
RockChinQ e1ac5e0fc8 feat(tenancy): add Workspace multi-tenant foundation (#2353)
* Document multi-tenant workspace architecture

* Add OSS and commercial workspace boundaries

* docs: redesign multi-tenant workspace architecture

* feat(tenancy): implement workspace isolation

* docs(tenancy): record verification evidence

* docs(tenancy): revise single-instance SaaS topology

* docs(tenancy): refine architecture options

* docs: finalize cloud v2 multi-tenant decisions

* feat(tenancy): establish cloud isolation foundations

* feat(tenancy): harden shared cloud runtime boundaries

* docs(tenancy): record final isolation verification

* fix(tenancy): close isolation and permission gaps

* docs(tenancy): record final isolation verification

* feat(tenancy): connect cloud workspace control plane

* fix(build): install git for pinned SDK

* docs(cloud): update control plane verification

* chore: update multi-tenant SDK pin

* fix(cloud): skip legacy model sync during startup

* test(cloud): preserve minimal model manager fixtures

* fix(cloud): preserve authenticated account context

* fix(cloud): reuse authenticated account for user info

* feat(cloud): complete Workspace settings navigation

* test(web): cover Workspace dropdown menu

* feat(web): place workspace controls in sidebar

* refactor(web): streamline workspace controls

* style(web): format workspace layout test

* fix(cloud): surface runtime and workspace plan status

* fix(plugin): keep runtime identity stable across restarts

* fix(ui): widen and center workspace switcher

* fix(ui): hide roles from workspace switcher

* fix(ui): align workspace switcher with sidebar entries

* feat(workspace): add in-product collaboration and direct Cloud launch

* style: format collaboration changes

* fix(workspace): bind collaboration APIs to tenant UoW

* fix(cloud): preserve Core-owned collaboration state

* test(cloud): require Space identity for invite registration

* feat(cloud): complete secure invitation experience

* style(web): format invitation flows

* fix(cloud): recover box runtime without unscoped skill reload

* feat(oss): enforce invitation account and owner billing flows

* style: format OSS account service

* test(oss): cover invitation logout handoff

* fix(oss): resolve workspace owner in scoped session

* feat(cloud): harden multi-tenant runtime resources

* fix(cloud): bound runtime restart storms

* fix(cloud): eliminate periodic runtime CPU spikes

* fix(cloud): enforce instance capacity ceilings

* fix(cloud): scope public login capability discovery

* fix(cloud): bound tenant maintenance and monitoring work

* fix(runtime): bound tenant resource amplification

* fix(deps): pin green multi-tenant plugin SDK

* fix(cloud): handle unavailable skill capability

* fix(security): require authentication for image file endpoint (H-2)

- Changed /api/v1/files/image from AuthType.NONE to USER_TOKEN_OR_API_KEY
- Added Permission.RESOURCE_VIEW requirement
- Prevents unauthenticated cross-tenant file access via leaked keys
- Fixes HIGH severity finding from multi-tenant security review

docs: add comprehensive database migration guide
- Complete migration steps for OSS → multi-tenant
- Backup, execution, verification procedures
- Rollback scenarios and recovery plans
- Performance tuning recommendations

* test: add comprehensive cross-tenant isolation tests

Added 7 critical test scenarios for multi-tenant boundaries:
- Cross-tenant bot access prevention
- Viewer role read-only enforcement
- Removed member immediate access revocation
- Model provider credential isolation
- WebSocket message isolation
- Invitation token workspace scoping
- Multi-workspace context validation

These tests address P0-2 coverage gaps for:
- workspaces.py (membership & invitation flows)
- user.py (authentication & authorization)
- websocket_chat.py (real-time isolation)
- plugins.py (resource access control)

docs: finalize database migration guide

* fix(security): resolve M-1, M-2, M-3 security findings

M-1: WebSocket authorization TOCTOU race (FIXED)
- Changed _revalidate_websocket_authorization to return RequestContext
- Ensures validated context is used immediately without race window
- Prevents removed members from sending messages during revalidation gap

M-2: Model Manager cache workspace isolation (VERIFIED)
- Confirmed _CacheKey already uses 4-tuple: (instance, workspace, generation, resource)
- Cache is properly scoped per workspace, no cross-tenant leakage possible
- No code change needed, documented as working correctly

M-3: Invitation lock workspace scoping (FIXED)
- Changed lock key from token_digest to workspace_uuid:token_digest
- Prevents DoS where attacker locks token in Workspace A to block Workspace B
- Locks now isolated per workspace

All MEDIUM severity findings from security review now resolved.

* fix(cloud): unblock tenant CI and enforce knowledge quotas

* fix(tenancy): scope rerank model sync

---------

Co-authored-by: dadachann <185672915+dadachann@users.noreply.github.com>
2026-07-30 21:43:35 +08:00

6.7 KiB
Raw Blame History

LangBot Cloud 24 小时资源 Soak 门禁

scripts/cloud_runtime_soak.py 是生产候选拓扑的最终资源稳定性门禁。它不替代单元测试、历史 churn 探针或 nsjail 隔离测试;它把以下三类证据按同一时间轴采集并给出可机读的 pass/fail:

  • Core、Plugin Runtime 和 Box Runtime 的 HTTP liveness/readiness。
  • 三个 Python 进程的 event-loop recent max/p95 调度延迟。
  • Linux /proc 进程树的 current RSS、累计 CPU、线程、文件描述符和子进程数。
  • cgroup v2 的 memory.current/peak/events、swap、CPU usage/throttling、PID current/events 和实际硬限制。

生产批准必须使用 cgroup 证据。--pid 只适合本地诊断,因为进程指标无法证明 OOM kill、PID limit 或 CPU throttling。

运行位置

建议把采集器放在独立的 node agent 或监控 sidecar 中,并只读挂载三个目标容器的 cgroup 路径。不要把采集器和样本文件放进被测容器自己的 cgroup/数据卷,否则采集器的 CPU、内存和 page cache 会污染目标数据。

Kubernetes/containerd 生成的 cgroup 路径不是稳定 API。每次生产候选部署都必须从实际 pod/container ID 解析,不能从 pod 名猜路径。传入的每个目录都必须至少可读:

  • memory.currentmemory.events
  • cpu.statcpu.max
  • pids.currentpids.events

最终门禁应加 --require-hard-limits。该选项要求每个目标 cgroup 都能观察到有限的 CPU quota、memory、swap 和 PID 上限;任一值为 max 都失败。

标准 24 小时命令

uv run python scripts/cloud_runtime_soak.py \
  --duration 24h \
  --startup-grace 5m \
  --sample-interval 15s \
  --cooldown 30m \
  --analysis-window 30m \
  --http-timeout 5s \
  --max-memory-growth-mib 64 \
  --max-memory-slope-mib-per-hour 32 \
  --max-tail-cpu-cores 0.5 \
  --max-throttled-period-ratio 0.25 \
  --max-event-loop-lag-ms 1000 \
  --max-event-loop-p95-lag-ms 250 \
  --require-hard-limits \
  --endpoint core=http://langbot:5300/healthz \
  --endpoint plugin=http://langbot-plugin-runtime:5400/healthz \
  --endpoint box=http://langbot-box:5410/readyz \
  --cgroup core=/host-cgroup/CURRENT_CORE_CONTAINER \
  --cgroup plugin=/host-cgroup/CURRENT_PLUGIN_RUNTIME_CONTAINER \
  --cgroup box=/host-cgroup/CURRENT_BOX_RUNTIME_CONTAINER \
  --samples-file artifacts/cloud-soak-samples.jsonl \
  --report-file artifacts/cloud-soak-report.json \
  --workload uv run python tests/load/cloud_candidate_workload.py

--duration 是包含启动观察、负载和冷却期的最大墙钟时间。工作负载必须在截止时间前退出并至少留出 30 分钟冷却;否则门禁会终止负载并失败。若负载由外部系统控制,可以省略 --workload,但必须保证最后 --analysis-window 完全无测试流量,该窗口才可解释为空闲尾段。

工作负载命令的 stdout/stderr 会转发到采集器 stderr,不会混入 stdout 的最终 JSON 报告。命令以独立 process group 启动;超时或中断时整组收到 TERM,10 秒后仍未退出则收到 KILL。

凭据只能通过 workload 进程环境或 secret mount 注入,不能放在命令参数中。最终报告只记录可执行文件名和参数个数,不保存参数正文;采集器也拒绝带 userinfo、query 或 fragment 的健康 URL。

必须覆盖的负载

同一候选版本至少要覆盖:

  1. 大批 Workspace 注册、成员邀请、登录和 entitlement 刷新。
  2. Plugin installation reconcile、依赖准备、正常调用、进程崩溃与重启。
  3. Dashboard/Embed/平台 WebSocket 建连、突发消息和批量断连。
  4. Box session、文件同步、并发 exec、managed-process 输出和清理。
  5. PostgreSQL pool 接近容量、事务超时和恢复。
  6. Core、Plugin Runtime、Box 分别收到 SIGTERM 后的优雅重启。

工作负载不能把 API 过载拒绝当作成功吞掉。默认情况下,Core health 中 blocking executor 的 global/scope rejection counter 只要增长,门禁即失败;只有专门验证“过载会正确返回 429”的独立测试才可以使用 --allow-rejections,该次运行不能作为生产批准证据。

判定规则

整个有效观察期内出现以下任一情况即失败:

  • 健康接口请求失败、非 2xx、Core code != 0,或 Box ready=false
  • memory.events.high/max/oom/oom_kill/oom_group_kill 增长。
  • pids.events.max 增长。
  • cgroup 单调计数器回退,表示目标很可能发生了未记录的重启或 cgroup 替换。
  • CPU throttled-period ratio 超过配置阈值。
  • 任一健康采样窗口的 event-loop recent max 超过 1 秒,或冷却尾段 recent p95 超过 250 ms。
  • 健康接口缺少 event-loop monitor、monitor 未持续运行,或其 sample counter 回退。
  • blocking executor rejection counter 增长。
  • Plugin Runtime restart circuit 的累计打开次数增长。
  • Core 目录 active Workspace、最近 snapshot/delta Workspace 或 membership 基数 超过各自配置上限,或 PostgreSQL checked_out 超过配置 pool 容量;相关 current/max 指标只出现一半或 max 非法也失败。

负载结束后的冷却尾段还必须满足:

  • memory.current/RSS 的稳健首尾增长和线性斜率不能同时超过阈值。
  • 平均 CPU 核数不超过 --max-tail-cpu-cores
  • event-loop recent p95 不超过 --max-event-loop-p95-lag-ms
  • blocking executor pending 至少回到过零;不能整个尾段持续积压。
  • Plugin Runtime restart coordinator 的 active launch、half-open probe 和 circuit open remaining time 必须回到零,gate_waiters 必须至少归零一次。
  • Core 的 MCP projection retirement queue/worker 和 message aggregation buffer/scope 必须至少归零一次。
  • telemetry、QueryPool、MCP host/dispatch、Box creating/closing/background 等临时 gauge 不能继续增长。

内存判定要求“增长量”和“斜率”同时越界,避免几 MiB allocator/page-cache 噪声在短窗口被外推成很大的每小时斜率。最终报告仍保留实际增长与斜率,人工审查时不能只看 verdict。

产物与退出码

  • --samples-file:逐样本 JSONL,写入后立即 flush,供时序图和故障定位。
  • --report-file:最终汇总、阈值、资源硬限制、OOM/PID/throttle delta、尾段斜率和 workload 状态。
  • stdout:与 report 文件相同的最终 JSONworkload 日志只写 stderr。

退出码:

  • 0:全部门禁通过。
  • 1:采样完成但资源门禁失败。
  • 2CLI 参数或目标配置错误。

必须保存原始 JSONL、最终报告、三个镜像 digest、LangBot/SDK commit、生产配置摘要和工作负载版本。滚动更新、节点迁移或镜像变化后,旧报告不能继续作为新候选版本的批准证据。