scan_skill_directory is disabled for Cloud-managed sandboxes, but the
register_skill tool was still advertised to the model whenever a backend was
available, so every call failed closed without a usable explanation. Advertise
register_skill only when the runtime may scan sandbox directories, and keep
activate available in both modes.
The workspace file tools (read/write/edit/glob/grep) ran a script through
`python`, which a sandbox built from a host rootfs may not provide: those hosts
ship `python3` only, and a read-only /usr prevents creating a `python` shim
inside the jail. Every file tool then failed with `/bin/sh: 1: python: not
found`. Resolve the interpreter with `command -v python3 || command -v python`,
matching the interpreter the Box service already uses for its own scripts.
Passkeys and passwords belong to the Account, but those routes required a Workspace -- which the WebUI never sends for them -- and then resolved the caller by looking the login name up as an email address. Every Account whose login name is not its email answered 404, so the passkey panel could neither add nor list a key.
The routes now authenticate with the account token and address the Account the caller authenticated as, matching the self-service second-factor routes in the same file. /space/bind-authorize-url carried the same Workspace requirement.
The login page also surfaced a TypeError instead of the passkey login failure message when an assertion was rejected.
Cloud runtime refuses unscoped database access and workspace_operation_logs is row-level-secured on the bound Workspace. The background writer, the retention pass and the metadata reads all run outside a request lifetime, so each of them failed and the failure was swallowed by the best-effort handler: enabling tracing recorded nothing at all.
The route wrapper also traces after the handler released its scope, so level resolution has to carry the Workspace itself. Without it the level read back as 0 and the request was dropped before it was even queued.
The startup gate now enumerates the Workspaces through this instance's execution bindings, because a cross-Workspace metadata scan cannot see them under RLS, and it runs after the Workspace service exists. The writer starts with a clean context so it cannot inherit the scope of the request that first enqueued a trace, and a dropped write is reported instead of being logged only at debug.
The second-factor listing was instance-wide and the three mutating endpoints never checked that the target Account belongs to the caller's Workspace, so an owner or admin could see and reset another tenant's account. The listing now resolves the Workspace members and the mutations answer 404 for a foreign account, which does not disclose that it exists.
Also fixes the TOTP write path on Postgres: six write points passed timezone-aware datetimes into timestamp-without-time-zone columns, which asyncpg rejects, so every enroll, confirm, revoke and recovery-code consumption raised a 500 there. SQLite tolerated the same values.
The panel derives its oversight view from the API authorization instead of the length of the account list, which would have hidden it in a single-member Workspace.
The panel answered "who, when, what" poorly: the actual before → after diff
was hidden behind a click, the actor was a muted line, and a whole row went to
the route while duration and hash sat in a permanent right column. The data
already arrives with the page, so the panel now reads top-down in the order an
operator actually looks:
* who and when first, then the action and resource, then the diff spelled out
inline (long values collapse to their size so a pipeline config does not
render as two near-identical blobs), with route / method / status / duration
/ evidence hash moved into the expanded diagnostics line;
* the whole card is the expand toggle with a rotating chevron, which removes
the ambiguity of a small text link whose expanded state looked identical;
* records are bucketed into today / yesterday / dated groups so a long history
reads as a timeline;
* a "changes only" toggle (reusing the level filter) answers "who changed
something" in one click -- in the live log 489 of 499 rows are pure views,
so this is the difference between a haystack and the ten rows that matter.
Tracing fixes found while reading the live log:
* PUT /pipelines/<uuid>/extensions was classified as a plugin config change
because the generic ('/extensions') rule ran first; it now maps to a
pipeline-specific action;
* that handler recorded no diff at all, so "modified extension config" could
never say what changed. It now captures the previous bindings and reports the
before → after.
Cache: the integrity result is now keyed by (Workspace, listing filters) so
each view reuses its own verification, and a delete invalidates all views for
the Workspace. Fourteen tests cover the projection, the verifier, the cache
lifecycle and the failure modes.
The integrity scan selected whole rows, so a cold pass paged in the
changes/detail payloads and the client fingerprint for up to
MAX_INTEGRITY_SCAN_ROWS rows the verifier never reads. Project only the hash
columns, and add a lightweight verifier so the scan no longer builds a full
display dict per row.
Repair the read cache so it is a latency shield, not a correctness shortcut:
a result computed within INTEGRITY_CACHE_TTL_SECONDS is served from the
per-Workspace cache (opening, refreshing and paging all land inside that
window and pay nothing), while a cache miss re-verifies the whole window. An
incremental scan that skips previously verified ids can never see an edit to
an already-cached row -- exactly the tampering this feature exists to expose.
Verified against a live 414-row log: 50 content edits and 7 re-signed links
are all detected, including an edit to a row verified on a previous pass.
Also fix two correctness gaps and one maintenance bug:
- age-based prune deleted rows without invalidating the cached prefix;
- the boundary baseline is now read only when the history exceeds the scan
window, which a bounded scan makes the rare case;
- the operation-log retention block had drifted outside the per-binding loop
in the maintenance task, so only the last discovered Workspace was ever
pruned while the others grew unbounded.
Tests: scan projection, verifier parity, TTL cache reuse, cache-miss
re-verification, edit to an already-verified row, hash mismatch, chain break
and prune invalidation.
A saved pipeline may carry the post-migration shape
{'ai': {'runner': {'id': ''}, 'runner_config': {}}} while no runner has been
selected yet. The planner matched the empty id against the certified
'plugin:author/name/component' form, failed, and returned a malformed-id
blocker, so the whole 'migrate all' batch stopped with both pipelines reported
as blocked even though there was no legacy section and nothing to convert.
A blank id with no legacy runner section now reports not_legacy: no work is
needed and no target is synthesized. A blank id that still coexists with a
legacy section remains an explicit invalid_runner_id blocker, because that
ambiguous state must be resolved by the operator rather than guessed.
The certified-archive admission gate treated any declared certificate it could
not resolve as an untrusted archive and rejected the install with
CERTIFIED_PLUGIN_OSS_FORCE_REQUIRED. OSS ships an empty
plugin.certification.trusted_public_keys ring and the marketplace signs every
package with its own issuer key, so all certified marketplace packages (for
example langbot-team/RunnerDemo and langbot-team/LocalAgent) failed at the
'validating plugin package' step before artifact storage.
Admission now distinguishes an unresolvable declaration from a configured trust
decision that fails:
- record certificate_id only when the key_id is actually present in the
configured ring, so an empty ring yields an unresolvable declaration;
- OSS degrades an unresolvable declaration to the existing oss_dev dedicated
profile with CERTIFIED_PLUGIN_OSS_UNTRUSTED_DEDICATED instead of blocking;
- a resolvable declaration that still fails keeps requiring the explicit
administrator force (CERTIFIED_PLUGIN_OSS_FORCE_REQUIRED);
- Cloud stays fail-closed and rejects before storage.
No shared-runtime privilege is granted when the issuer is not trusted, so this
withholds isolation rather than escalating it.
Resolve the revision-graph conflict introduced by the release line.
Conflicts
- tests/integration/persistence/test_rag_document_identity.py: master had
independently introduced the same dynamic-head helper (`current_head`) plus
a `DOCUMENT_IDENTITY_REVISION` constant, and it explicitly upgrades to that
revision. Keep master's semantics (the explicit revision matters because
upgrading to the head now also traverses the TOTP branch) while keeping the
module-level alembic imports on this branch.
New migration
- 0030_merge_totp_into_release joins the TOTP branch's own merge revision
(0026_merge_totp_and_rag_identity) with the release head
(0029_merge_rag_identity). Both reached 0025_rag_document_identity without
including each other, which left Alembic with two heads and made
`upgrade head` fail with "Multiple head revisions are present".
Verification
- Single Alembic head confirmed (0030_merge_totp_into_release).
- tests/integration/persistence/test_rag_document_identity.py: 34 passed.
- Full fast integration suite: 356 passed, 84 skipped.
Address the migration and formatting failures reported on the TOTP branch.
Migrations
- Register totp_credentials and totp_recovery_codes in
_ALEMBIC_TENANT_TABLES. On a legacy PostgreSQL install these two tables
reference users.uuid, which only exists after 0009, so create_all() must
not run ahead of Alembic the way it did for the other tenant tables.
Without this the PostgreSQL migration test failed with
"column uuid referenced in foreign key constraint does not exist".
Tests
- Resolve the Alembic head dynamically in the RAG document identity
regression instead of pinning 0025_rag_document_identity. The TOTP and
RAG branches now meet at a merge revision, so the pinned value was no
longer the head. This matches the convention already used by
test_migrations_postgres.
Formatting
- Apply ruff format to the new backend modules and prettier to the locale
files and TOTP components, so the lint jobs pass.
* fix(pipeline): stop emitting empty final assistant messages
An empty final streaming chunk was appended as an empty message chain when no sandbox outbox attachment was present, so platforms received an empty reply. Only append the chain when attachments were actually collected, and gate the "Call ..." tool notice behind output.misc.track-function-calls so it no longer becomes the final chunk of a stream.
* fix(wecombot): do not close stream on a blank final snapshot
A blank final snapshot closed the WeCom stream with an empty bubble and pushed the real answer into a separate reply_text message. Skip blank final snapshots and keep the session open so the following non-blank chunk can finalize it. Covered by new regression tests.
Preserve the published timezone-less schema and epoch API contract across run lifecycle, deadlines, leases, heartbeat registry, events, transcripts and retention cutoffs. Exercise the real Host journal with asyncpg under a non-superuser role on metadata and published-migration schemas; 32 PostgreSQL regressions fail before the fix and pass after it.