scan_skill_directory is disabled for Cloud-managed sandboxes, but the
register_skill tool was still advertised to the model whenever a backend was
available, so every call failed closed without a usable explanation. Advertise
register_skill only when the runtime may scan sandbox directories, and keep
activate available in both modes.
The workspace file tools (read/write/edit/glob/grep) ran a script through
`python`, which a sandbox built from a host rootfs may not provide: those hosts
ship `python3` only, and a read-only /usr prevents creating a `python` shim
inside the jail. Every file tool then failed with `/bin/sh: 1: python: not
found`. Resolve the interpreter with `command -v python3 || command -v python`,
matching the interpreter the Box service already uses for its own scripts.
A full-text search sends no query vector by design, but search() validated
len([]) == 0 as the expected embedding dimension, so the request failed with
'Embedding dimension 0 is not enabled for this deployment' and hid the
backend capability error. An empty query vector now means no expected
dimension to verify.
The NDJSON debug stream emits its frames after the request handler returned,
so the request-scoped tenant scope is already closed by then. The deferred
execution therefore failed with TenantScopeRequiredError on its first
persistence access and the client only saw a generic runner_error frame with
no log entry.
Re-enter the Workspace tenant scope around the deferred execution and log
unexpected failures with their traceback.
Pick up the shared-worker startup and installation-binding fixes released as
langbot-plugin 0.7.5 (tag v0.7.5), so Core deployments stop installing the
0.7.4 wheel that cannot launch shared pool workers and drops the installation
binding on plugin dispatches and relays.
Passkeys and passwords belong to the Account, but those routes required a Workspace -- which the WebUI never sends for them -- and then resolved the caller by looking the login name up as an email address. Every Account whose login name is not its email answered 404, so the passkey panel could neither add nor list a key.
The routes now authenticate with the account token and address the Account the caller authenticated as, matching the self-service second-factor routes in the same file. /space/bind-authorize-url carried the same Workspace requirement.
The login page also surfaced a TypeError instead of the passkey login failure message when an assertion was rejected.
Cloud runtime refuses unscoped database access and workspace_operation_logs is row-level-secured on the bound Workspace. The background writer, the retention pass and the metadata reads all run outside a request lifetime, so each of them failed and the failure was swallowed by the best-effort handler: enabling tracing recorded nothing at all.
The route wrapper also traces after the handler released its scope, so level resolution has to carry the Workspace itself. Without it the level read back as 0 and the request was dropped before it was even queued.
The startup gate now enumerates the Workspaces through this instance's execution bindings, because a cross-Workspace metadata scan cannot see them under RLS, and it runs after the Workspace service exists. The writer starts with a clean context so it cannot inherit the scope of the request that first enqueued a trace, and a dropped write is reported instead of being logged only at debug.
The second-factor listing was instance-wide and the three mutating endpoints never checked that the target Account belongs to the caller's Workspace, so an owner or admin could see and reset another tenant's account. The listing now resolves the Workspace members and the mutations answer 404 for a foreign account, which does not disclose that it exists.
Also fixes the TOTP write path on Postgres: six write points passed timezone-aware datetimes into timestamp-without-time-zone columns, which asyncpg rejects, so every enroll, confirm, revoke and recovery-code consumption raised a 500 there. SQLite tolerated the same values.
The panel derives its oversight view from the API authorization instead of the length of the account list, which would have hidden it in a single-member Workspace.
Rename the traceability surface to "operation log", move its entry from the sidebar into the settings dialog, and collect the capture level and the retention policy in one dialog. The retention fields are shown inline instead of behind a disclosure; edits stay a draft until Save, and Reset restores the persisted policy. Copy updated across the eight locales.
- render the banner through the Alert API so its grid keeps one full-width text column instead of wrapping inside the empty icon column
- point the cloud link at space.langbot.app/cloud?environment=dedicated and the self-hosted link at the GitHub repository
- keep the last known state and retry the probe (1/2/4/8s plus a re-check on window focus) so a restarting backend cannot hide the banner for the session
- publish --beta-banner-height so the fixed sidebar container, the sidebar wrapper and the wizard shell start below the banner
Systematically cross-checked all 263 route/method pairs against the classifier
and fixed every case where a real change could be mislabelled:
- POST /skills and PUT /skills/<name> were classified as skill_view (read
bucket), so creating or editing a skill was silently unlogged at the mutation
level. They now record skill_update.
- The public webhook ingress (/bots/<uuid>) and the embedded chat widget
(/embed/*) are unauthenticated visitor traffic, not Workspace changes; they
are dropped instead of logged as create/resource.
- Uploading files and indexing documents (data flowing into a resource) are
offered as observations so they never fill the log; deleting a knowledge-base
file is destructive and now records file_delete.
- A plugin's own config-file edit/delete is a plugin change, not an opaque
resource delete; ordered before the /config rule that also matched it.
- The bucket now comes from the action table (plus 'a read verb is always a
read'), so a write action can no longer be forced into the read bucket.
- Added agents and files resource families; localized every new action/resource
type in all nine locales.
- Split the route rules on the HTTP method so a DELETE is classified as a
destructive action instead of falling into the read/write bucket. Uninstalling
a plugin or skill and deleting a knowledge base or MCP server were recorded as
a view/update and left no readable trace.
- Add plugin_uninstall, skill_uninstall, knowledge_base_delete and mcp_delete.
- Skip the assistant's own conversation traffic: a chat session produced opaque
create/resource rows that answered nothing.
- Localize the new actions and the assistant_conversation resource type in all locales.
- Sync the assistant tool action list and apply ruff formatting.
- Collapse the retention policy section by default, showing a one-line summary when collapsed.
- Add the owner/admin-only list_operation_logs assistant tool backed by audit.view.
- Enforce per-tool permissions on both tool listing and invocation.
- Add a substring search filter (search) to query_logs and the tool so callers can
ask narrow questions (plugin name, account, verb) instead of dumping the full history.
- Project records to a compact shape so a page stays under the assistant result budget.
- Localize the new assistant tool label and retention summary across all locales.
The panel answered "who, when, what" poorly: the actual before → after diff
was hidden behind a click, the actor was a muted line, and a whole row went to
the route while duration and hash sat in a permanent right column. The data
already arrives with the page, so the panel now reads top-down in the order an
operator actually looks:
* who and when first, then the action and resource, then the diff spelled out
inline (long values collapse to their size so a pipeline config does not
render as two near-identical blobs), with route / method / status / duration
/ evidence hash moved into the expanded diagnostics line;
* the whole card is the expand toggle with a rotating chevron, which removes
the ambiguity of a small text link whose expanded state looked identical;
* records are bucketed into today / yesterday / dated groups so a long history
reads as a timeline;
* a "changes only" toggle (reusing the level filter) answers "who changed
something" in one click -- in the live log 489 of 499 rows are pure views,
so this is the difference between a haystack and the ten rows that matter.
Tracing fixes found while reading the live log:
* PUT /pipelines/<uuid>/extensions was classified as a plugin config change
because the generic ('/extensions') rule ran first; it now maps to a
pipeline-specific action;
* that handler recorded no diff at all, so "modified extension config" could
never say what changed. It now captures the previous bindings and reports the
before → after.
Cache: the integrity result is now keyed by (Workspace, listing filters) so
each view reuses its own verification, and a delete invalidates all views for
the Workspace. Fourteen tests cover the projection, the verifier, the cache
lifecycle and the failure modes.
The integrity scan selected whole rows, so a cold pass paged in the
changes/detail payloads and the client fingerprint for up to
MAX_INTEGRITY_SCAN_ROWS rows the verifier never reads. Project only the hash
columns, and add a lightweight verifier so the scan no longer builds a full
display dict per row.
Repair the read cache so it is a latency shield, not a correctness shortcut:
a result computed within INTEGRITY_CACHE_TTL_SECONDS is served from the
per-Workspace cache (opening, refreshing and paging all land inside that
window and pay nothing), while a cache miss re-verifies the whole window. An
incremental scan that skips previously verified ids can never see an edit to
an already-cached row -- exactly the tampering this feature exists to expose.
Verified against a live 414-row log: 50 content edits and 7 re-signed links
are all detected, including an edit to a row verified on a previous pass.
Also fix two correctness gaps and one maintenance bug:
- age-based prune deleted rows without invalidating the cached prefix;
- the boundary baseline is now read only when the history exceeds the scan
window, which a bounded scan makes the rare case;
- the operation-log retention block had drifted outside the per-binding loop
in the maintenance task, so only the last discovered Workspace was ever
pruned while the others grew unbounded.
Tests: scan projection, verifier parity, TTL cache reuse, cache-miss
re-verification, edit to an already-verified row, hash mismatch, chain break
and prune invalidation.
Reading the panel re-hashed up to MAX_INTEGRITY_SCAN_ROWS (20000) rows on
every open, refresh and page turn -- measured ~280ms for 385 rows, growing
linearly with history -- which is what made the log feel slow. The chain is
append-only, so a verified prefix stays valid: keep a per-Workspace cache and,
once warm, verify only the rows appended since (bounded window), reusing the
prior result across reads. Deletions (prune / row budget) invalidate it.
Also fix the cache-hit path (it referenced a field that was never stored) and
the "scanned" count (it double-counted reused rows).
Measured on 385 rows: cold ~280ms -> warm ~0.7ms -> incremental ~15ms.