Passkeys and passwords belong to the Account, but those routes required a Workspace -- which the WebUI never sends for them -- and then resolved the caller by looking the login name up as an email address. Every Account whose login name is not its email answered 404, so the passkey panel could neither add nor list a key.
The routes now authenticate with the account token and address the Account the caller authenticated as, matching the self-service second-factor routes in the same file. /space/bind-authorize-url carried the same Workspace requirement.
The login page also surfaced a TypeError instead of the passkey login failure message when an assertion was rejected.
The second-factor listing was instance-wide and the three mutating endpoints never checked that the target Account belongs to the caller's Workspace, so an owner or admin could see and reset another tenant's account. The listing now resolves the Workspace members and the mutations answer 404 for a foreign account, which does not disclose that it exists.
Also fixes the TOTP write path on Postgres: six write points passed timezone-aware datetimes into timestamp-without-time-zone columns, which asyncpg rejects, so every enroll, confirm, revoke and recovery-code consumption raised a 500 there. SQLite tolerated the same values.
The panel derives its oversight view from the API authorization instead of the length of the account list, which would have hidden it in a single-member Workspace.
Rename the traceability surface to "operation log", move its entry from the sidebar into the settings dialog, and collect the capture level and the retention policy in one dialog. The retention fields are shown inline instead of behind a disclosure; edits stay a draft until Save, and Reset restores the persisted policy. Copy updated across the eight locales.
- render the banner through the Alert API so its grid keeps one full-width text column instead of wrapping inside the empty icon column
- point the cloud link at space.langbot.app/cloud?environment=dedicated and the self-hosted link at the GitHub repository
- keep the last known state and retry the probe (1/2/4/8s plus a re-check on window focus) so a restarting backend cannot hide the banner for the session
- publish --beta-banner-height so the fixed sidebar container, the sidebar wrapper and the wizard shell start below the banner
Systematically cross-checked all 263 route/method pairs against the classifier
and fixed every case where a real change could be mislabelled:
- POST /skills and PUT /skills/<name> were classified as skill_view (read
bucket), so creating or editing a skill was silently unlogged at the mutation
level. They now record skill_update.
- The public webhook ingress (/bots/<uuid>) and the embedded chat widget
(/embed/*) are unauthenticated visitor traffic, not Workspace changes; they
are dropped instead of logged as create/resource.
- Uploading files and indexing documents (data flowing into a resource) are
offered as observations so they never fill the log; deleting a knowledge-base
file is destructive and now records file_delete.
- A plugin's own config-file edit/delete is a plugin change, not an opaque
resource delete; ordered before the /config rule that also matched it.
- The bucket now comes from the action table (plus 'a read verb is always a
read'), so a write action can no longer be forced into the read bucket.
- Added agents and files resource families; localized every new action/resource
type in all nine locales.
- Split the route rules on the HTTP method so a DELETE is classified as a
destructive action instead of falling into the read/write bucket. Uninstalling
a plugin or skill and deleting a knowledge base or MCP server were recorded as
a view/update and left no readable trace.
- Add plugin_uninstall, skill_uninstall, knowledge_base_delete and mcp_delete.
- Skip the assistant's own conversation traffic: a chat session produced opaque
create/resource rows that answered nothing.
- Localize the new actions and the assistant_conversation resource type in all locales.
- Sync the assistant tool action list and apply ruff formatting.
- Collapse the retention policy section by default, showing a one-line summary when collapsed.
- Add the owner/admin-only list_operation_logs assistant tool backed by audit.view.
- Enforce per-tool permissions on both tool listing and invocation.
- Add a substring search filter (search) to query_logs and the tool so callers can
ask narrow questions (plugin name, account, verb) instead of dumping the full history.
- Project records to a compact shape so a page stays under the assistant result budget.
- Localize the new assistant tool label and retention summary across all locales.
The panel answered "who, when, what" poorly: the actual before → after diff
was hidden behind a click, the actor was a muted line, and a whole row went to
the route while duration and hash sat in a permanent right column. The data
already arrives with the page, so the panel now reads top-down in the order an
operator actually looks:
* who and when first, then the action and resource, then the diff spelled out
inline (long values collapse to their size so a pipeline config does not
render as two near-identical blobs), with route / method / status / duration
/ evidence hash moved into the expanded diagnostics line;
* the whole card is the expand toggle with a rotating chevron, which removes
the ambiguity of a small text link whose expanded state looked identical;
* records are bucketed into today / yesterday / dated groups so a long history
reads as a timeline;
* a "changes only" toggle (reusing the level filter) answers "who changed
something" in one click -- in the live log 489 of 499 rows are pure views,
so this is the difference between a haystack and the ten rows that matter.
Tracing fixes found while reading the live log:
* PUT /pipelines/<uuid>/extensions was classified as a plugin config change
because the generic ('/extensions') rule ran first; it now maps to a
pipeline-specific action;
* that handler recorded no diff at all, so "modified extension config" could
never say what changed. It now captures the previous bindings and reports the
before → after.
Cache: the integrity result is now keyed by (Workspace, listing filters) so
each view reuses its own verification, and a delete invalidates all views for
the Workspace. Fourteen tests cover the projection, the verifier, the cache
lifecycle and the failure modes.
- pageOf was inserted under the wrong namespace (the locale files already had a
hideDetails in toolCalls), so the UI rendered the raw key. Removed it and
replaced the dropdown with a numbered pager (first/last always visible, a
one-page window around the current page, gaps as an ellipsis).
- Restore the right-hand column (duration + record hash) that the previous
declutter removed; the digest is still opt-in, but the hash stays visible.
- Hide the status code on success: a 200 is the default, so only >=400 shows.
The list was noisy and hard to read: a constant "verified" badge and the
"L1/L2" level badge on every row carried no information, the same "view X"
line repeated until it buried real changes, and each row rendered its full
payload diff inline.
- Drop the per-row level badge and the "verified" badge; keep only an
exception badge, so a healthy row shows nothing extra.
- Collapse consecutive identical rows into one line with an "xN" count.
- Make the change digest opt-in per row (one expanded at a time) and show a
"<n> changes" toggle instead of dumping the diff.
- Show the registered route under each row so "Resource / System" is no
longer ambiguous about what actually happened.
- Replace the arrows-only pager with a page selector ("Page x / y") plus the
arrows, for logs with hundreds of records.
Verified: tsc, eslint, prettier and check-i18n all pass.
- Chain integrity: the "read newest hash, then insert" pair in record() could be
interleaved by concurrent traces (the panel opens a dozen parallel GETs), so two
rows linked to the same predecessor and one looked like a broken chain even
though nobody edited anything. Serialize the pair with a per-service lock.
- Resource classification: generic verbs fell back to resource_type='resource', so
almost every read showed up as "View resource / Resource" (pipelines, bots,
providers, users, monitoring...). Derive the family from the registered route
identity (pipeline, bot, model_provider, llm_model, monitoring, user, workspace,
webhook, api_key, system).
- Frontend: render the resource family for the new types (8 locales) and keep the
resource_id badge added earlier.
Verified: backend integration 358 passed, ruff check/format clean; frontend tsc,
eslint, prettier and check-i18n all pass.
- Move the operation-traceability service and its HTTP routes into a standalone
pkg/operation_trace package; the Core controller package no longer depends on it
(routes stay auto-discovered by the controller package scan).
- Gate the hot path on one module-level boolean exposed as ap.operation_trace_active:
while no Workspace has tracing enabled the route wrapper performs zero trace work
and zero database round trips. The gate opens when a Workspace enables tracing and
is primed once at startup for already-enabled Workspaces.
- Record the installed extension identity for GitHub / marketplace / local / upload
installs (owner/repo, author/name, filename) so the trace no longer shows a
nameless "a plugin was installed" event and the card names the extension.
- Merge migrations 0032 + 0033 into a single 0032_operation_traceability revision
carrying the full table shape and an index set that matches the ORM.
Tests: unit 5165 passed / integration 358 passed; ruff check + format clean.
Granularity follow-up: a record used to say only "installed an extension".
- Add resolve_resource_identity(): derive the acted-on resource from trusted route
params or the install-style payload (plugin_author/plugin_name), so every
resource family names its target. Unknown keys are ignored and sensitive keys
are skipped, keeping the audit trail bounded and secret-free.
- The route wrapper applies it to every traced request; the identity only fills
in when a handler did not publish a more precise one.
- Publish field-level before/after diffs for plugin config, skill, knowledge
base, MCP server, pipeline, provider and model updates, reusing
changed_fields()/build_summary() so secret-looking keys stay redacted and a
masked round-trip never shows up as a phantom change.
- Count hash mismatches / broken links over the whole filtered history instead
of only the returned page, so the badges no longer reset to zero on page 2.
- Classify failing record ids during the same scan and expose them through an
integrity=all|issues|hash_mismatch|chain_broken listing filter.
- Report scanned_count / scan_truncated so an approximate (row-capped) summary
is never mistaken for a clean one.
- Resolve the acted-on resource identity from route params or the request body,
giving every resource family a "which plugin/skill/knowledge base" trace.
- Render resource types and the two counters through i18n, keeping all eight
locale files key- and line-aligned.
- Export download failed because the URL builder appended the path to a
"/" base, producing a protocol-relative "//api/..." URL that the browser
read as host "api"; the request never reached the backend. Resolve the
base to the current origin, mirroring the other URL builders.
- Remove the orphan ``clear_prune`` action: the branch has no prune/clear
route and no route rule maps to it, so it could never be produced. Its
i18n key is removed too.
- Remove ``tamperedCount``, now unreferenced after the verification badge
was split into integrity/chain counters. The operationTrace catalogue is
now free of unproduced keys.
- Correct the settings controller docstring: only governance, operation
logs, filters and export routes exist; there is no on-demand delete.
Verified: no producer for clear_prune, classify() can no longer emit it,
operationTrace orphan scan returns none, tsc/check-i18n/prettier/py_compile
all pass.
- Group the toolbar actions into a single right-aligned cluster so
"Download logs" and "Refresh" sit together instead of being spread apart
by the toolbar's space-between layout.
- Align the record header's filter row with the rest of the panel: replace
the ``ml-auto`` push (which left the filters misaligned once they wrapped
onto a second line) with ``justify-between``.
- Capture scope: treat the ``audit`` bucket (viewing the log / tracing
settings) as a read-level observation instead of recording it from the
mutation level. At "mutations only" the log no longer fills with GET page
views; views are recorded only at the read level.
- Coverage: classify plugin lifecycle (install / upgrade / config / page),
skills, knowledge bases and MCP servers, with dedicated actions and
resource types, so those operations are traced instead of falling into the
generic resource bucket. Route rules now support multi-fragment AND
matching and per-verb read/write actions.
- Panel: move "Download logs" next to "Refresh" in the toolbar.
- Verification: count hash mismatches and broken links separately instead of
collapsing both into one "tampered" signal.
- i18n: add the new action labels plus the split verification labels in all
eight locales.
Performance of the operation traceability surface:
- Cache the retention / row-budget / dedupe policy per Workspace and
invalidate it on governance change, so the hot write path no longer
issues three metadata round trips per recorded operation.
- Resolve the governance payload with one metadata query instead of four.
- Turn the dedupe probe into an existence check (LIMIT 1) instead of a
full COUNT over the dedupe window.
- Defer the row-budget COUNT to every N inserts; prune still enforces the
exact budget, and the maintenance loop keeps it bounded.
- Panel: fetch only the records page when paging or filtering instead of
re-requesting governance and the filter catalogue every time.
- Export: include the effective filters (and limit) in the artifact so a
downloaded CSV is self-describing.
Fixes and cleanup:
- Drop the unused OperationLogPruneReport import that broke `tsc`.
- i18n: remove the 9 redundant blank lines unique to en-US and format the
six non-compliant locale files with the project Prettier config; all
eight locales keep identical key sets, key order and placeholders.
The bot session monitor paging controls use `common.next` and
`common.previous`, but those keys never existed in any locale. They were
only rendered because of the hardcoded `defaultValue: 'Next'`/`'Previous'`.
Removing those fallbacks (previous commit) made i18next render the raw key
name, so the Playwright smoke tests could no longer find the "Next" button.
Add `common.previous` and `common.next` to all 8 locales (reusing each
language's existing `guidedTour.previous`/`next` wording) and re-align the
locale key order with the `en-US` reference.
Verified locally:
- scripts/check-i18n.mjs passes for all locales
- every t('...') key referenced by BotSessionMonitor.tsx now exists in en-US
- tests/e2e/bot-session-tool-timeline.spec.ts: 13/13 passed (including the
two request-recovery cases that failed in CI)
- prettier / tsc / eslint clean
The bot session monitor referenced several translation keys that did not
exist in any locale file. Because i18next `fallbackLng` is `zh-Hans`, the
missing keys fell back to the hardcoded English `defaultValue`, leaking
untranslated strings (e.g. "0 sessions", "User ID or name", "Search") into
every non-English UI.
- Add the missing keys to all 8 locales with proper translations:
`common.search`, `bots.sessionMonitor.{totalSessions,userSearch,startDate,endDate}`
and `monitoring.toolCalls.{showDetails,hideDetails}`.
- Remove the hardcoded `defaultValue` fallbacks in `BotSessionMonitor.tsx`
so translations are always driven by the locale files.
- Reorder nested keys in every locale file to match the `en-US.ts` reference
order, keeping values untouched. This removes all structural drift between
locale files (previously `guidedTour.*`, `plugins.*`, etc. were ordered
differently).
Verified: `scripts/check-i18n.mjs` reports matching keys for all locales,
every key/value pair is preserved, Prettier passes, and `tsc --noEmit` is clean.
The invitation error copy also renders as a transient sonner toast, so the
plain text locator resolved to two elements and tripped Playwright strict
mode while the toast was on screen.
Tag the inline error region with data-testid="invitation-error" and assert
against it, keeping the check deterministic under CI parallelism.
The workspace assistant added 39 keys under `assistant.*` to en-US,
zh-Hans and ja-JP, but es-ES, ru-RU, th-TH, vi-VN and zh-Hant were never
updated. The i18n key consistency check fails on all five, blocking the
release PR.
Add the full assistant namespace with translated copy so every locale
matches the en-US reference. No source changes; keys and placeholders
(`{{count}}`) mirror en-US exactly.
Run prettier over the assistant files and the new dock unit test so the
web lint job passes, and remove the unused ASSISTANT_BUTTON_SIZE import
flagged by code quality review.
Opening the panel left the user at the oldest message because the scroll
effect did not depend on `open`. Add it and defer the scroll to the next
frame so the popover has laid out before the sentinel is measured.
A click or an aborted swipe never arms the long press, so no drop handler
ran and a hover-expanded button stayed expanded after the pointer left.
Track real pointer presence and re-sync the rail on release.
`shouldCollapseRail` was keyed on `!railExpanded`, mirroring the expand
helper instead of negating it. Once a hover revealed the button it could
never collapse back into the rail, so the collapse looked broken after the
first restore.
Key the predicate on `railExpanded` being true, restore the mirror
relationship with `shouldExpandRail`, and add regression coverage for the
expand -> leave -> collapse round trip. Also lengthen the rail strip.
Resolve conflicts in the workspace assistant integration:
- DynamicFormItemComponent: keep HEAD's compact model selector and
`disabled` field support while adopting master's `sortModelsByCatalog`
ordering and `MODEL_SELECT_TRIGGER_CLASS`.
- i18n (en-US, ja-JP, zh-Hans): keep both the `assistant` namespace from
HEAD and master's `sidebarGuide` / `pipelineMigration` additions.
Add 0030_merge_assistant to join the assistant conversations branch with
the released chain so the migration graph converges on a single head.
Rename it from 0030_merge_assistant_conversations to stay within the 32
character revision limit enforced by test_migrations.
Also add assistant button docking (long-press drag, edge collapse to a
short blue rail, hover restore) with pure helpers in assistant-dock.ts
and unit coverage, plus auto-collapse for finished tool result cards.
Resolve the revision-graph conflict introduced by the release line.
Conflicts
- tests/integration/persistence/test_rag_document_identity.py: master had
independently introduced the same dynamic-head helper (`current_head`) plus
a `DOCUMENT_IDENTITY_REVISION` constant, and it explicitly upgrades to that
revision. Keep master's semantics (the explicit revision matters because
upgrading to the head now also traverses the TOTP branch) while keeping the
module-level alembic imports on this branch.
New migration
- 0030_merge_totp_into_release joins the TOTP branch's own merge revision
(0026_merge_totp_and_rag_identity) with the release head
(0029_merge_rag_identity). Both reached 0025_rag_document_identity without
including each other, which left Alembic with two heads and made
`upgrade head` fail with "Multiple head revisions are present".
Verification
- Single Alembic head confirmed (0030_merge_totp_into_release).
- tests/integration/persistence/test_rag_document_identity.py: 34 passed.
- Full fast integration suite: 356 passed, 84 skipped.
resetPassword() unconditionally added `method` and an empty `totp_code` to
every request, which broke the Playwright smoke test that asserts the
recovery-key flow posts exactly {user, recovery_key, new_password}.
Send only the second-factor fields that apply to the selected method, so the
default recovery-key flow stays byte-compatible with existing callers while
the totp / recovery_code methods still carry their inputs.
Address the migration and formatting failures reported on the TOTP branch.
Migrations
- Register totp_credentials and totp_recovery_codes in
_ALEMBIC_TENANT_TABLES. On a legacy PostgreSQL install these two tables
reference users.uuid, which only exists after 0009, so create_all() must
not run ahead of Alembic the way it did for the other tenant tables.
Without this the PostgreSQL migration test failed with
"column uuid referenced in foreign key constraint does not exist".
Tests
- Resolve the Alembic head dynamically in the RAG document identity
regression instead of pinning 0025_rag_document_identity. The TOTP and
RAG branches now meet at a merge revision, so the pinned value was no
longer the head. This matches the convention already used by
test_migrations_postgres.
Formatting
- Apply ruff format to the new backend modules and prettier to the locale
files and TOTP components, so the lint jobs pass.
Add a per-Account time-based one-time password (TOTP) second sign-in
factor, plus the owner/admin tooling needed to operate it.
Backend
- TotpService: enrolment, constant-time verification with a +/- one step
drift window, single-use recovery codes and disablement.
* shared secret is only ever persisted as a Fernet token whose key is
derived (HKDF-SHA256) from the instance JWT secret and gated by a
key_version epoch;
* recovery codes are only ever persisted as salted
PBKDF2-HMAC-SHA256 digests (600k iterations) and are single-use;
* the consumed counter advances monotonically so a captured code cannot
be replayed inside the same time window.
- New persistence entities and alembic migrations for credentials and
recovery codes.
- Login second-factor challenge, bound to the Account that passed the
password step.
- Owner/admin oversight endpoints to inspect and re-bind the second
factor of any Account.
Frontend
- Account settings panel for enrolling, managing, re-binding and
disabling the second factor.
- Login second-factor step and the matching client methods.
- Strings for all eight locales.
The shared secret never crosses the API boundary: enrolment returns a
server-rendered QR code (as a data: URL) and the plaintext secret is
discarded as soon as the image is produced.
Addresses review feedback on the marketplace installed-state change.
- Progress: byte-derived progress no longer has time-based drift layered on
top, and fallback drift is clamped to the current stage range, so the bar
cannot exceed the download band. 90/100 bytes at 40s elapsed used to report
61% against a declared 5-45% band; it now reports 41%. Drift is measured
from when the stage was entered rather than from task start.
- Skills: a marketplace skill is no longer reported installed from a bare
skill name. The backend names skills from their own SKILL.md and records no
publisher, so alice/review and bob/review install identically; matching the
bare name marked every publisher's skill as installed. Skill cards now
resolve to not-installed until the installed skill carries a publisher.
- Stages: "installing plugin dependencies" and "launching plugin" described
work this task context cannot observe (installation persistence ran under
the former; the runtime installs dependencies and starts the plugin inside
apply_plugin_installation under the latter). They become "persisting the
installation" and "installing or starting plugin", and the frontend maps
that combined step to the dependency stage rather than the launch stage.
- Removed the per-dependency progress fields: the backend never populated
them, so the UI could never have displayed them.
The stage mapping and progress maths move to install-progress.ts, and the
installed-state matching to a React-free marketplace-installed.ts, so both
are covered by executable tests (+9).
Verified: ruff, tsc --noEmit, prettier --check, eslint (0 errors), 98/98 unit
tests.