- render the banner through the Alert API so its grid keeps one full-width text column instead of wrapping inside the empty icon column
- point the cloud link at space.langbot.app/cloud?environment=dedicated and the self-hosted link at the GitHub repository
- keep the last known state and retry the probe (1/2/4/8s plus a re-check on window focus) so a restarting backend cannot hide the banner for the session
- publish --beta-banner-height so the fixed sidebar container, the sidebar wrapper and the wizard shell start below the banner
- Collapse the retention policy section by default, showing a one-line summary when collapsed.
- Add the owner/admin-only list_operation_logs assistant tool backed by audit.view.
- Enforce per-tool permissions on both tool listing and invocation.
- Add a substring search filter (search) to query_logs and the tool so callers can
ask narrow questions (plugin name, account, verb) instead of dumping the full history.
- Project records to a compact shape so a page stays under the assistant result budget.
- Localize the new assistant tool label and retention summary across all locales.
The panel answered "who, when, what" poorly: the actual before → after diff
was hidden behind a click, the actor was a muted line, and a whole row went to
the route while duration and hash sat in a permanent right column. The data
already arrives with the page, so the panel now reads top-down in the order an
operator actually looks:
* who and when first, then the action and resource, then the diff spelled out
inline (long values collapse to their size so a pipeline config does not
render as two near-identical blobs), with route / method / status / duration
/ evidence hash moved into the expanded diagnostics line;
* the whole card is the expand toggle with a rotating chevron, which removes
the ambiguity of a small text link whose expanded state looked identical;
* records are bucketed into today / yesterday / dated groups so a long history
reads as a timeline;
* a "changes only" toggle (reusing the level filter) answers "who changed
something" in one click -- in the live log 489 of 499 rows are pure views,
so this is the difference between a haystack and the ten rows that matter.
Tracing fixes found while reading the live log:
* PUT /pipelines/<uuid>/extensions was classified as a plugin config change
because the generic ('/extensions') rule ran first; it now maps to a
pipeline-specific action;
* that handler recorded no diff at all, so "modified extension config" could
never say what changed. It now captures the previous bindings and reports the
before → after.
Cache: the integrity result is now keyed by (Workspace, listing filters) so
each view reuses its own verification, and a delete invalidates all views for
the Workspace. Fourteen tests cover the projection, the verifier, the cache
lifecycle and the failure modes.
- pageOf was inserted under the wrong namespace (the locale files already had a
hideDetails in toolCalls), so the UI rendered the raw key. Removed it and
replaced the dropdown with a numbered pager (first/last always visible, a
one-page window around the current page, gaps as an ellipsis).
- Restore the right-hand column (duration + record hash) that the previous
declutter removed; the digest is still opt-in, but the hash stays visible.
- Hide the status code on success: a 200 is the default, so only >=400 shows.
The list was noisy and hard to read: a constant "verified" badge and the
"L1/L2" level badge on every row carried no information, the same "view X"
line repeated until it buried real changes, and each row rendered its full
payload diff inline.
- Drop the per-row level badge and the "verified" badge; keep only an
exception badge, so a healthy row shows nothing extra.
- Collapse consecutive identical rows into one line with an "xN" count.
- Make the change digest opt-in per row (one expanded at a time) and show a
"<n> changes" toggle instead of dumping the diff.
- Show the registered route under each row so "Resource / System" is no
longer ambiguous about what actually happened.
- Replace the arrows-only pager with a page selector ("Page x / y") plus the
arrows, for logs with hundreds of records.
Verified: tsc, eslint, prettier and check-i18n all pass.
- Move the operation-traceability service and its HTTP routes into a standalone
pkg/operation_trace package; the Core controller package no longer depends on it
(routes stay auto-discovered by the controller package scan).
- Gate the hot path on one module-level boolean exposed as ap.operation_trace_active:
while no Workspace has tracing enabled the route wrapper performs zero trace work
and zero database round trips. The gate opens when a Workspace enables tracing and
is primed once at startup for already-enabled Workspaces.
- Record the installed extension identity for GitHub / marketplace / local / upload
installs (owner/repo, author/name, filename) so the trace no longer shows a
nameless "a plugin was installed" event and the card names the extension.
- Merge migrations 0032 + 0033 into a single 0032_operation_traceability revision
carrying the full table shape and an index set that matches the ORM.
Tests: unit 5165 passed / integration 358 passed; ruff check + format clean.
Granularity follow-up: a record used to say only "installed an extension".
- Add resolve_resource_identity(): derive the acted-on resource from trusted route
params or the install-style payload (plugin_author/plugin_name), so every
resource family names its target. Unknown keys are ignored and sensitive keys
are skipped, keeping the audit trail bounded and secret-free.
- The route wrapper applies it to every traced request; the identity only fills
in when a handler did not publish a more precise one.
- Publish field-level before/after diffs for plugin config, skill, knowledge
base, MCP server, pipeline, provider and model updates, reusing
changed_fields()/build_summary() so secret-looking keys stay redacted and a
masked round-trip never shows up as a phantom change.
- Count hash mismatches / broken links over the whole filtered history instead
of only the returned page, so the badges no longer reset to zero on page 2.
- Classify failing record ids during the same scan and expose them through an
integrity=all|issues|hash_mismatch|chain_broken listing filter.
- Report scanned_count / scan_truncated so an approximate (row-capped) summary
is never mistaken for a clean one.
- Resolve the acted-on resource identity from route params or the request body,
giving every resource family a "which plugin/skill/knowledge base" trace.
- Render resource types and the two counters through i18n, keeping all eight
locale files key- and line-aligned.
- Export download failed because the URL builder appended the path to a
"/" base, producing a protocol-relative "//api/..." URL that the browser
read as host "api"; the request never reached the backend. Resolve the
base to the current origin, mirroring the other URL builders.
- Remove the orphan ``clear_prune`` action: the branch has no prune/clear
route and no route rule maps to it, so it could never be produced. Its
i18n key is removed too.
- Remove ``tamperedCount``, now unreferenced after the verification badge
was split into integrity/chain counters. The operationTrace catalogue is
now free of unproduced keys.
- Correct the settings controller docstring: only governance, operation
logs, filters and export routes exist; there is no on-demand delete.
Verified: no producer for clear_prune, classify() can no longer emit it,
operationTrace orphan scan returns none, tsc/check-i18n/prettier/py_compile
all pass.
- Group the toolbar actions into a single right-aligned cluster so
"Download logs" and "Refresh" sit together instead of being spread apart
by the toolbar's space-between layout.
- Align the record header's filter row with the rest of the panel: replace
the ``ml-auto`` push (which left the filters misaligned once they wrapped
onto a second line) with ``justify-between``.
- Capture scope: treat the ``audit`` bucket (viewing the log / tracing
settings) as a read-level observation instead of recording it from the
mutation level. At "mutations only" the log no longer fills with GET page
views; views are recorded only at the read level.
- Coverage: classify plugin lifecycle (install / upgrade / config / page),
skills, knowledge bases and MCP servers, with dedicated actions and
resource types, so those operations are traced instead of falling into the
generic resource bucket. Route rules now support multi-fragment AND
matching and per-verb read/write actions.
- Panel: move "Download logs" next to "Refresh" in the toolbar.
- Verification: count hash mismatches and broken links separately instead of
collapsing both into one "tampered" signal.
- i18n: add the new action labels plus the split verification labels in all
eight locales.
Performance of the operation traceability surface:
- Cache the retention / row-budget / dedupe policy per Workspace and
invalidate it on governance change, so the hot write path no longer
issues three metadata round trips per recorded operation.
- Resolve the governance payload with one metadata query instead of four.
- Turn the dedupe probe into an existence check (LIMIT 1) instead of a
full COUNT over the dedupe window.
- Defer the row-budget COUNT to every N inserts; prune still enforces the
exact budget, and the maintenance loop keeps it bounded.
- Panel: fetch only the records page when paging or filtering instead of
re-requesting governance and the filter catalogue every time.
- Export: include the effective filters (and limit) in the artifact so a
downloaded CSV is self-describing.
Fixes and cleanup:
- Drop the unused OperationLogPruneReport import that broke `tsc`.
- i18n: remove the 9 redundant blank lines unique to en-US and format the
six non-compliant locale files with the project Prettier config; all
eight locales keep identical key sets, key order and placeholders.
The bot session monitor referenced several translation keys that did not
exist in any locale file. Because i18next `fallbackLng` is `zh-Hans`, the
missing keys fell back to the hardcoded English `defaultValue`, leaking
untranslated strings (e.g. "0 sessions", "User ID or name", "Search") into
every non-English UI.
- Add the missing keys to all 8 locales with proper translations:
`common.search`, `bots.sessionMonitor.{totalSessions,userSearch,startDate,endDate}`
and `monitoring.toolCalls.{showDetails,hideDetails}`.
- Remove the hardcoded `defaultValue` fallbacks in `BotSessionMonitor.tsx`
so translations are always driven by the locale files.
- Reorder nested keys in every locale file to match the `en-US.ts` reference
order, keeping values untouched. This removes all structural drift between
locale files (previously `guidedTour.*`, `plugins.*`, etc. were ordered
differently).
Verified: `scripts/check-i18n.mjs` reports matching keys for all locales,
every key/value pair is preserved, Prettier passes, and `tsc --noEmit` is clean.
The invitation error copy also renders as a transient sonner toast, so the
plain text locator resolved to two elements and tripped Playwright strict
mode while the toast was on screen.
Tag the inline error region with data-testid="invitation-error" and assert
against it, keeping the check deterministic under CI parallelism.
Run prettier over the assistant files and the new dock unit test so the
web lint job passes, and remove the unused ASSISTANT_BUTTON_SIZE import
flagged by code quality review.
Opening the panel left the user at the oldest message because the scroll
effect did not depend on `open`. Add it and defer the scroll to the next
frame so the popover has laid out before the sentinel is measured.
A click or an aborted swipe never arms the long press, so no drop handler
ran and a hover-expanded button stayed expanded after the pointer left.
Track real pointer presence and re-sync the rail on release.
`shouldCollapseRail` was keyed on `!railExpanded`, mirroring the expand
helper instead of negating it. Once a hover revealed the button it could
never collapse back into the rail, so the collapse looked broken after the
first restore.
Key the predicate on `railExpanded` being true, restore the mirror
relationship with `shouldExpandRail`, and add regression coverage for the
expand -> leave -> collapse round trip. Also lengthen the rail strip.
Resolve conflicts in the workspace assistant integration:
- DynamicFormItemComponent: keep HEAD's compact model selector and
`disabled` field support while adopting master's `sortModelsByCatalog`
ordering and `MODEL_SELECT_TRIGGER_CLASS`.
- i18n (en-US, ja-JP, zh-Hans): keep both the `assistant` namespace from
HEAD and master's `sidebarGuide` / `pipelineMigration` additions.
Add 0030_merge_assistant to join the assistant conversations branch with
the released chain so the migration graph converges on a single head.
Rename it from 0030_merge_assistant_conversations to stay within the 32
character revision limit enforced by test_migrations.
Also add assistant button docking (long-press drag, edge collapse to a
short blue rail, hover restore) with pure helpers in assistant-dock.ts
and unit coverage, plus auto-collapse for finished tool result cards.
Resolve the revision-graph conflict introduced by the release line.
Conflicts
- tests/integration/persistence/test_rag_document_identity.py: master had
independently introduced the same dynamic-head helper (`current_head`) plus
a `DOCUMENT_IDENTITY_REVISION` constant, and it explicitly upgrades to that
revision. Keep master's semantics (the explicit revision matters because
upgrading to the head now also traverses the TOTP branch) while keeping the
module-level alembic imports on this branch.
New migration
- 0030_merge_totp_into_release joins the TOTP branch's own merge revision
(0026_merge_totp_and_rag_identity) with the release head
(0029_merge_rag_identity). Both reached 0025_rag_document_identity without
including each other, which left Alembic with two heads and made
`upgrade head` fail with "Multiple head revisions are present".
Verification
- Single Alembic head confirmed (0030_merge_totp_into_release).
- tests/integration/persistence/test_rag_document_identity.py: 34 passed.
- Full fast integration suite: 356 passed, 84 skipped.
resetPassword() unconditionally added `method` and an empty `totp_code` to
every request, which broke the Playwright smoke test that asserts the
recovery-key flow posts exactly {user, recovery_key, new_password}.
Send only the second-factor fields that apply to the selected method, so the
default recovery-key flow stays byte-compatible with existing callers while
the totp / recovery_code methods still carry their inputs.
Address the migration and formatting failures reported on the TOTP branch.
Migrations
- Register totp_credentials and totp_recovery_codes in
_ALEMBIC_TENANT_TABLES. On a legacy PostgreSQL install these two tables
reference users.uuid, which only exists after 0009, so create_all() must
not run ahead of Alembic the way it did for the other tenant tables.
Without this the PostgreSQL migration test failed with
"column uuid referenced in foreign key constraint does not exist".
Tests
- Resolve the Alembic head dynamically in the RAG document identity
regression instead of pinning 0025_rag_document_identity. The TOTP and
RAG branches now meet at a merge revision, so the pinned value was no
longer the head. This matches the convention already used by
test_migrations_postgres.
Formatting
- Apply ruff format to the new backend modules and prettier to the locale
files and TOTP components, so the lint jobs pass.
Add a per-Account time-based one-time password (TOTP) second sign-in
factor, plus the owner/admin tooling needed to operate it.
Backend
- TotpService: enrolment, constant-time verification with a +/- one step
drift window, single-use recovery codes and disablement.
* shared secret is only ever persisted as a Fernet token whose key is
derived (HKDF-SHA256) from the instance JWT secret and gated by a
key_version epoch;
* recovery codes are only ever persisted as salted
PBKDF2-HMAC-SHA256 digests (600k iterations) and are single-use;
* the consumed counter advances monotonically so a captured code cannot
be replayed inside the same time window.
- New persistence entities and alembic migrations for credentials and
recovery codes.
- Login second-factor challenge, bound to the Account that passed the
password step.
- Owner/admin oversight endpoints to inspect and re-bind the second
factor of any Account.
Frontend
- Account settings panel for enrolling, managing, re-binding and
disabling the second factor.
- Login second-factor step and the matching client methods.
- Strings for all eight locales.
The shared secret never crosses the API boundary: enrolment returns a
server-rendered QR code (as a data: URL) and the plaintext secret is
discarded as soon as the image is produced.
Addresses review feedback on the marketplace installed-state change.
- Progress: byte-derived progress no longer has time-based drift layered on
top, and fallback drift is clamped to the current stage range, so the bar
cannot exceed the download band. 90/100 bytes at 40s elapsed used to report
61% against a declared 5-45% band; it now reports 41%. Drift is measured
from when the stage was entered rather than from task start.
- Skills: a marketplace skill is no longer reported installed from a bare
skill name. The backend names skills from their own SKILL.md and records no
publisher, so alice/review and bob/review install identically; matching the
bare name marked every publisher's skill as installed. Skill cards now
resolve to not-installed until the installed skill carries a publisher.
- Stages: "installing plugin dependencies" and "launching plugin" described
work this task context cannot observe (installation persistence ran under
the former; the runtime installs dependencies and starts the plugin inside
apply_plugin_installation under the latter). They become "persisting the
installation" and "installing or starting plugin", and the frontend maps
that combined step to the dependency stage rather than the launch stage.
- Removed the per-dependency progress fields: the backend never populated
them, so the UI could never have displayed them.
The stage mapping and progress maths move to install-progress.ts, and the
installed-state matching to a React-free marketplace-installed.ts, so both
are covered by executable tests (+9).
Verified: ruff, tsc --noEmit, prettier --check, eslint (0 errors), 98/98 unit
tests.
Marketplace cards now reflect whether an extension is already installed in
the current workspace, and the installed-extension list gains a search box.
Backend (stream install progress):
- _read_httpx_response_limited gains an optional task_context: it publishes
download_total from Content-Length before the first chunk and updates
download_current / download_speed per chunk. The marketplace download path
previously had no progress reporting; it now matches the GitHub path.
- _marketplace_get forwards task_context to that helper.
- install_plugin resets the per-install counters so re-installing the same
plugin cannot inherit stale metadata, and reports human-readable stages:
preparing -> downloading -> inspecting -> storing -> installing
dependencies -> launching -> waiting for plugin to become ready.
Frontend (installed state):
- New marketplace-installed helper normalises the sidebar identities
(plugin: author/name, mcp: author__name, skill: bare name) into one
type:author/name index and resolves a card's installed state from it.
useMarketplaceInstalledIndex memoises on the sidebar lists, so a finished
install (which refreshes the sidebar) re-evaluates the cards automatically.
- PluginMarketCardVO carries installed / hasUpdate. An installed extension
turns its download affordance into a hollow green ring with a green check
in place, instead of adding a separate badge; the count slot switches to
the installed label. Cards with an available update use amber.
- PluginMarketComponent derives the annotated list and shares the index with
RecommendationLists.
Frontend (install task UI):
- mapActionToStage matches the new connector stage strings. The pre-download
stages are checked before the generic "install" match, because
"preparing plugin install" also contains "install".
- Stage progress ranges are non-overlapping; overall progress interpolates on
real byte counts while downloading and drifts monotonically elsewhere,
capped at 99%.
- The progress dialog and task queue expose the launching stage.
Frontend (installed list search):
- The installed list had no search at all. A query box in the page header
filters by label / name / author / description, case-insensitively, applied
before grouping so grouped and flat views both honour it.
- Search misses and an empty list now show distinct empty states, with a
clear action on a search miss.
- AsyncTask entity gains the optional created_at field.
i18n: new marketplace / install / search strings across all 8 locales.
Verified: ruff format + check, tsc --noEmit, prettier --check, eslint
(0 errors), and 89/89 frontend unit tests.