The panel answered "who, when, what" poorly: the actual before → after diff
was hidden behind a click, the actor was a muted line, and a whole row went to
the route while duration and hash sat in a permanent right column. The data
already arrives with the page, so the panel now reads top-down in the order an
operator actually looks:
* who and when first, then the action and resource, then the diff spelled out
inline (long values collapse to their size so a pipeline config does not
render as two near-identical blobs), with route / method / status / duration
/ evidence hash moved into the expanded diagnostics line;
* the whole card is the expand toggle with a rotating chevron, which removes
the ambiguity of a small text link whose expanded state looked identical;
* records are bucketed into today / yesterday / dated groups so a long history
reads as a timeline;
* a "changes only" toggle (reusing the level filter) answers "who changed
something" in one click -- in the live log 489 of 499 rows are pure views,
so this is the difference between a haystack and the ten rows that matter.
Tracing fixes found while reading the live log:
* PUT /pipelines/<uuid>/extensions was classified as a plugin config change
because the generic ('/extensions') rule ran first; it now maps to a
pipeline-specific action;
* that handler recorded no diff at all, so "modified extension config" could
never say what changed. It now captures the previous bindings and reports the
before → after.
Cache: the integrity result is now keyed by (Workspace, listing filters) so
each view reuses its own verification, and a delete invalidates all views for
the Workspace. Fourteen tests cover the projection, the verifier, the cache
lifecycle and the failure modes.
The integrity scan selected whole rows, so a cold pass paged in the
changes/detail payloads and the client fingerprint for up to
MAX_INTEGRITY_SCAN_ROWS rows the verifier never reads. Project only the hash
columns, and add a lightweight verifier so the scan no longer builds a full
display dict per row.
Repair the read cache so it is a latency shield, not a correctness shortcut:
a result computed within INTEGRITY_CACHE_TTL_SECONDS is served from the
per-Workspace cache (opening, refreshing and paging all land inside that
window and pay nothing), while a cache miss re-verifies the whole window. An
incremental scan that skips previously verified ids can never see an edit to
an already-cached row -- exactly the tampering this feature exists to expose.
Verified against a live 414-row log: 50 content edits and 7 re-signed links
are all detected, including an edit to a row verified on a previous pass.
Also fix two correctness gaps and one maintenance bug:
- age-based prune deleted rows without invalidating the cached prefix;
- the boundary baseline is now read only when the history exceeds the scan
window, which a bounded scan makes the rare case;
- the operation-log retention block had drifted outside the per-binding loop
in the maintenance task, so only the last discovered Workspace was ever
pruned while the others grew unbounded.
Tests: scan projection, verifier parity, TTL cache reuse, cache-miss
re-verification, edit to an already-verified row, hash mismatch, chain break
and prune invalidation.
Reading the panel re-hashed up to MAX_INTEGRITY_SCAN_ROWS (20000) rows on
every open, refresh and page turn -- measured ~280ms for 385 rows, growing
linearly with history -- which is what made the log feel slow. The chain is
append-only, so a verified prefix stays valid: keep a per-Workspace cache and,
once warm, verify only the rows appended since (bounded window), reusing the
prior result across reads. Deletions (prune / row budget) invalidate it.
Also fix the cache-hit path (it referenced a field that was never stored) and
the "scanned" count (it double-counted reused rows).
Measured on 385 rows: cold ~280ms -> warm ~0.7ms -> incremental ~15ms.
The pull request merge produced two Alembic heads: this branch's
0032_operation_traceability and master's 0032_cert_artifact_digest both
descended from 0031_merge_totp_assistant, so the integration and migration
jobs failed with "Multiple head revisions". Renumber the traceability
migration to 0033 and point it at master's head (0032_cert_artifact_digest),
restoring a single linear head.
Also replace the noqa'd unused import of the traceability routes with an
importlib bootstrap (addresses the CodeQL unused-import finding) and merge
origin/master to pick up the certification digest backfill.
- pageOf was inserted under the wrong namespace (the locale files already had a
hideDetails in toolCalls), so the UI rendered the raw key. Removed it and
replaced the dropdown with a numbered pager (first/last always visible, a
one-page window around the current page, gaps as an ellipsis).
- Restore the right-hand column (duration + record hash) that the previous
declutter removed; the digest is still opt-in, but the hash stays visible.
- Hide the status code on success: a 200 is the default, so only >=400 shows.
The list was noisy and hard to read: a constant "verified" badge and the
"L1/L2" level badge on every row carried no information, the same "view X"
line repeated until it buried real changes, and each row rendered its full
payload diff inline.
- Drop the per-row level badge and the "verified" badge; keep only an
exception badge, so a healthy row shows nothing extra.
- Collapse consecutive identical rows into one line with an "xN" count.
- Make the change digest opt-in per row (one expanded at a time) and show a
"<n> changes" toggle instead of dumping the diff.
- Show the registered route under each row so "Resource / System" is no
longer ambiguous about what actually happened.
- Replace the arrows-only pager with a page selector ("Page x / y") plus the
arrows, for logs with hundreds of records.
Verified: tsc, eslint, prettier and check-i18n all pass.
Tracing ran inline with the request: every traced call did several database
round trips (dedupe probe, actor-name lookup, newest-hash read) and a
serialized insert. With tracing on, the WebUI's burst of parallel requests
(login alone touches a dozen endpoints) saturated the write path and stalled
the whole service.
Now the request path only builds a bounded dict and pushes it onto a queue
(measured ~0.005ms per trace, zero database work). A single background writer
coroutine drains the queue and persists sequentially, which also removes the
read-newest-hash race that used to produce false "chain broken" reports; the
previous asyncio lock is gone. The actor display-name lookup and the dedupe
key hash move into the writer as well. The queue is bounded: when traffic
outruns the writer the oldest pending trace is dropped instead of slowing the
request down.
Verified: migration tests pass, ruff clean, a focused writer test shows
enqueue is non-blocking, rows link with zero chain breaks and no internal
key leaks into the stored row.
- Chain integrity: the "read newest hash, then insert" pair in record() could be
interleaved by concurrent traces (the panel opens a dozen parallel GETs), so two
rows linked to the same predecessor and one looked like a broken chain even
though nobody edited anything. Serialize the pair with a per-service lock.
- Resource classification: generic verbs fell back to resource_type='resource', so
almost every read showed up as "View resource / Resource" (pipelines, bots,
providers, users, monitoring...). Derive the family from the registered route
identity (pipeline, bot, model_provider, llm_model, monitoring, user, workspace,
webhook, api_key, system).
- Frontend: render the resource family for the new types (8 locales) and keep the
resource_id badge added earlier.
Verified: backend integration 358 passed, ruff check/format clean; frontend tsc,
eslint, prettier and check-i18n all pass.
- Move the operation-traceability service and its HTTP routes into a standalone
pkg/operation_trace package; the Core controller package no longer depends on it
(routes stay auto-discovered by the controller package scan).
- Gate the hot path on one module-level boolean exposed as ap.operation_trace_active:
while no Workspace has tracing enabled the route wrapper performs zero trace work
and zero database round trips. The gate opens when a Workspace enables tracing and
is primed once at startup for already-enabled Workspaces.
- Record the installed extension identity for GitHub / marketplace / local / upload
installs (owner/repo, author/name, filename) so the trace no longer shows a
nameless "a plugin was installed" event and the card names the extension.
- Merge migrations 0032 + 0033 into a single 0032_operation_traceability revision
carrying the full table shape and an index set that matches the ORM.
Tests: unit 5165 passed / integration 358 passed; ruff check + format clean.
- Remove the unused request_summary local: the stored summary is derived by the
service from changes + the classified action, so handlers publishing
quart.g.operation_log_summary was a dead channel that never reached the row.
- Drop the unused datetime import in the operation log model.
- Apply ruff format to the touched files so `ruff format --check` passes.
Granularity follow-up: a record used to say only "installed an extension".
- Add resolve_resource_identity(): derive the acted-on resource from trusted route
params or the install-style payload (plugin_author/plugin_name), so every
resource family names its target. Unknown keys are ignored and sensitive keys
are skipped, keeping the audit trail bounded and secret-free.
- The route wrapper applies it to every traced request; the identity only fills
in when a handler did not publish a more precise one.
- Publish field-level before/after diffs for plugin config, skill, knowledge
base, MCP server, pipeline, provider and model updates, reusing
changed_fields()/build_summary() so secret-looking keys stay redacted and a
masked round-trip never shows up as a phantom change.
- Count hash mismatches / broken links over the whole filtered history instead
of only the returned page, so the badges no longer reset to zero on page 2.
- Classify failing record ids during the same scan and expose them through an
integrity=all|issues|hash_mismatch|chain_broken listing filter.
- Report scanned_count / scan_truncated so an approximate (row-capped) summary
is never mistaken for a clean one.
- Resolve the acted-on resource identity from route params or the request body,
giving every resource family a "which plugin/skill/knowledge base" trace.
- Render resource types and the two counters through i18n, keeping all eight
locale files key- and line-aligned.
- Export download failed because the URL builder appended the path to a
"/" base, producing a protocol-relative "//api/..." URL that the browser
read as host "api"; the request never reached the backend. Resolve the
base to the current origin, mirroring the other URL builders.
- Remove the orphan ``clear_prune`` action: the branch has no prune/clear
route and no route rule maps to it, so it could never be produced. Its
i18n key is removed too.
- Remove ``tamperedCount``, now unreferenced after the verification badge
was split into integrity/chain counters. The operationTrace catalogue is
now free of unproduced keys.
- Correct the settings controller docstring: only governance, operation
logs, filters and export routes exist; there is no on-demand delete.
Verified: no producer for clear_prune, classify() can no longer emit it,
operationTrace orphan scan returns none, tsc/check-i18n/prettier/py_compile
all pass.
Tracing is now opt-in: a Workspace records nothing until an owner or admin
explicitly turns it on, so auditing stays off the hot path by default.
View and configure remain limited to owners and admins: the audit.view
permission is granted to those roles only, non-owner/admin callers receive a
403 from the governance write route, and the panel already gates both on the
owner/admin membership role.
- Group the toolbar actions into a single right-aligned cluster so
"Download logs" and "Refresh" sit together instead of being spread apart
by the toolbar's space-between layout.
- Align the record header's filter row with the rest of the panel: replace
the ``ml-auto`` push (which left the filters misaligned once they wrapped
onto a second line) with ``justify-between``.
- Capture scope: treat the ``audit`` bucket (viewing the log / tracing
settings) as a read-level observation instead of recording it from the
mutation level. At "mutations only" the log no longer fills with GET page
views; views are recorded only at the read level.
- Coverage: classify plugin lifecycle (install / upgrade / config / page),
skills, knowledge bases and MCP servers, with dedicated actions and
resource types, so those operations are traced instead of falling into the
generic resource bucket. Route rules now support multi-fragment AND
matching and per-verb read/write actions.
- Panel: move "Download logs" next to "Refresh" in the toolbar.
- Verification: count hash mismatches and broken links separately instead of
collapsing both into one "tampered" signal.
- i18n: add the new action labels plus the split verification labels in all
eight locales.
Performance of the operation traceability surface:
- Cache the retention / row-budget / dedupe policy per Workspace and
invalidate it on governance change, so the hot write path no longer
issues three metadata round trips per recorded operation.
- Resolve the governance payload with one metadata query instead of four.
- Turn the dedupe probe into an existence check (LIMIT 1) instead of a
full COUNT over the dedupe window.
- Defer the row-budget COUNT to every N inserts; prune still enforces the
exact budget, and the maintenance loop keeps it bounded.
- Panel: fetch only the records page when paging or filtering instead of
re-requesting governance and the filter catalogue every time.
- Export: include the effective filters (and limit) in the artifact so a
downloaded CSV is self-describing.
Fixes and cleanup:
- Drop the unused OperationLogPruneReport import that broke `tsc`.
- i18n: remove the 9 redundant blank lines unique to en-US and format the
six non-compliant locale files with the project Prettier config; all
eight locales keep identical key sets, key order and placeholders.
A saved pipeline may carry the post-migration shape
{'ai': {'runner': {'id': ''}, 'runner_config': {}}} while no runner has been
selected yet. The planner matched the empty id against the certified
'plugin:author/name/component' form, failed, and returned a malformed-id
blocker, so the whole 'migrate all' batch stopped with both pipelines reported
as blocked even though there was no legacy section and nothing to convert.
A blank id with no legacy runner section now reports not_legacy: no work is
needed and no target is synthesized. A blank id that still coexists with a
legacy section remains an explicit invalid_runner_id blocker, because that
ambiguous state must be resolved by the operator rather than guessed.
The certified-archive admission gate treated any declared certificate it could
not resolve as an untrusted archive and rejected the install with
CERTIFIED_PLUGIN_OSS_FORCE_REQUIRED. OSS ships an empty
plugin.certification.trusted_public_keys ring and the marketplace signs every
package with its own issuer key, so all certified marketplace packages (for
example langbot-team/RunnerDemo and langbot-team/LocalAgent) failed at the
'validating plugin package' step before artifact storage.
Admission now distinguishes an unresolvable declaration from a configured trust
decision that fails:
- record certificate_id only when the key_id is actually present in the
configured ring, so an empty ring yields an unresolvable declaration;
- OSS degrades an unresolvable declaration to the existing oss_dev dedicated
profile with CERTIFIED_PLUGIN_OSS_UNTRUSTED_DEDICATED instead of blocking;
- a resolvable declaration that still fails keeps requiring the explicit
administrator force (CERTIFIED_PLUGIN_OSS_FORCE_REQUIRED);
- Cloud stays fail-closed and rejects before storage.
No shared-runtime privilege is granted when the issuer is not trusted, so this
withholds isolation rather than escalating it.
The bot session monitor paging controls use `common.next` and
`common.previous`, but those keys never existed in any locale. They were
only rendered because of the hardcoded `defaultValue: 'Next'`/`'Previous'`.
Removing those fallbacks (previous commit) made i18next render the raw key
name, so the Playwright smoke tests could no longer find the "Next" button.
Add `common.previous` and `common.next` to all 8 locales (reusing each
language's existing `guidedTour.previous`/`next` wording) and re-align the
locale key order with the `en-US` reference.
Verified locally:
- scripts/check-i18n.mjs passes for all locales
- every t('...') key referenced by BotSessionMonitor.tsx now exists in en-US
- tests/e2e/bot-session-tool-timeline.spec.ts: 13/13 passed (including the
two request-recovery cases that failed in CI)
- prettier / tsc / eslint clean
The bot session monitor referenced several translation keys that did not
exist in any locale file. Because i18next `fallbackLng` is `zh-Hans`, the
missing keys fell back to the hardcoded English `defaultValue`, leaking
untranslated strings (e.g. "0 sessions", "User ID or name", "Search") into
every non-English UI.
- Add the missing keys to all 8 locales with proper translations:
`common.search`, `bots.sessionMonitor.{totalSessions,userSearch,startDate,endDate}`
and `monitoring.toolCalls.{showDetails,hideDetails}`.
- Remove the hardcoded `defaultValue` fallbacks in `BotSessionMonitor.tsx`
so translations are always driven by the locale files.
- Reorder nested keys in every locale file to match the `en-US.ts` reference
order, keeping values untouched. This removes all structural drift between
locale files (previously `guidedTour.*`, `plugins.*`, etc. were ordered
differently).
Verified: `scripts/check-i18n.mjs` reports matching keys for all locales,
every key/value pair is preserved, Prettier passes, and `tsc --noEmit` is clean.
The invitation error copy also renders as a transient sonner toast, so the
plain text locator resolved to two elements and tripped Playwright strict
mode while the toast was on screen.
Tag the inline error region with data-testid="invitation-error" and assert
against it, keeping the check deterministic under CI parallelism.