The panel answered "who, when, what" poorly: the actual before → after diff
was hidden behind a click, the actor was a muted line, and a whole row went to
the route while duration and hash sat in a permanent right column. The data
already arrives with the page, so the panel now reads top-down in the order an
operator actually looks:
* who and when first, then the action and resource, then the diff spelled out
inline (long values collapse to their size so a pipeline config does not
render as two near-identical blobs), with route / method / status / duration
/ evidence hash moved into the expanded diagnostics line;
* the whole card is the expand toggle with a rotating chevron, which removes
the ambiguity of a small text link whose expanded state looked identical;
* records are bucketed into today / yesterday / dated groups so a long history
reads as a timeline;
* a "changes only" toggle (reusing the level filter) answers "who changed
something" in one click -- in the live log 489 of 499 rows are pure views,
so this is the difference between a haystack and the ten rows that matter.
Tracing fixes found while reading the live log:
* PUT /pipelines/<uuid>/extensions was classified as a plugin config change
because the generic ('/extensions') rule ran first; it now maps to a
pipeline-specific action;
* that handler recorded no diff at all, so "modified extension config" could
never say what changed. It now captures the previous bindings and reports the
before → after.
Cache: the integrity result is now keyed by (Workspace, listing filters) so
each view reuses its own verification, and a delete invalidates all views for
the Workspace. Fourteen tests cover the projection, the verifier, the cache
lifecycle and the failure modes.
The pull request merge produced two Alembic heads: this branch's
0032_operation_traceability and master's 0032_cert_artifact_digest both
descended from 0031_merge_totp_assistant, so the integration and migration
jobs failed with "Multiple head revisions". Renumber the traceability
migration to 0033 and point it at master's head (0032_cert_artifact_digest),
restoring a single linear head.
Also replace the noqa'd unused import of the traceability routes with an
importlib bootstrap (addresses the CodeQL unused-import finding) and merge
origin/master to pick up the certification digest backfill.
- Move the operation-traceability service and its HTTP routes into a standalone
pkg/operation_trace package; the Core controller package no longer depends on it
(routes stay auto-discovered by the controller package scan).
- Gate the hot path on one module-level boolean exposed as ap.operation_trace_active:
while no Workspace has tracing enabled the route wrapper performs zero trace work
and zero database round trips. The gate opens when a Workspace enables tracing and
is primed once at startup for already-enabled Workspaces.
- Record the installed extension identity for GitHub / marketplace / local / upload
installs (owner/repo, author/name, filename) so the trace no longer shows a
nameless "a plugin was installed" event and the card names the extension.
- Merge migrations 0032 + 0033 into a single 0032_operation_traceability revision
carrying the full table shape and an index set that matches the ORM.
Tests: unit 5165 passed / integration 358 passed; ruff check + format clean.
- Remove the unused request_summary local: the stored summary is derived by the
service from changes + the classified action, so handlers publishing
quart.g.operation_log_summary was a dead channel that never reached the row.
- Drop the unused datetime import in the operation log model.
- Apply ruff format to the touched files so `ruff format --check` passes.
Granularity follow-up: a record used to say only "installed an extension".
- Add resolve_resource_identity(): derive the acted-on resource from trusted route
params or the install-style payload (plugin_author/plugin_name), so every
resource family names its target. Unknown keys are ignored and sensitive keys
are skipped, keeping the audit trail bounded and secret-free.
- The route wrapper applies it to every traced request; the identity only fills
in when a handler did not publish a more precise one.
- Publish field-level before/after diffs for plugin config, skill, knowledge
base, MCP server, pipeline, provider and model updates, reusing
changed_fields()/build_summary() so secret-looking keys stay redacted and a
masked round-trip never shows up as a phantom change.
- Count hash mismatches / broken links over the whole filtered history instead
of only the returned page, so the badges no longer reset to zero on page 2.
- Classify failing record ids during the same scan and expose them through an
integrity=all|issues|hash_mismatch|chain_broken listing filter.
- Report scanned_count / scan_truncated so an approximate (row-capped) summary
is never mistaken for a clean one.
- Resolve the acted-on resource identity from route params or the request body,
giving every resource family a "which plugin/skill/knowledge base" trace.
- Render resource types and the two counters through i18n, keeping all eight
locale files key- and line-aligned.
- Export download failed because the URL builder appended the path to a
"/" base, producing a protocol-relative "//api/..." URL that the browser
read as host "api"; the request never reached the backend. Resolve the
base to the current origin, mirroring the other URL builders.
- Remove the orphan ``clear_prune`` action: the branch has no prune/clear
route and no route rule maps to it, so it could never be produced. Its
i18n key is removed too.
- Remove ``tamperedCount``, now unreferenced after the verification badge
was split into integrity/chain counters. The operationTrace catalogue is
now free of unproduced keys.
- Correct the settings controller docstring: only governance, operation
logs, filters and export routes exist; there is no on-demand delete.
Verified: no producer for clear_prune, classify() can no longer emit it,
operationTrace orphan scan returns none, tsc/check-i18n/prettier/py_compile
all pass.
Performance of the operation traceability surface:
- Cache the retention / row-budget / dedupe policy per Workspace and
invalidate it on governance change, so the hot write path no longer
issues three metadata round trips per recorded operation.
- Resolve the governance payload with one metadata query instead of four.
- Turn the dedupe probe into an existence check (LIMIT 1) instead of a
full COUNT over the dedupe window.
- Defer the row-budget COUNT to every N inserts; prune still enforces the
exact budget, and the maintenance loop keeps it bounded.
- Panel: fetch only the records page when paging or filtering instead of
re-requesting governance and the filter catalogue every time.
- Export: include the effective filters (and limit) in the artifact so a
downloaded CSV is self-describing.
Fixes and cleanup:
- Drop the unused OperationLogPruneReport import that broke `tsc`.
- i18n: remove the 9 redundant blank lines unique to en-US and format the
six non-compliant locale files with the project Prettier config; all
eight locales keep identical key sets, key order and placeholders.
Resolve the revision-graph conflict introduced by the release line.
Conflicts
- tests/integration/persistence/test_rag_document_identity.py: master had
independently introduced the same dynamic-head helper (`current_head`) plus
a `DOCUMENT_IDENTITY_REVISION` constant, and it explicitly upgrades to that
revision. Keep master's semantics (the explicit revision matters because
upgrading to the head now also traverses the TOTP branch) while keeping the
module-level alembic imports on this branch.
New migration
- 0030_merge_totp_into_release joins the TOTP branch's own merge revision
(0026_merge_totp_and_rag_identity) with the release head
(0029_merge_rag_identity). Both reached 0025_rag_document_identity without
including each other, which left Alembic with two heads and made
`upgrade head` fail with "Multiple head revisions are present".
Verification
- Single Alembic head confirmed (0030_merge_totp_into_release).
- tests/integration/persistence/test_rag_document_identity.py: 34 passed.
- Full fast integration suite: 356 passed, 84 skipped.
Address the migration and formatting failures reported on the TOTP branch.
Migrations
- Register totp_credentials and totp_recovery_codes in
_ALEMBIC_TENANT_TABLES. On a legacy PostgreSQL install these two tables
reference users.uuid, which only exists after 0009, so create_all() must
not run ahead of Alembic the way it did for the other tenant tables.
Without this the PostgreSQL migration test failed with
"column uuid referenced in foreign key constraint does not exist".
Tests
- Resolve the Alembic head dynamically in the RAG document identity
regression instead of pinning 0025_rag_document_identity. The TOTP and
RAG branches now meet at a merge revision, so the pinned value was no
longer the head. This matches the convention already used by
test_migrations_postgres.
Formatting
- Apply ruff format to the new backend modules and prettier to the locale
files and TOTP components, so the lint jobs pass.
Add a per-Account time-based one-time password (TOTP) second sign-in
factor, plus the owner/admin tooling needed to operate it.
Backend
- TotpService: enrolment, constant-time verification with a +/- one step
drift window, single-use recovery codes and disablement.
* shared secret is only ever persisted as a Fernet token whose key is
derived (HKDF-SHA256) from the instance JWT secret and gated by a
key_version epoch;
* recovery codes are only ever persisted as salted
PBKDF2-HMAC-SHA256 digests (600k iterations) and are single-use;
* the consumed counter advances monotonically so a captured code cannot
be replayed inside the same time window.
- New persistence entities and alembic migrations for credentials and
recovery codes.
- Login second-factor challenge, bound to the Account that passed the
password step.
- Owner/admin oversight endpoints to inspect and re-bind the second
factor of any Account.
Frontend
- Account settings panel for enrolling, managing, re-binding and
disabling the second factor.
- Login second-factor step and the matching client methods.
- Strings for all eight locales.
The shared secret never crosses the API boundary: enrolment returns a
server-rendered QR code (as a data: URL) and the plaintext secret is
discarded as soon as the image is produced.
- Add MediaCache using xxHash3-128 (with sha256 fallback) content-addressable storage
- Externalize message chain image payloads before recording monitoring and discarded messages
- Strip base64 payloads to null in SQLite monitoring_messages, dropping row size from megabytes to hundreds of bytes
- Add GET /api/v1/files/media/<filename> route with immutable HTTP cache headers to serve cached media
- Integrate age-based retention (default 30 days) and configurable disk quota with MaintenanceService cleanup loop
- Add defensive sanitizer in MonitoringService.record_message against oversized raw base64 payloads
- Add comprehensive unit tests and end-to-end verification covering CAS deduplication, route serving, and LRU pruning
* fix(runner): align SDK pin and workspace-aware integration fixtures
* fix(ci): format sources and resolve current migration head
* test(persistence): align standalone migration fixtures with current models
* test(web): align smoke fixtures with current processor UI
---------
Co-authored-by: dadachann <185672915+dadachann@users.noreply.github.com>
* fix(cloud): provision login workspace just in time
* fix(oauth): send callback URI during code exchange
* fix(oauth): preserve callback URI through browser exchange
* fix(oauth): negotiate redirect-bound codes
---------
Co-authored-by: dadachann <185672915+dadachann@users.noreply.github.com>
* fix(cloud): launch new accounts through Space
* style: format Cloud entry URL
* fix(cloud): wait for launch workspace projection
---------
Co-authored-by: dadachann <185672915+dadachann@users.noreply.github.com>