close_trace() called the coroutine TelemetryManager.start_send_task without
awaiting or scheduling it, so every payload closed on that path was silently
dropped and the runtime logged "coroutine ... was never awaited". The only
other caller awaits it correctly, so both call styles stay supported: call the
manager, schedule the returned coroutine on the running loop, and close it with
a debug log when no loop is available.
Regression case added in tests/unit_tests/telemetry/test_trace.py: it fails
before this change (payload never delivered) and passes after.
Execution telemetry already reported window counters. This adds one bounded,
content-free chain per inbound platform event, so a single event can be
followed from its source platform through routing and processing to the
platform API calls it caused.
- telemetry/trace.py: ContextVar trace identity with route and run scopes, and
a 32-stage bound per chain.
- telemetry/execution.py: keeps at most 64 chains for 120s, decides sampling
from the observed outcome, and sends one payload per chain keyed by
query_id = trace_id, so Space fetches a chain by primary identity. Window
counters are unchanged and remain the source of coverage statistics.
- botmgr / pipelinemgr / orchestrator: bind the trace at the ingress boundary,
scope routing identity per dispatch, and attach the run identity, so nested
lanes (event -> route -> pipeline/runner -> platform API) reuse one chain. A
run reached without an ingress (WebUI debug, service API) owns its own chain.
- space.execution_trace selects off | failures | sampled | all (default
sampled: failures, WebUI debug runs and every N-th success).
Chains carry only code-defined identifiers (adapter and runner types, route
identity as type:uuid, run id); never message content, tool arguments or
platform user identifiers.
scan_skill_directory is disabled for Cloud-managed sandboxes, but the
register_skill tool was still advertised to the model whenever a backend was
available, so every call failed closed without a usable explanation. Advertise
register_skill only when the runtime may scan sandbox directories, and keep
activate available in both modes.
The workspace file tools (read/write/edit/glob/grep) ran a script through
`python`, which a sandbox built from a host rootfs may not provide: those hosts
ship `python3` only, and a read-only /usr prevents creating a `python` shim
inside the jail. Every file tool then failed with `/bin/sh: 1: python: not
found`. Resolve the interpreter with `command -v python3 || command -v python`,
matching the interpreter the Box service already uses for its own scripts.
Cloud runtime refuses unscoped database access and workspace_operation_logs is row-level-secured on the bound Workspace. The background writer, the retention pass and the metadata reads all run outside a request lifetime, so each of them failed and the failure was swallowed by the best-effort handler: enabling tracing recorded nothing at all.
The route wrapper also traces after the handler released its scope, so level resolution has to carry the Workspace itself. Without it the level read back as 0 and the request was dropped before it was even queued.
The startup gate now enumerates the Workspaces through this instance's execution bindings, because a cross-Workspace metadata scan cannot see them under RLS, and it runs after the Workspace service exists. The writer starts with a clean context so it cannot inherit the scope of the request that first enqueued a trace, and a dropped write is reported instead of being logged only at debug.
The panel answered "who, when, what" poorly: the actual before → after diff
was hidden behind a click, the actor was a muted line, and a whole row went to
the route while duration and hash sat in a permanent right column. The data
already arrives with the page, so the panel now reads top-down in the order an
operator actually looks:
* who and when first, then the action and resource, then the diff spelled out
inline (long values collapse to their size so a pipeline config does not
render as two near-identical blobs), with route / method / status / duration
/ evidence hash moved into the expanded diagnostics line;
* the whole card is the expand toggle with a rotating chevron, which removes
the ambiguity of a small text link whose expanded state looked identical;
* records are bucketed into today / yesterday / dated groups so a long history
reads as a timeline;
* a "changes only" toggle (reusing the level filter) answers "who changed
something" in one click -- in the live log 489 of 499 rows are pure views,
so this is the difference between a haystack and the ten rows that matter.
Tracing fixes found while reading the live log:
* PUT /pipelines/<uuid>/extensions was classified as a plugin config change
because the generic ('/extensions') rule ran first; it now maps to a
pipeline-specific action;
* that handler recorded no diff at all, so "modified extension config" could
never say what changed. It now captures the previous bindings and reports the
before → after.
Cache: the integrity result is now keyed by (Workspace, listing filters) so
each view reuses its own verification, and a delete invalidates all views for
the Workspace. Fourteen tests cover the projection, the verifier, the cache
lifecycle and the failure modes.
The integrity scan selected whole rows, so a cold pass paged in the
changes/detail payloads and the client fingerprint for up to
MAX_INTEGRITY_SCAN_ROWS rows the verifier never reads. Project only the hash
columns, and add a lightweight verifier so the scan no longer builds a full
display dict per row.
Repair the read cache so it is a latency shield, not a correctness shortcut:
a result computed within INTEGRITY_CACHE_TTL_SECONDS is served from the
per-Workspace cache (opening, refreshing and paging all land inside that
window and pay nothing), while a cache miss re-verifies the whole window. An
incremental scan that skips previously verified ids can never see an edit to
an already-cached row -- exactly the tampering this feature exists to expose.
Verified against a live 414-row log: 50 content edits and 7 re-signed links
are all detected, including an edit to a row verified on a previous pass.
Also fix two correctness gaps and one maintenance bug:
- age-based prune deleted rows without invalidating the cached prefix;
- the boundary baseline is now read only when the history exceeds the scan
window, which a bounded scan makes the rare case;
- the operation-log retention block had drifted outside the per-binding loop
in the maintenance task, so only the last discovered Workspace was ever
pruned while the others grew unbounded.
Tests: scan projection, verifier parity, TTL cache reuse, cache-miss
re-verification, edit to an already-verified row, hash mismatch, chain break
and prune invalidation.
A saved pipeline may carry the post-migration shape
{'ai': {'runner': {'id': ''}, 'runner_config': {}}} while no runner has been
selected yet. The planner matched the empty id against the certified
'plugin:author/name/component' form, failed, and returned a malformed-id
blocker, so the whole 'migrate all' batch stopped with both pipelines reported
as blocked even though there was no legacy section and nothing to convert.
A blank id with no legacy runner section now reports not_legacy: no work is
needed and no target is synthesized. A blank id that still coexists with a
legacy section remains an explicit invalid_runner_id blocker, because that
ambiguous state must be resolved by the operator rather than guessed.
The certified-archive admission gate treated any declared certificate it could
not resolve as an untrusted archive and rejected the install with
CERTIFIED_PLUGIN_OSS_FORCE_REQUIRED. OSS ships an empty
plugin.certification.trusted_public_keys ring and the marketplace signs every
package with its own issuer key, so all certified marketplace packages (for
example langbot-team/RunnerDemo and langbot-team/LocalAgent) failed at the
'validating plugin package' step before artifact storage.
Admission now distinguishes an unresolvable declaration from a configured trust
decision that fails:
- record certificate_id only when the key_id is actually present in the
configured ring, so an empty ring yields an unresolvable declaration;
- OSS degrades an unresolvable declaration to the existing oss_dev dedicated
profile with CERTIFIED_PLUGIN_OSS_UNTRUSTED_DEDICATED instead of blocking;
- a resolvable declaration that still fails keeps requiring the explicit
administrator force (CERTIFIED_PLUGIN_OSS_FORCE_REQUIRED);
- Cloud stays fail-closed and rejects before storage.
No shared-runtime privilege is granted when the issuer is not trusted, so this
withholds isolation rather than escalating it.
* fix(pipeline): stop emitting empty final assistant messages
An empty final streaming chunk was appended as an empty message chain when no sandbox outbox attachment was present, so platforms received an empty reply. Only append the chain when attachments were actually collected, and gate the "Call ..." tool notice behind output.misc.track-function-calls so it no longer becomes the final chunk of a stream.
* fix(wecombot): do not close stream on a blank final snapshot
A blank final snapshot closed the WeCom stream with an empty bubble and pushed the real answer into a separate reply_text message. Skip blank final snapshots and keep the session open so the following non-blank chunk can finalize it. Covered by new regression tests.
- Add MediaCache using xxHash3-128 (with sha256 fallback) content-addressable storage
- Externalize message chain image payloads before recording monitoring and discarded messages
- Strip base64 payloads to null in SQLite monitoring_messages, dropping row size from megabytes to hundreds of bytes
- Add GET /api/v1/files/media/<filename> route with immutable HTTP cache headers to serve cached media
- Integrate age-based retention (default 30 days) and configurable disk quota with MaintenanceService cleanup loop
- Add defensive sanitizer in MonitoringService.record_message against oversized raw base64 payloads
- Add comprehensive unit tests and end-to-end verification covering CAS deduplication, route serving, and LRU pruning