mirror of
https://github.com/langbot-app/LangBot.git
synced 2026-08-12 22:00:57 +00:00
fix(cloud): bound tenant maintenance and monitoring work
This commit is contained in:
@@ -16,7 +16,7 @@
|
|||||||
|
|
||||||
## 1. 当前已形成的交付基线
|
## 1. 当前已形成的交付基线
|
||||||
|
|
||||||
- LangBot Core 全量 `2833 passed, 33 skipped`,Plugin SDK 全量
|
- LangBot Core 全量 `2839 passed, 33 skipped`,Plugin SDK 全量
|
||||||
`1325 passed`,闭源适配器 `40 passed`,Space Go 全量测试通过;三仓格式、
|
`1325 passed`,闭源适配器 `40 passed`,Space Go 全量测试通过;三仓格式、
|
||||||
静态检查和 `git diff --check` 已通过。
|
静态检查和 `git diff --check` 已通过。
|
||||||
- Plugin Runtime 和 Box Runtime 的公开健康接口、event-loop lag 与有界
|
- Plugin Runtime 和 Box Runtime 的公开健康接口、event-loop lag 与有界
|
||||||
@@ -49,6 +49,12 @@
|
|||||||
- Core Cloud manager 已连接一次性 PostgreSQL 16,并从 `pg_settings` 读回
|
- Core Cloud manager 已连接一次性 PostgreSQL 16,并从 `pg_settings` 读回
|
||||||
`statement_timeout=60000ms`、`lock_timeout=5000ms` 和
|
`statement_timeout=60000ms`、`lock_timeout=5000ms` 和
|
||||||
`idle_in_transaction_session_timeout=60000ms`;测试结束后引擎已显式 dispose。
|
`idle_in_transaction_session_timeout=60000ms`;测试结束后引擎已显式 dispose。
|
||||||
|
- 独立异常路径复核已补齐 HTTPX 超限/取消时的底层流关闭;Monitoring 查询、导出和
|
||||||
|
detail 物化量均有实例上限与绝对上限,detail 统计使用数据库聚合。Token statistics
|
||||||
|
不再拉取全部历史 LLM call 在 Python 中分桶,而由 PostgreSQL/SQLite 聚合并只返回
|
||||||
|
有界的最新时间桶和模型分组,截断状态在响应中显式可见。邀请、Monitoring 和 Storage
|
||||||
|
周期清理已合并为一个先等待首个 interval 的调度器,同一周期只进行一次 Workspace
|
||||||
|
discovery;数据库删除批次和本地/S3 文件候选也有每轮硬上限。
|
||||||
|
|
||||||
以上结果是进入生产候选验证的前提,不是 SaaS 上线批准。
|
以上结果是进入生产候选验证的前提,不是 SaaS 上线批准。
|
||||||
|
|
||||||
@@ -197,6 +203,9 @@ PostgreSQL/pgvector 和代表性 Workspace 配置分布,测量:
|
|||||||
- 启动、目录重放、批量 reconcile 和故障恢复的耗时与峰值;
|
- 启动、目录重放、批量 reconcile 和故障恢复的耗时与峰值;
|
||||||
- remote MCP 数量增加及目录 generation 批量切换时的数据库 QPS、回收队列和
|
- remote MCP 数量增加及目录 generation 批量切换时的数据库 QPS、回收队列和
|
||||||
event-loop lag,确认不存在与 session 数量成比例的空闲轮询;
|
event-loop lag,确认不存在与 session 数量成比例的空闲轮询;
|
||||||
|
- 在最大 retention/backlog 和并发 Dashboard 请求下执行不带时间范围的 Monitoring
|
||||||
|
overview/token statistics,验证 SQL 分桶、statement timeout、响应截断和 cleanup
|
||||||
|
追赶不会形成 PostgreSQL CPU 尖峰或 Core RSS 增长;
|
||||||
- 单实例可批准的 Workspace、活跃 Bot、plugin worker 和 sandbox 上限。
|
- 单实例可批准的 Workspace、活跃 Bot、plugin worker 和 sandbox 上限。
|
||||||
|
|
||||||
容量上限必须写入生产配置与告警,不能只保留在测试报告中。
|
容量上限必须写入生产配置与告警,不能只保留在测试报告中。
|
||||||
|
|||||||
@@ -52,6 +52,9 @@ SDK 已先行发布到分支提交 `7c0b9827ed8597a1c84151b83fcf6307934fd944`,
|
|||||||
- Platform bot reload/remove/shutdown 统一串行化,旧 bot、代理、adapter 任务和进程会先停止再从注册表移除。
|
- Platform bot reload/remove/shutdown 统一串行化,旧 bot、代理、adapter 任务和进程会先停止再从注册表移除。
|
||||||
- Model provider requester 新增异步关闭契约;provider reload/remove、Workspace generation 替换、全量 reload 和 Application shutdown 都会确定性关闭旧 requester,允许第三方 requester 安全持有自己的 HTTP client 或连接池。
|
- Model provider requester 新增异步关闭契约;provider reload/remove、Workspace generation 替换、全量 reload 和 Application shutdown 都会确定性关闭旧 requester,允许第三方 requester 安全持有自己的 HTTP client 或连接池。
|
||||||
- Plugin Runtime、Box Runtime、stdio transport、adapter 连接和共享 HTTP client 均补齐 close/cancel/await。
|
- Plugin Runtime、Box Runtime、stdio transport、adapter 连接和共享 HTTP client 均补齐 close/cancel/await。
|
||||||
|
- HTTPX 有界流在超限异常或消费者取消时会立即关闭底层响应流;原来的 response hook
|
||||||
|
只在正常读完后由 HTTPX 自动关闭,持久客户端反复收到超大响应时可能积累未释放连接。
|
||||||
|
已消费响应在超限分支也会先 `aclose()` 再传播错误。
|
||||||
- `Application.dispose()` 只允许一个可追踪 shutdown task;重复的信号、窗口关闭或调用方清理不会铺开多个并行停机流程。
|
- `Application.dispose()` 只允许一个可追踪 shutdown task;重复的信号、窗口关闭或调用方清理不会铺开多个并行停机流程。
|
||||||
- Lark、微信、钉钉、企业微信和 QQ Official 的凭证交换后台任务统一进入 Application TaskManager,受全局/单 Workspace admission 约束并随应用停机取消;容量满时关闭尚未调度的 coroutine 并返回 429,不留下游离 task。
|
- Lark、微信、钉钉、企业微信和 QQ Official 的凭证交换后台任务统一进入 Application TaskManager,受全局/单 Workspace admission 约束并随应用停机取消;容量满时关闭尚未调度的 coroutine 并返回 429,不留下游离 task。
|
||||||
- `TaskCapacityError` 已下沉到无 Application/controller 依赖的纯错误模块。原来的 HTTP 过载异常路径会在特定冷启动导入顺序下触发 TaskManager/controller 循环导入,把应返回的 429 变成框架 500。
|
- `TaskCapacityError` 已下沉到无 Application/controller 依赖的纯错误模块。原来的 HTTP 过载异常路径会在特定冷启动导入顺序下触发 TaskManager/controller 循环导入,把应返回的 429 变成框架 500。
|
||||||
@@ -110,6 +113,17 @@ SDK 已先行发布到分支提交 `7c0b9827ed8597a1c84151b83fcf6307934fd944`,
|
|||||||
目标的签名 delta 中作为 tombstone 返回。注册创建新个人 Workspace 前通过
|
目标的签名 delta 中作为 tombstone 返回。注册创建新个人 Workspace 前通过
|
||||||
PostgreSQL transaction advisory lock 串行执行全局 active 数量准入;达到上限
|
PostgreSQL transaction advisory lock 串行执行全局 active 数量准入;达到上限
|
||||||
返回 503,避免多个 Space 副本同时观察到最后一个空位。
|
返回 503,避免多个 Space 副本同时观察到最后一个空位。
|
||||||
|
- Monitoring 分页、offset、CSV export 和 session/message detail 均在 service
|
||||||
|
边界执行实例配置上限与不可放大的绝对上限;detail 的完整统计改为 SQL aggregate,
|
||||||
|
只物化有界的 tool/LLM/error 明细并显式返回 `detail_truncated`。默认分页 1,000、
|
||||||
|
export 10,000、detail 2,000,绝对上限分别为 5,000、50,000、10,000。
|
||||||
|
- Token statistics 的时间序列不再把筛选范围内的全部 LLM call 拉回 Python 分桶;
|
||||||
|
PostgreSQL 使用 `date_trunc`、SQLite 使用 `strftime` 在数据库中聚合,并只返回
|
||||||
|
最近 1,000 个时间桶(绝对上限 10,000)。模型分组复用分页上限并在 SQL 中按 token
|
||||||
|
排序、限制;两类结果都返回显式的 `*_truncated` 标志。
|
||||||
|
- Monitoring 过期数据每表每轮默认最多删除 4 个批次、绝对最多 100 个批次;本地/S3
|
||||||
|
过期上传文件候选和每轮删除默认最多 1,000、绝对最多 10,000。单个历史数据量异常的
|
||||||
|
Workspace 不再能让一次维护循环无限物化候选或持续清空全部 backlog。
|
||||||
|
|
||||||
### CPU 和事件循环保护
|
### CPU 和事件循环保护
|
||||||
|
|
||||||
@@ -134,6 +148,11 @@ SDK 已先行发布到分支提交 `7c0b9827ed8597a1c84151b83fcf6307934fd944`,
|
|||||||
- Box session 枚举、旧 generation 回收和 admission 计数均通过 Workspace 索引执行;admission 过期回收通过最小堆执行,不再在每次 RPC 上产生 O(实例总 session/grant 数) 的扫描。
|
- Box session 枚举、旧 generation 回收和 admission 计数均通过 Workspace 索引执行;admission 过期回收通过最小堆执行,不再在每次 RPC 上产生 O(实例总 session/grant 数) 的扫描。
|
||||||
- Model、Pipeline、RAG 和 Platform manager 均维护 Workspace 到运行时 key 的二级索引。Workspace generation 更新只清理目标 Workspace 的缓存和运行时,不再扫描实例内所有租户的 provider/model、pipeline、knowledge runtime 或 bot;回归测试使用禁止全局迭代的映射验证该边界。
|
- Model、Pipeline、RAG 和 Platform manager 均维护 Workspace 到运行时 key 的二级索引。Workspace generation 更新只清理目标 Workspace 的缓存和运行时,不再扫描实例内所有租户的 provider/model、pipeline、knowledge runtime 或 bot;回归测试使用禁止全局迭代的映射验证该边界。
|
||||||
- Cloud heartbeat 直接读取已加载且有容量边界的 Pipeline、MCP、KnowledgeBase 和 Bot registry 计数,不再为每个活跃 Workspace 依次打开 Tenant UoW、执行四类 COUNT 查询;这消除了租户数增长后每日周期性形成的串行 SQL/CPU 尖峰。OSS 模式仍保留数据库统计语义。
|
- Cloud heartbeat 直接读取已加载且有容量边界的 Pipeline、MCP、KnowledgeBase 和 Bot registry 计数,不再为每个活跃 Workspace 依次打开 Tenant UoW、执行四类 COUNT 查询;这消除了租户数增长后每日周期性形成的串行 SQL/CPU 尖峰。OSS 模式仍保留数据库统计语义。
|
||||||
|
- 邀请、Monitoring 和 Storage 的三个周期清理 task 合并为一个
|
||||||
|
`resource-maintenance` 调度器。调度器先等待首个 interval,不与启动加载争抢资源;
|
||||||
|
同一到期周期只执行一次 active Workspace discovery,然后按 Workspace 串行运行
|
||||||
|
有界 job,单 Workspace 失败不跳过其他 Workspace。默认相同的一小时周期由此从
|
||||||
|
三次全租户发现和三个同时唤醒的任务收敛为一次发现和一个任务。
|
||||||
- Cloud 启动阶段先生成一份经过部署适配器和目录投影校验的 Workspace binding 快照,Model、Platform、Pipeline、RAG 和 Plugin 初始化共用该快照,初始化完成后立即释放;避免启动期间为每个 manager 重复执行整批租户发现和投影校验。
|
- Cloud 启动阶段先生成一份经过部署适配器和目录投影校验的 Workspace binding 快照,Model、Platform、Pipeline、RAG 和 Plugin 初始化共用该快照,初始化完成后立即释放;避免启动期间为每个 manager 重复执行整批租户发现和投影校验。
|
||||||
- Platform、Pipeline 和 RAG 的资源加载在使用已验证启动快照时不再为每个 Bot/Pipeline/KnowledgeBase 重新查询同一个 execution binding;常规请求和动态更新路径仍保留数据库 generation fence。
|
- Platform、Pipeline 和 RAG 的资源加载在使用已验证启动快照时不再为每个 Bot/Pipeline/KnowledgeBase 重新查询同一个 execution binding;常规请求和动态更新路径仍保留数据库 generation fence。
|
||||||
- MCP 初始 host 和 shutdown burst 由实例级 semaphore/批次限制;默认 `mcp.lifecycle_concurrency=16`,支持 `MCP__LIFECYCLE_CONCURRENCY` 覆写并硬性限制最大 128。初始加载不再先为每个 server 创建一个等待 semaphore 的 task,而是由一个可取消 dispatcher 每批最多物化 `lifecycle_concurrency` 个子 task;同时去掉了 ORM server/config 的双份临时列表,避免大量租户启动时集中占用 CPU、内存、socket 和文件句柄。
|
- MCP 初始 host 和 shutdown burst 由实例级 semaphore/批次限制;默认 `mcp.lifecycle_concurrency=16`,支持 `MCP__LIFECYCLE_CONCURRENCY` 覆写并硬性限制最大 128。初始加载不再先为每个 server 创建一个等待 semaphore 的 task,而是由一个可取消 dispatcher 每批最多物化 `lifecycle_concurrency` 个子 task;同时去掉了 ORM server/config 的双份临时列表,避免大量租户启动时集中占用 CPU、内存、socket 和文件句柄。
|
||||||
@@ -187,6 +206,10 @@ SDK 已先行发布到分支提交 `7c0b9827ed8597a1c84151b83fcf6307934fd944`,
|
|||||||
- Core 与 SDK 各进程的通用阻塞 executor 默认使用 8 个 worker、128 个 pending 槽位、每 Workspace 4 个在途槽位;它是实例/进程级共享背压,不由 Workspace 或插件 manifest 调高,单 Workspace 配置硬性不得超过 worker 的一半。生产值应按容器 CPU 和上游阻塞时延校准,不能把 pending 当吞吐配置无限放大。
|
- Core 与 SDK 各进程的通用阻塞 executor 默认使用 8 个 worker、128 个 pending 槽位、每 Workspace 4 个在途槽位;它是实例/进程级共享背压,不由 Workspace 或插件 manifest 调高,单 Workspace 配置硬性不得超过 worker 的一半。生产值应按容器 CPU 和上游阻塞时延校准,不能把 pending 当吞吐配置无限放大。
|
||||||
- 插件包下载上限 64 MiB,pip stdout/stderr 保留上限各 1 MiB;这不会限制安装进程实际输出,只限制父进程内存中的诊断副本。
|
- 插件包下载上限 64 MiB,pip stdout/stderr 保留上限各 1 MiB;这不会限制安装进程实际输出,只限制父进程内存中的诊断副本。
|
||||||
- 通用远程响应和媒体默认上限 10 MiB;错误诊断正文只保留 4 KiB。Plugin binary storage 默认 10 MiB、绝对上限 64 MiB;Skill 文本、Plugin UI 和 host edit 分别限制为 1 MiB、4 MiB 和 1 MiB。
|
- 通用远程响应和媒体默认上限 10 MiB;错误诊断正文只保留 4 KiB。Plugin binary storage 默认 10 MiB、绝对上限 64 MiB;Skill 文本、Plugin UI 和 host edit 分别限制为 1 MiB、4 MiB 和 1 MiB。
|
||||||
|
- Monitoring 查询上限由 `monitoring.query_limits` 配置并支持原生环境变量覆写,但始终
|
||||||
|
受代码绝对上限约束;cleanup 的每表批次数和 Storage 每轮文件数同样采用实例配置加
|
||||||
|
绝对上限。时间序列默认/绝对上限为 1,000/10,000 个数据库聚合桶,模型分组复用分页
|
||||||
|
上限。提高这些值必须计入 V-08/V-09 的数据库 CPU 与 Core RSS 容量曲线。
|
||||||
- Managed-process relay 保留 stdout 的原始换行,并按 64 KiB WebSocket frame 分块;不再承诺“一行对应一个 frame”。这是为无换行输出提供确定内存边界所需的协议收敛。
|
- Managed-process relay 保留 stdout 的原始换行,并按 64 KiB WebSocket frame 分块;不再承诺“一行对应一个 frame”。这是为无换行输出提供确定内存边界所需的协议收敛。
|
||||||
- 本轮没有把 Pipeline、Model、KnowledgeBase 等合法租户资源改成 lazy runtime。该改动会改变启动和请求语义,留到 Workspace placement/释放机制一起设计。
|
- 本轮没有把 Pipeline、Model、KnowledgeBase 等合法租户资源改成 lazy runtime。该改动会改变启动和请求语义,留到 Workspace placement/释放机制一起设计。
|
||||||
- 本轮没有为普通 nsjail 声称伪硬盘配额;严格 Cloud readiness 保持失败关闭。
|
- 本轮没有为普通 nsjail 声称伪硬盘配额;严格 Cloud readiness 保持失败关闭。
|
||||||
@@ -197,7 +220,7 @@ SDK 已先行发布到分支提交 `7c0b9827ed8597a1c84151b83fcf6307934fd944`,
|
|||||||
| --- | --- |
|
| --- | --- |
|
||||||
| LangBot Ruff + `git diff --check` | 通过 |
|
| LangBot Ruff + `git diff --check` | 通过 |
|
||||||
| Plugin SDK Ruff + `git diff --check` | 通过 |
|
| Plugin SDK Ruff + `git diff --check` | 通过 |
|
||||||
| LangBot 全量测试(使用远端精确钉住的新 SDK,含 unit/integration/Box/E2E) | `2833 passed, 33 skipped` |
|
| LangBot 全量测试(使用远端精确钉住的新 SDK,含 unit/integration/Box/E2E) | `2839 passed, 33 skipped` |
|
||||||
| Plugin SDK 全量测试 | `1325 passed` |
|
| Plugin SDK 全量测试 | `1325 passed` |
|
||||||
| Space Go 全量测试与闭源 Cloud Adapter 测试 | Go `go test ./...` 通过;Adapter `40 passed` |
|
| Space Go 全量测试与闭源 Cloud Adapter 测试 | Go `go test ./...` 通过;Adapter `40 passed` |
|
||||||
| Space PostgreSQL 16 Cloud v2 目录与并发容量准入 | 通过;两个注册并发争用最后一个槽位时 `1 success / 1 capacity rejection / 1 active Workspace` |
|
| Space PostgreSQL 16 Cloud v2 目录与并发容量准入 | 通过;两个注册并发争用最后一个槽位时 `1 success / 1 capacity rejection / 1 active Workspace` |
|
||||||
@@ -271,7 +294,7 @@ Plugin SDK audit 每个阶段执行 25,000 次 loopback RPC、5,000 次安装 bi
|
|||||||
|
|
||||||
探针要求第二阶段的结构状态与第一阶段精确相等,并对第二阶段 RSS 与 tracemalloc 增长设置失败阈值。macOS 的 RSS 来源是 `getrusage` peak,因此这里验证的是峰值增量边界而非“当前 RSS 回落”;最终 Linux 24 小时 soak 仍需采集 current RSS/PSS 和 cgroup `memory.current`。
|
探针要求第二阶段的结构状态与第一阶段精确相等,并对第二阶段 RSS 与 tracemalloc 增长设置失败阈值。macOS 的 RSS 来源是 `getrusage` peak,因此这里验证的是峰值增量边界而非“当前 RSS 回落”;最终 Linux 24 小时 soak 仍需采集 current RSS/PSS 和 cgroup `memory.current`。
|
||||||
|
|
||||||
LangBot 全量测试的 33 个 skip 中,22 个是默认全量运行未提供 PostgreSQL/pgvector 而跳过的集成用例,10 个是未提供 Valkey,另 1 个是可选环境的 collection skip;真实 PostgreSQL 相关路径已由上表单独运行覆盖。Plugin SDK 的 26 个 warning 为现有 Pydantic v2 deprecation 与 aiohttp AppKey 建议;没有失败、未关闭资源或资源上限降级。Core 当前全量产生 193 个既有第三方/兼容性 warning;`ResourceWarning` 和 `PytestUnraisableExceptionWarning` 仍由 pytest 配置提升为错误,本轮没有此类泄漏告警。
|
LangBot 全量测试的 33 个 skip 中,22 个是默认全量运行未提供 PostgreSQL/pgvector 而跳过的集成用例,10 个是未提供 Valkey,另 1 个是可选环境的 collection skip;真实 PostgreSQL 相关路径已由上表单独运行覆盖。Plugin SDK 的 26 个 warning 为现有 Pydantic v2 deprecation 与 aiohttp AppKey 建议;没有失败、未关闭资源或资源上限降级。Core 当前全量产生 194 个既有第三方/兼容性 warning;`ResourceWarning` 和 `PytestUnraisableExceptionWarning` 仍由 pytest 配置提升为错误,本轮没有此类泄漏告警。
|
||||||
|
|
||||||
Linux Runtime 探针使用上述镜像并只读挂载本地最新 SDK 源码:
|
Linux Runtime 探针使用上述镜像并只读挂载本地最新 SDK 源码:
|
||||||
|
|
||||||
|
|||||||
@@ -21,6 +21,8 @@ from .tenant import TenantContext, require_workspace_uuid
|
|||||||
LOG_FILE_PATTERN = re.compile(r'^langbot-(\d{4}-\d{2}-\d{2})\.log(?:\.\d+)?$')
|
LOG_FILE_PATTERN = re.compile(r'^langbot-(\d{4}-\d{2}-\d{2})\.log(?:\.\d+)?$')
|
||||||
DEFAULT_UPLOAD_FILE_RETENTION_DAYS = 7
|
DEFAULT_UPLOAD_FILE_RETENTION_DAYS = 7
|
||||||
DEFAULT_LOG_RETENTION_DAYS = 3
|
DEFAULT_LOG_RETENTION_DAYS = 3
|
||||||
|
DEFAULT_MAX_FILES_PER_RUN = 1000
|
||||||
|
HARD_MAX_FILES_PER_RUN = 10000
|
||||||
UPLOAD_OWNER_TYPES = ('upload_image', 'upload_document', 'upload')
|
UPLOAD_OWNER_TYPES = ('upload_image', 'upload_document', 'upload')
|
||||||
|
|
||||||
|
|
||||||
@@ -51,6 +53,17 @@ class MaintenanceService:
|
|||||||
def __init__(self, ap: app.Application) -> None:
|
def __init__(self, ap: app.Application) -> None:
|
||||||
self.ap = ap
|
self.ap = ap
|
||||||
|
|
||||||
|
def _max_files_per_run(self) -> int:
|
||||||
|
cleanup_cfg = (
|
||||||
|
getattr(getattr(self.ap, 'instance_config', None), 'data', {}).get('storage', {}).get('cleanup', {})
|
||||||
|
)
|
||||||
|
value = self._positive_int(
|
||||||
|
cleanup_cfg.get('max_files_per_run', DEFAULT_MAX_FILES_PER_RUN),
|
||||||
|
DEFAULT_MAX_FILES_PER_RUN,
|
||||||
|
'storage.cleanup.max_files_per_run',
|
||||||
|
)
|
||||||
|
return min(value, HARD_MAX_FILES_PER_RUN)
|
||||||
|
|
||||||
@_workspace_scope
|
@_workspace_scope
|
||||||
async def cleanup_expired_files(self, context: ExecutionContext) -> dict[str, int]:
|
async def cleanup_expired_files(self, context: ExecutionContext) -> dict[str, int]:
|
||||||
if not isinstance(context, ExecutionContext):
|
if not isinstance(context, ExecutionContext):
|
||||||
@@ -252,6 +265,7 @@ class MaintenanceService:
|
|||||||
provider = self.ap.storage_mgr.storage_provider
|
provider = self.ap.storage_mgr.storage_provider
|
||||||
cutoff = datetime.datetime.now(datetime.timezone.utc) - datetime.timedelta(days=retention_days)
|
cutoff = datetime.datetime.now(datetime.timezone.utc) - datetime.timedelta(days=retention_days)
|
||||||
candidates = []
|
candidates = []
|
||||||
|
max_candidates = self._max_files_per_run()
|
||||||
paginator = provider.s3_client.get_paginator('list_objects_v2')
|
paginator = provider.s3_client.get_paginator('list_objects_v2')
|
||||||
|
|
||||||
seen_prefixes: set[str] = set()
|
seen_prefixes: set[str] = set()
|
||||||
@@ -274,6 +288,8 @@ class MaintenanceService:
|
|||||||
'modified_at': last_modified.isoformat(),
|
'modified_at': last_modified.isoformat(),
|
||||||
}
|
}
|
||||||
)
|
)
|
||||||
|
if len(candidates) >= max_candidates:
|
||||||
|
return candidates
|
||||||
|
|
||||||
return candidates
|
return candidates
|
||||||
|
|
||||||
@@ -310,6 +326,7 @@ class MaintenanceService:
|
|||||||
storage_root = Path('data/storage')
|
storage_root = Path('data/storage')
|
||||||
cutoff = datetime.datetime.now().timestamp() - retention_days * 86400
|
cutoff = datetime.datetime.now().timestamp() - retention_days * 86400
|
||||||
candidates = []
|
candidates = []
|
||||||
|
max_candidates = self._max_files_per_run()
|
||||||
seen_roots: set[Path] = set()
|
seen_roots: set[Path] = set()
|
||||||
for owner_type in UPLOAD_OWNER_TYPES:
|
for owner_type in UPLOAD_OWNER_TYPES:
|
||||||
scoped_root = storage_root / self.ap.storage_mgr.scoped_prefix(context, owner_type=owner_type)
|
scoped_root = storage_root / self.ap.storage_mgr.scoped_prefix(context, owner_type=owner_type)
|
||||||
@@ -335,6 +352,8 @@ class MaintenanceService:
|
|||||||
if include_paths:
|
if include_paths:
|
||||||
item['path'] = str(entry)
|
item['path'] = str(entry)
|
||||||
candidates.append(item)
|
candidates.append(item)
|
||||||
|
if len(candidates) >= max_candidates:
|
||||||
|
return candidates
|
||||||
return candidates
|
return candidates
|
||||||
|
|
||||||
def _expired_log_candidates(self, retention_days: int, include_paths: bool = False) -> list[dict[str, Any]]:
|
def _expired_log_candidates(self, retention_days: int, include_paths: bool = False) -> list[dict[str, Any]]:
|
||||||
|
|||||||
@@ -15,6 +15,20 @@ from ..context import ExecutionContext
|
|||||||
from .tenant import TenantContext, require_workspace_uuid
|
from .tenant import TenantContext, require_workspace_uuid
|
||||||
|
|
||||||
|
|
||||||
|
_DEFAULT_MONITORING_PAGE_ROWS = 1000
|
||||||
|
_DEFAULT_MONITORING_EXPORT_ROWS = 10000
|
||||||
|
_DEFAULT_MONITORING_DETAIL_ROWS = 2000
|
||||||
|
_DEFAULT_MONITORING_TIMESERIES_BUCKETS = 1000
|
||||||
|
_DEFAULT_MONITORING_MAX_OFFSET = 1000000
|
||||||
|
_HARD_MAX_MONITORING_PAGE_ROWS = 5000
|
||||||
|
_HARD_MAX_MONITORING_EXPORT_ROWS = 50000
|
||||||
|
_HARD_MAX_MONITORING_DETAIL_ROWS = 10000
|
||||||
|
_HARD_MAX_MONITORING_TIMESERIES_BUCKETS = 10000
|
||||||
|
_HARD_MAX_MONITORING_OFFSET = 10000000
|
||||||
|
_DEFAULT_CLEANUP_BATCHES_PER_TABLE = 4
|
||||||
|
_HARD_MAX_CLEANUP_BATCHES_PER_TABLE = 100
|
||||||
|
|
||||||
|
|
||||||
def _workspace_transaction(method):
|
def _workspace_transaction(method):
|
||||||
"""Run an explicit service entrypoint in one Workspace transaction."""
|
"""Run an explicit service entrypoint in one Workspace transaction."""
|
||||||
|
|
||||||
@@ -38,6 +52,88 @@ class MonitoringService:
|
|||||||
def __init__(self, ap: app.Application) -> None:
|
def __init__(self, ap: app.Application) -> None:
|
||||||
self.ap = ap
|
self.ap = ap
|
||||||
|
|
||||||
|
def _configured_query_limit(self, name: str, default: int, hard_max: int) -> int:
|
||||||
|
config = (
|
||||||
|
getattr(getattr(self.ap, 'instance_config', None), 'data', {}).get('monitoring', {}).get('query_limits', {})
|
||||||
|
)
|
||||||
|
try:
|
||||||
|
value = int(config.get(name, default))
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
value = default
|
||||||
|
return min(max(value, 1), hard_max)
|
||||||
|
|
||||||
|
def normalize_page_window(self, limit: int, offset: int = 0) -> tuple[int, int]:
|
||||||
|
"""Clamp tenant-controlled pagination before constructing a DB query."""
|
||||||
|
|
||||||
|
page_cap = self._configured_query_limit(
|
||||||
|
'page_rows',
|
||||||
|
_DEFAULT_MONITORING_PAGE_ROWS,
|
||||||
|
_HARD_MAX_MONITORING_PAGE_ROWS,
|
||||||
|
)
|
||||||
|
offset_cap = self._configured_query_limit(
|
||||||
|
'max_offset',
|
||||||
|
_DEFAULT_MONITORING_MAX_OFFSET,
|
||||||
|
_HARD_MAX_MONITORING_OFFSET,
|
||||||
|
)
|
||||||
|
try:
|
||||||
|
normalized_limit = int(limit)
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
normalized_limit = 100
|
||||||
|
try:
|
||||||
|
normalized_offset = int(offset)
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
normalized_offset = 0
|
||||||
|
return (
|
||||||
|
min(max(normalized_limit, 1), page_cap),
|
||||||
|
min(max(normalized_offset, 0), offset_cap),
|
||||||
|
)
|
||||||
|
|
||||||
|
def normalize_export_limit(self, limit: int) -> int:
|
||||||
|
"""Clamp exports that are currently materialized as an in-memory list."""
|
||||||
|
|
||||||
|
export_cap = self._configured_query_limit(
|
||||||
|
'export_rows',
|
||||||
|
_DEFAULT_MONITORING_EXPORT_ROWS,
|
||||||
|
_HARD_MAX_MONITORING_EXPORT_ROWS,
|
||||||
|
)
|
||||||
|
try:
|
||||||
|
normalized = int(limit)
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
normalized = _DEFAULT_MONITORING_EXPORT_ROWS
|
||||||
|
return min(max(normalized, 1), export_cap)
|
||||||
|
|
||||||
|
def _detail_limit(self) -> int:
|
||||||
|
return self._configured_query_limit(
|
||||||
|
'detail_rows',
|
||||||
|
_DEFAULT_MONITORING_DETAIL_ROWS,
|
||||||
|
_HARD_MAX_MONITORING_DETAIL_ROWS,
|
||||||
|
)
|
||||||
|
|
||||||
|
def _timeseries_bucket_limit(self) -> int:
|
||||||
|
return self._configured_query_limit(
|
||||||
|
'timeseries_buckets',
|
||||||
|
_DEFAULT_MONITORING_TIMESERIES_BUCKETS,
|
||||||
|
_HARD_MAX_MONITORING_TIMESERIES_BUCKETS,
|
||||||
|
)
|
||||||
|
|
||||||
|
@staticmethod
|
||||||
|
def _token_bucket_expression(
|
||||||
|
timestamp_column: sqlalchemy.Column,
|
||||||
|
*,
|
||||||
|
bucket: str,
|
||||||
|
dialect_name: str,
|
||||||
|
):
|
||||||
|
"""Build a server-side hour/day bucket for supported business databases."""
|
||||||
|
|
||||||
|
if bucket not in {'hour', 'day'}:
|
||||||
|
bucket = 'hour'
|
||||||
|
if dialect_name == 'postgresql':
|
||||||
|
return sqlalchemy.func.date_trunc(bucket, timestamp_column)
|
||||||
|
if dialect_name == 'sqlite':
|
||||||
|
bucket_format = '%Y-%m-%d %H:00' if bucket == 'hour' else '%Y-%m-%d'
|
||||||
|
return sqlalchemy.func.strftime(bucket_format, timestamp_column)
|
||||||
|
raise RuntimeError(f'Unsupported monitoring database dialect: {dialect_name}')
|
||||||
|
|
||||||
@staticmethod
|
@staticmethod
|
||||||
def _require_write_context(context: ExecutionContext | None) -> str:
|
def _require_write_context(context: ExecutionContext | None) -> str:
|
||||||
"""Reject background/runtime writes that lost their execution fence."""
|
"""Reject background/runtime writes that lost their execution fence."""
|
||||||
@@ -57,6 +153,7 @@ class MonitoringService:
|
|||||||
context: ExecutionContext,
|
context: ExecutionContext,
|
||||||
retention_days: int,
|
retention_days: int,
|
||||||
batch_size: int = 1000,
|
batch_size: int = 1000,
|
||||||
|
max_batches_per_table: int | None = None,
|
||||||
) -> dict[str, int]:
|
) -> dict[str, int]:
|
||||||
"""Delete monitoring records older than the specified retention period.
|
"""Delete monitoring records older than the specified retention period.
|
||||||
|
|
||||||
@@ -72,6 +169,24 @@ class MonitoringService:
|
|||||||
raise ValueError('retention_days must be >= 1')
|
raise ValueError('retention_days must be >= 1')
|
||||||
if batch_size < 1:
|
if batch_size < 1:
|
||||||
raise ValueError('batch_size must be >= 1')
|
raise ValueError('batch_size must be >= 1')
|
||||||
|
if max_batches_per_table is None:
|
||||||
|
cleanup_config = (
|
||||||
|
getattr(getattr(self.ap, 'instance_config', None), 'data', {})
|
||||||
|
.get('monitoring', {})
|
||||||
|
.get('auto_cleanup', {})
|
||||||
|
)
|
||||||
|
max_batches_per_table = cleanup_config.get(
|
||||||
|
'max_batches_per_table_per_run',
|
||||||
|
_DEFAULT_CLEANUP_BATCHES_PER_TABLE,
|
||||||
|
)
|
||||||
|
try:
|
||||||
|
max_batches_per_table = int(max_batches_per_table)
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
max_batches_per_table = _DEFAULT_CLEANUP_BATCHES_PER_TABLE
|
||||||
|
max_batches_per_table = min(
|
||||||
|
max(max_batches_per_table, 1),
|
||||||
|
_HARD_MAX_CLEANUP_BATCHES_PER_TABLE,
|
||||||
|
)
|
||||||
|
|
||||||
cutoff = datetime.datetime.now(datetime.timezone.utc).replace(tzinfo=None) - datetime.timedelta(
|
cutoff = datetime.datetime.now(datetime.timezone.utc).replace(tzinfo=None) - datetime.timedelta(
|
||||||
days=retention_days
|
days=retention_days
|
||||||
@@ -132,6 +247,7 @@ class MonitoringService:
|
|||||||
pk_column=pk_column,
|
pk_column=pk_column,
|
||||||
cutoff=cutoff,
|
cutoff=cutoff,
|
||||||
batch_size=batch_size,
|
batch_size=batch_size,
|
||||||
|
max_batches=max_batches_per_table,
|
||||||
)
|
)
|
||||||
return deleted_counts
|
return deleted_counts
|
||||||
|
|
||||||
@@ -157,11 +273,12 @@ class MonitoringService:
|
|||||||
pk_column: sqlalchemy.Column,
|
pk_column: sqlalchemy.Column,
|
||||||
cutoff: datetime.datetime,
|
cutoff: datetime.datetime,
|
||||||
batch_size: int,
|
batch_size: int,
|
||||||
|
max_batches: int,
|
||||||
) -> int:
|
) -> int:
|
||||||
workspace_uuid = self._require_write_context(context)
|
workspace_uuid = self._require_write_context(context)
|
||||||
deleted_total = 0
|
deleted_total = 0
|
||||||
|
|
||||||
while True:
|
for _batch_number in range(max_batches):
|
||||||
|
|
||||||
async def delete_batch() -> tuple[int, int]:
|
async def delete_batch() -> tuple[int, int]:
|
||||||
select_result = await self.ap.persistence_mgr.execute_async(
|
select_result = await self.ap.persistence_mgr.execute_async(
|
||||||
@@ -739,6 +856,8 @@ class MonitoringService:
|
|||||||
"""
|
"""
|
||||||
LLMCall = persistence_monitoring.MonitoringLLMCall
|
LLMCall = persistence_monitoring.MonitoringLLMCall
|
||||||
workspace_uuid = require_workspace_uuid(context)
|
workspace_uuid = require_workspace_uuid(context)
|
||||||
|
if bucket not in {'hour', 'day'}:
|
||||||
|
bucket = 'hour'
|
||||||
|
|
||||||
conditions = [LLMCall.workspace_uuid == workspace_uuid]
|
conditions = [LLMCall.workspace_uuid == workspace_uuid]
|
||||||
if bot_ids:
|
if bot_ids:
|
||||||
@@ -812,21 +931,29 @@ class MonitoringService:
|
|||||||
}
|
}
|
||||||
|
|
||||||
# ---- Per-model breakdown ----
|
# ---- Per-model breakdown ----
|
||||||
by_model_query = _apply(
|
model_total_tokens = sqlalchemy.func.coalesce(sqlalchemy.func.sum(LLMCall.total_tokens), 0)
|
||||||
sqlalchemy.select(
|
model_limit, _unused_offset = self.normalize_page_window(_HARD_MAX_MONITORING_PAGE_ROWS)
|
||||||
LLMCall.model_name,
|
by_model_query = (
|
||||||
sqlalchemy.func.count(LLMCall.id),
|
_apply(
|
||||||
sqlalchemy.func.coalesce(sqlalchemy.func.sum(LLMCall.input_tokens), 0),
|
sqlalchemy.select(
|
||||||
sqlalchemy.func.coalesce(sqlalchemy.func.sum(LLMCall.output_tokens), 0),
|
LLMCall.model_name,
|
||||||
sqlalchemy.func.coalesce(sqlalchemy.func.sum(LLMCall.total_tokens), 0),
|
sqlalchemy.func.count(LLMCall.id),
|
||||||
sqlalchemy.func.coalesce(sqlalchemy.func.sum(LLMCall.duration), 0),
|
sqlalchemy.func.coalesce(sqlalchemy.func.sum(LLMCall.input_tokens), 0),
|
||||||
sqlalchemy.func.coalesce(sqlalchemy.func.sum(LLMCall.cost), 0.0),
|
sqlalchemy.func.coalesce(sqlalchemy.func.sum(LLMCall.output_tokens), 0),
|
||||||
sqlalchemy.func.sum(sqlalchemy.case((LLMCall.status == 'error', 1), else_=0)),
|
model_total_tokens,
|
||||||
).group_by(LLMCall.model_name)
|
sqlalchemy.func.coalesce(sqlalchemy.func.sum(LLMCall.duration), 0),
|
||||||
|
sqlalchemy.func.coalesce(sqlalchemy.func.sum(LLMCall.cost), 0.0),
|
||||||
|
sqlalchemy.func.sum(sqlalchemy.case((LLMCall.status == 'error', 1), else_=0)),
|
||||||
|
).group_by(LLMCall.model_name)
|
||||||
|
)
|
||||||
|
.order_by(model_total_tokens.desc())
|
||||||
|
.limit(model_limit + 1)
|
||||||
)
|
)
|
||||||
by_model_result = await self.ap.persistence_mgr.execute_async(by_model_query)
|
by_model_result = await self.ap.persistence_mgr.execute_async(by_model_query)
|
||||||
|
by_model_rows = by_model_result.all()
|
||||||
|
by_model_truncated = len(by_model_rows) > model_limit
|
||||||
by_model = []
|
by_model = []
|
||||||
for mrow in by_model_result.all():
|
for mrow in by_model_rows[:model_limit]:
|
||||||
(
|
(
|
||||||
model_name,
|
model_name,
|
||||||
m_calls,
|
m_calls,
|
||||||
@@ -851,44 +978,59 @@ class MonitoringService:
|
|||||||
'avg_duration_ms': int((m_duration or 0) / m_calls) if m_calls > 0 else 0,
|
'avg_duration_ms': int((m_duration or 0) / m_calls) if m_calls > 0 else 0,
|
||||||
}
|
}
|
||||||
)
|
)
|
||||||
by_model.sort(key=lambda x: x['total_tokens'], reverse=True)
|
|
||||||
|
|
||||||
# ---- Time-bucketed series ----
|
# ---- Time-bucketed series ----
|
||||||
# Use a DB-agnostic bucketing approach: fetch (timestamp, tokens) rows and
|
# Aggregate before materialization. Requests may omit their time window,
|
||||||
# aggregate in Python. The window is bounded by the time filter, so this is
|
# so fetching every historical call and bucketing in Python is unsafe.
|
||||||
# cheap for typical dashboard ranges (hours/days).
|
engine = self.ap.persistence_mgr.get_db_engine()
|
||||||
series_query = _apply(
|
bucket_expression = self._token_bucket_expression(
|
||||||
sqlalchemy.select(
|
LLMCall.timestamp,
|
||||||
LLMCall.timestamp,
|
bucket=bucket,
|
||||||
LLMCall.input_tokens,
|
dialect_name=engine.dialect.name,
|
||||||
LLMCall.output_tokens,
|
)
|
||||||
LLMCall.total_tokens,
|
bucket_limit = self._timeseries_bucket_limit()
|
||||||
).order_by(LLMCall.timestamp.asc())
|
series_query = (
|
||||||
|
_apply(
|
||||||
|
sqlalchemy.select(
|
||||||
|
bucket_expression.label('bucket'),
|
||||||
|
sqlalchemy.func.coalesce(sqlalchemy.func.sum(LLMCall.input_tokens), 0),
|
||||||
|
sqlalchemy.func.coalesce(sqlalchemy.func.sum(LLMCall.output_tokens), 0),
|
||||||
|
sqlalchemy.func.coalesce(sqlalchemy.func.sum(LLMCall.total_tokens), 0),
|
||||||
|
sqlalchemy.func.count(LLMCall.id),
|
||||||
|
).group_by(bucket_expression)
|
||||||
|
)
|
||||||
|
.order_by(bucket_expression.desc())
|
||||||
|
.limit(bucket_limit + 1)
|
||||||
)
|
)
|
||||||
series_result = await self.ap.persistence_mgr.execute_async(series_query)
|
series_result = await self.ap.persistence_mgr.execute_async(series_query)
|
||||||
|
|
||||||
bucket_fmt = '%Y-%m-%d %H:00' if bucket == 'hour' else '%Y-%m-%d'
|
bucket_fmt = '%Y-%m-%d %H:00' if bucket == 'hour' else '%Y-%m-%d'
|
||||||
buckets: dict[str, dict] = {}
|
series_rows = series_result.all()
|
||||||
for srow in series_result.all():
|
timeseries_truncated = len(series_rows) > bucket_limit
|
||||||
ts, s_in, s_out, s_total = srow
|
timeseries = []
|
||||||
if ts is None:
|
for bucket_value, s_in, s_out, s_total, calls in reversed(series_rows[:bucket_limit]):
|
||||||
|
if bucket_value is None:
|
||||||
continue
|
continue
|
||||||
key = ts.strftime(bucket_fmt)
|
bucket_key = (
|
||||||
b = buckets.setdefault(
|
bucket_value.strftime(bucket_fmt)
|
||||||
key,
|
if isinstance(bucket_value, (datetime.datetime, datetime.date))
|
||||||
{'bucket': key, 'input_tokens': 0, 'output_tokens': 0, 'total_tokens': 0, 'calls': 0},
|
else str(bucket_value)
|
||||||
|
)
|
||||||
|
timeseries.append(
|
||||||
|
{
|
||||||
|
'bucket': bucket_key,
|
||||||
|
'input_tokens': int(s_in or 0),
|
||||||
|
'output_tokens': int(s_out or 0),
|
||||||
|
'total_tokens': int(s_total or 0),
|
||||||
|
'calls': int(calls or 0),
|
||||||
|
}
|
||||||
)
|
)
|
||||||
b['input_tokens'] += int(s_in or 0)
|
|
||||||
b['output_tokens'] += int(s_out or 0)
|
|
||||||
b['total_tokens'] += int(s_total or 0)
|
|
||||||
b['calls'] += 1
|
|
||||||
|
|
||||||
timeseries = [buckets[k] for k in sorted(buckets.keys())]
|
|
||||||
|
|
||||||
return {
|
return {
|
||||||
'summary': summary,
|
'summary': summary,
|
||||||
'by_model': by_model,
|
'by_model': by_model,
|
||||||
|
'by_model_truncated': by_model_truncated,
|
||||||
'timeseries': timeseries,
|
'timeseries': timeseries,
|
||||||
|
'timeseries_truncated': timeseries_truncated,
|
||||||
'bucket': bucket,
|
'bucket': bucket,
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -904,6 +1046,7 @@ class MonitoringService:
|
|||||||
offset: int = 0,
|
offset: int = 0,
|
||||||
) -> tuple[list[dict], int]:
|
) -> tuple[list[dict], int]:
|
||||||
"""Get messages with filters"""
|
"""Get messages with filters"""
|
||||||
|
limit, offset = self.normalize_page_window(limit, offset)
|
||||||
workspace_uuid = require_workspace_uuid(context)
|
workspace_uuid = require_workspace_uuid(context)
|
||||||
conditions = [persistence_monitoring.MonitoringMessage.workspace_uuid == workspace_uuid]
|
conditions = [persistence_monitoring.MonitoringMessage.workspace_uuid == workspace_uuid]
|
||||||
|
|
||||||
@@ -958,6 +1101,7 @@ class MonitoringService:
|
|||||||
offset: int = 0,
|
offset: int = 0,
|
||||||
) -> tuple[list[dict], int]:
|
) -> tuple[list[dict], int]:
|
||||||
"""Get LLM calls with filters"""
|
"""Get LLM calls with filters"""
|
||||||
|
limit, offset = self.normalize_page_window(limit, offset)
|
||||||
workspace_uuid = require_workspace_uuid(context)
|
workspace_uuid = require_workspace_uuid(context)
|
||||||
conditions = [persistence_monitoring.MonitoringLLMCall.workspace_uuid == workspace_uuid]
|
conditions = [persistence_monitoring.MonitoringLLMCall.workspace_uuid == workspace_uuid]
|
||||||
|
|
||||||
@@ -1012,6 +1156,7 @@ class MonitoringService:
|
|||||||
offset: int = 0,
|
offset: int = 0,
|
||||||
) -> tuple[list[dict], int]:
|
) -> tuple[list[dict], int]:
|
||||||
"""Get tool calls with filters"""
|
"""Get tool calls with filters"""
|
||||||
|
limit, offset = self.normalize_page_window(limit, offset)
|
||||||
workspace_uuid = require_workspace_uuid(context)
|
workspace_uuid = require_workspace_uuid(context)
|
||||||
conditions = [persistence_monitoring.MonitoringToolCall.workspace_uuid == workspace_uuid]
|
conditions = [persistence_monitoring.MonitoringToolCall.workspace_uuid == workspace_uuid]
|
||||||
|
|
||||||
@@ -1064,6 +1209,7 @@ class MonitoringService:
|
|||||||
offset: int = 0,
|
offset: int = 0,
|
||||||
) -> tuple[list[dict], int]:
|
) -> tuple[list[dict], int]:
|
||||||
"""Get embedding calls with filters"""
|
"""Get embedding calls with filters"""
|
||||||
|
limit, offset = self.normalize_page_window(limit, offset)
|
||||||
workspace_uuid = require_workspace_uuid(context)
|
workspace_uuid = require_workspace_uuid(context)
|
||||||
conditions = [persistence_monitoring.MonitoringEmbeddingCall.workspace_uuid == workspace_uuid]
|
conditions = [persistence_monitoring.MonitoringEmbeddingCall.workspace_uuid == workspace_uuid]
|
||||||
|
|
||||||
@@ -1116,6 +1262,7 @@ class MonitoringService:
|
|||||||
offset: int = 0,
|
offset: int = 0,
|
||||||
) -> tuple[list[dict], int]:
|
) -> tuple[list[dict], int]:
|
||||||
"""Get sessions with filters"""
|
"""Get sessions with filters"""
|
||||||
|
limit, offset = self.normalize_page_window(limit, offset)
|
||||||
workspace_uuid = require_workspace_uuid(context)
|
workspace_uuid = require_workspace_uuid(context)
|
||||||
conditions = [persistence_monitoring.MonitoringSession.workspace_uuid == workspace_uuid]
|
conditions = [persistence_monitoring.MonitoringSession.workspace_uuid == workspace_uuid]
|
||||||
|
|
||||||
@@ -1171,6 +1318,7 @@ class MonitoringService:
|
|||||||
offset: int = 0,
|
offset: int = 0,
|
||||||
) -> tuple[list[dict], int]:
|
) -> tuple[list[dict], int]:
|
||||||
"""Get errors with filters"""
|
"""Get errors with filters"""
|
||||||
|
limit, offset = self.normalize_page_window(limit, offset)
|
||||||
workspace_uuid = require_workspace_uuid(context)
|
workspace_uuid = require_workspace_uuid(context)
|
||||||
conditions = [persistence_monitoring.MonitoringError.workspace_uuid == workspace_uuid]
|
conditions = [persistence_monitoring.MonitoringError.workspace_uuid == workspace_uuid]
|
||||||
|
|
||||||
@@ -1218,8 +1366,9 @@ class MonitoringService:
|
|||||||
context: TenantContext,
|
context: TenantContext,
|
||||||
session_id: str,
|
session_id: str,
|
||||||
) -> dict:
|
) -> dict:
|
||||||
"""Get detailed analysis for a specific session"""
|
"""Get bounded session details with full statistics computed in SQL."""
|
||||||
workspace_uuid = require_workspace_uuid(context)
|
workspace_uuid = require_workspace_uuid(context)
|
||||||
|
detail_limit = self._detail_limit()
|
||||||
# Get session info
|
# Get session info
|
||||||
session_query = sqlalchemy.select(persistence_monitoring.MonitoringSession).where(
|
session_query = sqlalchemy.select(persistence_monitoring.MonitoringSession).where(
|
||||||
persistence_monitoring.MonitoringSession.workspace_uuid == workspace_uuid,
|
persistence_monitoring.MonitoringSession.workspace_uuid == workspace_uuid,
|
||||||
@@ -1236,60 +1385,99 @@ class MonitoringService:
|
|||||||
|
|
||||||
session = session_row[0] if isinstance(session_row, tuple) else session_row
|
session = session_row[0] if isinstance(session_row, tuple) else session_row
|
||||||
|
|
||||||
# Get messages for this session
|
message_stats_result = await self.ap.persistence_mgr.execute_async(
|
||||||
messages_query = (
|
sqlalchemy.select(
|
||||||
sqlalchemy.select(persistence_monitoring.MonitoringMessage)
|
sqlalchemy.func.count(persistence_monitoring.MonitoringMessage.id).label('total'),
|
||||||
.where(
|
sqlalchemy.func.sum(
|
||||||
|
sqlalchemy.case(
|
||||||
|
(persistence_monitoring.MonitoringMessage.status == 'success', 1),
|
||||||
|
else_=0,
|
||||||
|
)
|
||||||
|
).label('success'),
|
||||||
|
sqlalchemy.func.sum(
|
||||||
|
sqlalchemy.case(
|
||||||
|
(persistence_monitoring.MonitoringMessage.status == 'error', 1),
|
||||||
|
else_=0,
|
||||||
|
)
|
||||||
|
).label('error'),
|
||||||
|
sqlalchemy.func.sum(
|
||||||
|
sqlalchemy.case(
|
||||||
|
(persistence_monitoring.MonitoringMessage.status == 'pending', 1),
|
||||||
|
else_=0,
|
||||||
|
)
|
||||||
|
).label('pending'),
|
||||||
|
sqlalchemy.func.min(persistence_monitoring.MonitoringMessage.timestamp).label('first_timestamp'),
|
||||||
|
sqlalchemy.func.max(persistence_monitoring.MonitoringMessage.timestamp).label('last_timestamp'),
|
||||||
|
).where(
|
||||||
persistence_monitoring.MonitoringMessage.workspace_uuid == workspace_uuid,
|
persistence_monitoring.MonitoringMessage.workspace_uuid == workspace_uuid,
|
||||||
persistence_monitoring.MonitoringMessage.session_id == session_id,
|
persistence_monitoring.MonitoringMessage.session_id == session_id,
|
||||||
)
|
)
|
||||||
.order_by(persistence_monitoring.MonitoringMessage.timestamp.asc())
|
|
||||||
)
|
)
|
||||||
messages_result = await self.ap.persistence_mgr.execute_async(messages_query)
|
message_stats = message_stats_result.one()
|
||||||
messages_rows = messages_result.all()
|
|
||||||
|
|
||||||
# Count messages by status
|
llm_stats_result = await self.ap.persistence_mgr.execute_async(
|
||||||
success_messages = 0
|
sqlalchemy.select(
|
||||||
error_messages = 0
|
sqlalchemy.func.count(persistence_monitoring.MonitoringLLMCall.id).label('total_calls'),
|
||||||
pending_messages = 0
|
sqlalchemy.func.coalesce(
|
||||||
for row in messages_rows:
|
sqlalchemy.func.sum(persistence_monitoring.MonitoringLLMCall.input_tokens),
|
||||||
msg = row[0] if isinstance(row, tuple) else row
|
0,
|
||||||
if msg.status == 'success':
|
).label('total_input_tokens'),
|
||||||
success_messages += 1
|
sqlalchemy.func.coalesce(
|
||||||
elif msg.status == 'error':
|
sqlalchemy.func.sum(persistence_monitoring.MonitoringLLMCall.output_tokens),
|
||||||
error_messages += 1
|
0,
|
||||||
elif msg.status == 'pending':
|
).label('total_output_tokens'),
|
||||||
pending_messages += 1
|
sqlalchemy.func.coalesce(
|
||||||
|
sqlalchemy.func.sum(persistence_monitoring.MonitoringLLMCall.total_tokens),
|
||||||
# Get LLM calls for this session
|
0,
|
||||||
llm_query = sqlalchemy.select(persistence_monitoring.MonitoringLLMCall).where(
|
).label('total_tokens'),
|
||||||
persistence_monitoring.MonitoringLLMCall.workspace_uuid == workspace_uuid,
|
sqlalchemy.func.coalesce(
|
||||||
persistence_monitoring.MonitoringLLMCall.session_id == session_id,
|
sqlalchemy.func.sum(persistence_monitoring.MonitoringLLMCall.duration),
|
||||||
|
0,
|
||||||
|
).label('total_duration'),
|
||||||
|
sqlalchemy.func.sum(
|
||||||
|
sqlalchemy.case(
|
||||||
|
(persistence_monitoring.MonitoringLLMCall.status == 'success', 1),
|
||||||
|
else_=0,
|
||||||
|
)
|
||||||
|
).label('success_calls'),
|
||||||
|
sqlalchemy.func.sum(
|
||||||
|
sqlalchemy.case(
|
||||||
|
(persistence_monitoring.MonitoringLLMCall.status != 'success', 1),
|
||||||
|
else_=0,
|
||||||
|
)
|
||||||
|
).label('error_calls'),
|
||||||
|
).where(
|
||||||
|
persistence_monitoring.MonitoringLLMCall.workspace_uuid == workspace_uuid,
|
||||||
|
persistence_monitoring.MonitoringLLMCall.session_id == session_id,
|
||||||
|
)
|
||||||
)
|
)
|
||||||
llm_result = await self.ap.persistence_mgr.execute_async(llm_query)
|
llm_stats = llm_stats_result.one()
|
||||||
llm_rows = llm_result.all()
|
|
||||||
|
|
||||||
# Calculate LLM statistics
|
tool_stats_result = await self.ap.persistence_mgr.execute_async(
|
||||||
total_llm_calls = len(llm_rows)
|
sqlalchemy.select(
|
||||||
total_input_tokens = 0
|
sqlalchemy.func.count(persistence_monitoring.MonitoringToolCall.id).label('total_calls'),
|
||||||
total_output_tokens = 0
|
sqlalchemy.func.coalesce(
|
||||||
total_tokens = 0
|
sqlalchemy.func.sum(persistence_monitoring.MonitoringToolCall.duration),
|
||||||
total_duration = 0
|
0,
|
||||||
success_llm_calls = 0
|
).label('total_duration'),
|
||||||
error_llm_calls = 0
|
sqlalchemy.func.sum(
|
||||||
|
sqlalchemy.case(
|
||||||
for row in llm_rows:
|
(persistence_monitoring.MonitoringToolCall.status == 'success', 1),
|
||||||
llm_call = row[0] if isinstance(row, tuple) else row
|
else_=0,
|
||||||
total_input_tokens += llm_call.input_tokens
|
)
|
||||||
total_output_tokens += llm_call.output_tokens
|
).label('success_calls'),
|
||||||
total_tokens += llm_call.total_tokens
|
sqlalchemy.func.sum(
|
||||||
total_duration += llm_call.duration
|
sqlalchemy.case(
|
||||||
if llm_call.status == 'success':
|
(persistence_monitoring.MonitoringToolCall.status != 'success', 1),
|
||||||
success_llm_calls += 1
|
else_=0,
|
||||||
else:
|
)
|
||||||
error_llm_calls += 1
|
).label('error_calls'),
|
||||||
|
).where(
|
||||||
# Get tool calls for this session
|
persistence_monitoring.MonitoringToolCall.workspace_uuid == workspace_uuid,
|
||||||
|
persistence_monitoring.MonitoringToolCall.session_id == session_id,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
tool_stats = tool_stats_result.one()
|
||||||
tool_query = (
|
tool_query = (
|
||||||
sqlalchemy.select(persistence_monitoring.MonitoringToolCall)
|
sqlalchemy.select(persistence_monitoring.MonitoringToolCall)
|
||||||
.where(
|
.where(
|
||||||
@@ -1297,9 +1485,12 @@ class MonitoringService:
|
|||||||
persistence_monitoring.MonitoringToolCall.session_id == session_id,
|
persistence_monitoring.MonitoringToolCall.session_id == session_id,
|
||||||
)
|
)
|
||||||
.order_by(persistence_monitoring.MonitoringToolCall.timestamp.asc())
|
.order_by(persistence_monitoring.MonitoringToolCall.timestamp.asc())
|
||||||
|
.limit(detail_limit + 1)
|
||||||
)
|
)
|
||||||
tool_result = await self.ap.persistence_mgr.execute_async(tool_query)
|
tool_result = await self.ap.persistence_mgr.execute_async(tool_query)
|
||||||
tool_rows = tool_result.all()
|
tool_rows = tool_result.all()
|
||||||
|
tool_calls_truncated = len(tool_rows) > detail_limit
|
||||||
|
tool_rows = tool_rows[:detail_limit]
|
||||||
|
|
||||||
tool_calls = [
|
tool_calls = [
|
||||||
self.ap.persistence_mgr.serialize_model(
|
self.ap.persistence_mgr.serialize_model(
|
||||||
@@ -1308,19 +1499,6 @@ class MonitoringService:
|
|||||||
for row in tool_rows
|
for row in tool_rows
|
||||||
]
|
]
|
||||||
|
|
||||||
total_tool_calls = len(tool_rows)
|
|
||||||
success_tool_calls = 0
|
|
||||||
error_tool_calls = 0
|
|
||||||
total_tool_duration = 0
|
|
||||||
for row in tool_rows:
|
|
||||||
tool_call = row[0] if isinstance(row, tuple) else row
|
|
||||||
total_tool_duration += tool_call.duration
|
|
||||||
if tool_call.status == 'success':
|
|
||||||
success_tool_calls += 1
|
|
||||||
else:
|
|
||||||
error_tool_calls += 1
|
|
||||||
|
|
||||||
# Get errors for this session
|
|
||||||
error_query = (
|
error_query = (
|
||||||
sqlalchemy.select(persistence_monitoring.MonitoringError)
|
sqlalchemy.select(persistence_monitoring.MonitoringError)
|
||||||
.where(
|
.where(
|
||||||
@@ -1328,9 +1506,12 @@ class MonitoringService:
|
|||||||
persistence_monitoring.MonitoringError.session_id == session_id,
|
persistence_monitoring.MonitoringError.session_id == session_id,
|
||||||
)
|
)
|
||||||
.order_by(persistence_monitoring.MonitoringError.timestamp.desc())
|
.order_by(persistence_monitoring.MonitoringError.timestamp.desc())
|
||||||
|
.limit(detail_limit + 1)
|
||||||
)
|
)
|
||||||
error_result = await self.ap.persistence_mgr.execute_async(error_query)
|
error_result = await self.ap.persistence_mgr.execute_async(error_query)
|
||||||
error_rows = error_result.all()
|
error_rows = error_result.all()
|
||||||
|
errors_truncated = len(error_rows) > detail_limit
|
||||||
|
error_rows = error_rows[:detail_limit]
|
||||||
|
|
||||||
errors = [
|
errors = [
|
||||||
self.ap.persistence_mgr.serialize_model(
|
self.ap.persistence_mgr.serialize_model(
|
||||||
@@ -1339,42 +1520,49 @@ class MonitoringService:
|
|||||||
for row in error_rows
|
for row in error_rows
|
||||||
]
|
]
|
||||||
|
|
||||||
# Calculate session duration
|
if message_stats.first_timestamp is not None and message_stats.last_timestamp is not None:
|
||||||
if messages_rows:
|
session_duration_seconds = int(
|
||||||
first_msg = messages_rows[0][0] if isinstance(messages_rows[0], tuple) else messages_rows[0]
|
(message_stats.last_timestamp - message_stats.first_timestamp).total_seconds()
|
||||||
last_msg = messages_rows[-1][0] if isinstance(messages_rows[-1], tuple) else messages_rows[-1]
|
)
|
||||||
session_duration_seconds = int((last_msg.timestamp - first_msg.timestamp).total_seconds())
|
|
||||||
else:
|
else:
|
||||||
session_duration_seconds = 0
|
session_duration_seconds = 0
|
||||||
|
total_llm_calls = int(llm_stats.total_calls or 0)
|
||||||
|
total_tool_calls = int(tool_stats.total_calls or 0)
|
||||||
|
|
||||||
return {
|
return {
|
||||||
'session_id': session_id,
|
'session_id': session_id,
|
||||||
'found': True,
|
'found': True,
|
||||||
'session': self.ap.persistence_mgr.serialize_model(persistence_monitoring.MonitoringSession, session),
|
'session': self.ap.persistence_mgr.serialize_model(persistence_monitoring.MonitoringSession, session),
|
||||||
'message_stats': {
|
'message_stats': {
|
||||||
'total': len(messages_rows),
|
'total': int(message_stats.total or 0),
|
||||||
'success': success_messages,
|
'success': int(message_stats.success or 0),
|
||||||
'error': error_messages,
|
'error': int(message_stats.error or 0),
|
||||||
'pending': pending_messages,
|
'pending': int(message_stats.pending or 0),
|
||||||
},
|
},
|
||||||
'llm_stats': {
|
'llm_stats': {
|
||||||
'total_calls': total_llm_calls,
|
'total_calls': total_llm_calls,
|
||||||
'success_calls': success_llm_calls,
|
'success_calls': int(llm_stats.success_calls or 0),
|
||||||
'error_calls': error_llm_calls,
|
'error_calls': int(llm_stats.error_calls or 0),
|
||||||
'total_input_tokens': total_input_tokens,
|
'total_input_tokens': int(llm_stats.total_input_tokens or 0),
|
||||||
'total_output_tokens': total_output_tokens,
|
'total_output_tokens': int(llm_stats.total_output_tokens or 0),
|
||||||
'total_tokens': total_tokens,
|
'total_tokens': int(llm_stats.total_tokens or 0),
|
||||||
'average_duration_ms': int(total_duration / total_llm_calls) if total_llm_calls > 0 else 0,
|
'average_duration_ms': (int(llm_stats.total_duration / total_llm_calls) if total_llm_calls > 0 else 0),
|
||||||
},
|
},
|
||||||
'tool_calls': tool_calls,
|
'tool_calls': tool_calls,
|
||||||
'tool_stats': {
|
'tool_stats': {
|
||||||
'total_calls': total_tool_calls,
|
'total_calls': total_tool_calls,
|
||||||
'success_calls': success_tool_calls,
|
'success_calls': int(tool_stats.success_calls or 0),
|
||||||
'error_calls': error_tool_calls,
|
'error_calls': int(tool_stats.error_calls or 0),
|
||||||
'total_duration_ms': total_tool_duration,
|
'total_duration_ms': int(tool_stats.total_duration or 0),
|
||||||
'average_duration_ms': int(total_tool_duration / total_tool_calls) if total_tool_calls > 0 else 0,
|
'average_duration_ms': (
|
||||||
|
int(tool_stats.total_duration / total_tool_calls) if total_tool_calls > 0 else 0
|
||||||
|
),
|
||||||
},
|
},
|
||||||
'errors': errors,
|
'errors': errors,
|
||||||
|
'detail_truncated': {
|
||||||
|
'tool_calls': tool_calls_truncated,
|
||||||
|
'errors': errors_truncated,
|
||||||
|
},
|
||||||
'session_duration_seconds': session_duration_seconds,
|
'session_duration_seconds': session_duration_seconds,
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -1383,8 +1571,9 @@ class MonitoringService:
|
|||||||
context: TenantContext,
|
context: TenantContext,
|
||||||
message_id: str,
|
message_id: str,
|
||||||
) -> dict:
|
) -> dict:
|
||||||
"""Get detailed information for a specific message including associated LLM calls and errors"""
|
"""Get bounded message details with full statistics computed in SQL."""
|
||||||
workspace_uuid = require_workspace_uuid(context)
|
workspace_uuid = require_workspace_uuid(context)
|
||||||
|
detail_limit = self._detail_limit()
|
||||||
# Get message info
|
# Get message info
|
||||||
message_query = sqlalchemy.select(persistence_monitoring.MonitoringMessage).where(
|
message_query = sqlalchemy.select(persistence_monitoring.MonitoringMessage).where(
|
||||||
persistence_monitoring.MonitoringMessage.workspace_uuid == workspace_uuid,
|
persistence_monitoring.MonitoringMessage.workspace_uuid == workspace_uuid,
|
||||||
@@ -1401,7 +1590,31 @@ class MonitoringService:
|
|||||||
|
|
||||||
message = message_row[0] if isinstance(message_row, tuple) else message_row
|
message = message_row[0] if isinstance(message_row, tuple) else message_row
|
||||||
|
|
||||||
# Get LLM calls for this message
|
llm_stats_result = await self.ap.persistence_mgr.execute_async(
|
||||||
|
sqlalchemy.select(
|
||||||
|
sqlalchemy.func.count(persistence_monitoring.MonitoringLLMCall.id).label('total_calls'),
|
||||||
|
sqlalchemy.func.coalesce(
|
||||||
|
sqlalchemy.func.sum(persistence_monitoring.MonitoringLLMCall.input_tokens),
|
||||||
|
0,
|
||||||
|
).label('total_input_tokens'),
|
||||||
|
sqlalchemy.func.coalesce(
|
||||||
|
sqlalchemy.func.sum(persistence_monitoring.MonitoringLLMCall.output_tokens),
|
||||||
|
0,
|
||||||
|
).label('total_output_tokens'),
|
||||||
|
sqlalchemy.func.coalesce(
|
||||||
|
sqlalchemy.func.sum(persistence_monitoring.MonitoringLLMCall.total_tokens),
|
||||||
|
0,
|
||||||
|
).label('total_tokens'),
|
||||||
|
sqlalchemy.func.coalesce(
|
||||||
|
sqlalchemy.func.sum(persistence_monitoring.MonitoringLLMCall.duration),
|
||||||
|
0,
|
||||||
|
).label('total_duration'),
|
||||||
|
).where(
|
||||||
|
persistence_monitoring.MonitoringLLMCall.workspace_uuid == workspace_uuid,
|
||||||
|
persistence_monitoring.MonitoringLLMCall.message_id == message_id,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
llm_stats = llm_stats_result.one()
|
||||||
llm_query = (
|
llm_query = (
|
||||||
sqlalchemy.select(persistence_monitoring.MonitoringLLMCall)
|
sqlalchemy.select(persistence_monitoring.MonitoringLLMCall)
|
||||||
.where(
|
.where(
|
||||||
@@ -1409,9 +1622,12 @@ class MonitoringService:
|
|||||||
persistence_monitoring.MonitoringLLMCall.message_id == message_id,
|
persistence_monitoring.MonitoringLLMCall.message_id == message_id,
|
||||||
)
|
)
|
||||||
.order_by(persistence_monitoring.MonitoringLLMCall.timestamp.asc())
|
.order_by(persistence_monitoring.MonitoringLLMCall.timestamp.asc())
|
||||||
|
.limit(detail_limit + 1)
|
||||||
)
|
)
|
||||||
llm_result = await self.ap.persistence_mgr.execute_async(llm_query)
|
llm_result = await self.ap.persistence_mgr.execute_async(llm_query)
|
||||||
llm_rows = llm_result.all()
|
llm_rows = llm_result.all()
|
||||||
|
llm_calls_truncated = len(llm_rows) > detail_limit
|
||||||
|
llm_rows = llm_rows[:detail_limit]
|
||||||
|
|
||||||
llm_calls = [
|
llm_calls = [
|
||||||
self.ap.persistence_mgr.serialize_model(
|
self.ap.persistence_mgr.serialize_model(
|
||||||
@@ -1420,13 +1636,6 @@ class MonitoringService:
|
|||||||
for row in llm_rows
|
for row in llm_rows
|
||||||
]
|
]
|
||||||
|
|
||||||
# Calculate LLM statistics
|
|
||||||
total_input_tokens = sum(call.input_tokens for call in llm_rows)
|
|
||||||
total_output_tokens = sum(call.output_tokens for call in llm_rows)
|
|
||||||
total_tokens = sum(call.total_tokens for call in llm_rows)
|
|
||||||
total_duration = sum(call.duration for call in llm_rows)
|
|
||||||
|
|
||||||
# Get errors for this message
|
|
||||||
error_query = (
|
error_query = (
|
||||||
sqlalchemy.select(persistence_monitoring.MonitoringError)
|
sqlalchemy.select(persistence_monitoring.MonitoringError)
|
||||||
.where(
|
.where(
|
||||||
@@ -1434,9 +1643,12 @@ class MonitoringService:
|
|||||||
persistence_monitoring.MonitoringError.message_id == message_id,
|
persistence_monitoring.MonitoringError.message_id == message_id,
|
||||||
)
|
)
|
||||||
.order_by(persistence_monitoring.MonitoringError.timestamp.asc())
|
.order_by(persistence_monitoring.MonitoringError.timestamp.asc())
|
||||||
|
.limit(detail_limit + 1)
|
||||||
)
|
)
|
||||||
error_result = await self.ap.persistence_mgr.execute_async(error_query)
|
error_result = await self.ap.persistence_mgr.execute_async(error_query)
|
||||||
error_rows = error_result.all()
|
error_rows = error_result.all()
|
||||||
|
errors_truncated = len(error_rows) > detail_limit
|
||||||
|
error_rows = error_rows[:detail_limit]
|
||||||
|
|
||||||
errors = [
|
errors = [
|
||||||
self.ap.persistence_mgr.serialize_model(
|
self.ap.persistence_mgr.serialize_model(
|
||||||
@@ -1444,6 +1656,7 @@ class MonitoringService:
|
|||||||
)
|
)
|
||||||
for row in error_rows
|
for row in error_rows
|
||||||
]
|
]
|
||||||
|
total_llm_calls = int(llm_stats.total_calls or 0)
|
||||||
|
|
||||||
return {
|
return {
|
||||||
'message_id': message_id,
|
'message_id': message_id,
|
||||||
@@ -1451,14 +1664,18 @@ class MonitoringService:
|
|||||||
'message': self.ap.persistence_mgr.serialize_model(persistence_monitoring.MonitoringMessage, message),
|
'message': self.ap.persistence_mgr.serialize_model(persistence_monitoring.MonitoringMessage, message),
|
||||||
'llm_calls': llm_calls,
|
'llm_calls': llm_calls,
|
||||||
'llm_stats': {
|
'llm_stats': {
|
||||||
'total_calls': len(llm_rows),
|
'total_calls': total_llm_calls,
|
||||||
'total_input_tokens': total_input_tokens,
|
'total_input_tokens': int(llm_stats.total_input_tokens or 0),
|
||||||
'total_output_tokens': total_output_tokens,
|
'total_output_tokens': int(llm_stats.total_output_tokens or 0),
|
||||||
'total_tokens': total_tokens,
|
'total_tokens': int(llm_stats.total_tokens or 0),
|
||||||
'total_duration_ms': total_duration,
|
'total_duration_ms': int(llm_stats.total_duration or 0),
|
||||||
'average_duration_ms': int(total_duration / len(llm_rows)) if len(llm_rows) > 0 else 0,
|
'average_duration_ms': (int(llm_stats.total_duration / total_llm_calls) if total_llm_calls > 0 else 0),
|
||||||
},
|
},
|
||||||
'errors': errors,
|
'errors': errors,
|
||||||
|
'detail_truncated': {
|
||||||
|
'llm_calls': llm_calls_truncated,
|
||||||
|
'errors': errors_truncated,
|
||||||
|
},
|
||||||
}
|
}
|
||||||
|
|
||||||
# ========== Export Methods ==========
|
# ========== Export Methods ==========
|
||||||
@@ -1548,6 +1765,7 @@ class MonitoringService:
|
|||||||
limit: int = 100000,
|
limit: int = 100000,
|
||||||
) -> list[dict]:
|
) -> list[dict]:
|
||||||
"""Export messages as list of dictionaries for CSV conversion"""
|
"""Export messages as list of dictionaries for CSV conversion"""
|
||||||
|
limit = self.normalize_export_limit(limit)
|
||||||
workspace_uuid = require_workspace_uuid(context)
|
workspace_uuid = require_workspace_uuid(context)
|
||||||
conditions = [persistence_monitoring.MonitoringMessage.workspace_uuid == workspace_uuid]
|
conditions = [persistence_monitoring.MonitoringMessage.workspace_uuid == workspace_uuid]
|
||||||
|
|
||||||
@@ -1603,6 +1821,7 @@ class MonitoringService:
|
|||||||
limit: int = 100000,
|
limit: int = 100000,
|
||||||
) -> list[dict]:
|
) -> list[dict]:
|
||||||
"""Export LLM calls as list of dictionaries for CSV conversion"""
|
"""Export LLM calls as list of dictionaries for CSV conversion"""
|
||||||
|
limit = self.normalize_export_limit(limit)
|
||||||
workspace_uuid = require_workspace_uuid(context)
|
workspace_uuid = require_workspace_uuid(context)
|
||||||
conditions = [persistence_monitoring.MonitoringLLMCall.workspace_uuid == workspace_uuid]
|
conditions = [persistence_monitoring.MonitoringLLMCall.workspace_uuid == workspace_uuid]
|
||||||
|
|
||||||
@@ -1657,6 +1876,7 @@ class MonitoringService:
|
|||||||
limit: int = 100000,
|
limit: int = 100000,
|
||||||
) -> list[dict]:
|
) -> list[dict]:
|
||||||
"""Export embedding calls as list of dictionaries for CSV conversion"""
|
"""Export embedding calls as list of dictionaries for CSV conversion"""
|
||||||
|
limit = self.normalize_export_limit(limit)
|
||||||
workspace_uuid = require_workspace_uuid(context)
|
workspace_uuid = require_workspace_uuid(context)
|
||||||
conditions = [persistence_monitoring.MonitoringEmbeddingCall.workspace_uuid == workspace_uuid]
|
conditions = [persistence_monitoring.MonitoringEmbeddingCall.workspace_uuid == workspace_uuid]
|
||||||
|
|
||||||
@@ -1708,6 +1928,7 @@ class MonitoringService:
|
|||||||
limit: int = 100000,
|
limit: int = 100000,
|
||||||
) -> list[dict]:
|
) -> list[dict]:
|
||||||
"""Export errors as list of dictionaries for CSV conversion"""
|
"""Export errors as list of dictionaries for CSV conversion"""
|
||||||
|
limit = self.normalize_export_limit(limit)
|
||||||
workspace_uuid = require_workspace_uuid(context)
|
workspace_uuid = require_workspace_uuid(context)
|
||||||
conditions = [persistence_monitoring.MonitoringError.workspace_uuid == workspace_uuid]
|
conditions = [persistence_monitoring.MonitoringError.workspace_uuid == workspace_uuid]
|
||||||
|
|
||||||
@@ -1758,6 +1979,7 @@ class MonitoringService:
|
|||||||
limit: int = 100000,
|
limit: int = 100000,
|
||||||
) -> list[dict]:
|
) -> list[dict]:
|
||||||
"""Export sessions as list of dictionaries for CSV conversion"""
|
"""Export sessions as list of dictionaries for CSV conversion"""
|
||||||
|
limit = self.normalize_export_limit(limit)
|
||||||
workspace_uuid = require_workspace_uuid(context)
|
workspace_uuid = require_workspace_uuid(context)
|
||||||
conditions = [persistence_monitoring.MonitoringSession.workspace_uuid == workspace_uuid]
|
conditions = [persistence_monitoring.MonitoringSession.workspace_uuid == workspace_uuid]
|
||||||
|
|
||||||
@@ -2011,6 +2233,7 @@ class MonitoringService:
|
|||||||
offset: int = 0,
|
offset: int = 0,
|
||||||
) -> tuple[list[dict], int]:
|
) -> tuple[list[dict], int]:
|
||||||
"""Get feedback list with filters."""
|
"""Get feedback list with filters."""
|
||||||
|
limit, offset = self.normalize_page_window(limit, offset)
|
||||||
workspace_uuid = require_workspace_uuid(context)
|
workspace_uuid = require_workspace_uuid(context)
|
||||||
conditions = [persistence_monitoring.MonitoringFeedback.workspace_uuid == workspace_uuid]
|
conditions = [persistence_monitoring.MonitoringFeedback.workspace_uuid == workspace_uuid]
|
||||||
|
|
||||||
@@ -2063,6 +2286,7 @@ class MonitoringService:
|
|||||||
limit: int = 100000,
|
limit: int = 100000,
|
||||||
) -> list[dict]:
|
) -> list[dict]:
|
||||||
"""Export feedback as list of dictionaries for CSV conversion."""
|
"""Export feedback as list of dictionaries for CSV conversion."""
|
||||||
|
limit = self.normalize_export_limit(limit)
|
||||||
workspace_uuid = require_workspace_uuid(context)
|
workspace_uuid = require_workspace_uuid(context)
|
||||||
conditions = [persistence_monitoring.MonitoringFeedback.workspace_uuid == workspace_uuid]
|
conditions = [persistence_monitoring.MonitoringFeedback.workspace_uuid == workspace_uuid]
|
||||||
|
|
||||||
|
|||||||
+96
-69
@@ -309,13 +309,6 @@ class Application:
|
|||||||
name='cloud-manifest-refresh',
|
name='cloud-manifest-refresh',
|
||||||
scopes=[core_entities.LifecycleControlScope.APPLICATION],
|
scopes=[core_entities.LifecycleControlScope.APPLICATION],
|
||||||
)
|
)
|
||||||
if self.workspace_collaboration_service is not None:
|
|
||||||
self.task_mgr.create_task(
|
|
||||||
self.workspace_collaboration_service.run_expired_invitation_cleanup(),
|
|
||||||
name='workspace-invitation-cleanup',
|
|
||||||
scopes=[core_entities.LifecycleControlScope.APPLICATION],
|
|
||||||
)
|
|
||||||
|
|
||||||
await self.plugin_connector.initialize_plugins()
|
await self.plugin_connector.initialize_plugins()
|
||||||
|
|
||||||
# 后续可能会允许动态重启其他任务
|
# 后续可能会允许动态重启其他任务
|
||||||
@@ -354,73 +347,68 @@ class Application:
|
|||||||
scopes=[core_entities.LifecycleControlScope.APPLICATION],
|
scopes=[core_entities.LifecycleControlScope.APPLICATION],
|
||||||
)
|
)
|
||||||
|
|
||||||
# Start monitoring data cleanup task if enabled
|
|
||||||
monitoring_cfg = self.instance_config.data.get('monitoring', {})
|
monitoring_cfg = self.instance_config.data.get('monitoring', {})
|
||||||
auto_cleanup_cfg = monitoring_cfg.get('auto_cleanup', {})
|
auto_cleanup_cfg = monitoring_cfg.get('auto_cleanup', {})
|
||||||
if auto_cleanup_cfg.get('enabled', True):
|
monitoring_enabled = auto_cleanup_cfg.get('enabled', True)
|
||||||
retention_days = self._get_positive_int_config(
|
retention_days = self._get_positive_int_config(
|
||||||
auto_cleanup_cfg.get('retention_days', 30),
|
auto_cleanup_cfg.get('retention_days', 30),
|
||||||
default=30,
|
default=30,
|
||||||
name='monitoring.auto_cleanup.retention_days',
|
name='monitoring.auto_cleanup.retention_days',
|
||||||
)
|
)
|
||||||
delete_batch_size = self._get_positive_int_config(
|
delete_batch_size = self._get_positive_int_config(
|
||||||
auto_cleanup_cfg.get('delete_batch_size', 1000),
|
auto_cleanup_cfg.get('delete_batch_size', 1000),
|
||||||
default=1000,
|
default=1000,
|
||||||
name='monitoring.auto_cleanup.delete_batch_size',
|
name='monitoring.auto_cleanup.delete_batch_size',
|
||||||
)
|
)
|
||||||
check_interval_hours = self._get_positive_float_config(
|
monitoring_interval_seconds = (
|
||||||
|
self._get_positive_float_config(
|
||||||
auto_cleanup_cfg.get('check_interval_hours', 1),
|
auto_cleanup_cfg.get('check_interval_hours', 1),
|
||||||
default=1,
|
default=1,
|
||||||
name='monitoring.auto_cleanup.check_interval_hours',
|
name='monitoring.auto_cleanup.check_interval_hours',
|
||||||
)
|
)
|
||||||
|
* 3600
|
||||||
|
)
|
||||||
|
|
||||||
async def monitoring_cleanup_loop():
|
|
||||||
check_interval_seconds = check_interval_hours * 3600
|
|
||||||
while True:
|
|
||||||
try:
|
|
||||||
bindings = await self.workspace_service.list_active_execution_bindings()
|
|
||||||
for binding in bindings:
|
|
||||||
context = ExecutionContext(
|
|
||||||
instance_uuid=binding.instance_uuid,
|
|
||||||
workspace_uuid=binding.workspace_uuid,
|
|
||||||
placement_generation=binding.placement_generation,
|
|
||||||
trigger_principal=PrincipalContext(PrincipalType.SYSTEM),
|
|
||||||
)
|
|
||||||
deleted = await self.monitoring_service.cleanup_expired_records(
|
|
||||||
context,
|
|
||||||
retention_days,
|
|
||||||
batch_size=delete_batch_size,
|
|
||||||
)
|
|
||||||
total_deleted = sum(deleted.values())
|
|
||||||
if total_deleted > 0:
|
|
||||||
self.logger.info(
|
|
||||||
f'Monitoring auto-cleanup: deleted {total_deleted} expired records '
|
|
||||||
f'for Workspace {context.workspace_uuid} (retention={retention_days}d): {deleted}'
|
|
||||||
)
|
|
||||||
except Exception as e:
|
|
||||||
self.logger.warning(f'Monitoring auto-cleanup error: {e}')
|
|
||||||
await asyncio.sleep(check_interval_seconds)
|
|
||||||
|
|
||||||
self.task_mgr.create_task(
|
|
||||||
monitoring_cleanup_loop(),
|
|
||||||
name='monitoring-cleanup',
|
|
||||||
scopes=[core_entities.LifecycleControlScope.APPLICATION],
|
|
||||||
)
|
|
||||||
|
|
||||||
# Start storage/log maintenance task if enabled
|
|
||||||
storage_cleanup_cfg = self.instance_config.data.get('storage', {}).get('cleanup', {})
|
storage_cleanup_cfg = self.instance_config.data.get('storage', {}).get('cleanup', {})
|
||||||
if storage_cleanup_cfg.get('enabled', True) and self.maintenance_service is not None:
|
storage_enabled = storage_cleanup_cfg.get('enabled', True) and self.maintenance_service is not None
|
||||||
check_interval_hours = self._get_positive_float_config(
|
storage_interval_seconds = (
|
||||||
|
self._get_positive_float_config(
|
||||||
storage_cleanup_cfg.get('check_interval_hours', 1),
|
storage_cleanup_cfg.get('check_interval_hours', 1),
|
||||||
default=1,
|
default=1,
|
||||||
name='storage.cleanup.check_interval_hours',
|
name='storage.cleanup.check_interval_hours',
|
||||||
)
|
)
|
||||||
|
* 3600
|
||||||
|
)
|
||||||
|
|
||||||
async def storage_cleanup_loop():
|
maintenance_intervals: dict[str, float] = {}
|
||||||
check_interval_seconds = check_interval_hours * 3600
|
if monitoring_enabled:
|
||||||
|
maintenance_intervals['monitoring'] = monitoring_interval_seconds
|
||||||
|
if storage_enabled:
|
||||||
|
maintenance_intervals['storage'] = storage_interval_seconds
|
||||||
|
if self.workspace_collaboration_service is not None:
|
||||||
|
maintenance_intervals['invitations'] = 3600.0
|
||||||
|
|
||||||
|
if maintenance_intervals:
|
||||||
|
|
||||||
|
async def resource_maintenance_loop():
|
||||||
|
"""Share tenant discovery and serialize periodic maintenance."""
|
||||||
|
|
||||||
|
loop = asyncio.get_running_loop()
|
||||||
|
started_at = loop.time()
|
||||||
|
next_due = {name: started_at + interval for name, interval in maintenance_intervals.items()}
|
||||||
while True:
|
while True:
|
||||||
|
await asyncio.sleep(max(min(next_due.values()) - loop.time(), 0.0))
|
||||||
|
observed_at = loop.time()
|
||||||
|
due = {name for name, due_at in next_due.items() if due_at <= observed_at}
|
||||||
|
if not due:
|
||||||
|
continue
|
||||||
try:
|
try:
|
||||||
bindings = await self.workspace_service.list_active_execution_bindings()
|
bindings = await self.workspace_service.list_active_execution_bindings()
|
||||||
|
except asyncio.CancelledError:
|
||||||
|
raise
|
||||||
|
except Exception as exc:
|
||||||
|
self.logger.warning(f'Resource maintenance Workspace discovery failed: {exc}')
|
||||||
|
else:
|
||||||
for binding in bindings:
|
for binding in bindings:
|
||||||
context = ExecutionContext(
|
context = ExecutionContext(
|
||||||
instance_uuid=binding.instance_uuid,
|
instance_uuid=binding.instance_uuid,
|
||||||
@@ -428,20 +416,59 @@ class Application:
|
|||||||
placement_generation=binding.placement_generation,
|
placement_generation=binding.placement_generation,
|
||||||
trigger_principal=PrincipalContext(PrincipalType.SYSTEM),
|
trigger_principal=PrincipalContext(PrincipalType.SYSTEM),
|
||||||
)
|
)
|
||||||
deleted = await self.maintenance_service.cleanup_expired_files(context)
|
if 'monitoring' in due:
|
||||||
total_deleted = sum(deleted.values())
|
try:
|
||||||
if total_deleted > 0:
|
deleted = await self.monitoring_service.cleanup_expired_records(
|
||||||
self.logger.info(
|
context,
|
||||||
f'Storage maintenance for Workspace {context.workspace_uuid}: '
|
retention_days,
|
||||||
f'deleted expired files: {deleted}'
|
batch_size=delete_batch_size,
|
||||||
|
)
|
||||||
|
total_deleted = sum(deleted.values())
|
||||||
|
if total_deleted > 0:
|
||||||
|
self.logger.info(
|
||||||
|
f'Monitoring auto-cleanup: deleted {total_deleted} expired records '
|
||||||
|
f'for Workspace {context.workspace_uuid} '
|
||||||
|
f'(retention={retention_days}d): {deleted}'
|
||||||
|
)
|
||||||
|
except asyncio.CancelledError:
|
||||||
|
raise
|
||||||
|
except Exception as exc:
|
||||||
|
self.logger.warning(
|
||||||
|
f'Monitoring auto-cleanup failed for '
|
||||||
|
f'Workspace {context.workspace_uuid}: {exc}'
|
||||||
|
)
|
||||||
|
if 'storage' in due:
|
||||||
|
try:
|
||||||
|
deleted = await self.maintenance_service.cleanup_expired_files(context)
|
||||||
|
total_deleted = sum(deleted.values())
|
||||||
|
if total_deleted > 0:
|
||||||
|
self.logger.info(
|
||||||
|
f'Storage maintenance for Workspace {context.workspace_uuid}: '
|
||||||
|
f'deleted expired files: {deleted}'
|
||||||
|
)
|
||||||
|
except asyncio.CancelledError:
|
||||||
|
raise
|
||||||
|
except Exception as exc:
|
||||||
|
self.logger.warning(
|
||||||
|
f'Storage maintenance failed for Workspace {context.workspace_uuid}: {exc}'
|
||||||
|
)
|
||||||
|
if 'invitations' in due:
|
||||||
|
try:
|
||||||
|
await self.workspace_collaboration_service.cleanup_expired_invitations(
|
||||||
|
active_bindings=bindings,
|
||||||
)
|
)
|
||||||
except Exception as e:
|
except asyncio.CancelledError:
|
||||||
self.logger.warning(f'Storage maintenance error: {e}')
|
raise
|
||||||
await asyncio.sleep(check_interval_seconds)
|
except Exception as exc:
|
||||||
|
self.logger.warning(f'Expired Workspace invitation cleanup failed: {exc}')
|
||||||
|
|
||||||
|
completed_at = loop.time()
|
||||||
|
for name in due:
|
||||||
|
next_due[name] = completed_at + maintenance_intervals[name]
|
||||||
|
|
||||||
self.task_mgr.create_task(
|
self.task_mgr.create_task(
|
||||||
storage_cleanup_loop(),
|
resource_maintenance_loop(),
|
||||||
name='storage-maintenance',
|
name='resource-maintenance',
|
||||||
scopes=[core_entities.LifecycleControlScope.APPLICATION],
|
scopes=[core_entities.LifecycleControlScope.APPLICATION],
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|||||||
@@ -59,6 +59,17 @@ _RUNTIME_POLICY_DEFAULTS = {
|
|||||||
}
|
}
|
||||||
},
|
},
|
||||||
'mcp': {'stdio': {'enabled': True}},
|
'mcp': {'stdio': {'enabled': True}},
|
||||||
|
'monitoring': {
|
||||||
|
'query_limits': {
|
||||||
|
'page_rows': 1000,
|
||||||
|
'export_rows': 10000,
|
||||||
|
'detail_rows': 2000,
|
||||||
|
'timeseries_buckets': 1000,
|
||||||
|
'max_offset': 1000000,
|
||||||
|
},
|
||||||
|
'auto_cleanup': {'max_batches_per_table_per_run': 4},
|
||||||
|
},
|
||||||
|
'storage': {'cleanup': {'max_files_per_run': 1000}},
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
|
|||||||
@@ -34,11 +34,22 @@ class _LimitedHTTPXAsyncByteStream(httpx.AsyncByteStream):
|
|||||||
self._read_bytes = 0
|
self._read_bytes = 0
|
||||||
|
|
||||||
async def __aiter__(self):
|
async def __aiter__(self):
|
||||||
async for chunk in self._inner:
|
try:
|
||||||
self._read_bytes += len(chunk)
|
async for chunk in self._inner:
|
||||||
if self._read_bytes > self._max_bytes:
|
self._read_bytes += len(chunk)
|
||||||
raise RemoteResponseTooLargeError(f'Remote response exceeds the {self._max_bytes}-byte limit')
|
if self._read_bytes > self._max_bytes:
|
||||||
yield chunk
|
raise RemoteResponseTooLargeError(f'Remote response exceeds the {self._max_bytes}-byte limit')
|
||||||
|
yield chunk
|
||||||
|
except BaseException:
|
||||||
|
# HTTPX only closes a response after normal stream exhaustion. If
|
||||||
|
# this limiter raises (or its consumer is cancelled), explicitly
|
||||||
|
# release the underlying connection before propagating the original
|
||||||
|
# failure so persistent clients cannot accumulate stranded streams.
|
||||||
|
try:
|
||||||
|
await self._inner.aclose()
|
||||||
|
except BaseException:
|
||||||
|
pass
|
||||||
|
raise
|
||||||
|
|
||||||
async def aclose(self) -> None:
|
async def aclose(self) -> None:
|
||||||
await self._inner.aclose()
|
await self._inner.aclose()
|
||||||
@@ -64,6 +75,7 @@ def httpx_response_limit_hooks(
|
|||||||
|
|
||||||
if response.is_stream_consumed:
|
if response.is_stream_consumed:
|
||||||
if len(response.content) > max_bytes:
|
if len(response.content) > max_bytes:
|
||||||
|
await response.aclose()
|
||||||
raise RemoteResponseTooLargeError(f'Remote response exceeds the {max_bytes}-byte limit')
|
raise RemoteResponseTooLargeError(f'Remote response exceeds the {max_bytes}-byte limit')
|
||||||
return
|
return
|
||||||
response.stream = _LimitedHTTPXAsyncByteStream(response.stream, max_bytes)
|
response.stream = _LimitedHTTPXAsyncByteStream(response.stream, max_bytes)
|
||||||
|
|||||||
@@ -556,6 +556,7 @@ class WorkspaceCollaborationService:
|
|||||||
self,
|
self,
|
||||||
*,
|
*,
|
||||||
retention: datetime.timedelta = datetime.timedelta(0),
|
retention: datetime.timedelta = datetime.timedelta(0),
|
||||||
|
active_bindings: typing.Iterable[WorkspaceExecutionBinding] | None = None,
|
||||||
) -> int:
|
) -> int:
|
||||||
"""Delete expired invitation records without crossing Cloud tenant scopes."""
|
"""Delete expired invitation records without crossing Cloud tenant scopes."""
|
||||||
cutoff = self._utcnow() - retention
|
cutoff = self._utcnow() - retention
|
||||||
@@ -576,7 +577,8 @@ class WorkspaceCollaborationService:
|
|||||||
if not callable(list_bindings) or not callable(tenant_uow):
|
if not callable(list_bindings) or not callable(tenant_uow):
|
||||||
raise RuntimeError('Cloud invitation cleanup requires tenant units of work')
|
raise RuntimeError('Cloud invitation cleanup requires tenant units of work')
|
||||||
deleted = 0
|
deleted = 0
|
||||||
for binding in await list_bindings():
|
bindings = active_bindings if active_bindings is not None else await list_bindings()
|
||||||
|
for binding in bindings:
|
||||||
async with tenant_uow(binding.workspace_uuid) as uow:
|
async with tenant_uow(binding.workspace_uuid) as uow:
|
||||||
deleted += await cleanup_session(uow.session, binding.workspace_uuid)
|
deleted += await cleanup_session(uow.session, binding.workspace_uuid)
|
||||||
return deleted
|
return deleted
|
||||||
|
|||||||
@@ -217,6 +217,9 @@ storage:
|
|||||||
uploaded_file_retention_days: 7
|
uploaded_file_retention_days: 7
|
||||||
# LangBot log files older than this many days will be deleted
|
# LangBot log files older than this many days will be deleted
|
||||||
log_retention_days: 3
|
log_retention_days: 3
|
||||||
|
# Bound per-Workspace file cleanup and diagnostic candidate lists.
|
||||||
|
# Supports STORAGE__CLEANUP__MAX_FILES_PER_RUN (hard cap: 10000).
|
||||||
|
max_files_per_run: 1000
|
||||||
s3:
|
s3:
|
||||||
endpoint_url: ''
|
endpoint_url: ''
|
||||||
access_key_id: ''
|
access_key_id: ''
|
||||||
@@ -271,6 +274,22 @@ mcp:
|
|||||||
# MCP__STDIO__ENABLED=false even when Box Runtime is available.
|
# MCP__STDIO__ENABLED=false even when Box Runtime is available.
|
||||||
enabled: true
|
enabled: true
|
||||||
monitoring:
|
monitoring:
|
||||||
|
query_limits:
|
||||||
|
# Maximum records materialized by one paginated monitoring request.
|
||||||
|
# Supports MONITORING__QUERY_LIMITS__PAGE_ROWS (hard cap: 5000).
|
||||||
|
page_rows: 1000
|
||||||
|
# CSV exports are currently assembled in memory. Keep this lower than
|
||||||
|
# the historical 100000-row default (hard cap: 50000).
|
||||||
|
export_rows: 10000
|
||||||
|
# Maximum related records returned by one session/message detail view
|
||||||
|
# (hard cap: 10000). Aggregate statistics remain database-computed.
|
||||||
|
detail_rows: 2000
|
||||||
|
# Token charts are grouped in SQL and return only the newest buckets
|
||||||
|
# (hard cap: 10000). Supports an environment variable override.
|
||||||
|
timeseries_buckets: 1000
|
||||||
|
# Bound high-offset scans that can otherwise monopolize PostgreSQL CPU
|
||||||
|
# (hard cap: 10000000).
|
||||||
|
max_offset: 1000000
|
||||||
auto_cleanup:
|
auto_cleanup:
|
||||||
# Enable automatic cleanup of expired monitoring records
|
# Enable automatic cleanup of expired monitoring records
|
||||||
enabled: true
|
enabled: true
|
||||||
@@ -280,6 +299,9 @@ monitoring:
|
|||||||
check_interval_hours: 1
|
check_interval_hours: 1
|
||||||
# Number of expired rows to delete per table batch
|
# Number of expired rows to delete per table batch
|
||||||
delete_batch_size: 1000
|
delete_batch_size: 1000
|
||||||
|
# Prevent one large Workspace backlog from monopolizing PostgreSQL.
|
||||||
|
# Supports MONITORING__AUTO_CLEANUP__MAX_BATCHES_PER_TABLE_PER_RUN.
|
||||||
|
max_batches_per_table_per_run: 4
|
||||||
box:
|
box:
|
||||||
# Master switch for the Box sandbox runtime. When false, LangBot does NOT
|
# Master switch for the Box sandbox runtime. When false, LangBot does NOT
|
||||||
# attempt to connect to a remote Box runtime nor start a local stdio Box
|
# attempt to connect to a remote Box runtime nor start a local stdio Box
|
||||||
|
|||||||
@@ -905,6 +905,32 @@ class TestMaintenanceServiceExpiredLocalUploadCandidates:
|
|||||||
# Verify - path included
|
# Verify - path included
|
||||||
assert 'path' in result[0]
|
assert 'path' in result[0]
|
||||||
|
|
||||||
|
def test_expired_local_upload_candidates_respects_run_limit(self):
|
||||||
|
ap = SimpleNamespace(
|
||||||
|
logger=SimpleNamespace(warning=Mock()),
|
||||||
|
storage_mgr=_scoped_storage_manager(),
|
||||||
|
instance_config=SimpleNamespace(data={'storage': {'cleanup': {'max_files_per_run': 2}}}),
|
||||||
|
)
|
||||||
|
service = MaintenanceService(ap)
|
||||||
|
entries = []
|
||||||
|
for index in range(3):
|
||||||
|
entry = Mock(spec=Path)
|
||||||
|
entry.is_file = Mock(return_value=True)
|
||||||
|
entry.stat = Mock(return_value=SimpleNamespace(st_size=100, st_mtime=0))
|
||||||
|
entry.relative_to = Mock(return_value=Path(f'scoped/old-{index}.txt'))
|
||||||
|
entries.append(entry)
|
||||||
|
|
||||||
|
with patch.object(Path, 'exists', return_value=True):
|
||||||
|
with patch.object(Path, 'rglob', return_value=entries):
|
||||||
|
result = service._expired_local_upload_candidates(TEST_CONTEXT, 7)
|
||||||
|
|
||||||
|
assert [item['key'] for item in result] == [
|
||||||
|
'scoped/old-0.txt',
|
||||||
|
'scoped/old-1.txt',
|
||||||
|
]
|
||||||
|
ap.instance_config.data['storage']['cleanup']['max_files_per_run'] = 999999
|
||||||
|
assert service._max_files_per_run() == 10000
|
||||||
|
|
||||||
|
|
||||||
ISOLATION_WORKSPACE_A = '00000000-0000-0000-0000-00000000000a'
|
ISOLATION_WORKSPACE_A = '00000000-0000-0000-0000-00000000000a'
|
||||||
ISOLATION_WORKSPACE_B = '00000000-0000-0000-0000-00000000000b'
|
ISOLATION_WORKSPACE_B = '00000000-0000-0000-0000-00000000000b'
|
||||||
|
|||||||
@@ -11,7 +11,7 @@ from langbot.pkg.api.http.authz import WorkspaceRequiredError
|
|||||||
from langbot.pkg.api.http.context import ExecutionContext
|
from langbot.pkg.api.http.context import ExecutionContext
|
||||||
from langbot.pkg.api.http.service.monitoring import MonitoringService
|
from langbot.pkg.api.http.service.monitoring import MonitoringService
|
||||||
from langbot.pkg.entity.persistence.base import Base
|
from langbot.pkg.entity.persistence.base import Base
|
||||||
from langbot.pkg.entity.persistence.monitoring import MonitoringMessage
|
from langbot.pkg.entity.persistence.monitoring import MonitoringLLMCall, MonitoringMessage
|
||||||
from langbot.pkg.entity.persistence.workspace import Workspace
|
from langbot.pkg.entity.persistence.workspace import Workspace
|
||||||
from langbot.pkg.persistence.mgr import PersistenceManager
|
from langbot.pkg.persistence.mgr import PersistenceManager
|
||||||
|
|
||||||
@@ -176,6 +176,161 @@ async def test_feedback_upsert_and_cancel_are_workspace_scoped(service):
|
|||||||
assert (await service.get_feedback_stats(context_b))['total_feedback'] == 1
|
assert (await service.get_feedback_stats(context_b))['total_feedback'] == 1
|
||||||
|
|
||||||
|
|
||||||
|
async def test_monitoring_queries_and_detail_views_are_strictly_bounded(service):
|
||||||
|
context = _context(WORKSPACE_A)
|
||||||
|
service.ap.instance_config.data['monitoring'] = {
|
||||||
|
'query_limits': {
|
||||||
|
'page_rows': 2,
|
||||||
|
'export_rows': 2,
|
||||||
|
'detail_rows': 2,
|
||||||
|
'timeseries_buckets': 2,
|
||||||
|
'max_offset': 10,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
await service.record_session_start(
|
||||||
|
context,
|
||||||
|
session_id='same-session',
|
||||||
|
bot_id='same-bot',
|
||||||
|
bot_name='Same Bot',
|
||||||
|
pipeline_id='same-pipeline',
|
||||||
|
pipeline_name='Same Pipeline',
|
||||||
|
)
|
||||||
|
message_ids = [await _record_message(service, context, f'message-{index}') for index in range(4)]
|
||||||
|
for index in range(3):
|
||||||
|
await service.record_llm_call(
|
||||||
|
context,
|
||||||
|
bot_id='same-bot',
|
||||||
|
bot_name='Same Bot',
|
||||||
|
pipeline_id='same-pipeline',
|
||||||
|
pipeline_name='Same Pipeline',
|
||||||
|
session_id='same-session',
|
||||||
|
model_name='model',
|
||||||
|
input_tokens=1,
|
||||||
|
output_tokens=2,
|
||||||
|
duration=10,
|
||||||
|
message_id=message_ids[0],
|
||||||
|
)
|
||||||
|
await service.record_tool_call(
|
||||||
|
context,
|
||||||
|
tool_name=f'tool-{index}',
|
||||||
|
tool_source='native',
|
||||||
|
duration=5,
|
||||||
|
session_id='same-session',
|
||||||
|
message_id=message_ids[0],
|
||||||
|
)
|
||||||
|
await service.record_error(
|
||||||
|
context,
|
||||||
|
bot_id='same-bot',
|
||||||
|
bot_name='Same Bot',
|
||||||
|
pipeline_id='same-pipeline',
|
||||||
|
pipeline_name='Same Pipeline',
|
||||||
|
error_type='Failure',
|
||||||
|
error_message=f'error-{index}',
|
||||||
|
session_id='same-session',
|
||||||
|
message_id=message_ids[0],
|
||||||
|
)
|
||||||
|
|
||||||
|
page, total = await service.get_messages(context, limit=100000, offset=-5)
|
||||||
|
exported = await service.export_messages(context, limit=100000)
|
||||||
|
session_detail = await service.get_session_analysis(context, 'same-session')
|
||||||
|
message_detail = await service.get_message_details(context, message_ids[0])
|
||||||
|
|
||||||
|
assert total == 4
|
||||||
|
assert len(page) == 2
|
||||||
|
assert len(exported) == 2
|
||||||
|
assert session_detail['message_stats']['total'] == 4
|
||||||
|
assert session_detail['llm_stats']['total_calls'] == 3
|
||||||
|
assert session_detail['tool_stats']['total_calls'] == 3
|
||||||
|
assert len(session_detail['tool_calls']) == 2
|
||||||
|
assert len(session_detail['errors']) == 2
|
||||||
|
assert session_detail['detail_truncated'] == {
|
||||||
|
'tool_calls': True,
|
||||||
|
'errors': True,
|
||||||
|
}
|
||||||
|
assert message_detail['llm_stats']['total_calls'] == 3
|
||||||
|
assert len(message_detail['llm_calls']) == 2
|
||||||
|
assert len(message_detail['errors']) == 2
|
||||||
|
assert message_detail['detail_truncated'] == {
|
||||||
|
'llm_calls': True,
|
||||||
|
'errors': True,
|
||||||
|
}
|
||||||
|
|
||||||
|
service.ap.instance_config.data['monitoring']['query_limits'] = {
|
||||||
|
'page_rows': 999999,
|
||||||
|
'export_rows': 999999,
|
||||||
|
'detail_rows': 999999,
|
||||||
|
'timeseries_buckets': 999999,
|
||||||
|
'max_offset': 99999999,
|
||||||
|
}
|
||||||
|
assert service.normalize_page_window(999999, 99999999) == (5000, 10000000)
|
||||||
|
assert service.normalize_export_limit(999999) == 50000
|
||||||
|
assert service._detail_limit() == 10000
|
||||||
|
assert service._timeseries_bucket_limit() == 10000
|
||||||
|
|
||||||
|
|
||||||
|
async def test_token_statistics_aggregate_and_limit_groups_in_database(service):
|
||||||
|
context = _context(WORKSPACE_A)
|
||||||
|
service.ap.instance_config.data['monitoring'] = {
|
||||||
|
'query_limits': {
|
||||||
|
'page_rows': 1,
|
||||||
|
'timeseries_buckets': 2,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
first_hour = datetime.datetime(2026, 7, 28, 10, 0)
|
||||||
|
rows = [
|
||||||
|
{
|
||||||
|
'id': f'llm-{index}',
|
||||||
|
'workspace_uuid': WORKSPACE_A,
|
||||||
|
'timestamp': first_hour + datetime.timedelta(hours=hour, minutes=index),
|
||||||
|
'model_name': model,
|
||||||
|
'input_tokens': input_tokens,
|
||||||
|
'output_tokens': output_tokens,
|
||||||
|
'total_tokens': input_tokens + output_tokens,
|
||||||
|
'duration': 100,
|
||||||
|
'cost': 0.01,
|
||||||
|
'status': 'success',
|
||||||
|
'bot_id': 'same-bot',
|
||||||
|
'bot_name': 'Same Bot',
|
||||||
|
'pipeline_id': 'same-pipeline',
|
||||||
|
'pipeline_name': 'Same Pipeline',
|
||||||
|
'session_id': 'same-session',
|
||||||
|
}
|
||||||
|
for index, (hour, model, input_tokens, output_tokens) in enumerate(
|
||||||
|
[
|
||||||
|
(0, 'small-model', 1, 2),
|
||||||
|
(1, 'large-model', 3, 4),
|
||||||
|
(2, 'large-model', 5, 6),
|
||||||
|
(2, 'large-model', 7, 8),
|
||||||
|
]
|
||||||
|
)
|
||||||
|
]
|
||||||
|
await service.ap.persistence_mgr.execute_async(sqlalchemy.insert(MonitoringLLMCall), rows)
|
||||||
|
|
||||||
|
stats = await service.get_token_statistics(context, bucket='hour')
|
||||||
|
|
||||||
|
assert stats['summary']['total_calls'] == 4
|
||||||
|
assert stats['summary']['total_tokens'] == 36
|
||||||
|
assert stats['by_model_truncated'] is True
|
||||||
|
assert [model['model_name'] for model in stats['by_model']] == ['large-model']
|
||||||
|
assert stats['timeseries_truncated'] is True
|
||||||
|
assert stats['timeseries'] == [
|
||||||
|
{
|
||||||
|
'bucket': '2026-07-28 11:00',
|
||||||
|
'input_tokens': 3,
|
||||||
|
'output_tokens': 4,
|
||||||
|
'total_tokens': 7,
|
||||||
|
'calls': 1,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
'bucket': '2026-07-28 12:00',
|
||||||
|
'input_tokens': 12,
|
||||||
|
'output_tokens': 14,
|
||||||
|
'total_tokens': 26,
|
||||||
|
'calls': 2,
|
||||||
|
},
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
async def test_cleanup_commits_sqlite_delete_before_vacuum(tmp_path):
|
async def test_cleanup_commits_sqlite_delete_before_vacuum(tmp_path):
|
||||||
engine = create_async_engine(
|
engine = create_async_engine(
|
||||||
f'sqlite+aiosqlite:///{tmp_path / "monitoring-cleanup.db"}',
|
f'sqlite+aiosqlite:///{tmp_path / "monitoring-cleanup.db"}',
|
||||||
@@ -200,32 +355,38 @@ async def test_cleanup_commits_sqlite_delete_before_vacuum(tmp_path):
|
|||||||
)
|
)
|
||||||
)
|
)
|
||||||
await connection.execute(
|
await connection.execute(
|
||||||
sqlalchemy.insert(MonitoringMessage).values(
|
sqlalchemy.insert(MonitoringMessage),
|
||||||
id='expired-message',
|
[
|
||||||
workspace_uuid=WORKSPACE_A,
|
{
|
||||||
timestamp=datetime.datetime.now(datetime.timezone.utc).replace(tzinfo=None)
|
'id': f'expired-message-{index}',
|
||||||
- datetime.timedelta(days=30),
|
'workspace_uuid': WORKSPACE_A,
|
||||||
bot_id='bot',
|
'timestamp': datetime.datetime.now(datetime.timezone.utc).replace(tzinfo=None)
|
||||||
bot_name='Bot',
|
- datetime.timedelta(days=30),
|
||||||
pipeline_id='pipeline',
|
'bot_id': 'bot',
|
||||||
pipeline_name='Pipeline',
|
'bot_name': 'Bot',
|
||||||
message_content='expired',
|
'pipeline_id': 'pipeline',
|
||||||
session_id='session',
|
'pipeline_name': 'Pipeline',
|
||||||
status='success',
|
'message_content': 'expired',
|
||||||
level='info',
|
'session_id': 'session',
|
||||||
)
|
'status': 'success',
|
||||||
|
'level': 'info',
|
||||||
|
}
|
||||||
|
for index in range(5)
|
||||||
|
],
|
||||||
)
|
)
|
||||||
|
|
||||||
deleted = await MonitoringService(application).cleanup_expired_records(
|
deleted = await MonitoringService(application).cleanup_expired_records(
|
||||||
_context(WORKSPACE_A),
|
_context(WORKSPACE_A),
|
||||||
retention_days=1,
|
retention_days=1,
|
||||||
|
batch_size=2,
|
||||||
|
max_batches_per_table=1,
|
||||||
)
|
)
|
||||||
|
|
||||||
assert deleted['monitoring_messages'] == 1
|
assert deleted['monitoring_messages'] == 2
|
||||||
async with engine.connect() as connection:
|
async with engine.connect() as connection:
|
||||||
remaining = await connection.scalar(
|
remaining = await connection.scalar(
|
||||||
sqlalchemy.select(sqlalchemy.func.count()).select_from(MonitoringMessage)
|
sqlalchemy.select(sqlalchemy.func.count()).select_from(MonitoringMessage)
|
||||||
)
|
)
|
||||||
assert remaining == 0
|
assert remaining == 3
|
||||||
finally:
|
finally:
|
||||||
await engine.dispose()
|
await engine.dispose()
|
||||||
|
|||||||
@@ -0,0 +1,113 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import asyncio
|
||||||
|
from types import SimpleNamespace
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
from langbot.pkg.core.app import Application
|
||||||
|
|
||||||
|
|
||||||
|
pytestmark = pytest.mark.asyncio
|
||||||
|
|
||||||
|
|
||||||
|
class _TaskManager:
|
||||||
|
def __init__(self, stop: asyncio.Event) -> None:
|
||||||
|
self.stop = stop
|
||||||
|
self.tasks: list[asyncio.Task] = []
|
||||||
|
|
||||||
|
def create_task(self, coro, *, name='', **_kwargs):
|
||||||
|
task = asyncio.create_task(coro, name=name)
|
||||||
|
self.tasks.append(task)
|
||||||
|
return SimpleNamespace(task=task)
|
||||||
|
|
||||||
|
async def wait_all(self) -> None:
|
||||||
|
await self.stop.wait()
|
||||||
|
for task in self.tasks:
|
||||||
|
task.cancel()
|
||||||
|
await asyncio.gather(*self.tasks, return_exceptions=True)
|
||||||
|
|
||||||
|
|
||||||
|
async def _wait_forever() -> None:
|
||||||
|
await asyncio.Event().wait()
|
||||||
|
|
||||||
|
|
||||||
|
async def test_resource_maintenance_waits_and_shares_workspace_discovery() -> None:
|
||||||
|
stop = asyncio.Event()
|
||||||
|
completed = asyncio.Event()
|
||||||
|
discovery_calls = 0
|
||||||
|
job_calls: list[str] = []
|
||||||
|
|
||||||
|
async def list_bindings():
|
||||||
|
nonlocal discovery_calls
|
||||||
|
discovery_calls += 1
|
||||||
|
return [
|
||||||
|
SimpleNamespace(
|
||||||
|
instance_uuid='instance',
|
||||||
|
workspace_uuid='workspace',
|
||||||
|
placement_generation=1,
|
||||||
|
)
|
||||||
|
]
|
||||||
|
|
||||||
|
async def cleanup_monitoring(_context, _retention_days, *, batch_size):
|
||||||
|
assert batch_size == 10
|
||||||
|
job_calls.append('monitoring')
|
||||||
|
return {}
|
||||||
|
|
||||||
|
async def cleanup_storage(_context):
|
||||||
|
job_calls.append('storage')
|
||||||
|
completed.set()
|
||||||
|
return {}
|
||||||
|
|
||||||
|
application = Application()
|
||||||
|
application.event_loop = asyncio.get_running_loop()
|
||||||
|
application.event_loop_monitor = SimpleNamespace(start=lambda: None)
|
||||||
|
application.task_mgr = _TaskManager(stop)
|
||||||
|
application.plugin_connector = SimpleNamespace(initialize_plugins=lambda: asyncio.sleep(0))
|
||||||
|
application.platform_mgr = SimpleNamespace(run=_wait_forever)
|
||||||
|
application.ctrl = SimpleNamespace(run=_wait_forever)
|
||||||
|
application.http_ctrl = SimpleNamespace(run=_wait_forever)
|
||||||
|
application.telemetry = None
|
||||||
|
application.workspace_collaboration_service = None
|
||||||
|
application.workspace_service = SimpleNamespace(list_active_execution_bindings=list_bindings)
|
||||||
|
application.monitoring_service = SimpleNamespace(cleanup_expired_records=cleanup_monitoring)
|
||||||
|
application.maintenance_service = SimpleNamespace(cleanup_expired_files=cleanup_storage)
|
||||||
|
application.instance_config = SimpleNamespace(
|
||||||
|
data={
|
||||||
|
'monitoring': {
|
||||||
|
'auto_cleanup': {
|
||||||
|
'enabled': True,
|
||||||
|
'retention_days': 30,
|
||||||
|
'delete_batch_size': 10,
|
||||||
|
'check_interval_hours': 0.00002,
|
||||||
|
}
|
||||||
|
},
|
||||||
|
'storage': {
|
||||||
|
'cleanup': {
|
||||||
|
'enabled': True,
|
||||||
|
'check_interval_hours': 0.00002,
|
||||||
|
}
|
||||||
|
},
|
||||||
|
}
|
||||||
|
)
|
||||||
|
application.logger = SimpleNamespace(
|
||||||
|
info=lambda *_args, **_kwargs: None,
|
||||||
|
warning=lambda *_args, **_kwargs: None,
|
||||||
|
error=lambda *_args, **_kwargs: None,
|
||||||
|
debug=lambda *_args, **_kwargs: None,
|
||||||
|
)
|
||||||
|
|
||||||
|
async def no_web_info() -> None:
|
||||||
|
return None
|
||||||
|
|
||||||
|
application.print_web_access_info = no_web_info
|
||||||
|
run_task = asyncio.create_task(application.run())
|
||||||
|
try:
|
||||||
|
await asyncio.sleep(0.01)
|
||||||
|
assert discovery_calls == 0
|
||||||
|
await asyncio.wait_for(completed.wait(), timeout=1)
|
||||||
|
assert discovery_calls == 1
|
||||||
|
assert job_calls == ['monitoring', 'storage']
|
||||||
|
finally:
|
||||||
|
stop.set()
|
||||||
|
await asyncio.wait_for(run_task, timeout=1)
|
||||||
@@ -6,6 +6,7 @@ Tests session management, reuse, and cleanup.
|
|||||||
|
|
||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import asyncio
|
||||||
import threading
|
import threading
|
||||||
from types import SimpleNamespace
|
from types import SimpleNamespace
|
||||||
from unittest.mock import AsyncMock
|
from unittest.mock import AsyncMock
|
||||||
@@ -161,14 +162,18 @@ class TestReadLimited:
|
|||||||
|
|
||||||
async def test_httpx_hook_rejects_before_automatic_buffer_grows(self):
|
async def test_httpx_hook_rejects_before_automatic_buffer_grows(self):
|
||||||
class Source(httpx.AsyncByteStream):
|
class Source(httpx.AsyncByteStream):
|
||||||
|
def __init__(self):
|
||||||
|
self.closed = False
|
||||||
|
|
||||||
async def __aiter__(self):
|
async def __aiter__(self):
|
||||||
yield b'123'
|
yield b'123'
|
||||||
yield b'45'
|
yield b'45'
|
||||||
|
|
||||||
async def aclose(self):
|
async def aclose(self):
|
||||||
return None
|
self.closed = True
|
||||||
|
|
||||||
transport = httpx.MockTransport(lambda _request: httpx.Response(200, stream=Source()))
|
source = Source()
|
||||||
|
transport = httpx.MockTransport(lambda _request: httpx.Response(200, stream=source))
|
||||||
async with httpx.AsyncClient(
|
async with httpx.AsyncClient(
|
||||||
transport=transport,
|
transport=transport,
|
||||||
event_hooks=httpclient.httpx_response_limit_hooks(max_bytes=4),
|
event_hooks=httpclient.httpx_response_limit_hooks(max_bytes=4),
|
||||||
@@ -176,6 +181,37 @@ class TestReadLimited:
|
|||||||
with pytest.raises(httpclient.RemoteResponseTooLargeError, match='4-byte'):
|
with pytest.raises(httpclient.RemoteResponseTooLargeError, match='4-byte'):
|
||||||
await client.get('https://example.invalid')
|
await client.get('https://example.invalid')
|
||||||
|
|
||||||
|
assert source.closed
|
||||||
|
|
||||||
|
async def test_httpx_limited_stream_closes_source_when_consumer_is_cancelled(self):
|
||||||
|
class Source(httpx.AsyncByteStream):
|
||||||
|
def __init__(self):
|
||||||
|
self.closed = False
|
||||||
|
|
||||||
|
async def __aiter__(self):
|
||||||
|
yield b'123'
|
||||||
|
await asyncio.Event().wait()
|
||||||
|
|
||||||
|
async def aclose(self):
|
||||||
|
self.closed = True
|
||||||
|
|
||||||
|
source = Source()
|
||||||
|
stream = httpclient._LimitedHTTPXAsyncByteStream(source, max_bytes=4)
|
||||||
|
first_chunk_consumed = asyncio.Event()
|
||||||
|
|
||||||
|
async def consume():
|
||||||
|
async for _chunk in stream:
|
||||||
|
first_chunk_consumed.set()
|
||||||
|
|
||||||
|
task = asyncio.create_task(consume())
|
||||||
|
await first_chunk_consumed.wait()
|
||||||
|
task.cancel()
|
||||||
|
|
||||||
|
with pytest.raises(asyncio.CancelledError):
|
||||||
|
await task
|
||||||
|
|
||||||
|
assert source.closed
|
||||||
|
|
||||||
async def test_close_all_handles_already_closed(self):
|
async def test_close_all_handles_already_closed(self):
|
||||||
"""close_all handles already closed sessions gracefully."""
|
"""close_all handles already closed sessions gracefully."""
|
||||||
session = httpclient.get_session()
|
session = httpclient.get_session()
|
||||||
|
|||||||
Reference in New Issue
Block a user