feat(ai): 每个数据页面都有 AI 解读,靠一条带优先级的生产者/消费者队列

原来只有今日页有晨报、指标详情页有归因,其余页面一片空白。现在除设置外
的 10 个页面都有:健康、睡眠、运动、趋势、每日、身体成分、成绩预测、
身体年龄、挑战赛、运动详情。

不是给每个页面写一套,而是一个通用管线:
- services/scopes.py:一个页面一个 context builder,返回同一个信封。
  context["highlights"] 是已经算好的白话事实——模型负责解读它们,模型不
  可用时规则引擎原样渲染。两者引用同一批数字,所以降级读起来不像换了个 App。
  没数据的页面返回 None,宁可不出卡片,也不让模型对着空表格发挥。
- coach.scope_messages / parse_scope_insight:一套提示词吃所有页面,页面
  的差异全在 context 里,加页面 = 加一个 builder。
- 前端 <AiPanel scope="…">:一个组件渲染所有页面,轮询逻辑抽成
  lib/insight.ts 的 usePolledInsight,晨报卡也改用它。

## 队列

一次生成 40 秒到 4.5 分钟,所以什么都不能在请求里生成。页面只负责入队,
worker 负责消费(services/jobs.py)。

优先级才是用队列而不是后台线程的理由:同步完成后 prefetch 把所有页面按
背景优先级排进去,可能要跑半小时;而用户一打开某个页面,那个页面的任务
立刻提到队首、下一个就跑。你在看什么,队列就在算什么。

队列放在数据库而不是内存里,因为 gunicorn 有两个 worker:任务带 holder
声明后回读确认,和 scheduler.py 抢 tick 是同一套做法。id 由
user+kind+subject 推导,所以每几秒一次的轮询是幂等的入队,不会每几秒堆一
个任务。

## 网关中断时踩到的两个坑(当场修了)

写完正好赶上 oracle 那台机器不通,于是看到:
- 三次失败后任务被永久标 failed,网关恢复了也不会重试——一次瞬时中断就把
  那个页面的解读判了死刑,直到它的数据碰巧变化。加了冷却期,过期后重置
  尝试次数再排一次。
- 队列已经放弃了,页面还在 pending 转圈,要转满 8 分钟才停。meta.pending
  现在跟着队列状态走,并把失败原因带给卡片。

顺带把 BAND_SOURCES 从 routes/settings.py 下沉到 services/insights.py:
教练要拿它做参照,而 services 不该反向依赖 routes。

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
ericwyuan
2026-09-01 15:17:17 +08:00
parent 984a06a098
commit 241ae0d6a3
26 changed files with 1948 additions and 182 deletions

View File

@@ -302,6 +302,31 @@ CREATE TABLE IF NOT EXISTS ai_recommendations (
-- attribution, and `subject` is the day (briefing) or metric+range (trend).
-- Same reasoning as ai_recommendations above — a generation costs minutes, so
-- it can never sit inside a page load.
-- The coach's work queue. Generating one insight costs minutes against the
-- gateway, so nothing is produced inside a request: screens enqueue, a worker
-- consumes. `priority` is what makes the screen the user is actually looking
-- at jump ahead of the backfill queued after a sync (lower runs first).
--
-- `id` is derived from user+kind+subject, so enqueueing the same work twice
-- updates one row rather than piling up duplicates — which is what keeps a
-- poll every few seconds from queueing a job every few seconds.
CREATE TABLE IF NOT EXISTS ai_jobs (
id VARCHAR(160) PRIMARY KEY,
user_id VARCHAR(64) NOT NULL,
kind VARCHAR(32) NOT NULL,
subject VARCHAR(96) NOT NULL,
fingerprint VARCHAR(64),
priority INT NOT NULL DEFAULT 10,
status VARCHAR(16) NOT NULL DEFAULT 'pending',
attempts INT NOT NULL DEFAULT 0,
error TEXT,
holder VARCHAR(64),
claimed_at DATETIME,
created_at DATETIME,
updated_at DATETIME,
FOREIGN KEY (user_id) REFERENCES users(id)
);
CREATE TABLE IF NOT EXISTS ai_insights (
id VARCHAR(160) PRIMARY KEY,
user_id VARCHAR(64) NOT NULL,