原来每天只存 7 个指标,而 get_user_summary 一次就返回 60+ 字段, 另有睡眠分期、训练准备度、耐力分等独立端点从未被调用。 db.py: - health_data 新增 31 列(距离/活动卡路里/基础代谢/爬楼/强度分钟/ 久坐时长/最高最低心率/最大压力/身体电量四项/血氧/呼吸/ 睡眠深浅REM清醒分期/睡眠血氧/睡眠呼吸/睡眠压力/训练准备度/ VO2max/耐力分) - 新增 badges 与 personal_records 两张表,均以 (user_id, garmin_id) 为主键,重复同步更新而非累积 - 新增增量迁移: CREATE TABLE IF NOT EXISTS 对已存在的表不生效, 新列必须显式 ALTER,否则生产库上永远不会出现。按列名比对后 逐个补齐,SQLite 与 MariaDB 都幂等 services/garmin.py: - _extract_daily 改为汇总 user_summary + sleep + hrv + training_readiness + training_status + endurance_score 五个端点 - 每个可选端点用 _safe 包裹:某项设备不记录时留 NULL,不影响当天其余数据 - 新增 sync_badges / sync_personal_records(账号级,每次同步取一次) fix(garmin): 个人纪录整批写入失败 - Garmin 在同一份数据里混用 ISO 字符串和 Unix 毫秒时间戳, prStartTimeGmt 是 1570961412000,写进 DATETIME 列被 MariaDB 以 1292 拒绝,导致 11 项个人纪录一条都没存进去 - 新增 _to_datetime 统一处理 ISO / 毫秒 / 秒三种形状,并优先取 Garmin 自己提供的 *Formatted 字段 services/ai.py: - 送给模型的 CSV 从 7 列扩到 23 列,纳入身体电量、血氧、呼吸、 训练准备度、耐力分和睡眠分期 接口: GET /api/health/badges、/api/health/personal-records tests (+13, 共 292): - 徽章/纪录的往返、重复同步不累积、按用户隔离 - 两个用户可持有同一个 Garmin 徽章 id 而不冲突 - 时间戳三种形状的归一化及无效值不抛异常 NAS 实测: 7 天数据每天 31 项指标、65 个奖励、11 项个人纪录 Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
552 lines
20 KiB
Python
552 lines
20 KiB
Python
"""
|
||
Multi-provider LLM layer for health recommendations.
|
||
|
||
Design goals
|
||
------------
|
||
* **Switchable models** — every model lives in a catalog keyed by a short id
|
||
("gemini-flash", "llama-70b", ...). Callers pass an id; nothing else in the
|
||
codebase knows which vendor is behind it.
|
||
* **Large context** — daily metrics are serialised as compact CSV rather than
|
||
JSON, so a year of data costs a few thousand tokens instead of tens of
|
||
thousands. Each model declares its own window and the payload is trimmed to
|
||
fit the smallest of (model window, configured day budget).
|
||
* **Fallback** — if the preferred model errors or times out, the next healthy
|
||
model in the chain is tried before giving up. This mirrors the behaviour the
|
||
NAS deployment already relies on (Gemini primary, NVIDIA secondary).
|
||
|
||
Only text-in/text-out models are supported; no vision models are registered.
|
||
API keys are read from the environment — never hardcode them.
|
||
"""
|
||
import json
|
||
import os
|
||
import re
|
||
|
||
import requests
|
||
|
||
# Tunables are read per call rather than captured at import: module-level
|
||
# constants freeze whatever the environment held when the module first loaded,
|
||
# which both hides live config changes and leaks a developer's .env into tests.
|
||
FALLBACK_TIMEOUT = 60.0
|
||
FALLBACK_DAY_BUDGET = 365
|
||
|
||
|
||
def default_timeout():
|
||
return float(os.environ.get("AI_TIMEOUT_SECONDS") or FALLBACK_TIMEOUT)
|
||
|
||
|
||
def default_day_budget():
|
||
"""Max days of history to put in a prompt, before per-model trimming."""
|
||
return int(os.environ.get("AI_DAY_BUDGET") or FALLBACK_DAY_BUDGET)
|
||
|
||
|
||
# Output cap. Deliberately modest: a long generation is what blows past an
|
||
# upstream's own timeout (the self-hosted gateway allows its adapters only
|
||
# 30-45s), and the reply here is a short JSON list, not an essay.
|
||
FALLBACK_MAX_TOKENS = 1024
|
||
|
||
|
||
def default_max_tokens():
|
||
return int(os.environ.get("AI_MAX_TOKENS") or FALLBACK_MAX_TOKENS)
|
||
|
||
SYSTEM_PROMPT = (
|
||
"你是一名严谨的健康数据分析助手,负责解读用户的可穿戴设备(Garmin)数据。\n"
|
||
"要求:\n"
|
||
"1. 只依据给出的数据得出结论,数据不足时明确说明,不要编造数值。\n"
|
||
"2. 指出趋势、异常和相互关联(例如睡眠不足与静息心率升高的关系)。\n"
|
||
"3. 给出具体、可执行的建议,而不是泛泛而谈。\n"
|
||
"4. 你不是医生,不做诊断;发现明显异常时建议用户咨询专业医师。\n"
|
||
"5. 用简体中文回答。\n\n"
|
||
"输出严格为 JSON 数组,最多 5 条,每条 recommendation 不超过 120 字,\n"
|
||
"每个元素形如:\n"
|
||
'{"category": "睡眠", "recommendation": "……", "priority": "high|medium|low", '
|
||
'"basedOn": ["sleep_duration"]}\n'
|
||
"不要输出 JSON 以外的任何文字,不要用 markdown 代码块包裹。"
|
||
)
|
||
|
||
|
||
class AIError(Exception):
|
||
"""Raised when a provider cannot produce a completion."""
|
||
|
||
|
||
class Completion:
|
||
"""A model reply plus, where the endpoint reports it, the upstream that
|
||
actually served the request.
|
||
|
||
The self-hosted gateway multiplexes over nvidia/gemini/ollama and names
|
||
the winner in its response, so `upstream` is what makes a gateway-side
|
||
failover visible to the UI instead of silently invisible.
|
||
"""
|
||
|
||
__slots__ = ("text", "upstream")
|
||
|
||
def __init__(self, text, upstream=None):
|
||
self.text = text
|
||
self.upstream = upstream
|
||
|
||
|
||
# --- providers --------------------------------------------------------------
|
||
class Provider:
|
||
"""Base class. Subclasses turn a prompt into text.
|
||
|
||
`use_proxy` decides whether HTTP(S)_PROXY / ALL_PROXY from the environment
|
||
apply. It matters because the two kinds of endpoint want opposite answers:
|
||
overseas vendors (Gemini, NVIDIA) may only be reachable *through* a local
|
||
proxy, while a self-hosted box on a public IP is reachable directly and
|
||
breaks if forced through one.
|
||
"""
|
||
|
||
name = "base"
|
||
|
||
def __init__(
|
||
self, model_id, context_window, api_key_env, use_proxy=True, max_tokens=None
|
||
):
|
||
self.model_id = model_id
|
||
self.context_window = context_window
|
||
self.api_key_env = api_key_env
|
||
self.use_proxy = use_proxy
|
||
self._max_tokens = max_tokens
|
||
|
||
@property
|
||
def max_tokens(self):
|
||
"""Output cap for this endpoint.
|
||
|
||
Reasoning models emit a chain-of-thought *before* the answer, so a cap
|
||
sized for the answer alone gets spent on the thinking and truncates
|
||
before any JSON appears. Those endpoints therefore declare a larger
|
||
budget than the default.
|
||
"""
|
||
return self._max_tokens or default_max_tokens()
|
||
|
||
@property
|
||
def api_key(self):
|
||
return os.environ.get(self.api_key_env) or ""
|
||
|
||
def is_configured(self):
|
||
return bool(self.api_key)
|
||
|
||
def _session(self):
|
||
session = requests.Session()
|
||
# trust_env=False also drops netrc/CA-bundle env lookups, which is the
|
||
# intent here: talk to the host directly, exactly as configured.
|
||
session.trust_env = self.use_proxy
|
||
return session
|
||
|
||
def generate(self, prompt, timeout=None):
|
||
raise NotImplementedError
|
||
|
||
|
||
class GeminiProvider(Provider):
|
||
"""Google AI Studio (generativelanguage.googleapis.com)."""
|
||
|
||
name = "gemini"
|
||
BASE = "https://generativelanguage.googleapis.com/v1beta/models"
|
||
|
||
def generate(self, prompt, timeout=None):
|
||
if not self.is_configured():
|
||
raise AIError(f"{self.api_key_env} 未配置")
|
||
timeout = timeout or default_timeout()
|
||
url = f"{self.BASE}/{self.model_id}:generateContent"
|
||
payload = {
|
||
"contents": [{"parts": [{"text": prompt}]}],
|
||
"generationConfig": {
|
||
"temperature": 0.4,
|
||
"maxOutputTokens": self.max_tokens,
|
||
},
|
||
}
|
||
try:
|
||
resp = self._session().post(
|
||
url,
|
||
headers={
|
||
"Content-Type": "application/json",
|
||
"X-goog-api-key": self.api_key,
|
||
},
|
||
json=payload,
|
||
timeout=timeout,
|
||
)
|
||
except requests.RequestException as e:
|
||
raise AIError(f"gemini 请求失败: {e}") from e
|
||
|
||
if resp.status_code != 200:
|
||
raise AIError(f"gemini HTTP {resp.status_code}: {resp.text[:200]}")
|
||
|
||
try:
|
||
body = resp.json()
|
||
parts = body["candidates"][0]["content"]["parts"]
|
||
return Completion("".join(p.get("text", "") for p in parts))
|
||
except (ValueError, KeyError, IndexError) as e:
|
||
raise AIError(f"gemini 响应格式异常: {e}") from e
|
||
|
||
|
||
class OpenAICompatProvider(Provider):
|
||
"""Any endpoint speaking the OpenAI chat-completions schema (NVIDIA NIM,
|
||
Ollama, vLLM, ...).
|
||
|
||
`requires_key=False` covers self-hosted runtimes such as Ollama, which
|
||
authenticate by network reachability rather than by a token. Those are
|
||
opt-in: they count as configured only once their base URL is set, so an
|
||
unset OLLAMA_BASE_URL keeps the entry out of the fallback chain.
|
||
"""
|
||
|
||
name = "openai-compat"
|
||
|
||
def __init__(
|
||
self,
|
||
model_id,
|
||
context_window,
|
||
base_url_env,
|
||
default_base_url="",
|
||
api_key_env=None,
|
||
requires_key=True,
|
||
use_proxy=True,
|
||
max_tokens=None,
|
||
):
|
||
super().__init__(
|
||
model_id, context_window, api_key_env or "", use_proxy, max_tokens
|
||
)
|
||
self.base_url_env = base_url_env
|
||
self.default_base_url = default_base_url
|
||
self.requires_key = requires_key
|
||
|
||
@property
|
||
def base_url(self):
|
||
return os.environ.get(self.base_url_env) or self.default_base_url
|
||
|
||
def is_configured(self):
|
||
if not self.base_url:
|
||
return False
|
||
return bool(self.api_key) if self.requires_key else True
|
||
|
||
def generate(self, prompt, timeout=None):
|
||
if not self.is_configured():
|
||
raise AIError(
|
||
f"{self.api_key_env} 未配置" if self.requires_key
|
||
else f"{self.base_url_env} 未配置"
|
||
)
|
||
timeout = timeout or default_timeout()
|
||
url = f"{self.base_url.rstrip('/')}/chat/completions"
|
||
payload = {
|
||
"model": self.model_id,
|
||
"messages": [{"role": "user", "content": prompt}],
|
||
"temperature": 0.4,
|
||
"max_tokens": self.max_tokens,
|
||
}
|
||
headers = {"Content-Type": "application/json"}
|
||
if self.api_key:
|
||
headers["Authorization"] = f"Bearer {self.api_key}"
|
||
try:
|
||
resp = self._session().post(
|
||
url, headers=headers, json=payload, timeout=timeout
|
||
)
|
||
except requests.RequestException as e:
|
||
raise AIError(f"{self.model_id} 请求失败: {e}") from e
|
||
|
||
if resp.status_code != 200:
|
||
raise AIError(f"{self.model_id} HTTP {resp.status_code}: {resp.text[:200]}")
|
||
|
||
try:
|
||
body = resp.json()
|
||
# `provider` is a gateway extension, absent from stock OpenAI
|
||
# responses — hence the .get rather than an index.
|
||
return Completion(
|
||
body["choices"][0]["message"]["content"], body.get("provider")
|
||
)
|
||
except (ValueError, KeyError, IndexError) as e:
|
||
raise AIError(f"{self.model_id} 响应格式异常: {e}") from e
|
||
|
||
|
||
# --- catalog ----------------------------------------------------------------
|
||
NVIDIA_BASE = "https://integrate.api.nvidia.com/v1"
|
||
|
||
|
||
def _nvidia(model_id, context_window):
|
||
return OpenAICompatProvider(
|
||
model_id=model_id,
|
||
context_window=context_window,
|
||
api_key_env="NVIDIA_API_KEY",
|
||
base_url_env="NVIDIA_BASE_URL",
|
||
default_base_url=NVIDIA_BASE,
|
||
)
|
||
|
||
|
||
def _build_catalog():
|
||
"""Model id -> Provider. Text-only models with large context windows.
|
||
|
||
The NVIDIA model strings below were taken from that account's live
|
||
`GET /v1/models` listing. Do not guess them: ids that merely look
|
||
plausible (`qwen/qwen2.5-72b-instruct`, `deepseek-ai/deepseek-r1`)
|
||
return HTTP 404 from this endpoint.
|
||
"""
|
||
return {
|
||
# Preferred entry: the self-hosted gateway on the Oracle box. It
|
||
# multiplexes over nvidia/gemini/ollama behind one OpenAI-compatible
|
||
# endpoint and rotates several Gemini keys, so it absorbs the quota
|
||
# and timeout failures that a single upstream hits on its own. Its
|
||
# reply names the upstream that served the request.
|
||
"gateway": OpenAICompatProvider(
|
||
model_id=os.environ.get("AI_GATEWAY_MODEL") or "ai-gateway-auto",
|
||
context_window=128_000,
|
||
api_key_env="AI_GATEWAY_TOKEN",
|
||
base_url_env="AI_GATEWAY_BASE_URL",
|
||
# Self-hosted and directly reachable: a local proxy would only
|
||
# add a hop that times out.
|
||
use_proxy=False,
|
||
# Its primary upstream is a reasoning model that thinks out loud
|
||
# before answering; at the default cap the trace consumed the whole
|
||
# budget and the reply was truncated before the JSON began.
|
||
max_tokens=3000,
|
||
),
|
||
# Direct upstreams, for pinning one vendor or for running without the
|
||
# gateway. These need their own keys in this app's .env.
|
||
"gemini-flash": GeminiProvider(
|
||
model_id="gemini-flash-latest",
|
||
context_window=1_000_000,
|
||
api_key_env="GEMINI_API_KEY",
|
||
),
|
||
"llama-70b": _nvidia("meta/llama-3.3-70b-instruct", 128_000),
|
||
"nemotron-49b": _nvidia("nvidia/llama-3.3-nemotron-super-49b-v1.5", 128_000),
|
||
"mistral-large": _nvidia("mistralai/mistral-large-2-instruct", 128_000),
|
||
}
|
||
|
||
|
||
CATALOG = _build_catalog()
|
||
|
||
# Preference order used when no model is requested, and for fallback.
|
||
FALLBACK_CHAIN = "gateway,gemini-flash,llama-70b"
|
||
|
||
|
||
def default_chain():
|
||
"""Preference order, read from the environment on every call.
|
||
|
||
Deliberately not a module-level constant: it is read at request time so a
|
||
changed AI_MODEL_CHAIN takes effect without a restart, and so tests can
|
||
set it without reaching into module internals.
|
||
"""
|
||
raw = os.environ.get("AI_MODEL_CHAIN") or FALLBACK_CHAIN
|
||
return [m.strip() for m in raw.split(",") if m.strip()]
|
||
|
||
|
||
def list_models():
|
||
"""Catalog entries plus whether each one currently has credentials."""
|
||
chain = default_chain()
|
||
head = chain[0] if chain else None
|
||
return [
|
||
{
|
||
"id": mid,
|
||
"model": p.model_id,
|
||
"provider": p.name,
|
||
"contextWindow": p.context_window,
|
||
"configured": p.is_configured(),
|
||
"default": mid == head,
|
||
}
|
||
for mid, p in CATALOG.items()
|
||
]
|
||
|
||
|
||
def resolve_chain(preferred=None):
|
||
"""Ordered list of model ids to attempt, configured ones only."""
|
||
chain = []
|
||
if preferred:
|
||
if preferred not in CATALOG:
|
||
raise AIError(f"未知模型: {preferred}")
|
||
chain.append(preferred)
|
||
for mid in default_chain():
|
||
if mid in CATALOG and mid not in chain:
|
||
chain.append(mid)
|
||
configured = [m for m in chain if CATALOG[m].is_configured()]
|
||
if not configured:
|
||
raise AIError(
|
||
"没有可用的模型:请在 backend/.env 中配置 GEMINI_API_KEY 或 NVIDIA_API_KEY"
|
||
)
|
||
return configured
|
||
|
||
|
||
# --- prompt construction ----------------------------------------------------
|
||
# Kept deliberately short: every extra column multiplies by the number of
|
||
# days sent, and the column names double as the vocabulary the model cites
|
||
# back in `basedOn`.
|
||
_CSV_COLUMNS = [
|
||
("date", "date"),
|
||
("steps", "steps"),
|
||
("distanceMeters", "dist_m"),
|
||
("heartRate", "rest_hr"),
|
||
("heartRateMax", "max_hr"),
|
||
("heartRateVariability", "hrv"),
|
||
("stress", "stress"),
|
||
("stressMax", "stress_max"),
|
||
("bodyBatteryHigh", "bb_high"),
|
||
("bodyBatteryLow", "bb_low"),
|
||
("spo2Avg", "spo2"),
|
||
("respirationAvg", "resp"),
|
||
("intensityMinutes", "intensity_min"),
|
||
("caloriesBurned", "kcal"),
|
||
("activeCalories", "active_kcal"),
|
||
("floorsAscended", "floors"),
|
||
("trainingReadiness", "readiness"),
|
||
("enduranceScore", "endurance"),
|
||
]
|
||
|
||
|
||
def build_prompt(summary, activities=None, day_budget=None):
|
||
"""Render health history as a compact CSV prompt.
|
||
|
||
CSV rather than JSON: roughly 4x fewer tokens for the same numbers, which
|
||
is what makes a full year of history practical to send.
|
||
"""
|
||
day_budget = day_budget if day_budget is not None else default_day_budget()
|
||
rows = summary[-day_budget:] if day_budget else summary
|
||
header = (
|
||
",".join(label for _, label in _CSV_COLUMNS)
|
||
+ ",sleep_h,sleep_q,sleep_deep_s,sleep_rem_s,sleep_awake_s"
|
||
)
|
||
lines = [header]
|
||
for r in rows:
|
||
cells = []
|
||
for key, _ in _CSV_COLUMNS:
|
||
value = r.get(key)
|
||
cells.append("" if value is None else str(value))
|
||
sleep = r.get("sleep") or {}
|
||
for key in ("duration", "quality", "deepSeconds", "remSeconds", "awakeSeconds"):
|
||
value = sleep.get(key)
|
||
cells.append("" if value is None else str(value))
|
||
lines.append(",".join(cells))
|
||
|
||
sections = [
|
||
SYSTEM_PROMPT,
|
||
f"\n## 每日健康数据(共 {len(rows)} 天,CSV)\n" + "\n".join(lines),
|
||
]
|
||
|
||
if activities:
|
||
act_lines = ["type,start,duration_s,distance_km,kcal,avg_hr,max_hr"]
|
||
for a in activities[:200]:
|
||
act_lines.append(
|
||
",".join(
|
||
str(a.get(k) if a.get(k) is not None else "")
|
||
for k in (
|
||
"activity_type", "start_time", "duration",
|
||
"distance", "calories", "heart_rate_average",
|
||
"heart_rate_max",
|
||
)
|
||
)
|
||
)
|
||
sections.append(
|
||
f"\n## 运动记录(共 {min(len(activities), 200)} 条,CSV)\n"
|
||
+ "\n".join(act_lines)
|
||
)
|
||
|
||
return "\n".join(sections)
|
||
|
||
|
||
# --- response parsing -------------------------------------------------------
|
||
_VALID_PRIORITIES = {"high", "medium", "low"}
|
||
_FENCE = re.compile(r"^\s*```(?:json)?\s*|\s*```\s*$", re.MULTILINE)
|
||
|
||
|
||
def parse_recommendations(text):
|
||
"""Coerce a model reply into the same shape the rule engine returns.
|
||
|
||
Models routinely wrap JSON in markdown fences or add a sentence before it,
|
||
despite instructions, so both are tolerated here.
|
||
"""
|
||
if not text or not text.strip():
|
||
raise AIError("模型返回空响应")
|
||
|
||
cleaned = _FENCE.sub("", text).strip()
|
||
try:
|
||
data = json.loads(cleaned)
|
||
except ValueError:
|
||
start, end = cleaned.find("["), cleaned.rfind("]")
|
||
if start == -1 or end <= start:
|
||
raise AIError(f"模型未返回 JSON 数组: {text[:200]}")
|
||
try:
|
||
data = json.loads(cleaned[start : end + 1])
|
||
except ValueError as e:
|
||
raise AIError(f"模型返回的 JSON 无法解析: {e}") from e
|
||
|
||
if isinstance(data, dict):
|
||
data = [data]
|
||
if not isinstance(data, list):
|
||
raise AIError("模型返回的不是 JSON 数组")
|
||
|
||
recs = []
|
||
for i, item in enumerate(data):
|
||
if not isinstance(item, dict):
|
||
continue
|
||
text_value = (item.get("recommendation") or "").strip()
|
||
if not text_value:
|
||
continue
|
||
priority = str(item.get("priority", "medium")).lower()
|
||
if priority not in _VALID_PRIORITIES:
|
||
priority = "medium"
|
||
based_on = item.get("basedOn")
|
||
if not isinstance(based_on, list):
|
||
based_on = []
|
||
recs.append(
|
||
{
|
||
"id": f"ai-{i}",
|
||
"category": (item.get("category") or "综合").strip(),
|
||
"recommendation": text_value,
|
||
"priority": priority,
|
||
"basedOn": [str(b) for b in based_on],
|
||
"source": "ai",
|
||
}
|
||
)
|
||
|
||
if not recs:
|
||
raise AIError("模型未返回任何有效建议")
|
||
|
||
order = {"high": 0, "medium": 1, "low": 2}
|
||
recs.sort(key=lambda r: order[r["priority"]])
|
||
return recs
|
||
|
||
|
||
# --- entry point ------------------------------------------------------------
|
||
# One CSV day is ~40 characters ≈ 10 tokens. Half the window is left for the
|
||
# system prompt, the activity table and the model's own answer.
|
||
_TOKENS_PER_DAY = 10
|
||
_WINDOW_UTILISATION = 0.5
|
||
|
||
|
||
def max_days_for(provider, day_budget=None):
|
||
"""How many days of history fit in this model's context window.
|
||
|
||
Models in the chain have windows that differ by more than an order of
|
||
magnitude (32k for a local Ollama vs 1M for Gemini), so the payload has to
|
||
be sized per model — a prompt that fits Gemini would overflow Ollama.
|
||
"""
|
||
day_budget = day_budget if day_budget is not None else default_day_budget()
|
||
fits = int(provider.context_window * _WINDOW_UTILISATION / _TOKENS_PER_DAY)
|
||
return max(1, min(day_budget, fits)) if day_budget else max(1, fits)
|
||
|
||
|
||
def generate(summary, activities=None, preferred_model=None, day_budget=None):
|
||
"""Ask the first healthy model in the chain for recommendations.
|
||
|
||
Returns (recommendations, meta). `meta` records which model answered, how
|
||
much history it actually saw, and every model that failed on the way —
|
||
the failures are kept even on success so a silent degradation to a weaker
|
||
model is still visible.
|
||
"""
|
||
chain = resolve_chain(preferred_model)
|
||
errors = []
|
||
|
||
for model_id in chain:
|
||
provider = CATALOG[model_id]
|
||
days = max_days_for(provider, day_budget)
|
||
prompt = build_prompt(summary, activities, days)
|
||
try:
|
||
completion = provider.generate(prompt)
|
||
recs = parse_recommendations(completion.text)
|
||
return recs, {
|
||
"model": model_id,
|
||
"provider": provider.name,
|
||
"upstream": completion.upstream,
|
||
"days": min(len(summary), days),
|
||
"fallbackFrom": [e["model"] for e in errors],
|
||
"errors": errors,
|
||
}
|
||
except AIError as e:
|
||
errors.append({"model": model_id, "error": str(e)})
|
||
|
||
detail = "; ".join(f"{e['model']}: {e['error']}" for e in errors)
|
||
raise AIError(f"所有模型均失败 -> {detail}")
|