背景:文档停在 Node.js 时代或甲骨文 8123 部署,与生产(NAS :8124 + Flask + auth-hub + ai-gateway)严重脱节,曾导致凭旧记忆误判'无线上环境'。 - CLAUDE.md 重写:技术栈/结构/命令/部署事实/关键坑(F7 button、UTC 日期、 429 退避以 DB 为准、迁移幂等、AI 生成耗时) - docs/ARCHITECTURE.md 重写为 Flask 蓝图+services+可插拔数据层 + NAS 部署 - docs/DEVELOPMENT.md 重写为 Flask/CRA 开发指南 + push.sh 部署流程 - docs/REQUIREMENTS.md:部署条目改 NAS 8124;补 auth-hub/AI 教练/新修复 - docs/AUTH_HUB_INTEGRATION.md 新增(补 .env.example 悬空引用) - README.md:技术栈/DB/auth-hub/API 清单/部署节修正 - backend/config.py 与 .env.example:AUTH_HUB_REDIRECT_URI 默认 8123→8124, MariaDB 注释 Oracle→NAS - tests:GatewayCourtesy 并发测试对齐 MAX_CONCURRENT(AI_JOB_CONCURRENCY=2); conftest 禁用 create_app 后台队列线程,修整库测试 flaky(585 passed)
104 lines
4.5 KiB
Plaintext
104 lines
4.5 KiB
Plaintext
# --- Server ---
|
|
# BACKEND_PORT takes precedence over PORT. Prefer it: many tools inject PORT
|
|
# for the frontend, and Flask would otherwise take the React dev server's port.
|
|
BACKEND_PORT=5000
|
|
|
|
# --- Database: sqlite (default) or mariadb ---
|
|
DB_TYPE=sqlite
|
|
# SQLite file (used when DB_TYPE=sqlite)
|
|
DATABASE_PATH=./data/health.db
|
|
|
|
# MariaDB (used when DB_TYPE=mariadb) — production DB on the NAS
|
|
# (192.168.50.64, MariaDB 10.11). Connection is over the socket
|
|
# /run/mysqld/mysqld10.sock (or TCP 127.0.0.1:3306) as root; the socket path
|
|
# only matters when TCP auth is disabled for the app user.
|
|
# MARIADB_SOCKET=/run/mysqld/mysqld10.sock
|
|
# MARIADB_HOST=127.0.0.1
|
|
# MARIADB_PORT=3306
|
|
# MARIADB_USER=root
|
|
# MARIADB_PASSWORD=your_production_mariadb_password
|
|
# MARIADB_DATABASE=garmin_health_lab
|
|
|
|
# --- Auth ---
|
|
# CHANGE THIS in production! Used to sign JWTs (7-day expiry by default).
|
|
JWT_SECRET=dev_secret_change_me
|
|
JWT_EXPIRY_DAYS=7
|
|
|
|
# --- auth-hub OAuth2 provider (centralized SSO) ---
|
|
# See docs/AUTH_HUB_INTEGRATION.md for setup instructions.
|
|
#
|
|
# Base URL of the auth-hub service
|
|
AUTH_HUB_BASE_URL=http://129.146.26.249:5300
|
|
#
|
|
# OAuth2 client credentials (obtain from auth-hub.manage_clients create)
|
|
# NOTE: these must be registered against the auth-hub instance AUTH_HUB_BASE_URL
|
|
# actually points to (dev vs prod are separate databases with separate clients).
|
|
# Put the REAL values in backend/.env (gitignored) — never here.
|
|
AUTH_HUB_CLIENT_ID=your_client_id
|
|
AUTH_HUB_CLIENT_SECRET=your_client_secret
|
|
#
|
|
# Callback URL (must exactly match what's registered in auth-hub). Production
|
|
# registers both the public frp address (129.146.26.249:8124) and the LAN
|
|
# address (192.168.50.64:8124).
|
|
AUTH_HUB_REDIRECT_URI=http://129.146.26.249:8124/auth/callback
|
|
|
|
# --- CORS (comma-separated allowed front-end origins) ---
|
|
# localhost stays in the production list on purpose: CORS is not an auth
|
|
# boundary — every data route requires a valid JWT — so allowing a developer's
|
|
# dev server costs nothing and saves toggling this on every session.
|
|
CORS_ORIGIN=http://localhost:3000,http://localhost:5173
|
|
|
|
# --- AI models (text-only, large context) ---
|
|
# Put REAL keys in backend/.env — that file is gitignored. Never commit keys.
|
|
# Any model whose credentials are absent is skipped automatically.
|
|
|
|
# Self-hosted AI gateway (model id "gateway"). OpenAI-compatible; it fans out
|
|
# over nvidia/gemini/ollama itself and rotates several Gemini keys, so it
|
|
# absorbs single-vendor quota limits. Reached directly, bypassing any local
|
|
# HTTP proxy. NOTE: its NVIDIA upstream is a large reasoning model — replies
|
|
# can take 2-3 minutes, so set AI_TIMEOUT_SECONDS accordingly.
|
|
# HTTPS (Caddy, strips the /ai prefix) rather than http://…:5100 — the token
|
|
# rides in an Authorization header and should not cross the internet in clear.
|
|
AI_GATEWAY_BASE_URL=https://ai.zichuan.xyz/v1
|
|
AI_GATEWAY_TOKEN=
|
|
AI_GATEWAY_MODEL=ai-gateway-auto
|
|
|
|
# Google AI Studio -> "gemini-flash". Free-tier quota is small; 429s are common.
|
|
GEMINI_API_KEY=
|
|
|
|
# NVIDIA NIM -> "llama-70b", "nemotron-49b", "mistral-large".
|
|
# Model ids come from that account's live GET /v1/models — do not guess them.
|
|
NVIDIA_API_KEY=
|
|
# NVIDIA_BASE_URL=https://integrate.api.nvidia.com/v1
|
|
|
|
# Preference order. The first configured model answers; if it fails or times
|
|
# out, the next is tried. Read per request, so changes need no restart.
|
|
AI_MODEL_CHAIN=gateway,gemini-flash,llama-70b
|
|
|
|
# Max days of history sent (CSV-encoded). Trimmed further per model so the
|
|
# payload always fits that model's own context window.
|
|
AI_DAY_BUDGET=365
|
|
|
|
# Measured against the gateway, not guessed: a trivial prompt took 138s end to
|
|
# end, because its primary upstream emits a full chain of thought before the
|
|
# answer. Nothing user-facing blocks on this (the briefing generates in a
|
|
# background thread), but the timeout still has to clear the real latency.
|
|
AI_TIMEOUT_SECONDS=300
|
|
# Output cap. Reasoning models spend part of it thinking before they answer;
|
|
# entries that need more declare their own budget in services/ai.py.
|
|
AI_MAX_TOKENS=1024
|
|
|
|
# Output cap for the AI coach (晨报 / 趋势归因 / Copilot). Larger than
|
|
# AI_MAX_TOKENS above: the same reasoning trace is spent from this budget
|
|
# before the answer starts, and at 1024 the reply was all thinking with the
|
|
# JSON truncated away.
|
|
AI_COACH_MAX_TOKENS=4000
|
|
|
|
# --- AI coach job queue ---
|
|
# Three projects share the gateway, so going too high causes 502s.
|
|
AI_JOB_CONCURRENCY=2
|
|
AI_JOB_GAP_SECONDS=5
|
|
# Set AI_JOBS=false to stop consuming entirely (screens then show the computed
|
|
# figures with no model reading).
|
|
AI_JOBS=true
|