Files
GarminHealthLab/deploy/push.sh
ericwyuan 74f2721401 fix(deploy): S99garmin.sh 丢了可执行位,重启两天没人发现服务已经死了
09-06 07:08 NAS 重启(DHCP 顺带把它的局域网地址从 .64 换成了 .65),佳明健康
服务再也没起来——两天后才被发现。

根因:`deploy/S99garmin.sh` 本地就没有 +x(`-rw-r--r--`,另外三个部署脚本都是
`-rwxr-xr-x`)。DSM 的 rc.d 启动器碰到一个没有执行权限的软链接目标,既不执行
也不报错,静默跳过。push.sh 用 tar 同步 deploy/ 目录会原样带走权限位,所以自
09-01 起只要重新部署一次,就会把 NAS 上(原本可能是手工修过的)可执行权限
再次覆盖回不可执行——这颗雷从那天就埋下了,直到这次重启才被踩到。

- 本地补上 S99garmin.sh 的 +x
- push.sh 同步完 deploy/ 之后显式 chmod +x 三个会被开机脚本或部署流程直接
  执行的文件,并断言生效——以后这个位再丢,部署会报错退出,不会再悄悄失效
  到下次重启才现形
- push.sh 的默认目标 IP 也是这次连带发现的坑:硬编码的 192.168.50.64 已经
  证明会被 DHCP 换掉,改成当前地址 .65 并加注释——真正的解法是给 NAS 做
  DHCP 保留(MAC 90:09:d0:22:a6:33),不在这次改动范围内

现场同时处理:手动 chmod +x 后 sh start.sh 拉起服务;frpc 配置里的
localIP 硬编码着 192.168.50.64(同样的病),改成 127.0.0.1 后不再受局域网
IP 变化影响,公网已验证恢复。frpc 改动是 NAS 系统配置,不在这个仓库里,
旧文件备份在 NAS 的 /etc/frp/frpc.toml.bak-20260908。

遗留:auth-hub 注册的 LAN 回调地址还是 http://192.168.50.64:8124/auth/callback
(CLAUDE.md:99),局域网内直接用 .65 访问会在登录环节被拒;公网入口不受影响。
要不要把它也换成 DHCP 保留后的固定地址,留给用户决定。

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-08 07:27:22 +08:00

137 lines
6.2 KiB
Bash
Executable File

#!/bin/sh
# Push this working tree to the NAS and restart it.
#
# The NAS only accepts password auth, so one ssh master connection is opened
# up front and every later step rides on it: you type the password once, not
# five times. (macOS's bundled rsync 2.6.9 cannot carry a password through
# -e at all — hence tar over ssh.)
#
# ./deploy/push.sh [user@host] [port]
#
# The password is never stored here — ssh asks for it on your terminal. To
# stop being asked at all, run `ssh-copy-id -p 2222 ericwyuan@192.168.50.64`
# once; after that this script runs unattended.
#
# Never touches .env, .venv or the database on the far side.
set -e
# DHCP-assigned and has already changed once (192.168.50.64 -> .65 on the
# 2026-09-06 reboot) — this default will go stale again after the NAS's next
# reboot unless a DHCP reservation is set for it on the router. Until then,
# pass the current IP explicitly: `./deploy/push.sh ericwyuan@<ip>`.
HOST="${1:-ericwyuan@192.168.50.65}"
PORT="${2:-2222}"
# Needed for the sudo restart below. Prompted for rather than stored, and only
# used over the already-authenticated ssh master connection.
PASS="${NAS_PASSWORD:-}"
REPO="$(cd "$(dirname "$0")/.." && pwd)"
CTL="$(mktemp -u /tmp/garmin-deploy-XXXXXX)"
sh_() { ssh -S "$CTL" -o BatchMode=yes "$HOST" "$@"; }
cleanup() { ssh -S "$CTL" -O exit "$HOST" 2>/dev/null || true; }
trap cleanup EXIT
echo "==> connecting to $HOST:$PORT (password prompt follows, once)"
ssh -M -S "$CTL" -fN -p "$PORT" -o ControlPersist=300 "$HOST"
if [ -z "$PASS" ]; then
# Same password as the ssh login; asked for separately because ssh consumed
# the first one itself and sudo on the far side needs it on stdin.
printf 'sudo password for %s (needed to restart the root-owned service): ' "$HOST" >&2
stty -echo 2>/dev/null; read PASS; stty echo 2>/dev/null; echo >&2
fi
# The app dir has moved before; find it rather than assume it.
APP=$(sh_ 'for d in ~/apps/garmin-health-lab /volume1/web/garmin-health-lab; do
[ -d "$d/backend" ] && { echo "$d"; break; }; done')
[ -n "$APP" ] || { echo "cannot find the app dir on $HOST" >&2; exit 1; }
echo "==> app dir: $APP"
# STATIC_DIR is ./static relative to backend/, which is where start.sh cds to.
STATIC="$APP/backend/static"
# `cmd && VAR=x` would trip `set -e` when cmd fails, so spell it out.
if sh_ "[ -f '$APP/static/index.html' ]" 2>/dev/null; then
STATIC="$APP/static"
fi
echo "==> static dir: $STATIC"
# macOS bsdtar writes com.apple.provenance xattrs and ._ resource forks that
# the NAS's tar cannot read; it warns once per file and copies nothing useful.
TAR="tar czf - --no-xattrs"
export COPYFILE_DISABLE=1
echo "==> backend"
$TAR --exclude .venv --exclude .env --exclude __pycache__ \
--exclude '*.db' --exclude tests --exclude .pytest_cache \
-C "$REPO/backend" . | sh_ "tar xzf - -C '$APP/backend'"
# The restart below runs the NAS's own copies of stop.sh / start.sh, so they
# have to travel with the code — otherwise a fix to them lands in git, the
# deploy reports success, and the far side keeps running the old ones. That is
# exactly how the broken pgrep guards survived a deploy that was meant to fix
# them.
echo "==> deploy scripts"
$TAR -C "$REPO/deploy" . | sh_ "tar xzf - -C '$APP/deploy'"
# tar carries whatever bit the local file has, and it silently had none:
# S99garmin.sh lost +x at some point before this repo existed, every deploy
# since then re-clobbered the NAS's copy back to non-executable, and nothing
# noticed until a 2026-09-06 reboot left the service down for two days —
# DSM's rc.d runner does not execute a script it cannot execute, and does not
# warn either. Asserted here so a future loss of the bit fails the deploy
# instead of failing silently at the next reboot.
sh_ "chmod +x '$APP/deploy/S99garmin.sh' '$APP/deploy/start.sh' '$APP/deploy/stop.sh'"
if ! sh_ "[ -x '$APP/deploy/S99garmin.sh' ]"; then
echo "S99garmin.sh is not executable after chmod — the boot script will not run on the next reboot" >&2
exit 1
fi
echo "==> static (cleared first, so stale JS chunks do not pile up)"
if [ ! -f "$REPO/client/build/index.html" ]; then
echo "client/build is missing — run 'npm run build' first" >&2
exit 1
fi
sh_ "rm -rf '$STATIC' && mkdir -p '$STATIC'"
$TAR -C "$REPO/client/build" . | sh_ "tar xzf - -C '$STATIC'"
# The service is started at boot as root (DSM Task Scheduler -> S99garmin.sh),
# so logs/error.log and logs/access.log are root-owned. Restarting as
# ericwyuan therefore fails instantly — gunicorn cannot open its own error log
# — and this went unnoticed for a whole deploy: stop.sh could not kill a root
# process either, so the OLD master kept serving :8124, start.sh's health
# check saw a 200 and reported "started", and the new code was never loaded.
# Hence sudo, matching how the service actually runs.
SUDO="echo '$PASS' | sudo -S"
echo "==> restart (sudo: the service runs as root, started at boot)"
# DSM has no pgrep, and an unprivileged `ps` cannot see a root-owned process
# — either one silently returns nothing, which would make the pid comparison
# below always "pass" and put the check right back where it started.
pids_() {
sh_ "echo '$PASS' | sudo -S ps -eo pid,args 2>/dev/null \
| grep '$APP/backend/.venv/bin/gunicorn' | grep -v grep \
| awk '{print \$1}' | sort -n | tr '\n' ','"
}
BEFORE=$(pids_ || true)
sh_ "cd '$APP' && $SUDO sh deploy/stop.sh >/dev/null 2>&1; sleep 3; $SUDO sh deploy/start.sh" \
|| { echo "restart failed" >&2; exit 1; }
echo "==> health"
sh_ "curl -sf -m 5 -o /dev/null -w 'local api: %{http_code}\n' \
http://127.0.0.1:8124/api/health/status" || echo "local api: unreachable"
# A 200 alone proves nothing: it is exactly what a surviving old master
# returns. The master pid must have changed for the new code to be loaded.
AFTER=$(pids_ || true)
if [ -z "$AFTER" ]; then
echo "cannot see any gunicorn process - the pid check could not run." >&2
echo "Verify by hand before trusting this deploy." >&2
exit 1
fi
if [ "$BEFORE" = "$AFTER" ]; then
echo "the gunicorn pids did not change ($AFTER) - the old process is still" >&2
echo "serving and your changes are NOT live. Check $APP/logs/error.log." >&2
exit 1
fi
echo "==> gunicorn restarted: $BEFORE -> $AFTER"
echo "done. The public URL takes a few seconds longer (frp reconnecting)."