Web Day 34 健康檢查與監控
執行需求:CPU 可跑。上線的最後一步是「出事時要知道」。今天要把 Day 30 寫的 `/healthz` 端點升級成完整的健康檢查系統:liveness 與 readiness 兩個端點、Prometheus metrics 端點、以及用 Alertmanager 觸發告警。我們也會寫幾支 Python 小工具幫忙做 uptime monitoring、log 解析、alert routing。沒有 Prometheus 也可以用純 Python 寫的輕量替代方案,整套設計對 side project 與中型服務都適用。
引言
一個系統上線之後,「還活著嗎」這件事不會自己告訴你。應用程式可能因為 OOM 被 kill、Postgres 可能因為磁碟滿了而拒絕連線、Redis 可能因為網路不穩定而 timeout——這些問題都要在「使用者打客服電話抱怨」之前先被發現。今天的目標是建立一整套可觀察性(observability)機制:知道系統現在的狀態、保留歷史資料、出問題時主動通知。
可觀察性由三個支柱組成:metrics(指標)、logs(日誌)、traces(追蹤)。這個系列涵蓋前兩個:metrics 用 Prometheus 抓、logs 用 Python logging 結構化輸出。traces 在 side project 階段暫時不需要,Day 43 壓測章節會再簡短介紹。讀完今天,你會拿到一組 health endpoints、Prometheus metrics 端點、以及兩支用 Python 寫的監控工具腳本,能在沒有重型監控系統的情況下維持基本可觀察性。
今天的學習地圖分五段:第一,理解 liveness 與 readiness 的差別與正確用法;第二,把 /healthz 與 /readyz 兩個端點用 FastAPI 實作出來;第三,加上 Prometheus /metrics 端點與 JSON 結構化 logging;第四,寫兩支監控腳本(uptime_monitor、alert_watcher)做離線版的健康檢查;第五,了解沒有 Prometheus 時的替代方案與常見踩雷。學完之後你會有一個「出問題時能在 5 分鐘內知道」的系統,這是上線的最低標準。
為什麼監控是上線的最後一步
Day 30 到 Day 33 我們把 Docker、PostgreSQL、CI/CD、Caddy 都接好了,技術上這個服務已經可以對外提供。但「可以對外提供」跟「可以在正式環境穩定運行」是兩件事。正式環境會遇到開發環境看不到的問題:流量突然暴增、磁碟空間被 log 寫滿、DNS 過期、憑證忘了續期、依賴的第三方 API 故障。沒有監控的服務就像沒有儀表板的飛機——你可以飛,但出事時你不知道發生了什麼。
監控的價值不只是「出問題時收到通知」,更是「出問題時能在最短時間內定位問題」。當使用者回報「網站壞了」,你要在 5 分鐘內回答:是 API 掛了還是資料庫?是一台掛了還是多台?最近一次部署到現在有什麼改動?沒有 metrics 與 logs 的人只能亂猜;有了完整可觀察性的人可以一眼看出根因。今天我們把這個基礎建設蓋起來。
健康檢查的兩種角色
很多人會把 health check 當成「伺服器有回應就好」,但 Kubernetes 與現代容器平台區分兩種角色:liveness(活著)與 readiness(準備好)。前者回答「這個 process 是不是還在跑」,後者回答「這個 process 能不能接 request」。兩者答案不同的時候,就是系統正在「半死不活」的中間狀態——process 還在但暫時不該收新流量。
把這兩個概念延伸到你自己的應用:liveness 端點只檢查 process 本身還能回應(檢查 in-memory 物件、必要的常數是否還在);readiness 端點則額外檢查外部依賴(Postgres、Redis、外部 API)是否還連得上。差別在於 liveness 失敗時應該重啟 container、readiness 失敗時應該暫時從 load balancer 移除但保留 process。今天我們把這兩個端點都實作。
一個常見的設計錯誤是把 liveness 跟 readiness 合在一起。合在一起的結果是:當 Postgres 短暫不穩時,liveness 也跟著失敗,平台就重啟 container;但 container 重啟並不會讓 Postgres 變穩,反而讓 API 反覆重啟、connection pool 一直 reset。分開之後,Postgres 不穩時只會讓 readiness 失敗,平台把流量切走、process 留著等 Postgres 恢復,這才是正確的行為。
實作 liveness 與 readiness
FastAPI 的 dependency injection 系統很適合做這件事。我們把檢查邏輯寫成可重用的 dependency,再掛到對應的端點。先看一下 SQLModel 與 SQLAlchemy 的基礎版本:
"""app/health.py:liveness 與 readiness 檢查。"""
from collections.abc import Iterator
from datetime import datetime, timezone
from fastapi import APIRouter, Depends, status
from sqlalchemy import text
from sqlalchemy.engine import Engine
from sqlalchemy.orm import Session, sessionmaker
router = APIRouter(tags=["health"])
def get_engine() -> Engine:
"""在 lifespan 中建立的 engine 透過 app.state 注入。"""
from app.main import app
return app.state.engine
def get_session_factory() -> sessionmaker[Session]:
from app.main import app
return app.state.session_factory
def get_db_session(
factory: sessionmaker[Session] = Depends(get_session_factory),
) -> Iterator[Session]:
session = factory()
try:
yield session
finally:
session.close()
@router.get("/healthz")
def liveness() -> dict:
"""最輕量的活著檢查,不接觸任何外部依賴。"""
return {"status": "alive", "ts": datetime.now(timezone.utc).isoformat()}
@router.get("/readyz")
def readiness(
session: Session = Depends(get_db_session),
) -> dict:
"""檢查 Postgres 連線是否可用。"""
try:
session.execute(text("SELECT 1"))
return {"status": "ready", "db": "ok"}
except Exception as exc:
return {"status": "not_ready", "db": str(exc)}
這份程式碼有兩個關鍵設計:`/healthz` 不接觸任何外部依賴,只回應時間戳記,這樣即使資料庫全掛,liveness 仍然會通過、不會誤觸 container 重啟;`/readyz` 用 Depends 注入 session 並實際 query 一次,回應包含 db 狀態,讓監控系統能看到「資料庫是否還連得上」。`status_code` 在 failure 時可以改用 `response.status_code = status.HTTP_503_SERVICE_UNAVAILABLE`,讓 Caddy 與 Prometheus 能用 HTTP 狀態區分成功與失敗,這對自動告警非常有用。
另一個延伸是「深度健康檢查」:當 readiness 失敗時,可以進一步查 Postgres 的 `pg_stat_activity`、Redis 的 `INFO replication`、外部 API 的認證 token 是否過期。這些深度檢查通常做成獨立的 `/diag` 端點,不掛在 liveness/readiness 上,避免每次健康檢查都跑這些昂貴的查詢。
Prometheus metrics 端點
Prometheus 是 CNCF 畢業的開源 metrics 系統,透過 HTTP pull 模型抓 metrics:在應用程式上開一個 /metrics 端點,Prometheus 伺服器定期來拉。Python 端用 prometheus-client 函式庫把計數器、histogram 包好,FastAPI 起一個端點 expose 出去。
"""app/metrics.py:Prometheus metrics 設定與端點。"""
from prometheus_client import CONTENT_TYPE_LATEST, Counter, Histogram, generate_latest
REQUEST_COUNT = Counter(
"http_requests_total",
"HTTP 請求總數",
["method", "path", "status"],
)
REQUEST_LATENCY = Histogram(
"http_request_duration_seconds",
"HTTP 請求耗時(秒)",
["method", "path"],
buckets=(0.01, 0.05, 0.1, 0.25, 0.5, 1.0, 2.5, 5.0, 10.0),
)
DB_QUERY_LATENCY = Histogram(
"db_query_duration_seconds",
"資料庫查詢耗時(秒)",
["op"],
buckets=(0.001, 0.005, 0.01, 0.05, 0.1, 0.5, 1.0),
)
def metrics_response() -> tuple[bytes, str]:
return generate_latest(), CONTENT_TYPE_LATEST
這份程式碼定義了三組 metrics:`http_requests_total` 用 Counter 累加每個 method/path/status 組合的請求數;`http_request_duration_seconds` 用 Histogram 記錄耗時分布,buckets 設定覆蓋毫秒到十秒的常見範圍;`db_query_duration_seconds` 專門追蹤資料庫查詢耗時,這是效能調校最常看的一個指標。在路由或 middleware 用 time.perf_counter 量測每次請求,並把結果丟進 Histogram 與 Counter,metrics 端點就會自動累計。
FastAPI 端用 response_model 直接 expose 這份 metrics:
"""app/main.py 片段:註冊 metrics 端點。"""
from fastapi import FastAPI, Response
from app.metrics import REQUEST_COUNT, REQUEST_LATENCY, metrics_response
import time
app = FastAPI()
@app.middleware("http")
async def metrics_middleware(request, call_next):
started = time.perf_counter()
response = await call_next(request)
elapsed = time.perf_counter() - started
REQUEST_COUNT.labels(request.method, request.url.path, response.status_code).inc()
REQUEST_LATENCY.labels(request.method, request.url.path).observe(elapsed)
return response
@app.get("/metrics")
def metrics() -> Response:
payload, content_type = metrics_response()
return Response(content=payload, media_type=content_type)
這個 middleware 在每次請求時量測耗時、把結果推到 Counter 與 Histogram。`/metrics` 端點把累積的 metrics 序列化成 Prometheus 格式(純文字),讓 Prometheus server 來抓。對正式服務來說,這層包裝就足以看到每秒請求數、p95 回應時間、錯誤率等關鍵指標。
結構化 logging
metrics 告訴你「現在怎樣」,logs 告訴你「發生了什麼」。把 log 結構化(JSON 格式)能讓 log 聚合系統(Elasticsearch、Loki)做更好的查詢與聚合。Python 內建 logging 模組支援自訂 Formatter,輸出 JSON 格式的 log。
"""app/logging_config.py:JSON 格式的 logging 設定。"""
import json
import logging
import sys
from datetime import datetime, timezone
class JsonFormatter(logging.Formatter):
"""把 LogRecord 轉成 JSON 字串。"""
def format(self, record: logging.LogRecord) -> str:
payload = {
"ts": datetime.now(timezone.utc).isoformat(),
"level": record.levelname,
"logger": record.name,
"msg": record.getMessage(),
}
if record.exc_info:
payload["exc"] = self.formatException(record.exc_info)
for key, value in record.__dict__.items():
if key in payload or key.startswith("_"):
continue
try:
json.dumps(value)
payload[key] = value
except (TypeError, ValueError):
payload[key] = repr(value)
return json.dumps(payload, ensure_ascii=False)
def configure(level: str = "INFO") -> None:
handler = logging.StreamHandler(sys.stdout)
handler.setFormatter(JsonFormatter())
root = logging.getLogger()
root.handlers = [handler]
root.setLevel(level)
這份設定把所有 log 強制走 JSON Formatter,每行一個結構化記錄,包含時間戳記、層級、logger 名稱、訊息內容與例外資訊。`record.__dict__` 的部分會把自訂欄位(例如 `request_id`、`user_id`)自動納入,這對 tracing 很有用。實務上可以把 `JsonFormatter` 接上 Loki 或 Elasticsearch,就能用 SQL 風格的查詢(例如「找出 status_code=500 且 path=/api/bookings 的所有 log」)做除錯。
uptime monitoring 工具
Prometheus 雖然能監控 metrics,但「網站是否還活著」這種最基礎的監控,更適合用一支小工具每分鐘跑一次。我們寫一支輕量的 uptime monitor:對多個端點輪詢、記錄回應時間與狀態、寫進 SQLite 留下歷史紀錄。
"""scripts/uptime_monitor.py:每分鐘對多個端點做 health check 並寫入 SQLite。"""
import datetime as dt
import sqlite3
import sys
import time
import httpx
TARGETS = [
("api", "https://api.example.com/healthz"),
("admin", "https://admin.example.com/healthz"),
("db-via-api", "https://api.example.com/readyz"),
]
def init_db(path: str) -> None:
conn = sqlite3.connect(path)
conn.execute(
"CREATE TABLE IF NOT EXISTS uptime ("
"name TEXT, ts TEXT, status INTEGER, duration_ms REAL, msg TEXT)"
)
conn.commit()
return conn
def probe(name: str, url: str, conn: sqlite3.Connection) -> None:
started = time.perf_counter()
try:
res = httpx.get(url, timeout=5)
duration = (time.perf_counter() - started) * 1000
row = (name, dt.datetime.utcnow().isoformat(), res.status_code, duration, "ok" if res.status_code == 200 else "fail")
except httpx.HTTPError as exc:
duration = (time.perf_counter() - started) * 1000
row = (name, dt.datetime.utcnow().isoformat(), 0, duration, repr(exc))
conn.execute("INSERT INTO uptime VALUES (?, ?, ?, ?, ?)", row)
conn.commit()
print(row)
if __name__ == "__main__":
db_path = sys.argv[1] if len(sys.argv) > 1 else "uptime.db"
conn = init_db(db_path)
while True:
for name, url in TARGETS:
probe(name, url, conn)
time.sleep(60)
這支腳本示範了「最簡單的 uptime monitor」:用 SQLite 保存所有檢查結果、用 while True 與 sleep 60 跑成常駐程式。對 side project 來說,這比架一套 Prometheus + Grafana 簡單太多。要查「昨天凌晨 3 點 API 有沒有掛」,只要 `sqlite3 uptime.db "SELECT * FROM uptime WHERE name='api' AND ts > '2025-07-20 03:00'"` 就能看到。
alert 通知:當問題發生時
metrics 與 uptime 都只是「資料」,要變成「通知」還需要 alert routing。對 side project 來說,最簡單的做法是「連續失敗 N 次就寄信」。我們寫一支簡單的 alert watcher,讀 uptime.db,當某個端點連續失敗 3 次就發通知。
"""scripts/alert_watcher.py:讀 uptime.db,連續失敗 3 次就印出告警。"""
import os
import sqlite3
import sys
from collections import defaultdict
def recent_failures(conn: sqlite3.Connection, limit: int = 3) -> dict[str, list[tuple]]:
cur = conn.execute(
"SELECT name, status, msg FROM uptime ORDER BY id DESC LIMIT 200"
)
by_name: dict[str, list[tuple]] = defaultdict(list)
for name, status, msg in cur.fetchall():
by_name[name].append((status, msg))
return by_name
def report(by_name: dict[str, list[tuple]], threshold: int = 3) -> int:
fired = 0
for name, rows in by_name.items():
recent = rows[:threshold]
if all(status != 200 for status, _ in recent):
fired += 1
print(f"[ALERT] {name} 連續 {threshold} 次失敗:{recent[0][1]}")
return fired
if __name__ == "__main__":
db_path = sys.argv[1] if len(sys.argv) > 1 else "uptime.db"
webhook = os.environ.get("ALERT_WEBHOOK")
conn = sqlite3.connect(db_path)
fired = report(recent_failures(conn))
if fired and webhook:
import httpx
httpx.post(webhook, json={"text": f"alert watcher fired {fired}"})
sys.exit(0 if fired == 0 else 1)
這支腳本從 uptime.db 撈最近 200 筆、依名稱分組、檢查「最近 N 次是否全失敗」。實務上可以把 webhook 接到 Slack、Discord 或 LINE Notify。對三人以下的小團隊,這套「SQLite + 簡單 watcher」就能撐住基本的監控需求;等系統長大再升級到 Prometheus + Alertmanager。
沒有 Prometheus 時的替代流程
Prometheus + Grafana 是最完整的方案,但對 side project 來說有點重。如果你只是想「知道系統還活著」,前面那兩支腳本(uptime_monitor + alert_watcher)就夠了。如果想看到 metrics 圖表但不想架 Prometheus,可以用 hosted 服務:Grafana Cloud 免費層、New Relic free tier、UptimeRobot、Uptime Kuma 自架等。Uptime Kuma 是一個開源的 self-hosted 監控工具,介面漂亮、支援多種通訊協定(HTTP、TCP、Ping、DNS),是 side project 的好選擇。
另一個輕量替代是用 APScheduler 把 metrics 抓取排程跑在 API process 裡:每 30 秒把 counters 與 histograms 的當下值寫進 CSV,每天用 matplotlib 畫一張圖。這個做法對開發環境很方便,正式環境就略顯陽春。如果團隊規模小到只有一個人,又不想管 Prometheus 伺服器,這套是務實的選擇。
還有一個有趣的選擇是 OpenTelemetry(OTel):它把 metrics、logs、traces 三者整合在同一個 SDK,可以無痛從「純 metrics」升級到「加上 traces」。我們的 prometheus-client 設定可以保留,把 exporter 從 Prometheus 換成 OTel collector,後端就可以任意切換 Prometheus / Grafana / Datadog 等。今天先不展開 OTel,後續 Day 43 壓測章節會再簡短介紹。OTel 對 side project 來說負擔偏重,但如果你的團隊已經有 Datadog 或 New Relic 的合約,OTel 是最划算的選擇。
常見錯誤與踩雷
health 端點做太重的檢查。有些團隊會在 /healthz 裡面跑 migration 檢查、ping 全部外部 API、撈完整份 DB schema。這會讓 health 端點本身就吃掉大量資源,而且 Kubernetes 預設每 10 秒就打一次 health,過重的檢查會讓 API 變慢。正確做法是 health 只做最輕量的檢查(in-memory),重的檢查放到 readiness 或獨立的 /diag 端點。曾經有團隊把 S3 上傳當成 health 檢查的一部分,結果 S3 短暫不穩就把整個 K8s 叢集的 container 全部重啟,是非常慘痛的經驗。
metrics 卡住 Prometheus 抓取。Prometheus pull 模型假設 /metrics 端點在 1 到 2 秒內回應。如果你的 metrics 端點本身跑很慢(例如 histogram bucket 太多、labels 組合爆炸),會讓 Prometheus 抓 timeout、整個監控失效。常見解法是減少 label 維度、或在 middleware 之外用 background thread 把 metrics flush 到固定 buffer。另一個常見原因是把 user_id 當 label,這會讓 metrics 維度暴增,正確做法是把 user_id 寫進 log、不放進 metrics。
log 裡忘記脫敏。把 password、token、信用卡號直接寫進 log 是常見災難。結構化 logging 雖然方便,但更要把敏感欄位過濾掉。建議在 logging config 裡加一個 filter,看到 password、token、authorization 等關鍵字就把 value 換成 `***`。這層保護要在 logging 設定的最早期就加上,後面改會很痛苦,因為你可能根本不知道哪些 log 已經把敏感資料洩漏出去。
uptime monitor 跟 API 跑在同一台機器。如果 monitor 跟 API 在同一台主機,主機掛了 monitor 也跟著掛,就失去監控意義。正式的 monitor 應該跑在不同主機或不同雲端區域。今天的腳本可以包進 docker compose、放到獨立的監控主機跑,或者用雲端服務(UptimeRobot 等)。對 side project 來說,至少要把 monitor 跑在不同的 container、用不同的網路介面進入主機,這樣主機網卡故障時還能收到通知。
沒設定 metric 保留策略。metrics 與 log 都會持續長大。Prometheus 預設保留 15 天,SQLite 沒有內建 TTL,CSV 永遠不會自動刪。長期跑下來會把磁碟塞滿,正式部署一定要設定 retention:Prometheus 用 `--storage.tsdb.retention.time=30d`、SQLite 用定期 VACUUM 與分檔、log 用 logrotate 把檔案砍掉。曾經有團隊忘記設 retention,半年後 Prometheus 磁碟滿了連帶整個監控停擺,反而錯過當下的問題。
效能與實務提醒
Prometheus 的 scrape interval 預設 15 秒,對大部分應用已經足夠;如果你的指標變化很快(例如金融交易),可以縮短到 5 秒。但 scrape 太頻繁會在 metrics 端點造成 load,要權衡。另一個常見調校是 histogram bucket 的設計:bucket 數太多會讓 Prometheus 儲存爆炸、查詢變慢,bucket 太少又會失去解析度。一般建議從 5 到 7 個 bucket 開始,依需求再加。對 SLO 監控來說,p95 與 p99 是最常被看的兩個百分位數,bucket 至少要在 p99 附近有足夠解析度。
結構化 logging 的成本主要在 JSON 序列化。每行 log 多花 50 到 200 微秒看似不多,但一個請求若有 20 條 log,整體 latency 會多 1 到 4 毫秒。對高效能 API 來說要特別注意:可以用「debug 級別不序列化、生產級別才序列化」的方式減少成本。另一個技巧是把常見的 request_id、user_id 用 contextvars 存起來,logging filter 自動帶入,這樣既不用每條 log 都手動加參數,又能避免重複傳遞。
uptime monitor 與 alert watcher 跑成常駐程式時,記得用 supervisor(systemd、supervisord、docker compose)幫忙重啟,避免程式自己 crash 後就再也沒人監控了。這個「監控自己的監控」聽起來有點哲學,但在正式環境非常重要。Day 42 會再示範怎麼把這些監控工具包進 docker compose。另外建議監控腳本本身也要有 health endpoint,用 `curl http://localhost:9100/healthz` 確認它還在跑,這是「自己監控自己」的標準做法。
最後一個實務提醒:metrics 與 logs 是事後分析的工具,但「除錯」這件事常常需要「重現問題」。所以除了 metrics 與 logs,建議在 API 加上 request_id(每個請求一個 UUID),並把 request_id 同時寫進 log 與 response header。這樣當使用者回報問題時,工程師只要拿 request_id 就能在 log 系統裡查到完整軌跡,這比「我剛剛那個請求是哪個」省下大量時間。
小結
今天我們建立了完整的健康檢查與監控體系:liveness 與 readiness 兩種 health endpoint、Prometheus metrics 與 histogram、結構化 JSON logging、以及兩支 Python 監控腳本(uptime_monitor 與 alert_watcher)。重點觀念包括:liveness 與 readiness 的差別、metrics/log 的分工、以及沒有 Prometheus 時的替代方案。我們也說明了為什麼監控是上線的最後一步、為什麼 metrics 與 log 要分開設計、以及 request_id 在除錯流程中的關鍵角色。文末也提供了雲端服務與輕量自架選項。明天我們會進入貫穿專案的 Day 35,正式開始做「預約管理系統」的資料模型設計。
結語
健康檢查與監控是上線章節的收尾,也是貫穿專案的基礎。明天開始的「預約管理系統」會用到今天所有設計:liveness、readiness、metrics、log 都會接進專案的 main 程式,並在後面 Day 41 一鍵部署時一起被驗證。今天先把工具準備好,明天,我們會從資料模型開始,把這個虛構的服務從零搭建起來。今天介紹的所有 Python 程式(health endpoints、metrics middleware、JSON formatter、uptime monitor、alert watcher)都會沿用到專案篇,後續會依需求擴充與調整。從明天起所有的範例都會以「預約管理系統」為情境,所有設計決策都會先考慮到「在 production 能不能跑、能不能監控、能不能除錯」。
延伸資源
- Prometheus Python client:
https://github.com/prometheus/client_python - Kubernetes liveness 與 readiness 探針說明:
https://kubernetes.io/docs/concepts/configuration/liveness-readiness-startup-probes/ - 結構化 logging 12-factor app 原則:
https://12factor.net/logs - OpenTelemetry Python SDK:
https://opentelemetry.io/docs/languages/python/ - Uptime Kuma 自架監控工具:
https://github.com/louislam/uptime-kuma
留言
張貼留言