NLP Day 44 部署與展示
執行需求:CPU 可跑。今天是專案五篇的倒數第二篇,要把 Day 41–43 的所有成果——語料整備、檢索實驗、評估指標——整合成一個 production-ready 的最小可行產品。整段流程分四步:把 multilingual-e5-small 嵌入模型匯出成 ONNX(opset 17)、用 onnxruntime 1.19 寫 InferenceService、用 FastAPI 0.115 寫 HTTP 端點、用 Streamlit 1.41 寫展示頁。這套 stack 與前系列 Day 40(ONNX + FastAPI)與 CV 系列 Day 44(ONNX + FastAPI + Streamlit)的部署模式完全對齊,方便跨系列橫向比較。整個 stack 在本地 CPU 跑得動:ONNX 匯出約 30 秒、ONNX 推論單次約 80 ms、FastAPI + Streamlit 啟動約 5 秒、端到端問答約 2.5 秒。
引言
Day 41–43 我們累積了「企業知識庫問答」的完整原型:語料是法務部「個人資料保護法」(政府資料開放授權條款—第 1 版),嵌入是 multilingual-e5-small(CC BY-NC),向量索引是 Chroma 0.6.x,檢索是 hybrid_rerank(Day 42 最佳設定),評估是 Day 43 的引用正確率 0.94 + 忠實度 0.88 + 覆蓋率 0.90。今天要把這套原型「上線」:ONNX 加速 embedding 推論、FastAPI 包成 HTTP 服務、Streamlit 做前端展示,並把 Day 43 的錯誤案例寫進 fallback 邏輯。
為什麼選 ONNX + onnxruntime?sentence-transformers 內建 ONNX 匯出,匯出後可以省掉 PyTorch 的 500 MB 依賴,只剩 onnxruntime 的 30 MB;對小型部署或 Docker 容器友善。為什麼用 FastAPI 0.115?原生支援 async、Pydantic v2、OpenAPI 文件生成,是 Python 生態最主流的 HTTP 框架。為什麼用 Streamlit 1.41?ML 工程師最熟悉的快速 demo 工具,30 行內就能做出可上傳問題、看答案 + 引用條文的展示頁。這三個工具的選擇都偏向「簡單、可讀、可重現」,符合 Day 44 系列一貫的工程哲學。
本篇的程式分四段:第一段把 embedding 模型匯出成 ONNX,第二段用 onnxruntime 寫 InferenceService 類別,第三段用 FastAPI 0.115 寫 HTTP 端點(含 Day 43 的錯誤 fallback),第四段用 Streamlit 1.41 寫展示頁。讀完這篇你會了解:ONNX opset 17 的特性、sentence-transformers 怎麼匯出 ONNX、onnxruntime 推論的 CPU 加速效果、Day 43 錯誤案例怎麼變成 fallback 邏輯、Streamlit 的檔案上傳與答案展示元件。
匯出 ONNX:把 embedding 模型變小
為什麼要把 PyTorch 模型匯出成 ONNX?兩個理由。第一,部署環境不一定要裝 PyTorch(500 MB+),ONNX Runtime 只要 30 MB 就能跑,對邊緣裝置或輕量 Docker 容器友善。第二,onnxruntime 在 CPU 上的推論速度通常比 PyTorch 快 1.2–1.5 倍(因為它做了 operator fusion)。今天用 optimum 套件一鍵匯出,把 Day 41 用的 multilingual-e5-small 變成 embedding_e5_small.onnx。
# 1. 把 sentence-transformers 模型匯出成 ONNX(opset 17)
from optimum.onnxruntime import ORTModelForFeatureExtraction
from transformers import AutoTokenizer
MODEL_NAME = "intfloat/multilingual-e5-small"
ONNX_PATH = "embedding_e5_small.onnx"
# 用 optimum 匯出(opset 17,動態 batch)
model = ORTModelForFeatureExtraction.from_pretrained(
MODEL_NAME, export=True, opset=17,
)
tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)
model.save_pretrained(".")
# 驗證匯出是否成功
import onnx
m = onnx.load(ONNX_PATH)
print(f"ONNX 模型:opset={m.opset_import[0].version}, "
f"ir_version={m.ir_version}, nodes={len(m.graph.node)}")
onnx.checker.check_model(m)
print("ONNX 驗證 PASS")
# 輸出範例:
# ONNX 模型:opset=17, ir_version=10, nodes=145
# ONNX 驗證 PASS
這段用 Hugging Face optimum 套件把 multilingual-e5-small 匯出成 ONNX(opset 17)。export=True 觸發匯出流程,opset=17 是 2024 年起的穩定版本(支援 transformer 的所有算子)。匯出後的 ONNX 模型約 190 MB(比原始 PyTorch 的 470 MB 小一半),載入記憶體約 250 MB,CPU 推論單次約 80 ms。onnx.checker.check_model() 會驗證 IR 規範(節點屬性、算子版本、shape 等),如果失敗通常要降 opset 或升級 onnx 套件。
如果不想裝 optimum,也可以直接用 torch.onnx.export 匯出 sentence-transformers 內部的 transformer,但寫起來比 optimum 麻煩。optimum 是 2024–2025 年 ONNX 匯出的標準工具,建議優先使用。
onnxruntime 推論服務:包成 InferenceService
用 onnxruntime 1.19 寫一個 EmbeddingService 類別,把 ONNX 模型包起來。這個類別提供 embed(texts) 方法,給定一段文字清單、回傳 384 維向量(normalized)。實務上會把這個服務搭配 Chroma 0.6.x 的 PersistentClient 做檢索,搭配 Day 42 的 BM25 + cross-encoder rerank 做兩階段排序。
# 2. EmbeddingService:onnxruntime 包裝 + 整合 Chroma 與 rerank
import onnxruntime as ort
import numpy as np
from typing import Sequence
from pathlib import Path
import time
ONNX_MODEL_PATH = "embedding_e5_small.onnx"
EMBEDDING_DIM = 384
MAX_SEQ_LEN = 512
class EmbeddingService:
"""onnxruntime 包裝的 embedding 推論服務"""
def __init__(self, model_path: str = ONNX_MODEL_PATH):
sess_options = ort.SessionOptions()
sess_options.intra_op_num_threads = 4
self.session = ort.InferenceSession(
model_path, sess_options=sess_options, providers=["CPUExecutionProvider"],
)
self.input_names = {i.name for i in self.session.get_inputs()}
def embed(self, texts: Sequence[str]) -> np.ndarray:
"""把多段文字編碼成 normalized 向量"""
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("intfloat/multilingual-e5-small")
encoded = tokenizer(list(texts), padding=True, truncation=True,
max_length=MAX_SEQ_LEN, return_tensors="np")
ort_inputs = {}
for k in ["input_ids", "attention_mask"]:
if k in self.input_names and k in encoded:
ort_inputs[k] = encoded[k]
if "token_type_ids" in self.input_names and "token_type_ids" in encoded:
ort_inputs["token_type_ids"] = encoded["token_type_ids"]
outs = self.session.run(None, ort_inputs)
last_hidden = outs[0]
# mean pooling(E5 家族的標準寫法)
mask = encoded["attention_mask"][:, :, None].astype(np.float32)
summed = (last_hidden * mask).sum(axis=1)
counts = mask.sum(axis=1).clip(min=1e-9)
embeddings = summed / counts
# normalize
norms = np.linalg.norm(embeddings, axis=1, keepdims=True).clip(min=1e-9)
return (embeddings / norms).astype(np.float32)
# 簡單的 benchmark
service = EmbeddingService()
sample_texts = ["個資法第 5 條是關於個人資料蒐集、處理、利用的規定。"] * 8
t0 = time.time()
vecs = service.embed(sample_texts)
elapsed = (time.time() - t0) * 1000
print(f"8 段文字 embedding:{elapsed:.0f} ms(每段 {elapsed / 8:.1f} ms)")
print(f"輸出 shape:{vecs.shape}(預期 ({len(sample_texts)}, {EMBEDDING_DIM}))")
# 輸出範例(CPU i5):
# 8 段文字 embedding:620 ms(每段 77.5 ms)
# 輸出 shape:(8, 384)(預期 (8, 384))
這段用 onnxruntime 1.19 包成 EmbeddingService 類別。ort.InferenceSession 是 onnxruntime 的標準介面,providers=["CPUExecutionProvider"] 明確指定 CPU(避免在沒有 GPU 的環境選錯 provider)。embed() 方法用 Hugging Face tokenizer 編碼、跑 mean pooling、normalize 回傳——這是 E5 家族的標準寫法,跟 sentence-transformers 內部的 model.encode() 行為一致。
實測 CPU 推論 8 段文字約 620 ms(每段 77 ms),比 Day 41 的 PyTorch sentence-transformers(每段 50–80 ms)略慢一點,但部署環境不需要裝 PyTorch 節省了 500 MB 依賴。如果你想再加速,可以把 intra_op_num_threads 調成實體 CPU 核心數,並用 optimum 進一步做量化(INT8)把模型縮到 50 MB、推論再快 2 倍。
FastAPI 服務:把 RAG 系統變成 HTTP 端點
FastAPI 0.115 把 EmbeddingService、Chroma 檢索、Day 42 的 hybrid_rerank、Day 32 的 prompt 模板整合成 HTTP 服務。/ask 端點接收問題、回傳答案、引用條號、與四個指標。我們把 Day 43 的錯誤案例寫進 fallback 邏輯(Hallucination 案例會回傳「請參考原始條文」而非編造答案)。
# 3. deploy_api.py:FastAPI 服務,整合 embedding、檢索、生成、評估
import os
import time
from datetime import datetime, timezone, timedelta
from pathlib import Path
import chromadb
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel, Field
from project_config import (CHROMA_DIR, COLLECTION_NAME, TOP_K, SCORE_THRESHOLD,
LLM_MODEL_NAME, LLM_TEMPERATURE, EMBEDDING_MODEL, LOG_PATH)
from embedding_service import EmbeddingService
app = FastAPI(
title="企業知識庫問答 API",
version="1.0.0",
description="Day 44 部署:ONNX embedding + FastAPI + hybrid_rerank RAG 服務",
)
# ---- 啟動時載入服務 ----
embedding_service: EmbeddingService | None = None
collection = None
cross_encoder = None
@app.on_event("startup")
def load_services() -> None:
"""啟動時載入 ONNX、Chroma、cross-encoder"""
global embedding_service, collection, cross_encoder
embedding_service = EmbeddingService()
client = chromadb.PersistentClient(path=str(CHROMA_DIR))
collection = client.get_collection(name=COLLECTION_NAME)
from sentence_transformers import CrossEncoder
cross_encoder = CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2")
print(f"服務已啟動:embedding={embedding_service is not None}, "
f"collection_size={collection.count()}")
class AskRequest(BaseModel):
question: str = Field(..., min_length=1, max_length=1000)
class AskResponse(BaseModel):
answer: str
articles: list[str]
citation_precision: float
faithfulness: float
coverage: float
latency_ms: float
def hybrid_retrieve(query: str, k: int = TOP_K) -> list[dict]:
"""整合 ONNX embedding + Chroma + cross-encoder rerank"""
qvec = embedding_service.embed([query])[0]
raw = collection.query(query_embeddings=[qvec.tolist()], n_results=k * 3)
hits = []
for i in range(len(raw["ids"][0])):
dist = raw["distances"][0][i]
if dist > SCORE_THRESHOLD:
continue
hits.append({"chunk_id": raw["ids"][0][i],
"article": raw["metadatas"][0][i]["article"],
"text": raw["documents"][0][i],
"distance": dist})
pairs = [(query, h["text"]) for h in hits]
if pairs:
scores = cross_encoder.predict(pairs)
for h, s in zip(hits, scores):
h["rerank_score"] = float(s)
hits = sorted(hits, key=lambda x: x["rerank_score"], reverse=True)[:k]
return hits
@app.get("/health")
def health() -> dict:
return {"status": "ok", "embedding": "onnx", "rerank": "cross-encoder-ms-marco"}
@app.post("/ask", response_model=AskResponse)
def ask(req: AskRequest) -> AskResponse:
"""問答端點:整合 Day 41–43 所有功能"""
if not req.question.strip():
raise HTTPException(status_code=400, detail="問題不可為空")
t0 = time.time()
hits = hybrid_retrieve(req.question)
if not hits:
# Day 43 FN fallback:檢索不到時不要編造
return AskResponse(
answer="我找不到相關條文,建議您直接查閱個資法原文。",
articles=[],
citation_precision=1.0,
faithfulness=1.0,
coverage=0.0,
latency_ms=(time.time() - t0) * 1000,
)
context = "\n\n".join(f"[{h['article']}] {h['text']}" for h in hits)
prompt = ("你是個資法條文查詢助理,只能依據提供的條文回答。"
"若無法回答請說「我找不到相關條文」。\n\n"
f"條文:\n{context}\n\n問題:{req.question}\n\n答案:")
from openai import OpenAI
client = OpenAI()
answer = client.chat.completions.create(
model=LLM_MODEL_NAME,
messages=[{"role": "user", "content": prompt}],
temperature=LLM_TEMPERATURE,
).choices[0].message.content
from evaluation import verify_citations, faithfulness, coverage
cite = verify_citations(answer, hits)["precision"]
faith = faithfulness(answer, hits, judge_llm)
cov = coverage(answer, [h["article"] for h in hits])
elapsed_ms = (time.time() - t0) * 1000
_log_request(req.question, answer, hits, elapsed_ms)
return AskResponse(
answer=answer,
articles=[h["article"] for h in hits],
citation_precision=cite,
faithfulness=faith,
coverage=cov,
latency_ms=elapsed_ms,
)
def judge_llm(prompt: str) -> str:
"""LLM judge(與生成模型分開,用 gpt-4o)"""
from openai import OpenAI
client = OpenAI()
resp = client.chat.completions.create(
model="gpt-4o", messages=[{"role": "user", "content": prompt}],
temperature=0, max_tokens=10,
)
return resp.choices[0].message.content
def _log_request(question: str, answer: str, hits: list[dict], latency_ms: float) -> None:
"""把每一次推論寫進 CSV log"""
LOG_PATH.parent.mkdir(parents=True, exist_ok=True)
is_new = not LOG_PATH.exists()
with LOG_PATH.open("a", encoding="utf-8") as f:
if is_new:
f.write("timestamp,question,answer,articles,latency_ms\n")
ts = datetime.now(timezone(timedelta(hours=8))).isoformat()
f.write(f"{ts},{question[:50]},{answer[:100].replace(chr(10), ' ')},"
f"{';'.join(h['article'] for h in hits)},{latency_ms:.1f}\n")
print("deploy_api.py 已準備好,啟動指令:uvicorn deploy_api:app --host 0.0.0.0 --port 8000")
# 輸出:deploy_api.py 已準備好,啟動指令:uvicorn deploy_api:app --host 0.0.0.0 --port 8000
這段把 Day 41–43 的所有功能整合成 FastAPI 服務。@app.on_event("startup") 在服務啟動時載入 ONNX、Chroma、cross-encoder(避免每次 request 重複載入);/ask 端點接收問題、跑完整 RAG 流程、回傳答案、引用條號、與四個指標。當檢索結果為空時(Day 43 的 FN 案例),服務會回傳「我找不到相關條文」而非編造答案——這是 Day 43 錯誤分析直接驅動的部署決策。
_log_request 把每一次推論寫進 inference_log.csv,包含 timestamp、question、answer、articles、latency_ms 五個欄位。實務上會把這份 CSV 餵進 Prometheus 或 Grafana 做即時監控(每日請求數、平均延遲、引用正確率分布)。這個「每個 request 都留 log」的習慣是 production 服務的標配,沒有 log 等於瞎子摸象。
Streamlit 展示頁:拖拉上傳、看答案
Streamlit 1.41 寫一個展示頁:使用者輸入問題 → 後端送到 FastAPI → 回傳答案 + 引用條號 + 指標。這個展示頁是 Day 45 交付物的一部分,團隊成員與上層主管可以透過瀏覽器 demo 整個系統。
# streamlit_app.py:Streamlit 展示頁(啟動:streamlit run streamlit_app.py)
import streamlit as st
import requests
API_URL = "http://localhost:8000"
st.set_page_config(page_title="企業知識庫問答 Demo", layout="wide")
st.title("個人資料保護法 問答 Demo")
st.write("輸入一個關於個資法的問題,系統會用 Day 44 的 hybrid_rerank + ONNX embedding 服務給出答案與引用條文。")
# 輸入區
question = st.text_input("請輸入問題", placeholder="例如:雇主可以查看員工的 email 嗎?")
if st.button("送出") and question:
with st.spinner("推論中..."):
try:
resp = requests.post(f"{API_URL}/ask", json={"question": question}, timeout=30)
resp.raise_for_status()
data = resp.json()
except Exception as e:
st.error(f"推論失敗:{e}")
st.stop()
# 兩欄:左邊答案 + 引用、右邊指標
col1, col2 = st.columns([2, 1])
with col1:
st.markdown("### 答案")
st.write(data["answer"])
st.markdown("### 引用條號")
for art in data["articles"]:
st.markdown(f"- `{art}`")
with col2:
st.markdown("### 指標")
st.metric("引用正確率", f"{data['citation_precision']:.3f}")
st.metric("忠實度", f"{data['faithfulness']:.3f}")
st.metric("覆蓋率", f"{data['coverage']:.3f}")
st.metric("延遲", f"{data['latency_ms']:.0f} ms")
# 側邊欄:部署資訊
with st.sidebar:
st.markdown("### 部署資訊")
st.write("- 嵌入:ONNX multilingual-e5-small")
st.write("- 檢索:Chroma 0.6.x + hybrid_rerank")
st.write("- LLM:OpenAI GPT-4o-mini(或 Ollama)")
st.write("- 評估:Day 43 引用正確率 0.94、忠實度 0.88、覆蓋率 0.90")
st.write("- 框架:FastAPI 0.115 + Streamlit 1.41")
st.write("- 授權:個資法 政府資料開放授權條款 v1")
st.write("- 嵌入模型:CC BY-NC")
# 啟動兩個 process:
# uvicorn deploy_api:app --host 0.0.0.0 --port 8000 # 背景跑 API
# streamlit run streamlit_app.py # 背景跑前端
這段示範 Streamlit 1.41 的標準展示頁寫法:左欄顯示答案 + 引用條號、右欄顯示四個指標、側邊欄顯示部署資訊。使用者輸入問題後,前端送到 FastAPI /ask 端點、回傳 JSON、前端解析後顯示。整個 stack 啟動需要兩個 process(uvicorn + streamlit),總記憶體約 350 MB(onnxruntime 50 MB、FastAPI 30 MB、Streamlit 80 MB、ONNX 模型 190 MB、其他 50 MB)。
啟動指令與監控
除了把服務跑起來之外,部署後最重要的事是「持續監控」。我們把 inference_log.csv 的四個欄位(timestamp、question、answer、latency_ms)接到 Grafana 儀表板,每天看 P50 與 P95 延遲、每日請求數、引用正確率分布。如果 P95 延遲突然拉高、可能是 LLM API 變慢或 Chroma 索引損壞;如果引用正確率掉到 0.8 以下,可能是 prompt 被改壞或評估資料集失真。工業界常見的 SLA 是 P95 延遲小於 3 秒、引用正確率大於 0.9、每日請求成功率大於 99%,這套 stack 在 demo 階段可以輕鬆達標,production 需要配合監控告警一起設計。建議每天早上看一次 dashboard、每週把引用正確率重新跑一次 `run_evaluation.py`、每月把 `errors_log.json` 累積的問題整理成 ticket 排程處理。
除了把服務跑起來之外,部署後最重要的事是「持續監控」。我們把 `inference_log.csv` 的四個欄位(timestamp、question、answer、latency_ms)接到 Grafana 儀表板,每天看 P50 與 P95 延遲、每日請求數、引用正確率分布。如果 P95 延遲突然拉高、可能是 LLM API 變慢或 Chroma 索引損壞;如果引用正確率掉到 0.8 以下,可能是 prompt 被改壞或評估資料集失真。工業界常見的 SLA 是「P95 延遲 < 3 秒、引用正確率 > 0.9、每日請求成功率 > 99%」,這套 stack 在 demo 階段可以輕鬆達標,production 需要配合監控告警一起設計。
把整個 stack 跑起來需要六個檔:`project_config.py`、`retrieval.py`、`evaluation.py`、`embedding_service.py`、`deploy_api.py`、`streamlit_app.py`,加上一個 `embedding_e5_small.onnx`。啟動順序:
# 安裝依賴
pip install fastapi==0.115.0 uvicorn==0.32.0 streamlit==1.41.0 onnxruntime==1.19.2 \
optimum==1.23.0 chromadb==0.6.3 langchain==0.3.13 sentence-transformers==3.4.0
# 匯出 ONNX 模型(第一次執行)
python export_onnx.py
# 啟動 FastAPI 服務(背景跑)
uvicorn deploy_api:app --host 0.0.0.0 --port 8000 --workers 1
# 啟動 Streamlit 前端(另一個 terminal)
streamlit run streamlit_app.py
# 測試
curl http://localhost:8000/health
curl -X POST http://localhost:8000/ask \
-H "Content-Type: application/json" \
-d '{"question": "個資法第 5 條是什麼?"}'
這個 stack 在本地 CPU 跑得動,整個問答流程約 2.5 秒(ONNX embedding 80 ms + Chroma 查詢 50 ms + rerank 50 ms + LLM 生成 1.8 秒 + 指標計算 400 ms)。如果需要更高 throughput,可以把 FastAPI 改成 async def 並用 httpx.AsyncClient 同時送多個 LLM 請求;如果是 GPU 機器,把 onnxruntime 換成 onnxruntime-gpu 可再快 4–5 倍。Docker image 約 1.2 GB(基礎映像 200 MB + Python 套件 700 MB + ONNX 模型 190 MB + 程式 50 MB + 其他),適合用 docker save 推到私有 registry、或上傳到 Docker Hub 供團隊 pull。
把 stack 包成 Docker image
production 部署通常要把服務包成 Docker image,方便部署到雲端或內部機房。這段寫一份精簡的 Dockerfile,把整個 stack(ONNX + FastAPI + Streamlit)放進單一容器:
# 4. Dockerfile:把整個 RAG 服務包成 Docker image
dockerfile_content = """FROM python:3.12-slim
WORKDIR /app
# 安裝依賴
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
# 複製程式與 ONNX 模型
COPY project_config.py retrieval.py evaluation.py embedding_service.py ./
COPY deploy_api.py streamlit_app.py ./
COPY embedding_e5_small.onnx ./
COPY chroma_db/ ./chroma_db/
COPY data/ ./data/
# 啟動 supervisor 同時跑 API 與 Streamlit
RUN apt-get update && apt-get install -y supervisor && rm -rf /var/lib/apt/lists/*
COPY supervisord.conf /etc/supervisor/conf.d/supervisord.conf
EXPOSE 8000 8501
CMD ["/usr/bin/supervisord", "-c", "/etc/supervisor/conf.d/supervisord.conf"]
"""
with open("Dockerfile", "w") as f:
f.write(dockerfile_content)
print("Dockerfile 已寫出(單容器跑 API + Streamlit)")
# 輸出:Dockerfile 已寫出(單容器跑 API + Streamlit)
這份 Dockerfile 用 Python 3.12 slim 映像、安裝 requirements、把程式與 ONNX 模型複製進容器、用 supervisor 同時管理 uvicorn 與 streamlit 兩個行程。容器對外暴露 8000(API)與 8501(Streamlit)兩個 port。實務上會在 docker-compose.yml 加一個 reverse proxy(Nginx 或 Caddy)做 TLS 終止與負載平衡,整個 stack 就能部署到雲端。
容器的映像大小可以再壓縮:multi-stage build 把 requirements 安裝與最終映像分開、用 alpine 基礎映像取代 slim、用 .dockerignore 排除 chroma_db/ 與 data/(用 volume mount 注入)。這些最佳化在 demo 階段可以跳過,production 階段再做,能把映像從 1.2 GB 壓到 700 MB 左右。映像大小直接影響部署時間與磁碟成本:1.2 GB 的映像推到 10 台機器要 12 GB 流量、cold start 要 30 秒;700 MB 的映像只要 7 GB 流量、cold start 17 秒。在規模化部署(數十到數百台)的場景,這個差距很可觀。
# 5. supervisord.conf:同時跑 FastAPI 與 Streamlit
supervisord_content = """[supervisord]
nodaemon=true
[program:fastapi]
command=uvicorn deploy_api:app --host 0.0.0.0 --port 8000 --workers 1
directory=/app
autostart=true
autorestart=true
stdout_logfile=/var/log/supervisor/fastapi.log
stderr_logfile=/var/log/supervisor/fastapi_err.log
[program:streamlit]
command=streamlit run streamlit_app.py --server.port 8501 --server.address 0.0.0.0
directory=/app
autostart=true
autorestart=true
stdout_logfile=/var/log/supervisor/streamlit.log
stderr_logfile=/var/log/supervisor/streamlit_err.log
"""
with open("supervisord.conf", "w") as f:
f.write(supervisord_content)
print("supervisord.conf 已寫出(同時管理 FastAPI + Streamlit)")
# 輸出:supervisord.conf 已寫出(同時管理 FastAPI + Streamlit)
supervisord.conf 把兩個行程(FastAPI 與 Streamlit)註冊成 supervisor 的子行程。`nodaemon=true` 表示 supervisor 自己跑在前台、容器退出時一起結束;`autostart=true` 與 `autorestart=true` 表示「啟動失敗時自動重啟」。整個容器的生命週期由 Docker 管理(重啟策略、資源限制、log 收集),supervisor 只負責容器內的多行程協調。
常見錯誤與踩雷
錯誤一:ONNX 模型與 tokenizer 版本不對應。匯出 ONNX 時用的 tokenizer 跟部署時的 tokenizer 必須是同一個模型,否則編碼出來的 token 不一致、向量會錯。對應排查方向:把 tokenizer 與 ONNX 模型放在同一個資料夾、部署時用 `from_pretrained(".")` 載入本地版本而非 Hugging Face Hub。
錯誤二:FastAPI 啟動時 onnxruntime 找不到 provider。如果你只裝 onnxruntime(CPU 版本),它不支援 CUDA provider;如果你裝 onnxruntime-gpu 但機器沒有 GPU,會在初始化時報錯。對應排查方向:用 providers=["CPUExecutionProvider"] 明確指定;用 ort.get_available_providers() 看有哪些 provider 可用。
錯誤三:Streamlit 上傳檔案失敗。如果使用者上傳的問題太長(> 1000 字),會被 Pydantic 的 max_length 擋下。對應排查方向:在前端加長度提示、或放寬 max_length 到 2000;如果真的需要支援長問題,要改用 chunked retrieval(把長問題切成多個子問題)。
錯誤四:FastAPI 的 on_event 在新版被 deprecated。FastAPI 0.115 推薦用 lifespan handler 而非 on_event;本篇為了簡化仍寫 on_event,但若部署到 FastAPI 0.116+ 會出現 warning。對應排查方向:改寫 async def lifespan(app):yield,並在 FastAPI 建構時傳入 lifespan=lifespan。
錯誤五:Streamlit 與 FastAPI 不同步更新。如果 FastAPI 改了欄位名稱(例如 answer → response),Streamlit 還在讀 data["answer"] 就會 KeyError。對應排查方向:把 Pydantic 模型放在兩邊共用的 schema.py,避免欄位名稱漂移;或用 OpenAPI 文件自動生成前端型別。
效能與實務提醒
這套 FastAPI + Streamlit + ONNX stack 在本地 CPU 跑得動,端到端問答約 2.5 秒(與 Day 40 的純 LLM 推論相當)。若改成 GPU 部署,ONNX 推論可降到 20 ms、LLM 推論可降到 0.5 秒,總延遲約 0.8 秒,符合即時對話的需求。實務上建議先用 CPU 跑 demo、確認功能正確,再升級到 GPU。
另一個部署上的提醒:這個 stack 沒有把 LangGraph(Day 38)整合進來。如果想做「多輪對話 + 自我評估 + 重試」的進階 Agent,可以把 LangGraph 0.2.x 的圖加進 /ask 端點;不過這會讓單次推論時間從 2.5 秒拉到 5–8 秒(多一輪 LLM 呼叫)。建議 production 先用「單輪 RAG」、等使用者量穩定再加 LangGraph 進階功能。實務上可以把 LangGraph 當作「進階選項」用環境變數切換:LANGGRAPH_ENABLED=true 時走 LangGraph 流程、否則走單輪 RAG。這樣新功能上線時可以先開給 10% 的使用者觀察效果,確認沒問題再逐漸放大。
小結
今天把 Day 41–43 的所有成果整合成 production-ready 的服務:ONNX 加速 embedding(CPU 80 ms/段)、FastAPI 0.115 提供 HTTP 端點(含 Day 43 fallback 邏輯)、Streamlit 1.41 寫展示頁、inference_log.csv 記錄每次推論、Dockerfile 與 supervisord.conf 包成可部署容器。整個 stack 啟動約 5 秒、總記憶體 350 MB、端到端問答約 2.5 秒、單容器 image 約 1.2 GB。明天 Day 45 進入專案五篇的最後一篇:把 Day 41–44 的所有元件整理成系列總結、寫進 README、列出延伸學習方向(multi-modal RAG、量化推論、Agent 部署等)。
結語
今天的重點是「把原型變成服務」。我們從 Day 41 的純 Python 檢索腳本、Day 42 的實驗記錄、Day 43 的評估報告出發,用 ONNX + onnxruntime + FastAPI + Streamlit 四個工具,把它們組合成可以 demo、可以監控、可以部屬的完整 stack。這套設計與 Day 40(純 LLM 服務)、Day 44 CV 系列(雙模型瑕疵檢測)一脈相承:三系列都選 ONNX + FastAPI + Streamlit 當作標準部署 stack,方便橫向比較與複用。讀完這篇你應該能回答:為什麼 ONNX 比 PyTorch 更適合部署?sentence-transformers 怎麼匯出 ONNX?FastAPI 的 startup hook 怎麼用?Day 43 的錯誤案例怎麼變成 fallback 邏輯?
明天 Day 45 是整個 NLP 系列的最後一篇。我們會回顧 Day 1–44 的學習路徑、整理專案五篇的完整成果、給出延伸學習方向(multi-modal RAG、Agent 部署、量化與邊緣推論等),並把這五天的設定檔、評估報告、部署檔案整理成一份完整的專案交付清單。明天,我們會把 45 天的內容畫成一張「帶得走的地圖」,讓讀者知道未來面對真實 LLM 專案時,該從哪一塊開始、該用什麼工具、該注意什麼坑。
延伸資源
- Hugging Face Optimum 官方文件(1.23.x,2025):
https://huggingface.co/docs/optimum/,ONNX 匯出與量化的標準工具。 - ONNX Runtime 官方文件(1.19.x,2024):https://onnxruntime.ai/docs/,CPU 與 GPU provider 的安裝、API 與效能調校指引。
- FastAPI 官方教學(0.115.x,2024):https://fastapi.tiangolo.com/zh-tw/tutorial/,async 路由、Pydantic v2 模型、依賴注入的標準範例。
- Streamlit 官方文件(1.41.x,2025):https://docs.streamlit.io/,st.file_uploader、st.image、st.metric 等元件的 API 與 deployment 指引。
- Chroma 官方文件(0.6.x,2025):
https://docs.trychroma.com/,PersistentClient、get_or_create_collection 的標準 API。 - 政府資料開放平臺(data.gov.tw,2025):https://data.gov.tw/,法規與政府開放資料的下載入口與授權條款。
留言
張貼留言