Bài 6: CI/CD for LLM Apps¶
Tổng quan¶
CI/CD cho LLM app không có bước retrain, nhưng có hai thứ mà pipeline phần mềm thông thường thiếu: prompt testing và eval gate. Bài này dựng workflow GitHub Actions:
- CI: prompt linting + unit test tool/chain + eval gate (gọi pipeline từ Bài 2) chặn merge khi chất lượng tụt.
- CD: build image → push → deploy lên môi trường từ Bài 5 → smoke test → rollback được.
1. CI/CD cho LLM App khác gì¶
graph LR
subgraph "App thường"
C1[Code] --> U1[Unit / integration test] --> B1[Build] --> D1[Deploy] --> S1[Smoke test]
end
subgraph "LLM app - thêm 2 khối"
C2[Code + Prompt + Eval set] --> L2[Prompt lint] --> U2[Unit test tool/chain]
U2 --> E2[Eval gate - điểm không được tụt] --> B2[Build] --> D2[Deploy] --> S2[Smoke test thật]
end
| Thứ đáng chặn merge | Kiểm bằng |
|---|---|
| Prompt thiếu biến / sai cú pháp / version không hợp lệ | Prompt lint (mục 4) |
| Tool trả sai schema, chain vỡ | Unit test |
| Eval score tụt quá tolerance ở tổng hoặc slice quan trọng | Eval gate (Bài 2) |
| Cost/latency trung bình tăng bất thường | Guardrail metric trong eval report |
2. Góc nhìn ML để so sánh¶
Để hiểu vì sao LLM CI/CD nhẹ hơn MLOps ở một số mặt:
| Khía cạnh MLOps | LLM app tương ứng |
|---|---|
| Data validation (schema, phân bố, thiếu giá trị) | Validate golden set (case có đủ trường, slice hợp lệ); validate tài liệu nguồn trước khi ingest |
| Model validation (so metric với baseline trước khi promote) | Eval gate - so eval score với baseline |
| Training dài, tốn GPU | Không có (trừ nhánh fine-tuning phụ) → CI chạy trên runner thường, tính bằng phút |
| Continuous Training (CT): tự retrain khi data drift | Thường không cần - cải tiến = sửa prompt/RAG/eval, không phải weights |
| Model registry (weights + metadata) | Prompt registry (Bài 1) - prompt version + eval_score |
Khi nào LLM app cần nhánh giống MLOps
Chỉ khi bạn thật sự fine-tune (LoRA từ M1, Bài 3). Lúc đó thêm một pipeline tách biệt: validate data → train → eval → register model, chạy thủ công hoặc theo lịch, không nằm trên đường CI của mỗi PR.
3. GitHub Actions Nền Tảng¶
| Khái niệm | Vai trò |
|---|---|
Workflow (.github/workflows/*.yml) |
Một quy trình tự động; nhiều workflow cho CI, CD, nightly |
Trigger (on:) |
pull_request, push (branch), schedule (cron), workflow_dispatch (chạy tay) |
| Job | Nhóm step chạy trên một runner; job có thể phụ thuộc nhau (needs:) |
| Runner | Máy chạy job - ubuntu-latest (GitHub-hosted) hoặc self-hosted (khi cần GPU) |
| Secrets | ${{ secrets.OPENAI_API_KEY }} - lưu ở repo/environment settings, không in ra log |
Cache (actions/cache) |
Cache ~/.cache/pip, .eval_cache → job nhanh, rẻ |
| Environment | production với protection rule (required reviewer) trước bước deploy |
# .github/workflows/ci.yml - khung CI
name: ci
on:
pull_request:
branches: [main]
concurrency: # PR push liên tục -> huỷ run cũ
group: ci-${{ github.ref }}
cancel-in-progress: true
jobs:
lint-and-test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with: { python-version: "3.12" }
- uses: actions/cache@v4
with:
path: ~/.cache/pip
key: pip-${{ hashFiles('requirements.txt') }}
- run: pip install -r requirements.txt -r requirements-dev.txt
- run: python -m prompts.lint prompts/ # mục 4
- run: ruff check . && ruff format --check .
- run: pytest tests/unit -q
4. CI Patterns cho LLM¶
Prompt linting¶
# prompts/lint.py
import sys, glob, yaml
from string import Template
REQUIRED_META = {"name", "version", "model", "owner", "changelog"}
MAX_TEMPLATE_CHARS = 12000
def lint_file(path: str) -> list[str]:
errs = []
doc = yaml.safe_load(open(path, encoding="utf-8"))
missing = REQUIRED_META - doc.keys()
if missing:
errs.append(f"{path}: thiếu metadata {missing}")
if not isinstance(doc.get("version"), int) or doc["version"] < 1:
errs.append(f"{path}: version phải là số nguyên >= 1")
tpl = doc.get("template", "")
if len(tpl) > MAX_TEMPLATE_CHARS:
errs.append(f"{path}: template quá dài ({len(tpl)} ký tự)")
# biến trong template phải nằm trong danh sách khai báo
declared = set(doc.get("variables", []))
used = {m for m in Template(tpl).get_identifiers()} if hasattr(Template, "get_identifiers") else set()
if used - declared:
errs.append(f"{path}: biến dùng nhưng chưa khai báo: {used - declared}")
return errs
if __name__ == "__main__":
all_errs = [e for f in glob.glob(f"{sys.argv[1]}/**/*.yaml", recursive=True) for e in lint_file(f)]
for e in all_errs:
print("::error::" + e)
sys.exit(1 if all_errs else 0)
Lint bắt: thiếu metadata, version sai kiểu, template quá dài, biến dùng mà chưa khai báo (lỗi phổ biến nhất - xem Bài 1).
Unit test tool & chain¶
# tests/unit/test_tools.py
def test_search_law_returns_schema():
out = search_law.invoke({"query": "công ty TNHH"})
assert isinstance(out, list)
assert all({"article", "text", "score"} <= d.keys() for d in out)
def test_answer_chain_handles_empty_context(monkeypatch):
monkeypatch.setattr("src.rag.retrieve", lambda q: [])
resp = answer_chain("câu hỏi ngoài phạm vi")
assert "không tìm thấy" in resp.lower() # không được bịa khi rỗng
Đây là test deterministic - không gọi LLM thật (mock), chạy nhanh, luôn bật.
Eval gate trên PR¶
# .github/workflows/eval-gate.yml
name: eval-gate
on:
pull_request:
paths: ["prompts/**", "src/rag/**", "src/agent/**", "eval/**"]
jobs:
eval:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with: { python-version: "3.12" }
- uses: actions/cache@v4
with: { path: .eval_cache, key: eval-cache-${{ github.run_id }}, restore-keys: eval-cache- }
- run: pip install -r requirements.txt
- name: Eval subset
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
EVAL_CACHE_DIR: .eval_cache
run: python -m eval.run --dataset eval/datasets/legal_qa/v3.yaml --subset --report eval_report.md --json eval_report.json
- name: Gate
run: python -m eval.gate --run eval_report.json # exit 1 => check fail => chặn merge
- name: Comment on PR
if: always()
uses: actions/github-script@v7
with:
script: |
const fs = require('fs');
github.rest.issues.createComment({
...context.repo, issue_number: context.issue.number,
body: fs.readFileSync('eval_report.md','utf8'),
});
Bật branch protection: require check eval-gate / eval pass trước khi merge vào main.
Kiểm soát chi phí CI¶
| Kỹ thuật | Chi tiết |
|---|---|
| Subset thay full | 20–40 case trên PR; full 100–300 chạy nightly |
| Cache LLM response | Key (rendered_prompt, model, params) → chạy lại PR không tính tiền |
| Judge model rẻ trên PR | gpt-4o-mini cho subset; model mạnh cho nightly |
paths: filter |
Bỏ eval nếu PR chỉ đụng docs/CSS |
concurrency: cancel-in-progress |
Push liên tục huỷ run cũ |
5. CD Patterns¶
graph LR
MERGE[Merge vào main] --> BUILD[Build image + tag = git SHA]
BUILD --> PUSH[Push -> Artifact Registry]
PUSH --> DEPLOY[Deploy: VM pull tag mới + restart]
DEPLOY --> SMOKE[Smoke test: /health + 2 câu hỏi thật]
SMOKE -->|pass| DONE[Xong - lưu tag là last-known-good]
SMOKE -->|fail| RB[Rollback: pull last-known-good + restart]
Workflow CD¶
# .github/workflows/cd.yml
name: cd
on:
push:
branches: [main]
jobs:
deploy:
runs-on: ubuntu-latest
environment: production # có thể yêu cầu reviewer duyệt
steps:
- uses: actions/checkout@v4
- id: auth
uses: google-github-actions/auth@v2
with:
workload_identity_provider: ${{ secrets.WIF_PROVIDER }} # OIDC, không cần key file
service_account: deployer@PROJECT.iam.gserviceaccount.com
- uses: google-github-actions/setup-gcloud@v2
- name: Build & push
run: |
IMG=asia-southeast1-docker.pkg.dev/PROJECT/repo/llm-app
gcloud auth configure-docker asia-southeast1-docker.pkg.dev -q
docker build -t $IMG:${{ github.sha }} -t $IMG:latest .
docker push $IMG:${{ github.sha }}
docker push $IMG:latest
- name: Deploy to VM
run: |
gcloud compute ssh llm-app --zone=asia-southeast1-a --command '
docker pull asia-southeast1-docker.pkg.dev/PROJECT/repo/llm-app:latest &&
sudo systemctl restart llm-app'
- name: Smoke test
run: python -m deploy.smoke --url https://chat.example.com
- name: Rollback on failure
if: failure()
run: |
gcloud compute ssh llm-app --zone=asia-southeast1-a --command '
LKG=$(cat /var/run/llm-app.lkg) &&
docker pull asia-southeast1-docker.pkg.dev/PROJECT/repo/llm-app:$LKG &&
docker tag asia-southeast1-docker.pkg.dev/PROJECT/repo/llm-app:$LKG \
asia-southeast1-docker.pkg.dev/PROJECT/repo/llm-app:latest &&
sudo systemctl restart llm-app'
Smoke test¶
# deploy/smoke.py
import sys, httpx
def main(url: str):
assert httpx.get(f"{url}/health", timeout=10).json()["status"] == "ok"
r = httpx.post(f"{url}/chat", json={"question": "Công ty TNHH có tối đa bao nhiêu thành viên?"},
timeout=60)
assert r.status_code == 200
body = r.text
assert "50" in body, f"Câu trả lời không như mong đợi: {body[:200]}"
print("smoke OK")
if __name__ == "__main__":
try:
main(sys.argv[sys.argv.index("--url") + 1])
except Exception as e:
print("::error::smoke failed:", e); sys.exit(1)
Rollback - hai trục¶
| Đổi gì gây hồi quy | Rollback |
|---|---|
| Code / dependency | Deploy lại image tag last-known-good (như trên) |
| Prompt | Trỏ alias production về prompt version cũ (Bài 1) - không cần build lại nếu prompt load động; nếu git-based thì revert PR production.txt |
| Model / provider | Đổi config model về giá trị cũ (nên là biến môi trường / secret, không hardcode) |
Ghi lại last-known-good sau mỗi smoke test pass
echo ${{ github.sha }} | sudo tee /var/run/llm-app.lkg. Rollback chỉ có nghĩa khi bạn biết chính xác quay về đâu.
6. Hands-on: Workflow eval gate chặn deploy¶
Dùng eval pipeline từ Bài 2, prompt registry từ Bài 1, môi trường deploy từ Bài 5.
Bước 1 - CI: lint + unit test¶
prompts/lint.py+ step gọi nó trongci.yml.- 3–5 unit test cho tool/chain (mock LLM), chạy
pytest tests/unit. - Bật branch protection: require check
cipass.
Bước 2 - CI: eval gate¶
eval-gate.ymlchạyeval.run --subset+eval.gatetrên PR đụngprompts/**hoặcsrc/rag/**.- Cache
.eval_cache. Judge dùnggpt-4o-mini. - Comment
eval_report.mdlên PR. - Require check
eval-gate / eval.
Bước 3 - Chứng minh gate hoạt động¶
- Mở PR sửa
rag_answerprompt bỏ ràng buộc injection → eval sliceinjectiontụt → gateexit 1→ không merge được. Chụp lại PR check đỏ + comment report. - Sửa lại prompt cho đúng → gate xanh → merge được.
Bước 4 - CD: deploy + smoke + rollback¶
cd.yml: build image taggithub.sha, push, SSH vào VMdocker pull+systemctl restart.deploy/smoke.pygọi/health+ 1 câu hỏi có đáp án biết trước.- Bước
if: failure()rollback về.lkg. - Ghi
.lkgsau smoke pass.
Bước 5 - Diễn tập rollback¶
- Cố tình deploy một image lỗi (ví dụ sửa
/chattrả 500) → smoke fail → workflow tự rollback → xác nhậnhttps://.../healthvẫn ok và câu hỏi vẫn trả lời (bản cũ).
Tiêu chí hoàn thành¶
- PR không merge được khi eval subset tụt quá tolerance (có ảnh chụp)
- Eval report được comment tự động lên PR
- CI chạy < ~4 phút, không đốt tiền lặp lại nhờ cache
- Merge vào
main→ tự build + deploy + smoke test - Smoke fail → tự rollback về last-known-good, dịch vụ không gián đoạn
- Biết cách rollback prompt (trỏ alias) tách khỏi rollback code (image tag)
Tóm tắt¶
graph LR
PR[Pull Request] --> LINT[Prompt lint + ruff]
LINT --> UT[Unit test tool/chain - mock LLM]
UT --> EG{Eval gate<br/>subset, so baseline}
EG -->|tụt| BLOCK[Chặn merge + comment report]
EG -->|ổn| MERGE[Merge main]
MERGE --> CD[Build tag=SHA -> push -> deploy -> smoke]
CD -->|smoke pass| LKG[Ghi last-known-good]
CD -->|smoke fail| RB[Rollback image / prompt / model]
| Chủ đề | Cốt lõi |
|---|---|
| Khác MLOps | Không retrain; CI chạy phút không phải giờ; thêm prompt lint + eval gate |
| GitHub Actions | trigger · job/needs · runner · secrets · cache · environment protection |
| Prompt lint | Metadata đủ, version đúng kiểu, biến khai báo khớp template |
| Eval gate | Subset trên PR + branch protection; cache response; judge rẻ; paths: filter |
| CD | Tag = git SHA; deploy → smoke test thật → rollback tự động khi fail |
| Rollback | 3 trục độc lập: image tag (code) · prompt version alias · model config |