Bỏ qua

Bài 6: CI/CD for LLM Apps

Tổng quan

CI/CD cho LLM app không có bước retrain, nhưng có hai thứ mà pipeline phần mềm thông thường thiếu: prompt testing và eval gate. Bài này dựng workflow GitHub Actions:

  • CI: prompt linting + unit test tool/chain + eval gate (gọi pipeline từ Bài 2) chặn merge khi chất lượng tụt.
  • CD: build image → push → deploy lên môi trường từ Bài 5 → smoke test → rollback được.

1. CI/CD cho LLM App khác gì

graph LR
    subgraph "App thường"
        C1[Code] --> U1[Unit / integration test] --> B1[Build] --> D1[Deploy] --> S1[Smoke test]
    end
    subgraph "LLM app - thêm 2 khối"
        C2[Code + Prompt + Eval set] --> L2[Prompt lint] --> U2[Unit test tool/chain]
        U2 --> E2[Eval gate - điểm không được tụt] --> B2[Build] --> D2[Deploy] --> S2[Smoke test thật]
    end
Thứ đáng chặn merge Kiểm bằng
Prompt thiếu biến / sai cú pháp / version không hợp lệ Prompt lint (mục 4)
Tool trả sai schema, chain vỡ Unit test
Eval score tụt quá tolerance ở tổng hoặc slice quan trọng Eval gate (Bài 2)
Cost/latency trung bình tăng bất thường Guardrail metric trong eval report

2. Góc nhìn ML để so sánh

Để hiểu vì sao LLM CI/CD nhẹ hơn MLOps ở một số mặt:

Khía cạnh MLOps LLM app tương ứng
Data validation (schema, phân bố, thiếu giá trị) Validate golden set (case có đủ trường, slice hợp lệ); validate tài liệu nguồn trước khi ingest
Model validation (so metric với baseline trước khi promote) Eval gate - so eval score với baseline
Training dài, tốn GPU Không có (trừ nhánh fine-tuning phụ) → CI chạy trên runner thường, tính bằng phút
Continuous Training (CT): tự retrain khi data drift Thường không cần - cải tiến = sửa prompt/RAG/eval, không phải weights
Model registry (weights + metadata) Prompt registry (Bài 1) - prompt version + eval_score

Khi nào LLM app cần nhánh giống MLOps

Chỉ khi bạn thật sự fine-tune (LoRA từ M1, Bài 3). Lúc đó thêm một pipeline tách biệt: validate data → train → eval → register model, chạy thủ công hoặc theo lịch, không nằm trên đường CI của mỗi PR.


3. GitHub Actions Nền Tảng

Khái niệm Vai trò
Workflow (.github/workflows/*.yml) Một quy trình tự động; nhiều workflow cho CI, CD, nightly
Trigger (on:) pull_request, push (branch), schedule (cron), workflow_dispatch (chạy tay)
Job Nhóm step chạy trên một runner; job có thể phụ thuộc nhau (needs:)
Runner Máy chạy job - ubuntu-latest (GitHub-hosted) hoặc self-hosted (khi cần GPU)
Secrets ${{ secrets.OPENAI_API_KEY }} - lưu ở repo/environment settings, không in ra log
Cache (actions/cache) Cache ~/.cache/pip, .eval_cache → job nhanh, rẻ
Environment production với protection rule (required reviewer) trước bước deploy
# .github/workflows/ci.yml - khung CI
name: ci
on:
  pull_request:
    branches: [main]
concurrency:                      # PR push liên tục -> huỷ run cũ
  group: ci-${{ github.ref }}
  cancel-in-progress: true
jobs:
  lint-and-test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with: { python-version: "3.12" }
      - uses: actions/cache@v4
        with:
          path: ~/.cache/pip
          key: pip-${{ hashFiles('requirements.txt') }}
      - run: pip install -r requirements.txt -r requirements-dev.txt
      - run: python -m prompts.lint prompts/          # mục 4
      - run: ruff check . && ruff format --check .
      - run: pytest tests/unit -q

4. CI Patterns cho LLM

Prompt linting

# prompts/lint.py
import sys, glob, yaml
from string import Template

REQUIRED_META = {"name", "version", "model", "owner", "changelog"}
MAX_TEMPLATE_CHARS = 12000

def lint_file(path: str) -> list[str]:
    errs = []
    doc = yaml.safe_load(open(path, encoding="utf-8"))
    missing = REQUIRED_META - doc.keys()
    if missing:
        errs.append(f"{path}: thiếu metadata {missing}")
    if not isinstance(doc.get("version"), int) or doc["version"] < 1:
        errs.append(f"{path}: version phải là số nguyên >= 1")
    tpl = doc.get("template", "")
    if len(tpl) > MAX_TEMPLATE_CHARS:
        errs.append(f"{path}: template quá dài ({len(tpl)} ký tự)")
    # biến trong template phải nằm trong danh sách khai báo
    declared = set(doc.get("variables", []))
    used = {m for m in Template(tpl).get_identifiers()} if hasattr(Template, "get_identifiers") else set()
    if used - declared:
        errs.append(f"{path}: biến dùng nhưng chưa khai báo: {used - declared}")
    return errs

if __name__ == "__main__":
    all_errs = [e for f in glob.glob(f"{sys.argv[1]}/**/*.yaml", recursive=True) for e in lint_file(f)]
    for e in all_errs:
        print("::error::" + e)
    sys.exit(1 if all_errs else 0)

Lint bắt: thiếu metadata, version sai kiểu, template quá dài, biến dùng mà chưa khai báo (lỗi phổ biến nhất - xem Bài 1).

Unit test tool & chain

# tests/unit/test_tools.py
def test_search_law_returns_schema():
    out = search_law.invoke({"query": "công ty TNHH"})
    assert isinstance(out, list)
    assert all({"article", "text", "score"} <= d.keys() for d in out)

def test_answer_chain_handles_empty_context(monkeypatch):
    monkeypatch.setattr("src.rag.retrieve", lambda q: [])
    resp = answer_chain("câu hỏi ngoài phạm vi")
    assert "không tìm thấy" in resp.lower()      # không được bịa khi rỗng

Đây là test deterministic - không gọi LLM thật (mock), chạy nhanh, luôn bật.

Eval gate trên PR

# .github/workflows/eval-gate.yml
name: eval-gate
on:
  pull_request:
    paths: ["prompts/**", "src/rag/**", "src/agent/**", "eval/**"]
jobs:
  eval:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with: { python-version: "3.12" }
      - uses: actions/cache@v4
        with: { path: .eval_cache, key: eval-cache-${{ github.run_id }}, restore-keys: eval-cache- }
      - run: pip install -r requirements.txt
      - name: Eval subset
        env:
          OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
          EVAL_CACHE_DIR: .eval_cache
        run: python -m eval.run --dataset eval/datasets/legal_qa/v3.yaml --subset --report eval_report.md --json eval_report.json
      - name: Gate
        run: python -m eval.gate --run eval_report.json      # exit 1 => check fail => chặn merge
      - name: Comment on PR
        if: always()
        uses: actions/github-script@v7
        with:
          script: |
            const fs = require('fs');
            github.rest.issues.createComment({
              ...context.repo, issue_number: context.issue.number,
              body: fs.readFileSync('eval_report.md','utf8'),
            });

Bật branch protection: require check eval-gate / eval pass trước khi merge vào main.

Kiểm soát chi phí CI

Kỹ thuật Chi tiết
Subset thay full 20–40 case trên PR; full 100–300 chạy nightly
Cache LLM response Key (rendered_prompt, model, params) → chạy lại PR không tính tiền
Judge model rẻ trên PR gpt-4o-mini cho subset; model mạnh cho nightly
paths: filter Bỏ eval nếu PR chỉ đụng docs/CSS
concurrency: cancel-in-progress Push liên tục huỷ run cũ

5. CD Patterns

graph LR
    MERGE[Merge vào main] --> BUILD[Build image + tag = git SHA]
    BUILD --> PUSH[Push -> Artifact Registry]
    PUSH --> DEPLOY[Deploy: VM pull tag mới + restart]
    DEPLOY --> SMOKE[Smoke test: /health + 2 câu hỏi thật]
    SMOKE -->|pass| DONE[Xong - lưu tag là last-known-good]
    SMOKE -->|fail| RB[Rollback: pull last-known-good + restart]

Workflow CD

# .github/workflows/cd.yml
name: cd
on:
  push:
    branches: [main]
jobs:
  deploy:
    runs-on: ubuntu-latest
    environment: production          # có thể yêu cầu reviewer duyệt
    steps:
      - uses: actions/checkout@v4
      - id: auth
        uses: google-github-actions/auth@v2
        with:
          workload_identity_provider: ${{ secrets.WIF_PROVIDER }}   # OIDC, không cần key file
          service_account: deployer@PROJECT.iam.gserviceaccount.com
      - uses: google-github-actions/setup-gcloud@v2
      - name: Build & push
        run: |
          IMG=asia-southeast1-docker.pkg.dev/PROJECT/repo/llm-app
          gcloud auth configure-docker asia-southeast1-docker.pkg.dev -q
          docker build -t $IMG:${{ github.sha }} -t $IMG:latest .
          docker push $IMG:${{ github.sha }}
          docker push $IMG:latest
      - name: Deploy to VM
        run: |
          gcloud compute ssh llm-app --zone=asia-southeast1-a --command '
            docker pull asia-southeast1-docker.pkg.dev/PROJECT/repo/llm-app:latest &&
            sudo systemctl restart llm-app'
      - name: Smoke test
        run: python -m deploy.smoke --url https://chat.example.com
      - name: Rollback on failure
        if: failure()
        run: |
          gcloud compute ssh llm-app --zone=asia-southeast1-a --command '
            LKG=$(cat /var/run/llm-app.lkg) &&
            docker pull asia-southeast1-docker.pkg.dev/PROJECT/repo/llm-app:$LKG &&
            docker tag  asia-southeast1-docker.pkg.dev/PROJECT/repo/llm-app:$LKG \
                        asia-southeast1-docker.pkg.dev/PROJECT/repo/llm-app:latest &&
            sudo systemctl restart llm-app'

Smoke test

# deploy/smoke.py
import sys, httpx

def main(url: str):
    assert httpx.get(f"{url}/health", timeout=10).json()["status"] == "ok"
    r = httpx.post(f"{url}/chat", json={"question": "Công ty TNHH có tối đa bao nhiêu thành viên?"},
                   timeout=60)
    assert r.status_code == 200
    body = r.text
    assert "50" in body, f"Câu trả lời không như mong đợi: {body[:200]}"
    print("smoke OK")

if __name__ == "__main__":
    try:
        main(sys.argv[sys.argv.index("--url") + 1])
    except Exception as e:
        print("::error::smoke failed:", e); sys.exit(1)

Rollback - hai trục

Đổi gì gây hồi quy Rollback
Code / dependency Deploy lại image tag last-known-good (như trên)
Prompt Trỏ alias production về prompt version cũ (Bài 1) - không cần build lại nếu prompt load động; nếu git-based thì revert PR production.txt
Model / provider Đổi config model về giá trị cũ (nên là biến môi trường / secret, không hardcode)

Ghi lại last-known-good sau mỗi smoke test pass

echo ${{ github.sha }} | sudo tee /var/run/llm-app.lkg. Rollback chỉ có nghĩa khi bạn biết chính xác quay về đâu.


6. Hands-on: Workflow eval gate chặn deploy

Dùng eval pipeline từ Bài 2, prompt registry từ Bài 1, môi trường deploy từ Bài 5.

Bước 1 - CI: lint + unit test

  • prompts/lint.py + step gọi nó trong ci.yml.
  • 3–5 unit test cho tool/chain (mock LLM), chạy pytest tests/unit.
  • Bật branch protection: require check ci pass.

Bước 2 - CI: eval gate

  • eval-gate.yml chạy eval.run --subset + eval.gate trên PR đụng prompts/** hoặc src/rag/**.
  • Cache .eval_cache. Judge dùng gpt-4o-mini.
  • Comment eval_report.md lên PR.
  • Require check eval-gate / eval.

Bước 3 - Chứng minh gate hoạt động

  • Mở PR sửa rag_answer prompt bỏ ràng buộc injection → eval slice injection tụt → gate exit 1 → không merge được. Chụp lại PR check đỏ + comment report.
  • Sửa lại prompt cho đúng → gate xanh → merge được.

Bước 4 - CD: deploy + smoke + rollback

  • cd.yml: build image tag github.sha, push, SSH vào VM docker pull + systemctl restart.
  • deploy/smoke.py gọi /health + 1 câu hỏi có đáp án biết trước.
  • Bước if: failure() rollback về .lkg.
  • Ghi .lkg sau smoke pass.

Bước 5 - Diễn tập rollback

  • Cố tình deploy một image lỗi (ví dụ sửa /chat trả 500) → smoke fail → workflow tự rollback → xác nhận https://.../health vẫn ok và câu hỏi vẫn trả lời (bản cũ).

Tiêu chí hoàn thành

  • PR không merge được khi eval subset tụt quá tolerance (có ảnh chụp)
  • Eval report được comment tự động lên PR
  • CI chạy < ~4 phút, không đốt tiền lặp lại nhờ cache
  • Merge vào main → tự build + deploy + smoke test
  • Smoke fail → tự rollback về last-known-good, dịch vụ không gián đoạn
  • Biết cách rollback prompt (trỏ alias) tách khỏi rollback code (image tag)

Tóm tắt

graph LR
    PR[Pull Request] --> LINT[Prompt lint + ruff]
    LINT --> UT[Unit test tool/chain - mock LLM]
    UT --> EG{Eval gate<br/>subset, so baseline}
    EG -->|tụt| BLOCK[Chặn merge + comment report]
    EG -->|ổn| MERGE[Merge main]
    MERGE --> CD[Build tag=SHA -> push -> deploy -> smoke]
    CD -->|smoke pass| LKG[Ghi last-known-good]
    CD -->|smoke fail| RB[Rollback image / prompt / model]
Chủ đề Cốt lõi
Khác MLOps Không retrain; CI chạy phút không phải giờ; thêm prompt lint + eval gate
GitHub Actions trigger · job/needs · runner · secrets · cache · environment protection
Prompt lint Metadata đủ, version đúng kiểu, biến khai báo khớp template
Eval gate Subset trên PR + branch protection; cache response; judge rẻ; paths: filter
CD Tag = git SHA; deploy → smoke test thật → rollback tự động khi fail
Rollback 3 trục độc lập: image tag (code) · prompt version alias · model config