Bỏ qua

Bài 4: Containerisation & LLM Serving

Tổng quan

Module I, Bài 2 đã dạy chạy model local với Ollama và quantization (GGUF, AWQ...). Bài này chuyển sang serving cho production:

  • Đóng gói LLM app bằng Docker đúng cách (multi-stage, không nhét weights vào image, GPU support).
  • vLLM / TGI cho throughput cao: PagedAttention và continuous batching giải quyết bài toán gì.
  • Chọn serving stack: Ollama vs vLLM/TGI vs API provider.
  • FastAPI patterns cho LLM: streaming SSE, async concurrency, xử lý rate limit.

1. Docker cho LLM App

Nguyên tắc

Nguyên tắc Vì sao
Multi-stage build Tách stage build (compiler, dev headers) khỏi stage runtime → image nhỏ, ít bề mặt tấn công
Model weights KHÔNG nằm trong image Weights hàng GB làm image phình, build chậm, push/pull tốn; version weights tách khỏi version code
Pin mọi version python:3.12.4-slim, hash trong requirements.txt → build reproducible
Layer theo tần suất đổi COPY requirements.txt + pip install trước COPY . . → sửa code không phá cache layer cài dependency
Chạy bằng non-root user Giảm rủi ro nếu container bị chiếm
.dockerignore Loại .git, .venv, data/, *.gguf, notebook → context build nhẹ

Dockerfile mẫu (LLM app gọi API + có RAG index)

# ---------- build stage ----------
FROM python:3.12.4-slim AS build
WORKDIR /app
RUN pip install --no-cache-dir uv
COPY requirements.txt .
RUN uv pip install --system --no-cache -r requirements.txt

# ---------- runtime stage ----------
FROM python:3.12.4-slim AS runtime
RUN useradd -m -u 1000 appuser
WORKDIR /app
COPY --from=build /usr/local/lib/python3.12/site-packages /usr/local/lib/python3.12/site-packages
COPY --from=build /usr/local/bin /usr/local/bin
COPY src/ ./src/
COPY prompts/ ./prompts/
# Weights / vector index KHÔNG copy vào đây - mount volume hoặc tải lúc khởi động
USER appuser
EXPOSE 8000
HEALTHCHECK --interval=30s --timeout=5s --retries=3 \
  CMD python -c "import urllib.request; urllib.request.urlopen('http://localhost:8000/health')"
CMD ["uvicorn", "src.main:app", "--host", "0.0.0.0", "--port", "8000", "--workers", "2"]

Model weights & index: đưa vào runtime thế nào

Cách Khi nào
Bind mount / named volume Self-host trên 1 máy; weights nằm trên disk host, container mount -v /data/models:/models:ro
Tải lúc khởi động (từ GCS/S3/HF Hub) Cluster, autoscale; entrypoint script download_if_missing() vào volume cache
Init container (K8s) GKE; container phụ tải weights vào emptyDir trước khi container chính chạy

GPU support

# Base image có sẵn CUDA runtime
FROM nvidia/cuda:12.4.1-runtime-ubuntu22.04
# ... cài python, torch bản CUDA, vllm ...
# Chạy với GPU (cần nvidia-container-toolkit trên host)
docker run --gpus all -p 8000:8000 \
  -v /data/models:/models:ro \
  my-llm-serve:latest

# Kiểm tra GPU nhìn thấy trong container
docker run --gpus all --rm nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi

Image GPU rất nặng

Base CUDA + torch + vLLM dễ vượt 8–10 GB. Dùng -runtime- chứ không -devel-, xoá cache pip/apt trong cùng layer, cân nhắc multi-stage kể cả cho image GPU.


2. Self-hosted Serving: vLLM & TGI

Khi bạn tự host model (không gọi API provider), throughput là bài toán chính: nhiều request đồng thời trên GPU có hạn.

Vấn đề 1: KV cache phân mảnh → PagedAttention

Mỗi token sinh ra cần lưu key/value của tất cả token trước đó (KV cache). Cách naive: cấp trước một khối bộ nhớ liên tục bằng max_length cho mỗi request → lãng phí lớn (request ngắn vẫn chiếm chỗ dài) và phân mảnh.

PagedAttention (vLLM) quản lý KV cache như phân trang bộ nhớ ảo của OS: chia thành các block nhỏ cố định, cấp phát theo nhu cầu, block không cần liên tục. Kết quả:

  • Tận dụng ~96% bộ nhớ KV (so với ~20–40% cách naive).
  • Cho phép nhiều request đồng thời hơn trên cùng GPU → throughput cao hơn nhiều lần.
  • Chia sẻ block giữa các request có chung prefix (prompt caching cấp GPU).

Vấn đề 2: batch tĩnh lãng phí → Continuous Batching

graph TB
    subgraph "Static batching"
        S1[Chờ đủ N request] --> S2[Chạy cả batch tới khi request DÀI NHẤT xong] --> S3[Request ngắn ngồi chờ vô ích]
    end
    subgraph "Continuous batching"
        C1[Request xong -> rời batch ngay] --> C2[Chèn request mới vào slot trống ngay lập tức] --> C3[GPU luôn bận]
    end

Continuous batching (còn gọi in-flight batching): lập lịch ở mức iteration (mỗi bước sinh 1 token) thay vì mức request. Request kết thúc sẽ nhường slot cho request đang chờ ngay lập tức → GPU không idle.

vLLM - chạy nhanh

# OpenAI-compatible server
python -m vllm.entrypoints.openai.api_server \
  --model Qwen/Qwen2.5-7B-Instruct \
  --dtype bfloat16 \
  --gpu-memory-utilization 0.90 \
  --max-num-seqs 128 \
  --max-model-len 8192
# App gọi vào như gọi OpenAI (chỉ đổi base_url)
from openai import OpenAI
client = OpenAI(base_url="http://vllm:8000/v1", api_key="dummy")

TGI (Text Generation Inference - HuggingFace)

Tương đương vLLM về ý tưởng (paged/flash attention, continuous batching), đóng gói Docker sẵn, tích hợp tốt hệ sinh thái HF.

docker run --gpus all -p 8080:80 \
  -v /data/models:/data \
  ghcr.io/huggingface/text-generation-inference:latest \
  --model-id Qwen/Qwen2.5-7B-Instruct --max-batch-prefill-tokens 8192

Tham số throughput vs latency

Tham số (vLLM) Tăng lên thì Đánh đổi
--max-num-seqs Nhiều request đồng thời hơn → throughput ↑ Latency mỗi request ↑ khi bão hoà
--gpu-memory-utilization Nhiều KV cache hơn → chứa nhiều seq Sát 1.0 dễ OOM khi burst
--max-model-len Cho context dài hơn Ăn KV cache, giảm số seq đồng thời
Tensor parallel (--tensor-parallel-size) Model lớn chia nhiều GPU Overhead giao tiếp giữa GPU

3. Chọn Serving Stack

graph TD
    Q{Bạn cần gì?} --> API_Q[Không muốn quản GPU,<br/>scale co giãn]
    Q --> SELF_Q[Cần self-host: privacy,<br/>model riêng, chi phí token lớn]
    Q --> DEV_Q[Chỉ dev / prototype]
    API_Q --> API[API provider<br/>OpenAI / Anthropic / Vertex]
    SELF_Q --> VLLM[vLLM / TGI<br/>trên GPU]
    DEV_Q --> OLLAMA[Ollama]
Ollama vLLM / TGI API provider
Mục đích Dev, demo, máy cá nhân Production self-host, throughput cao Production, không quản hạ tầng
Throughput đồng thời Thấp Rất cao (PagedAttention + continuous batching) Cao (provider lo)
Cần GPU? Không bắt buộc (chạy CPU/GGUF được) Có (thực tế bắt buộc) Không
Chi phí Điện + phần cứng GPU (thuê/mua) - đáng khi volume lớn Per token - đáng khi volume vừa/nhỏ
Vận hành Rất nhẹ Nặng (GPU driver, OOM, autoscale) Nhẹ nhất
Model Open-weights Open-weights Model của provider (+ một số open qua Vertex/Bedrock)
Khi nào chọn Học, thử nghiệm > ~vài triệu token/ngày, hoặc bắt buộc on-prem Mặc định để bắt đầu; volume chưa lớn

Đừng self-host quá sớm

Chi phí thật của vLLM không phải tiền GPU mà là thời gian vận hành: driver, CUDA mismatch, OOM lúc burst, autoscale GPU. Bắt đầu bằng API provider; chuyển self-host khi phép tính token cho thấy tiết kiệm đủ bù chi phí vận hành.


4. FastAPI Patterns cho LLM

Streaming qua SSE

Người dùng muốn thấy token xuất hiện dần thay vì chờ 5 giây rồi hiện cả đoạn.

# src/main.py
from fastapi import FastAPI, Request
from fastapi.responses import StreamingResponse
import json, asyncio

app = FastAPI()

@app.post("/chat")
async def chat(req: Request):
    body = await req.json()
    question = body["question"]

    async def event_stream():
        try:
            async for token in llm_stream(question):        # async generator
                yield f"data: {json.dumps({'token': token}, ensure_ascii=False)}\n\n"
                if await req.is_disconnected():             # client đóng tab -> dừng, khỏi tốn tiền
                    break
            yield "data: [DONE]\n\n"
        except Exception as e:
            yield f"data: {json.dumps({'error': str(e)})}\n\n"

    return StreamingResponse(event_stream(), media_type="text/event-stream")

@app.get("/health")
async def health():
    return {"status": "ok"}

Async concurrency & backpressure

# Giới hạn số request LLM đồng thời -> tránh làm sập upstream / vượt rate limit
LLM_SEMAPHORE = asyncio.Semaphore(20)

async def llm_stream(question: str):
    async with LLM_SEMAPHORE:                    # request thứ 21 xếp hàng chờ
        async for tok in provider.stream(question):
            yield tok
  • Dùng client async (AsyncOpenAI, httpx.AsyncClient) - đừng gọi SDK sync trong endpoint async (block event loop).
  • Nhiều uvicorn --workers cho CPU-bound (parse, embed); async lo I/O-bound (chờ LLM).

Xử lý rate limit từ provider

# src/llm_client.py
import random, asyncio
from openai import AsyncOpenAI, RateLimitError, APITimeoutError

client = AsyncOpenAI(timeout=30.0, max_retries=0)   # tự quản retry để có jitter

async def call_with_retry(**kwargs):
    for attempt in range(5):
        try:
            return await client.chat.completions.create(**kwargs)
        except (RateLimitError, APITimeoutError) as e:
            if attempt == 4:
                raise
            backoff = min(2 ** attempt, 16) + random.uniform(0, 1)   # exp backoff + jitter
            await asyncio.sleep(backoff)
Vấn đề Xử lý
429 RateLimitError Exponential backoff + jitter; tôn trọng header Retry-After nếu có
Timeout Đặt timeout rõ ràng (30–60s); huỷ khi client disconnect
Burst vượt quota Semaphore giới hạn concurrency; hàng đợi có giới hạn + trả 503 khi đầy
Provider outage Fallback model/provider (xem Bài 7)

5. Hands-on: Dockerize Vietnamese LLM app + streaming FastAPI

Dùng RAG app từ Capstone Module I.

Bước 1 - FastAPI streaming endpoint

  • /chat (POST) trả StreamingResponse SSE, có nhả [DONE], dừng khi req.is_disconnected().
  • /health cho healthcheck.
  • Client LLM async + call_with_retry (backoff + jitter).
  • Semaphore(20) giới hạn concurrency.

Bước 2 - Dockerfile

  • Multi-stage (build → runtime), non-root user, HEALTHCHECK.
  • .dockerignore loại data/, .venv, *.gguf, .git.
  • Vector index không nằm trong image - mount -v ./data/index:/app/data/index:ro.
  • Build, kiểm tra image size (mục tiêu < ~400 MB cho app gọi API).

Bước 3 - Benchmark concurrency

  • Dùng hey / wrk / script asyncio bắn 50 request đồng thời vào /chat.
  • Đo: p50 / p95 latency, throughput (req/s), thời điểm bắt đầu xếp hàng do semaphore.
  • Thử Semaphore(5) vs Semaphore(50) → quan sát đánh đổi latency/throughput/lỗi 429.

Bước 4 (mở rộng) - Serve model local bằng vLLM

  • Nếu có GPU: chạy vllm ... --model Qwen/Qwen2.5-7B-Instruct trong container GPU.
  • Trỏ app vào base_url=http://vllm:8000/v1 (docker-compose 2 service).
  • So throughput đồng thời giữa Ollama và vLLM trên cùng câu hỏi.

Tiêu chí hoàn thành

  • /chat stream token thật, dừng khi client ngắt kết nối
  • Image multi-stage, non-root, có healthcheck, không chứa weights/index
  • Retry có exponential backoff + jitter cho 429/timeout
  • Có số benchmark p50/p95/throughput ở ít nhất 2 mức concurrency
  • (Mở rộng) so sánh Ollama vs vLLM về throughput đồng thời

Tóm tắt

graph LR
    CODE[App code + prompts] --> IMG[Docker image<br/>multi-stage, non-root, no weights]
    IMG --> RUN[Container]
    W[(Weights / index)] -.mount / tải lúc khởi động.-> RUN
    RUN --> FASTAPI[FastAPI: SSE stream + async + semaphore + retry]
    FASTAPI --> BACKEND{Backend}
    BACKEND -->|self-host| VLLM[vLLM / TGI<br/>PagedAttention + continuous batching]
    BACKEND -->|managed| API[API provider]
Chủ đề Cốt lõi
Docker Multi-stage; weights ngoài image; layer dependency trước code; non-root; healthcheck
GPU image Base nvidia/cuda:*-runtime, --gpus all, nvidia-container-toolkit; image rất nặng
PagedAttention KV cache như paging → tận dụng ~96% bộ nhớ, nhiều seq đồng thời
Continuous batching Lập lịch mức iteration; request xong nhường slot ngay → GPU không idle
Chọn stack Ollama (dev) · vLLM/TGI (self-host, volume lớn) · API (mặc định để bắt đầu)
FastAPI SSE streaming; async client; semaphore backpressure; retry backoff + jitter; dừng khi disconnect