Bài 4: Containerisation & LLM Serving¶
Tổng quan¶
Module I, Bài 2 đã dạy chạy model local với Ollama và quantization (GGUF, AWQ...). Bài này chuyển sang serving cho production:
- Đóng gói LLM app bằng Docker đúng cách (multi-stage, không nhét weights vào image, GPU support).
- vLLM / TGI cho throughput cao: PagedAttention và continuous batching giải quyết bài toán gì.
- Chọn serving stack: Ollama vs vLLM/TGI vs API provider.
- FastAPI patterns cho LLM: streaming SSE, async concurrency, xử lý rate limit.
1. Docker cho LLM App¶
Nguyên tắc¶
| Nguyên tắc | Vì sao |
|---|---|
| Multi-stage build | Tách stage build (compiler, dev headers) khỏi stage runtime → image nhỏ, ít bề mặt tấn công |
| Model weights KHÔNG nằm trong image | Weights hàng GB làm image phình, build chậm, push/pull tốn; version weights tách khỏi version code |
| Pin mọi version | python:3.12.4-slim, hash trong requirements.txt → build reproducible |
| Layer theo tần suất đổi | COPY requirements.txt + pip install trước COPY . . → sửa code không phá cache layer cài dependency |
| Chạy bằng non-root user | Giảm rủi ro nếu container bị chiếm |
.dockerignore |
Loại .git, .venv, data/, *.gguf, notebook → context build nhẹ |
Dockerfile mẫu (LLM app gọi API + có RAG index)¶
# ---------- build stage ----------
FROM python:3.12.4-slim AS build
WORKDIR /app
RUN pip install --no-cache-dir uv
COPY requirements.txt .
RUN uv pip install --system --no-cache -r requirements.txt
# ---------- runtime stage ----------
FROM python:3.12.4-slim AS runtime
RUN useradd -m -u 1000 appuser
WORKDIR /app
COPY --from=build /usr/local/lib/python3.12/site-packages /usr/local/lib/python3.12/site-packages
COPY --from=build /usr/local/bin /usr/local/bin
COPY src/ ./src/
COPY prompts/ ./prompts/
# Weights / vector index KHÔNG copy vào đây - mount volume hoặc tải lúc khởi động
USER appuser
EXPOSE 8000
HEALTHCHECK --interval=30s --timeout=5s --retries=3 \
CMD python -c "import urllib.request; urllib.request.urlopen('http://localhost:8000/health')"
CMD ["uvicorn", "src.main:app", "--host", "0.0.0.0", "--port", "8000", "--workers", "2"]
Model weights & index: đưa vào runtime thế nào¶
| Cách | Khi nào |
|---|---|
| Bind mount / named volume | Self-host trên 1 máy; weights nằm trên disk host, container mount -v /data/models:/models:ro |
| Tải lúc khởi động (từ GCS/S3/HF Hub) | Cluster, autoscale; entrypoint script download_if_missing() vào volume cache |
| Init container (K8s) | GKE; container phụ tải weights vào emptyDir trước khi container chính chạy |
GPU support¶
# Base image có sẵn CUDA runtime
FROM nvidia/cuda:12.4.1-runtime-ubuntu22.04
# ... cài python, torch bản CUDA, vllm ...
# Chạy với GPU (cần nvidia-container-toolkit trên host)
docker run --gpus all -p 8000:8000 \
-v /data/models:/models:ro \
my-llm-serve:latest
# Kiểm tra GPU nhìn thấy trong container
docker run --gpus all --rm nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi
Image GPU rất nặng
Base CUDA + torch + vLLM dễ vượt 8–10 GB. Dùng -runtime- chứ không -devel-, xoá cache pip/apt trong cùng layer, cân nhắc multi-stage kể cả cho image GPU.
2. Self-hosted Serving: vLLM & TGI¶
Khi bạn tự host model (không gọi API provider), throughput là bài toán chính: nhiều request đồng thời trên GPU có hạn.
Vấn đề 1: KV cache phân mảnh → PagedAttention¶
Mỗi token sinh ra cần lưu key/value của tất cả token trước đó (KV cache). Cách naive: cấp trước một khối bộ nhớ liên tục bằng max_length cho mỗi request → lãng phí lớn (request ngắn vẫn chiếm chỗ dài) và phân mảnh.
PagedAttention (vLLM) quản lý KV cache như phân trang bộ nhớ ảo của OS: chia thành các block nhỏ cố định, cấp phát theo nhu cầu, block không cần liên tục. Kết quả:
- Tận dụng ~96% bộ nhớ KV (so với ~20–40% cách naive).
- Cho phép nhiều request đồng thời hơn trên cùng GPU → throughput cao hơn nhiều lần.
- Chia sẻ block giữa các request có chung prefix (prompt caching cấp GPU).
Vấn đề 2: batch tĩnh lãng phí → Continuous Batching¶
graph TB
subgraph "Static batching"
S1[Chờ đủ N request] --> S2[Chạy cả batch tới khi request DÀI NHẤT xong] --> S3[Request ngắn ngồi chờ vô ích]
end
subgraph "Continuous batching"
C1[Request xong -> rời batch ngay] --> C2[Chèn request mới vào slot trống ngay lập tức] --> C3[GPU luôn bận]
end
Continuous batching (còn gọi in-flight batching): lập lịch ở mức iteration (mỗi bước sinh 1 token) thay vì mức request. Request kết thúc sẽ nhường slot cho request đang chờ ngay lập tức → GPU không idle.
vLLM - chạy nhanh¶
# OpenAI-compatible server
python -m vllm.entrypoints.openai.api_server \
--model Qwen/Qwen2.5-7B-Instruct \
--dtype bfloat16 \
--gpu-memory-utilization 0.90 \
--max-num-seqs 128 \
--max-model-len 8192
# App gọi vào như gọi OpenAI (chỉ đổi base_url)
from openai import OpenAI
client = OpenAI(base_url="http://vllm:8000/v1", api_key="dummy")
TGI (Text Generation Inference - HuggingFace)¶
Tương đương vLLM về ý tưởng (paged/flash attention, continuous batching), đóng gói Docker sẵn, tích hợp tốt hệ sinh thái HF.
docker run --gpus all -p 8080:80 \
-v /data/models:/data \
ghcr.io/huggingface/text-generation-inference:latest \
--model-id Qwen/Qwen2.5-7B-Instruct --max-batch-prefill-tokens 8192
Tham số throughput vs latency¶
| Tham số (vLLM) | Tăng lên thì | Đánh đổi |
|---|---|---|
--max-num-seqs |
Nhiều request đồng thời hơn → throughput ↑ | Latency mỗi request ↑ khi bão hoà |
--gpu-memory-utilization |
Nhiều KV cache hơn → chứa nhiều seq | Sát 1.0 dễ OOM khi burst |
--max-model-len |
Cho context dài hơn | Ăn KV cache, giảm số seq đồng thời |
Tensor parallel (--tensor-parallel-size) |
Model lớn chia nhiều GPU | Overhead giao tiếp giữa GPU |
3. Chọn Serving Stack¶
graph TD
Q{Bạn cần gì?} --> API_Q[Không muốn quản GPU,<br/>scale co giãn]
Q --> SELF_Q[Cần self-host: privacy,<br/>model riêng, chi phí token lớn]
Q --> DEV_Q[Chỉ dev / prototype]
API_Q --> API[API provider<br/>OpenAI / Anthropic / Vertex]
SELF_Q --> VLLM[vLLM / TGI<br/>trên GPU]
DEV_Q --> OLLAMA[Ollama]
| Ollama | vLLM / TGI | API provider | |
|---|---|---|---|
| Mục đích | Dev, demo, máy cá nhân | Production self-host, throughput cao | Production, không quản hạ tầng |
| Throughput đồng thời | Thấp | Rất cao (PagedAttention + continuous batching) | Cao (provider lo) |
| Cần GPU? | Không bắt buộc (chạy CPU/GGUF được) | Có (thực tế bắt buộc) | Không |
| Chi phí | Điện + phần cứng | GPU (thuê/mua) - đáng khi volume lớn | Per token - đáng khi volume vừa/nhỏ |
| Vận hành | Rất nhẹ | Nặng (GPU driver, OOM, autoscale) | Nhẹ nhất |
| Model | Open-weights | Open-weights | Model của provider (+ một số open qua Vertex/Bedrock) |
| Khi nào chọn | Học, thử nghiệm | > ~vài triệu token/ngày, hoặc bắt buộc on-prem | Mặc định để bắt đầu; volume chưa lớn |
Đừng self-host quá sớm
Chi phí thật của vLLM không phải tiền GPU mà là thời gian vận hành: driver, CUDA mismatch, OOM lúc burst, autoscale GPU. Bắt đầu bằng API provider; chuyển self-host khi phép tính token cho thấy tiết kiệm đủ bù chi phí vận hành.
4. FastAPI Patterns cho LLM¶
Streaming qua SSE¶
Người dùng muốn thấy token xuất hiện dần thay vì chờ 5 giây rồi hiện cả đoạn.
# src/main.py
from fastapi import FastAPI, Request
from fastapi.responses import StreamingResponse
import json, asyncio
app = FastAPI()
@app.post("/chat")
async def chat(req: Request):
body = await req.json()
question = body["question"]
async def event_stream():
try:
async for token in llm_stream(question): # async generator
yield f"data: {json.dumps({'token': token}, ensure_ascii=False)}\n\n"
if await req.is_disconnected(): # client đóng tab -> dừng, khỏi tốn tiền
break
yield "data: [DONE]\n\n"
except Exception as e:
yield f"data: {json.dumps({'error': str(e)})}\n\n"
return StreamingResponse(event_stream(), media_type="text/event-stream")
@app.get("/health")
async def health():
return {"status": "ok"}
Async concurrency & backpressure¶
# Giới hạn số request LLM đồng thời -> tránh làm sập upstream / vượt rate limit
LLM_SEMAPHORE = asyncio.Semaphore(20)
async def llm_stream(question: str):
async with LLM_SEMAPHORE: # request thứ 21 xếp hàng chờ
async for tok in provider.stream(question):
yield tok
- Dùng client async (
AsyncOpenAI,httpx.AsyncClient) - đừng gọi SDK sync trong endpoint async (block event loop). - Nhiều
uvicorn --workerscho CPU-bound (parse, embed); async lo I/O-bound (chờ LLM).
Xử lý rate limit từ provider¶
# src/llm_client.py
import random, asyncio
from openai import AsyncOpenAI, RateLimitError, APITimeoutError
client = AsyncOpenAI(timeout=30.0, max_retries=0) # tự quản retry để có jitter
async def call_with_retry(**kwargs):
for attempt in range(5):
try:
return await client.chat.completions.create(**kwargs)
except (RateLimitError, APITimeoutError) as e:
if attempt == 4:
raise
backoff = min(2 ** attempt, 16) + random.uniform(0, 1) # exp backoff + jitter
await asyncio.sleep(backoff)
| Vấn đề | Xử lý |
|---|---|
429 RateLimitError |
Exponential backoff + jitter; tôn trọng header Retry-After nếu có |
| Timeout | Đặt timeout rõ ràng (30–60s); huỷ khi client disconnect |
| Burst vượt quota | Semaphore giới hạn concurrency; hàng đợi có giới hạn + trả 503 khi đầy |
| Provider outage | Fallback model/provider (xem Bài 7) |
5. Hands-on: Dockerize Vietnamese LLM app + streaming FastAPI¶
Dùng RAG app từ Capstone Module I.
Bước 1 - FastAPI streaming endpoint¶
/chat(POST) trảStreamingResponseSSE, có nhả[DONE], dừng khireq.is_disconnected()./healthcho healthcheck.- Client LLM async +
call_with_retry(backoff + jitter). Semaphore(20)giới hạn concurrency.
Bước 2 - Dockerfile¶
- Multi-stage (build → runtime), non-root user,
HEALTHCHECK. .dockerignoreloạidata/,.venv,*.gguf,.git.- Vector index không nằm trong image - mount
-v ./data/index:/app/data/index:ro. - Build, kiểm tra image size (mục tiêu < ~400 MB cho app gọi API).
Bước 3 - Benchmark concurrency¶
- Dùng
hey/wrk/ scriptasynciobắn 50 request đồng thời vào/chat. - Đo: p50 / p95 latency, throughput (req/s), thời điểm bắt đầu xếp hàng do semaphore.
- Thử
Semaphore(5)vsSemaphore(50)→ quan sát đánh đổi latency/throughput/lỗi 429.
Bước 4 (mở rộng) - Serve model local bằng vLLM¶
- Nếu có GPU: chạy
vllm ... --model Qwen/Qwen2.5-7B-Instructtrong container GPU. - Trỏ app vào
base_url=http://vllm:8000/v1(docker-compose 2 service). - So throughput đồng thời giữa Ollama và vLLM trên cùng câu hỏi.
Tiêu chí hoàn thành¶
-
/chatstream token thật, dừng khi client ngắt kết nối - Image multi-stage, non-root, có healthcheck, không chứa weights/index
- Retry có exponential backoff + jitter cho 429/timeout
- Có số benchmark p50/p95/throughput ở ít nhất 2 mức concurrency
- (Mở rộng) so sánh Ollama vs vLLM về throughput đồng thời
Tóm tắt¶
graph LR
CODE[App code + prompts] --> IMG[Docker image<br/>multi-stage, non-root, no weights]
IMG --> RUN[Container]
W[(Weights / index)] -.mount / tải lúc khởi động.-> RUN
RUN --> FASTAPI[FastAPI: SSE stream + async + semaphore + retry]
FASTAPI --> BACKEND{Backend}
BACKEND -->|self-host| VLLM[vLLM / TGI<br/>PagedAttention + continuous batching]
BACKEND -->|managed| API[API provider]
| Chủ đề | Cốt lõi |
|---|---|
| Docker | Multi-stage; weights ngoài image; layer dependency trước code; non-root; healthcheck |
| GPU image | Base nvidia/cuda:*-runtime, --gpus all, nvidia-container-toolkit; image rất nặng |
| PagedAttention | KV cache như paging → tận dụng ~96% bộ nhớ, nhiều seq đồng thời |
| Continuous batching | Lập lịch mức iteration; request xong nhường slot ngay → GPU không idle |
| Chọn stack | Ollama (dev) · vLLM/TGI (self-host, volume lớn) · API (mặc định để bắt đầu) |
| FastAPI | SSE streaming; async client; semaphore backpressure; retry backoff + jitter; dừng khi disconnect |