LLM System Design
Designing LLM systems: method, case studies, inference infrastructure
Open "Foundations and Method"
Foundations and Method
Latency and cost, routing, caching, guardrails, evals and LLM observability
20 questions
Open "Case Studies"
Case Studies
End-to-end LLM case studies: support, documents, agents, data extraction, voice, search
20 questions
Open "Inference Infrastructure"
Inference Infrastructure
Inference engines, KV cache, LoRA, MoE, quantization, embeddings, GPU selection and cost per token
20 questions