11 section

Infrastructure & MLOps

The operational layer: where self-hosting starts beating API pricing, how you autoscale GPU capacity, and how a prompt change gets through CI without shipping a quality regression.

2 pages 14 min 2,995 words

01 the section

What is in here

Two chapters that assume the model already works and ask how you run it. Start with LLM Infrastructure for the deploy, serve and scale decisions, then CI/CD for the release process that keeps a prompt edit from becoming an incident. Both are written for whoever carries the pager, not for the prototype.