02 section
Model Landscape
Which model to use and what it will cost: the field as of March 2026, how to assess capability for your own task, and pricing you can forecast
What is in here
The taxonomy is the map; capability assessment is the method that stops you trusting a leaderboard. Pricing and the selection guide turn both into a number and a decision. Model names date fast — the frameworks here do not, so read for the method and swap in whatever shipped this quarter.
01
13 min
Model Taxonomy
A map of the March 2026 model field — frontier, reasoning, fast, open-weight, specialized — with the capability tiers and residency constraints that narrow your shortlist
model-selectionfundamentalsreasoning
02
9 min
Capability Assessment
Public benchmarks are contaminated and averaged; build a task-specific eval set, run internal Elo, and A/B the finalists on your own traffic instead
model-selectionevaluationproduction
03
11 min
Pricing and Costs
Where the money goes in a token-priced system: input/output asymmetry, cache discounts, batch tiers, and the crossover point where self-hosting beats an API
costmodel-selectioncaching
04
7 min
Model Selection Guide
A decision tree that turns latency budget, task difficulty, privacy posture and spend into one model choice — or a routed cascade of several
model-selectioncostlatency