The Optimization Ladder: Prompts, SFT, RL, and Routing
LLM optimization should climb from cheap control fixes to heavier training only when the eval proves the next step is worth it.
Read the essayNotes on measuring quality, learning from expert judgment, and finding the smallest change that makes a meaningful difference.
LLM optimization should climb from cheap control fixes to heavier training only when the eval proves the next step is worth it.
Read the essayReward signals encode product judgment. ML teams can build the harness, but domain experts know which errors matter, which tradeoffs are acceptable, and what good work looks like.
LLM optimization should climb from cheap control fixes to heavier training only when the eval proves the next step is worth it.
Frontier models are the right baseline for new workflows. Specialist models become attractive once the task repeats, the eval is stable, and the cost or latency curve starts limiting the product.
Production traces are not just logs. With the right capture, review, and holdout discipline, they become the evals that make model optimization safe.
Orchestra's operations benchmark showed that scaffolding and output control can make small models reliable before sparse fine-tuning work begins.
Selected from the Understudy research archive and adapted for Orchestra. Original publication dates and attribution are retained. Each essay links to its original source and relevant evidence.