Results / Operations workflows

A smaller model. A shorter path to the action.

Output control removed unnecessary generation before sparse fine-tuning repaired remaining failures. A separate Fireworks serving validation measured the resulting 8B route.

Median response latency · lower is better
Sonnet 4.6 · APIMean score 1.0000 · 100% strict pass
1,935 ms
Optimized Qwen3-8B · FireworksMean score 0.9630 · 92.22% strict pass
369 ms

Bounded workflows, measured at the action level.

The 30-task Operations holdout spans scheduling, compliance, safety, facilities, inventory, legal documents and cross-system execution. A local action-level scorer checks the workflow’s required behavior; it is not a rating of open-ended chat.

The methodology sweep uses 25 samples per task, or 750 trajectories per row. The Fireworks production-style serving check is separate: 30 tasks with three samples each, or 90 trajectories. Tinker evaluation latency and Fireworks serving latency are different runtime measurements.

Control the output before changing the weights.

Raw Qwen3-8B without prefill already scored 0.9560 but spent 1,390,498 evaluation tokens. JSON prefill reduced that to 74,008 tokens while scoring 0.9667. SFT plus the same prefill reached the same aggregate score: training did not explain the first improvement.

Qwen’s /no_think control cut generation further. Sparse supervised fine-tuning then improved the /no_think route from 0.9598 to 0.9733 on mean score, and from 91.2% to 94.67% strict pass. These are methodology-sweep results, not the separate 90-trajectory serving result.

Methodology sweep · 750 trajectories per row

InterventionMean score ± SDStrict passEval tokensp50 / p95
Qwen3-8B raw no prefill0.9560 ± 0.184794.00%1,390,49832,602 / 63,258 ms
Qwen3-8B + JSON prefill0.9667 ± 0.124893.33%74,0086,759 / 23,475 ms
SFT + JSON prefill0.9667 ± 0.124893.33%74,6306,629 / 26,958 ms
Qwen3-8B + /no_think0.9598 ± 0.131491.20%38,1135,343 / 21,394 ms
SFT + /no_think0.9733 ± 0.112494.67%39,0895,523 / 23,802 ms
All rows use the same 30-task holdout and 25 samples per task. ± is standard deviation. Latency here is the Tinker evaluation route; compare serving runtimes in the separate table below.

The serving comparison, with the quality difference visible.

At the 90-trajectory validation scale, the optimized Fireworks 8B route was about 5.2× faster at p50, with a mean score of 0.9630 versus Sonnet’s 1.0000. The smaller route preserved most of the reference score, not exact quality parity.

Fireworks used 33,111 observed tokens. At the reported $0.20 per million blended-token basis, that is $0.006617 for the validation slice. Sonnet’s recorded token/evaluation cost is $0.039969. The approximately 6.0× ratio compares these slice token costs.

The short-lived Fireworks deployment’s loaded validation cost was about $1.56 because cold-start GPU time dominated the small token cost. This is separate from the modeled token basis and must stay separate when considering production economics.

Serving validation · 90 trajectories

RouteMean scoreStrict passp50Token cost / slice
Sonnet 4.6 · Anthropic API1.0000100%1,935 ms$0.039969
SFT + /no_think · Fireworks 8B0.963092.22%369 ms$0.006617
Fireworks p95: 536 ms. Loaded Fireworks validation cost: approximately $1.56, including cold-start GPU time. This is not a claim about fully loaded steady-state production cost.

What the remaining errors tell us.

Raw /no_think repeatedly missed lease archival, Jira/Confluence incident work and Pipefy vendor onboarding. SFT mostly repaired the incident workflow and partially improved lease archival; vendor onboarding remained a stable miss.

The public snapshot records model families, intervention settings, trajectory counts, score distributions, tokens and runtime distinction. Exact harness versions, hardware details, confidence intervals, and an independently sealed final test are not reported there. The holdout was used repeatedly during method selection.

This page restates the original publication and its checked-in data. It does not rerun the experiment or establish that the short-lived Fireworks deployment is still serving.

Limitations and source record.

  • The optimized serving route scored below Sonnet; latency improvement is not evidence of equal quality.
  • 750-trajectory method comparisons and 90-trajectory serving validation have different runtimes and sample counts.
  • Reported token costs exclude idle capacity, deployment overhead and broader operational cost; the loaded validation cost is shown separately.
  • Repeated comparisons on the holdout limit claims about independent generalization.

Source: the original Understudy publication and its checked-in benchmark snapshot. Adapted for Orchestra on 5 September 2026. This page preserves historical evidence; it is not a new run, live telemetry or a forecast.

Read the original publication ↗

Own your intelligence

Talk to Orchestra