Results

Progress you can
inspect.

A result belongs beside its workload, baseline and method. Here are three different ways to improve the economics of real work.

The evidence, in context

These historical studies document specific tasks. Read the score definitions, sample counts and limitations before extending a result to your own workload.

01

CRM actions

Better CRM actions, from a better task contract.

A targeted prompt adapter improved Qwen’s mean score on seven write-heavy CRM tasks. The improvement came from reading failures and specifying the missing behavior.

Read the study

+13.1%

relative lift in mean score

0.630 ± 0.087 versus Sonnet 4.6 at 0.557 ± 0.066. Seven tasks; optimized n=10, baseline n=3.

02

Operations workflows

A smaller model. A shorter path to the action.

Output control removed unnecessary generation before sparse fine-tuning repaired remaining failures. A separate Fireworks serving validation measured the resulting 8B route.

Read the study

5.2×

lower median latency

369 ms versus 1,935 ms. Mean action-level score: 0.9630 versus 1.0000. 90 trajectories per serving comparison.

03

Warehouse labeling

Label the whole table. Inspect the difficult rows.

A post-trained 30B open route labeled 39,962 non-empty comments alongside Sonnet and Opus. The cost comparison is strong; the quality evidence needs a label-by-label reading.

Read the study

$2.82

reported cost for 39,962 rows

$2.82 open route · $12.48 Sonnet · $139.63 Opus. Agreement between models is not ground-truth accuracy.

Own your intelligence

Talk to Orchestra