Results
Progress you can inspect.
A result belongs beside its workload, baseline and method. Here are three different ways to improve the economics of real work.
01
CRM actions
Better CRM actions, from a better task contract.
A targeted prompt adapter improved Qwen's mean score on seven write-heavy CRM tasks. The improvement came from reading failures and specifying the missing behavior.
Read the study ->+13.1%
relative lift in mean score
0.630 ± 0.087 versus Sonnet 4.6 at 0.557 ± 0.066. Seven tasks; optimized n=10, baseline n=3.
02
Operations workflows
A smaller model. A shorter path to the action.
Output control removed unnecessary generation before sparse fine-tuning repaired remaining failures. A separate Fireworks serving validation measured the resulting 8B route.
Read the study ->5.2×
lower median latency
369 ms versus 1,935 ms. Mean action-level score: 0.9630 versus 1.0000. 90 trajectories per serving comparison.
03
Warehouse labeling
Label the whole table. Inspect the difficult rows.
A post-trained 30B open route labeled 39,962 non-empty comments alongside Sonnet and Opus. The cost comparison is strong; the quality evidence needs a label-by-label reading.
Read the study ->$2.82
reported cost for 39,962 rows
$2.82 open route · $12.48 Sonnet · $139.63 Opus. Agreement between models is not ground-truth accuracy.
