Results

Progress you can inspect.

A result belongs beside its workload, baseline and method. Here are three different ways to improve the economics of real work.

01

CRM actions

Better CRM actions, from a better task contract.

A targeted prompt adapter improved Qwen's mean score on seven write-heavy CRM tasks. The improvement came from reading failures and specifying the missing behavior.

Read the study ->

+13.1%

relative lift in mean score

0.630 ± 0.087 versus Sonnet 4.6 at 0.557 ± 0.066. Seven tasks; optimized n=10, baseline n=3.

02

Operations workflows

A smaller model. A shorter path to the action.

Output control removed unnecessary generation before sparse fine-tuning repaired remaining failures. A separate Fireworks serving validation measured the resulting 8B route.

Read the study ->

5.2×

lower median latency

369 ms versus 1,935 ms. Mean action-level score: 0.9630 versus 1.0000. 90 trajectories per serving comparison.

03

Warehouse labeling

Label the whole table. Inspect the difficult rows.

A post-trained 30B open route labeled 39,962 non-empty comments alongside Sonnet and Opus. The cost comparison is strong; the quality evidence needs a label-by-label reading.

Read the study ->

$2.82

reported cost for 39,962 rows

$2.82 open route · $12.48 Sonnet · $139.63 Opus. Agreement between models is not ground-truth accuracy.

Own your intelligence