Use cases
Start where the work repeats.
A clear task, a measurable outcome, and enough repetition to make learning it worthwhile. Find the part of your product that is ready for a cheaper model and a skill.
01
Compliance evidence packs
The pack auditors keep asking for, assembled from your systems of record.
The task · Assemble audit evidence in the standard four-section structure.
02
Coding sessions
The exploration and triage steps that recur inside every Claude Code or DeepSeek Harness session.
The task · Locate files, symbols and call sites; report paths, not dumps.
03
Scheduled digests
Recurring summaries with a stable shape and a fixed source of record.
The task · The week's incidents, with owners, status and follow-ups.
A useful starting point
Bring the task. And your definition of good.
You do not need a dataset to start. A workload that repeats, its current model, and examples of what still goes wrong are enough for a first conversation. We establish the evaluation first, then decide whether the next improvement belongs in the prompt, the route, or a learned skill.
Inspect the published results ->