In-path gateway for AI agents
One endpoint, every model, and a bill that keeps getting smaller.
Point your agents at Operant — one base URL, no code changes. It forwards every call, prices every turn, and then quietly adapts: caches what repeats, drops what is stale, and moves the work it has learned onto cheaper models carrying the right skill.
Without Operant this conversation would have cost $2.41; it cost $0.87 — −$1.54 (64%) saved.
| Call | Served by | Potential | Actual | Saved |
|---|---|---|---|---|
| #1 | claude-sonnet-5 | $0.31 | $0.31 | — |
| cachebreakpoints placed — the next turns read this prefix at 0.1× | ||||
| #2 | claude-sonnet-5 | $0.58 | $0.19 | −$0.39 |
| cache41,900 of 46,300 input tokens read from cache instead of re-sent | ||||
| #3 | claude-sonnet-5 | $0.74 | $0.22 | −$0.52 |
| compressionmasked 3 stale tool results (older than 2 turns) → 18,400 fewer tokens | ||||
| #4 | claude-haiku-4-5 learned:7 | $0.78 | $0.15 | −$0.63 |
| routingrule “Compliance evidence pack” served haiku + skill v2 — gate 100% | ||||
| Σ | 4 calls | $2.41 | $0.87 | −$1.54 |
Illustrative ledger. Counterfactuals are stated estimates: same tokens priced at the model the agent asked for, masked tokens restored, gateway-shaped cache unwound.