Order-status lookup
A bounded tool call with a repeatable outcome earned a lower-cost model.
- GPT-5 baseline
- $0.000441
- GPT-4.1 mini
- $0.000111
- Measured reduction
- 74.83%
Controlled benchmark · 12 July 2026
MargIQ routed a routine GPT-5 order lookup to GPT-4.1 mini at 74.8% lower measured model cost, while keeping GPT-5 when the request became safety-sensitive.
Report-only by default. No provider migration or production routing change.
Same workflow, different stakes
An AI workflow is a recurring production task. Inside it, request patterns can carry very different cost, quality, and risk requirements.
One recurring server-side task with multiple request paths.
Learn from repeated production traffic.
Compare cost, quality, latency, and risk.
Choose only from your available models.
A bounded tool call with a repeatable outcome earned a lower-cost model.
Physical harm, urgency, and reputational pressure required stronger judgment. MargIQ kept GPT-5.
Evaluation breadth
The suite covered recurring production-style work and adversarial neighbours designed to test when MargIQ should optimize, protect, or wait.
Second verified route
For a short, low-risk intent classification with a strict JSON output, MargIQ selected GPT-4o mini after candidate evaluation.
This is a path-level result, not a claim that MargIQ reduces every application's total AI spend by 90.7%.
Quality protection
Candidate outputs disagreed on decision-bearing fields and the benchmark had no authoritative quality definition to resolve them. MargIQ blocked optimization instead of guessing.
How to read the result
MargIQ learns recurring server-side AI work and the request patterns inside it, then selects a model for each path instead of assigning one model to everything.
No. The requested model remains the quality anchor and fallback. MargIQ routes only to models you make available, and only after evidence supports the change.
No. The measured percentages apply to the verified paths shown here. Account-wide savings depend on how much traffic is eligible for those paths.
Methodology and limits
The retained database contains complete transaction-level cost evidence for two lower-cost routes and one protected route.
OpenRouter traffic used an OpenAI-compatible client wrapped by MargIQ. GPT-5 was the requested model.
Policies needed repeated samples and evaluation before a lower-cost route could become active.
The published percentages apply only to the two verified routine paths with complete retained evidence.
Research record
The sanitized record includes the measured routes, quality protection outcomes, methodology, and limitations shown on this page.
Your workflows, your evidence
Start in report-only mode. MargIQ keeps your requested models active while it maps potential savings and quality boundaries.
Analyze my workflows free