Controlled benchmark · 12 July 2026

One support workflow. Two model decisions.

MargIQ routed a routine GPT-5 order lookup to GPT-4.1 mini at 74.8% lower measured model cost, while keeping GPT-5 when the request became safety-sensitive.

Report-only by default. No provider migration or production routing change.

Verified pathsEvidence changed the route
Routine order lookupGPT-5 GPT-4.1 mini
74.8% lower cost
Safety-sensitive refundGPT-5 retained
Quality protected
Controlled evidence, clearly labeled.This is a MargIQ-owned sandbox benchmark, not a customer case study.

Same workflow, different stakes

Routine work became cheaper. Risky work stayed protected.

An AI workflow is a recurring production task. Inside it, request patterns can carry very different cost, quality, and risk requirements.

WorkflowEcommerce support triage

One recurring server-side task with multiple request paths.

  1. 01Observe

    Learn from repeated production traffic.

  2. 02Evaluate

    Compare cost, quality, latency, and risk.

  3. 03Route automatically

    Choose only from your available models.

Routine pathLower-cost route

Order-status lookup

A bounded tool call with a repeatable outcome earned a lower-cost model.

GPT-5 baseline
$0.000441
GPT-4.1 mini
$0.000111
Measured reduction
74.83%
Higher-risk pathRequested model retained

Refund and safety escalation

Physical harm, urgency, and reputational pressure required stronger judgment. MargIQ kept GPT-5.

Requested model
GPT-5
Selected model
GPT-5
Saving claimed
$0.00

Evaluation breadth

Tested beyond one happy path.

The suite covered recurring production-style work and adversarial neighbours designed to test when MargIQ should optimize, protect, or wait.

15workflow families
21safety and generalization scenarios
58requests in the retained snapshot
ClassificationExtractionSummarizationWritingCustomer chatTool usePrivacySecurity-sensitive work

Second verified route

A strict classification path cost 90.7% less.

For a short, low-risk intent classification with a strict JSON output, MargIQ selected GPT-4o mini after candidate evaluation.

RequestedGPT-5$0.000259
SelectedGPT-4o mini$0.000024
Measured reduction90.73%

This is a path-level result, not a claim that MargIQ reduces every application's total AI spend by 90.7%.

Quality protection

Uncertainty did not become a savings claim.

One workflow remained on GPT-5.

Candidate outputs disagreed on decision-bearing fields and the benchmark had no authoritative quality definition to resolve them. MargIQ blocked optimization instead of guessing.

How to read the result

Workflow-level evidence, not a blanket downgrade.

What is workflow-aware model routing?

MargIQ learns recurring server-side AI work and the request patterns inside it, then selects a model for each path instead of assigning one model to everything.

Does MargIQ replace the requested model?

No. The requested model remains the quality anchor and fallback. MargIQ routes only to models you make available, and only after evidence supports the change.

Are these savings account-wide?

No. The measured percentages apply to the verified paths shown here. Account-wide savings depend on how much traffic is eligible for those paths.

Methodology and limits

What this benchmark proves, and what it does not.

The retained database contains complete transaction-level cost evidence for two lower-cost routes and one protected route.

Traffic

Real model calls in a controlled sandbox

OpenRouter traffic used an OpenAI-compatible client wrapped by MargIQ. GPT-5 was the requested model.

Decision basis

Recurring workflow evidence

Policies needed repeated samples and evaluation before a lower-cost route could become active.

Reporting

Measured routes kept separate from projections

The published percentages apply only to the two verified routine paths with complete retained evidence.

Limitations

  • This was a MargIQ-controlled sandbox benchmark, not customer production traffic.
  • The retained snapshot covers one active day and two recorded lower-cost decisions.
  • Short requests produced small absolute dollar savings despite large path-level reductions.
  • Results depend on traffic mix, token usage, available models, and provider pricing.
  • Latency observations were directional, not a controlled performance experiment.
  • Long-term drift, failure rates, and human-review impact were not measured.

Research record

Inspect and cite the benchmark evidence.

The sanitized record includes the measured routes, quality protection outcomes, methodology, and limitations shown on this page.

Suggested citationHegde, Rakshith. “MargIQ Workflow-Aware LLM Routing Benchmark.” MargIQ, 12 July 2026. Updated 27 July 2026.getmargiq.com/benchmarks/workflow-aware-llm-routing

Your workflows, your evidence

See where your model spend is actually necessary.

Start in report-only mode. MargIQ keeps your requested models active while it maps potential savings and quality boundaries.

Analyze my workflows free