# MargIQ Workflow-Aware LLM Routing Benchmark

Published 12 July 2026. Updated 27 July 2026.

Author: Rakshith Hegde, Founder of MargIQ.

Canonical URL: https://getmargiq.com/benchmarks/workflow-aware-llm-routing

Disclosure: This was a MargIQ-controlled sandbox benchmark, not customer production traffic.

## Direct answer

In a controlled benchmark, MargIQ routed a routine GPT-5 order lookup to GPT-4.1 mini at 74.83% lower measured model cost. A safety-sensitive refund request in the same ecommerce support workflow stayed on GPT-5. MargIQ therefore reduced cost only where path-level evidence supported the change.

## Coverage

- 15 workflow families evaluated
- 21 safety and generalization scenarios
- 58 requests in the retained snapshot
- 1 active benchmark day
- 2 recorded lower-cost routes
- 1 recorded protected route

The suite covered classification, extraction, summarization, writing, customer chat, tool use, privacy, and security-sensitive work.

## Verified route: routine order lookup

- Workflow: Ecommerce order-support triage
- Requested model: GPT-5
- Selected model: GPT-4.1 mini
- Requested-model baseline: $0.000441
- Selected-model cost: $0.000111
- Measured saving: $0.000330
- Measured model-cost reduction: 74.83%

## Protected route: safety-sensitive refund

- Workflow: Ecommerce order-support triage
- Requested model: GPT-5
- Selected model: GPT-5
- Saving claimed: $0
- Result: MargIQ retained the requested model because the request involved physical harm, urgency, and reputational pressure.

## Second verified route: intent classification

- Workflow: Customer-message intent routing
- Requested model: GPT-5
- Selected model: GPT-4o mini
- Requested-model baseline: $0.000259
- Selected-model cost: $0.000024
- Measured saving: $0.000235
- Measured model-cost reduction: 90.73%

## Methodology

The benchmark used controlled OpenRouter traffic through an OpenAI-compatible client wrapped by MargIQ. GPT-5 was the requested model. MargIQ observed recurring request patterns and evaluated model cost, output quality, latency, and risk before activating a lower-cost route.

The published percentages apply only to paths with complete transaction-level evidence in the retained snapshot.

## Limitations

- This was controlled sandbox traffic, not customer production traffic.
- The retained snapshot covers one active day and two recorded lower-cost decisions.
- Path-level reductions do not represent account-wide or guaranteed savings.
- Results depend on traffic mix, token usage, available models, and provider pricing.
- Latency observations were directional, not a controlled performance experiment.
- Long-term drift, failure rates, and human-review impact were not measured.

## Citation

Hegde, Rakshith. "MargIQ Workflow-Aware LLM Routing Benchmark." MargIQ, 12 July 2026. Updated 27 July 2026. https://getmargiq.com/benchmarks/workflow-aware-llm-routing
