Proof

Evidence first. Routing second.

A controlled benchmark shows MargIQ lowering model cost on routine paths while retaining GPT-5 when the same workflow becomes risky.

Verified routine path74.8% lower measured model cost.

GPT-5 moved to GPT-4.1 mini for an order lookup. A safety-sensitive refund stayed on GPT-5.

Read the controlled benchmark
Coverage15

Workflow families across classification, extraction, writing, chat, tool use, privacy, and security.

Scenarios21

Safety and generalization checks including paraphrases, multilingual inputs, and security neighbours.

Verified decisions3

Two lower-cost routes and one deliberate GPT-5 retention in the retained benchmark snapshot.

Methodology

Proof means showing what MargIQ changed and what it refused to change.

What we inspect

Recurring AI call sites, requested models, task shape, selected model, latency, usage, and policy source.

What we report

Workflow-level savings potential, request-level routing paths, safety guards, and model-pair summaries.

What we protect

Unknown, sensitive, or low-confidence work falls back to the requested model instead of being downgraded.

What we avoid

No raw workflow hashes, variant IDs, thresholds, embeddings, or internal policy JSON in customer-facing UI.

Honest limits

MargIQ should not overclaim savings before workflow evidence exists.

  • The published benchmark used controlled sandbox traffic, not customer production traffic.
  • Path-level savings do not represent a guaranteed reduction across all application traffic.
  • Results depend on workflow mix, token usage, available models, and current provider pricing.
  • Long-term drift and failure rates require production evidence over a longer period.