Evidence-backed model decisions

Measure cheaper models against what you shipped.

rightmodeler replays your real agent traces through cheaper models, judges each output against what you already shipped, and reports agreement, cost, evidence, sample size, and every abstention.

72%

cost reduction · PR summary · medium confidence · illustrative example

View on GitHub
Illustrative, not measured results
rightmodeler · per-step approval5 steps · quality floor 0.90
Approvedpr_summary
gpt-5.6gpt-5.4
save 72%quality 0.94reference+judge
Pending, cascade risktool_agentCASCADE
claude-opus-5llama-3.3-70b
save 41%quality 0.88trajectory
Approvedjson_extraction
gpt-5.4qwen3.7-flash
save 68%quality 1.00deterministic
Approvedsql_generation
gpt-5.4deepseek-chat
save 55%quality 0.91reference
Abstainedauth_code_editNO EVIDENCE · abstain
gpt-5.6·
save ·quality ·none

Reads the traces you already have.

Autodetects 10 trace formats into one per-step schema.

Reads traces from Claude Code, Codex, LangSmith / LangGraph, OpenAI SDK, Langfuse, Braintrust, Phoenix (OpenInference), OpenTelemetry GenAI, Helicone, W&B Weave.

What teams measured with rightmodeler

Chris Myers portrait

rightmodeler took AI Assist from brute force to precision routing. Costs fell 70.8%, responses got twice as fast, and quality held at 100%, measured, not assumed.

Chris Myers, CEO, B:Side Capital and Fund

Measured on a 20-query benchmark against the outputs B:Side had accepted.

Run it on your own traces.

Free until replay, then your own provider key. It’s a report, not a runtime gateway.