Evidence-backed model decisions

Measure cheaper models against what you shipped.

rightmodeler replays your real agent traces through cheaper models, judges each output against what you already shipped, and reports agreement, cost, evidence, sample size, and every abstention.

0%

cost reduction · PR summary · medium confidence · illustrative example

View on GitHub
Illustrative, not measured results
rightmodeler · per-step approval5 steps · quality floor 0.90
Approvedpr_summary
gpt-4.1gpt-4o-mini
save 72%quality 0.94reference+judge
Pending, cascade risktool_agentCASCADE
claude-opus-4llama-3.3-70b
save 41%quality 0.88trajectory
Approvedjson_extraction
gpt-4ogpt-4o-mini
save 68%quality 1.00deterministic
Approvedsql_generation
gpt-4odeepseek-chat
save 55%quality 0.91reference
Abstainedauth_code_editHIGH-RISK · abstain
gpt-4.1·
save ·quality ·none

Reads the traces you already have.

Autodetects 9 trace formats into one per-step schema.

Reads traces from Claude Code, Codex, LangSmith / LangGraph, OpenAI SDK, Langfuse, Braintrust, Phoenix (OpenInference), OpenTelemetry GenAI, LiteLLM StandardLoggingPayload.

Chris Myers portrait

rightmodeler took AI Assist from brute force to precision routing. Costs fell 70.8%, responses got twice as fast, and quality held at 100%, measured, not assumed.

Chris Myers, CEO, B:Side Capital and Fund

Measured on a 20-query benchmark against the outputs B:Side had accepted.

Run it on your own traces.

It’s a report, not a runtime gateway. Measure the savings on your own data first.