Evidence-backed model decisions
Measure cheaper models against what you shipped.
rightmodeler replays your real agent traces through cheaper models, judges each output against what you already shipped, and reports agreement, cost, evidence, sample size, and every abstention.
cost reduction · PR summary · medium confidence · illustrative example
Reads the traces you already have.
Autodetects 10 trace formats into one per-step schema.
Reads traces from Claude Code, Codex, LangSmith / LangGraph, OpenAI SDK, Langfuse, Braintrust, Phoenix (OpenInference), OpenTelemetry GenAI, Helicone, W&B Weave.
The platform
See everything. Measure candidates. Review the pull request.
What teams measured with rightmodeler

“rightmodeler took AI Assist from brute force to precision routing. Costs fell 70.8%, responses got twice as fast, and quality held at 100%, measured, not assumed.”
Measured on a 20-query benchmark against the outputs B:Side had accepted.
Run it on your own traces.
Free until replay, then your own provider key. It’s a report, not a runtime gateway.

