Comparison · LangSmith

rightmodeler vs LangSmith

LangSmith traces, evaluates, and deploys your agent, and its Engine now turns production traces into diagnosed issues and proposed fix pull requests. rightmodeler reads the runs LangSmith already recorded and settles a narrower question with evidence: can a cheaper model take this step without losing what you accepted?

Different job · failure fixes vs model right-sizingVisit LangSmith  (opens in a new tab)

TL;DR

Different job, same traces. You hire LangSmith to run the agent lifecycle: trace trees, dashboards and alerts, evaluators on production traces, deployment, and, as of 2026-09-22, Engine, which clusters recurring failures into issues, diagnoses each root cause against your traces and connected code, and proposes the fix as a pull request. You hire rightmodeler, a free MIT-licensed CLI, for model-cost substitution: it replays the LLM runs you export from LangSmith through cheaper candidates from your provider's live catalog, judges each one against the output you already accepted, and reports agreement, sample size, and abstentions per step. When a swap clears every gate, rightmodeler apply opens a draft pull request that changes model identifiers and nothing else, for a human to review and merge. Not observability. Not a runtime gateway.

What LangSmith Engine does, as of 2026-09-22

  • Per LangChain's docs, Engine is the LangSmith agent for agent engineering. It scans your tracing projects on a schedule, clusters recurring failures into prioritized issues, and diagnoses each root cause against your traces and, when you connect a GitHub repository, your source code. From an issue it opens a pull request with the proposed code or prompt fix in that repository, and its changelog for late August 2026 says those pull requests now open automatically. It keeps each issue current by attaching new traces that match the same failure pattern, reopens an issue when the failure comes back, and turns the traces behind an issue into ground-truth dataset examples so you can test the fix offline before it ships.
  • That is a full loop from production trace to reviewed fix, and for agent failures it is the right tool. Its scope is failures. Its documented issue categories run from Agent looping and Hallucination to Wrong tool and Tracing quality, and none of them is about which model a step should run on or whether a cheaper one would hold up there. You can tell Engine to prioritize Cost & Tokens, and its Context explosion category flags runaway token counts, but its docs describe no step replayed on other models. Engine's own reasoning runs on LangChain-managed inference, with no bring-your-own-key option, per its docs.
  • LangSmith also ships an LLM Gateway, in beta as of 2026-09-22 per its docs: one LangSmith key in front of several providers, with spend limits, rate limits, fallbacks, and redaction applied to live model calls. It governs the model you name. Which model to name at each step is still your call, and that call is the one rightmodeler gathers evidence for.
  • rightmodeler is narrower on purpose. It does not watch production, cluster failures, or edit prompts or code, and it cannot see a run until you export it. It answers one question per step with receipts, and its pull request changes model identifiers and nothing else.

Side by side

One note on words: reference agreement means how closely a candidate's output matches the one you already shipped for that step.

langsmith vs rightmodeler
starting point
LangSmith · production traces streaming into your LangSmith project
rightmodeler · LLM runs you export from LangSmith as JSON or JSONL
question answered
LangSmith · what keeps failing, why, and how to fix it (Engine)
rightmodeler · which model each step needs, and whether a cheaper one holds up
unit of work
LangSmith · the issue: a cluster of traces sharing one failure pattern
rightmodeler · the step family: every recorded LLM run that shares a run name
evidence
LangSmith · evaluator scores, feedback, and the traces linked to each issue
rightmodeler · reference agreement, sample size, abstentions, and a floor re-cleared on held-out cases
the pull request
LangSmith · a proposed code or prompt fix in your connected repository
rightmodeler · a draft that may change model identifiers only, with an evidence table in the body
inference and keys
LangSmith · Engine runs on LangChain-managed inference, no bring-your-own-key
rightmodeler · replays run on your own provider key; no rightmodeler server, account, or telemetry
seat at runtime
LangSmith · tracing beside every run; Deployment and the LLM Gateway (beta) can host or carry it
rightmodeler · none; it runs offline on files you exported

The situations that decide it

Pick the row that sounds like your week. Each names the honest winner.

The same failed tool call keeps turning up across your production runs, and you need the root cause and a fix before it recurs.

the right hire: LangSmith

This is what Engine is built for, per its docs: it clusters the failing traces into one issue, diagnoses the root cause against your traces and connected code, proposes the fix as a pull request, reopens the issue if the failure comes back, and turns the evidence traces into dataset examples for testing the fix. rightmodeler has no live view and never touches prompts or code.

Your LangGraph agent runs every node on a frontier model. The bill keeps climbing, and nobody can say which nodes need that model.

the right hire: rightmodeler

Every model call lands in the export as its own LLM run, and the audit groups runs into step families by run name, so give each node's model call its own run_name to get a verdict per node. It replays those runs on cheaper candidates from your provider's live catalog, judges them against the outputs you accepted, and keeps the frontier model wherever the evidence says to. Only a swap that clears the quality floor again on held-out cases becomes a draft pull request.

Your team already runs Engine on production traces and also wants the model bill down without adding risk.

the right hire: both, together

Keep Engine on failures and give the audit the cost question. Export the same project's LLM runs, and if you already grade with LangSmith scorers, pass --evaluator langsmith so those scorers, not a judge you did not write, decide each replay. Engine's fix pull requests and rightmodeler's model-only drafts arrive as separate reviews, so each change is judged on its own evidence.

LangSmith alone

  • Trace trees, dashboards and alerts on cost, latency, and errors, plus LLM-as-judge, code, and multi-turn evaluators run on production traces.
  • Engine, as of 2026-09-22: recurring issues found on a schedule, root causes traced into your connected code, fixes proposed as pull requests, and dataset examples built from the failing traces.
  • Deployment to run agents, and an LLM Gateway, in beta, that applies spend limits, rate limits, and fallbacks to live model calls.

LangSmith with rightmodeler beside it

  • Everything on the left stays. The audit reads the LLM runs LangSmith already recorded and adds a per-step answer to one question: does this step need the model it pays for?
  • Candidates come from your configured provider's live catalog and are judged against the output you accepted, either by a judge from a model family neither side belongs to or by your own LangSmith scorers.
  • A swap has to clear the quality floor again on held-out cases, multi-step swaps are confirmed end to end, and thin evidence ends in an abstention. What clears every gate becomes a draft pull request that a human reviews and merges.

How the audit runs on your LangSmith exports

No new instrumentation. Your LangSmith key is used to export runs, and by the CLI only if you choose LangSmith scorers or datasets.

Pull the LLM runs from your project with the LangSmith SDK, for example client.list_runs(project_name=..., run_type="llm"), and save them as JSON or JSONL with one run per record. The CLI detects the format from the trace_id, run_type, and dotted_order keys, orders each trace by dotted_order, and takes the model from extra.metadata.ls_model_name. It reads files only: there is no rightmodeler server, no account with us, and no telemetry.

The run writes report.md with a verdict per step family: the candidates tried on your recorded inputs, the agreement each earned against the output you accepted, the sample size behind it, and an abstention with a named reason wherever the evidence is thin. rightmodeler report --output json prints the same report as JSON. When a family clears every release gate, rightmodeler apply opens a draft pull request that changes model identifiers only and requests review from the owners of the files it touches. It never merges.

# preview the pipeline against the export, without spending

# replay and judge on your own provider key

# or let your own LangSmith scorers grade the replays

Setup guide: rightmodeler + LangSmith

The words the report uses

swap candidate
A cheaper model shortlisted from your configured provider's live catalog and replayed on one step's recorded inputs.
reference evidence
Scores measured against the output you actually shipped for that exact step. Evidence of agreement with shipped output, not proof of correctness: the production result is the reference, not ground truth.
quality floor
The minimum agreement a candidate must clear, and clear again on held-out cases, before the audit will recommend a swap. Configurable, and anything below it is a no.
cascade risk
The chance that a cheaper model at one step quietly degrades the later steps that build on its output. Swaps that could cascade are confirmed by running the pipeline end to end before they are recommended.
abstain
The audit's decision to recommend nothing for a step when the sample is small or the evidence is mixed, with the reason named. The current model stays.

Frequently asked questions

Doesn't LangSmith Engine already find cheaper models?

Not per its docs as of 2026-09-22. Engine's documented issue categories cover agent failures such as looping, hallucination, wrong tools, and truncated responses, and none of them is about which model a step should run on. You can ask it to prioritize Cost & Tokens, and it flags context explosion, but its docs describe no step replayed on other models and no measurement of whether a cheaper one holds up. That measurement is the whole of what rightmodeler does.

Both open pull requests. What is different about rightmodeler's?

Engine's pull request carries a proposed code or prompt fix for a diagnosed failure. rightmodeler apply opens a draft that may only change model identifiers: a diff check refuses anything else, the body carries the evidence table for each step family, and review is requested from the owners of the changed files. It opens only after the swap clears every release gate, and the CLI never merges; a person does.

Can't LangSmith experiments already compare models?

They can, honestly. Build a dataset, define evaluators, and an experiment will compare models across it, and Engine can now seed that dataset from failing traces. rightmodeler starts from the outputs you already accepted, so no dataset is required, and the verdict comes back per step with sample size and abstentions attached. If you keep curated LangSmith datasets and scorers, the audit can use both: rightmodeler corpus import --from langsmith:<dataset> builds the corpus, and --evaluator langsmith lets your scorers grade.

Does rightmodeler connect to my LangSmith account?

Only when you ask it to. The audit reads exported files on your machine, and replays run on your own provider key. The CLI calls the LangSmith API only on the evaluator path (--evaluator langsmith) or for corpus import, using the key in LANGSMITH_API_KEY unless you name another variable.

Is LangSmith open source?

The platform is not. LangSmith, including Engine and the LLM Gateway, is a commercial product that LangChain hosts, and self-hosting it takes a license from LangChain, per its docs. Its client SDK is MIT-licensed, and that SDK is all the export step needs. LangChain's open-source work includes the frameworks: LangChain, LangGraph, and Deep Agents. The rightmodeler CLI is MIT-licensed.

Does every audit end in recommended swaps?

No, and it is designed not to. Steps with thin evidence get an abstention rather than a swap, and a clean bill is a possible outcome: the report can come back saying your current model assignments already earn their cost.

Run the audit on your own traces

The CLI runs from npx, nothing to install, and your own traces settle the question.

View on GitHub