Aug 5, 2026

    Introducing AutoEvals: Automatically find the best model for your task

    A

    Amar Singh

    X

    It's the Summer of 2026. A new frontier model is released nearly every month. Each one comes with a new set of benchmark numbers, some hype on X, and the same nagging question for teams running LLMs in production: should we switch?

    Most teams never really answer it. Answering it honestly means building a dataset that represents your real traffic, standing up API access to every provider you want to test, writing a rubric, wiring up judging infrastructure, burning the tokens to run it all, and then interpreting the results. That's a week of work, per model, per task, every time something new ships. So the question gets shelved. The model that was chosen months ago, based on a price point and vibes, keeps running because swapping models in production is both hard and risky.

    And the risk is real. Benchmarks are averages over someone else's tasks. A model can top the leaderboards and still fumble your tool schemas, your output format, your edge cases. The only benchmark that actually matters is your own traffic, and until now, nobody had time to run it.

    We're introducing tooling to fix this. Introducing AutoEvals - the plug-and-play toolkit to find the right model for your task. Sign in or read the docs to get started.

    AutoEvals model comparison report showing a recommended model with quality, cost, and latency results

    Real production traffic is the best benchmark

    AutoEvals is designed to help developers find the best model for a given task or agent. Not the best on a public benchmark you never heard of - better on the exact requests your users sent last week.

    Integrate with Inference Gateway (in 2 minutes), collect data, click one button in the dashboard, and a few minutes later you have a definitive answer about which model you should use, with real numbers on the difference in quality, cost, and latency.

    Every model release becomes free upside instead of homework. When something new ships, you no longer weigh the announcement against a week of eval work. Run a comparison over lunch and know by the afternoon whether it matters for you.

    Switching stops being scary. The recommendation isn't a leaderboard rank. It's your own traffic, replayed and judged, with every sample inspectable. You can understand exactly how a new model handles requests compared to your current model before a single production request touches it.

    You stop paying for intelligence you don't need. Plenty of tasks are quietly overserved by expensive models. AutoEvals shows you the quality-cost frontier for your specific task, so you can see when a more affordable model offers the same level accuracy.

    AutoEvals recommendation card naming the winning model with the judge's reasoning
    The verdict: a concrete recommendation for your task, with the reasoning behind it

    How AutoEvals work in practice

    An AutoEval runs in three moves. It starts by sampling your recent production traffic, which is then replayed by a set of candidate models. Afterwards, LLM judges score each output for both quality, meaning how well the model did the task, and similarity, meaning how closely it behaves like the model you run today. It then summarizes these results into a model verdict, key findings, and per-model summaries.

    The two scores together tell you how safe a switch is. High quality with high similarity is a drop-in upgrade. High quality with low similarity means the model is good but behaves differently, so it earns a closer look first. When you do want to inspect the outputs individually, the Sample Viewer gives you side-by-side model responses and the judge's reasoning for each score.

    Sample Viewer showing side-by-side model responses with judge reasoning for each score
    The Sample Viewer: side-by-side model responses and the judge's reasoning for each score

    Getting your traffic in front of Catalyst is the only setup, and it's one integration: route your LLM calls through our Gateway (our CLI can instrument your project automatically), or report traces with our tracing SDK. From then on, comparing models is a button, and switching is a single line change.

    Running on our own agents

    We built AutoEvals for our own agents before we built it for anyone else, and two findings genuinely changed how we think about model selection.

    The gap between model generations is bigger than you think. One of our GTM agents had been running on the same well-regarded model for months. It felt fine, until we ran an AutoEval on it. Models at a fraction of the cost blew it out of the water with better long-horizon trajectories, tighter tool-schema adherence, and more precise results across the board. We left better results on the table because we didn't look.

    Some tasks have a skill ceiling. On that same task, the top frontier models landed on identical quality scores. Past a certain capability level, they were all simply taking the right action, and paying more bought nothing. Once you can see the ceiling, the model decision makes itself, and it's rarely the most expensive option.

    Quality versus cost scatter plot showing top models converging on the same quality score at very different prices
    Past a certain capability level, quality converges - paying more buys nothing

    Put those two findings together and the conclusion is hard to escape: if you're building with LLMs, your application should be getting more accurate and cheaper over time, just from the pace of model releases. The teams that capture that compounding advantage will be the ones for whom evaluation is effortless. That's what we built.

    Run your first comparison today

    AutoEvals are live now for every Inference account, free to start. Connect your traffic, click Run Comparison, and in a few minutes you'll know something most teams never find out: whether the model you're running is actually the best one for the job.

    Read the guide or sign up free to get started.

    The best model for your task might have shipped last week. Now you can actually find out.

    CONTACT

    Meet with our research team

    Schedule a call with our research team to learn more about how Specialized Language Models can cut costs and improve performance.