Latest Updates
AI engineering & research, product releases and upgrades, and thoughts about the industry from the Inference R&D, Inc. team.

Introducing AutoEvals: Automatically find the best model for your task
Automatically discover which model is best suited for your task, with the data you already generate. AutoEvals replays your production traffic across candidate models and tells you which one to run, with real numbers on quality, cost, and latency.
Aug 5, 2026
AAmar Singh

Schematron V2: Frontier HTML-to-JSON extraction at a fraction of the cost
Schematron is a family of small, purpose-built models that transform messy HTML into clean, structured JSON, delivering frontier-level extraction quality at a fraction of the cost and latency of large general-purpose LLMs.
Apr 16, 2026
AAmar Singh

Introducing Catalyst: Monitor, train, and deploy self-improving AI models
Build self-improving AI-systems from production data
Apr 14, 2026
SSam Hogan

How Inference.net trains Specialized Language Models that cut AI costs by up to 50x
Learn how Inference.net trains Specialized Language Models using the NVIDIA NeMO framework to delivery frontier accuracy at up to 50x lower cost.
Mar 11, 2026
SSam Hogan

Specialized LLMs: The model you need doesn't exist yet
Specialized LLMs trained on your own user data can match frontier quality for a fraction of the cost
Feb 5, 2026
SSam Hogan

Project OSSAS: Custom LLMs to process 100 Million Research Papers
Project OSSAS is a large-scale open-science initiative to make the world’s scientific knowledge accessible through AI-generated summaries of research papers.
Nov 11, 2025
SSam Hogan

LOGIC: Trustless Inference through Log-Probability Verification
A practical method for verifying LLM inference requests in trustless environments.
Nov 5, 2025
AAmar Singh

Hybrid-Attention models are the future for SLMs
Hybrid attention delivers up to 3x cost reduction compared to traditional transformer models
Nov 3, 2025
AAmar Singh

Announcing our $11.8M Series Seed
We raised $11.8 million in funding led by Multicoin Capital and a16z CSX to train and hosti custom language models that are faster, more affordable, and more accurate than what the Big Labs offer.
Oct 14, 2025
SSam Hogan

Schematron: An LLM trained for HTML -> JSON at scale
Schematron-8B and Schematron-3B deliver frontier-level extraction quality at 1-2% of the cost and 10x+ faster inference than large, general-purpose LLMs.
Sep 9, 2025
SSam Hogan