
New MIT and Sakana AI framework uses an LLM judge to cut evaluation costs for self-improving coding agents
MIT's SIFT framework uses a language model to cut coding agent evaluation costs, achieving 35.1% accuracy on Polyglot with reduced compute resources.
Ben Dickson
Google’s WikiSkill gives AI agents a memory of what went wrong — without putting it in the prompt
In tests, WikiSkill helped a 9B Qwen outscore a 27B model by turning past runs into reusable skills.
Ben Dickson
Stanford and Nvidia's open CLM-8B caches reusable agent actions and runs up to 9x faster than Jev in tests
Instead of writing out an answer, the open 8B model scores a fixed list of options. It ran faster than TypeSafe's Jev in the team's tests but gave up a few points of accuracy on tool calling.
Ben Dickson
Text handoffs slow AI models down. C2C lets them communicate through KV caches instead
C2C cuts out text handoffs between models — but only for teams that control their own inference stack.
Ben Dickson
Google’s open source EnvHarness lets AI agents train against environments that evolve with them
The approach presents an alternative to continuously building new simulators and training tasks from scratch: start with a trusted environment and dynamically reshape it around the agent’s current weaknesses.
Ben Dickson
Salesforce researchers took an AI agent from finishing 43.5% of browser tasks to 93% without touching the model
Fixing one agent failure tends to break another. Salesforce's answer is to evolve rival versions of the harness and promote only the ones that don't regress.
Ben Dickson
Frontier models can recover up to 65% of facts they can't directly recall — just by thinking longer
Google Research shows LLMs encode 95-98% of facts, with recall as the main bottleneck, aiding reliable app development without larger models.
Ben Dickson
Meta researchers taught an 8B AI model to match Claude Opus 4.5 — without the frontier price tag
Meta AI's EvoHarness-RL framework optimizes AI agent efficiency, improving success rates in complex workflows by teaching effective harness use.
Ben Dickson
Nvidia finds that simple linear math can replace costly AI model handoffs
No deep learning needed: Nvidia's linear math swaps AI models mid-task 25x faster, cutting compute costs on long agentic sessions.
Ben Dickson
One AI module faked 86% of a pipeline's accuracy gains by feeding another the answers
A decomposer module was caught planting fake answers to inflate its pipeline's accuracy score. Most of the 'progress' wasn't real.
Ben Dickson
Brex assumes its AI agents could do anything — so it watches the network, not the code
An LLM now decides which network requests Brex's AI agents are allowed to make — and it only has to step in about 2% of the time.
Ben Dickson
Four AI agents coordinating in real time outperformed Claude Opus 4.8 on enterprise coding tasks
Enterprise codebases stump AI agents working alone. Sharing discoveries with teammates in real time, instead of waiting for a review phase, nearly doubled task accuracy in new research.
Ben Dickson