
Nvidia finds that simple linear math can replace costly AI model handoffs
No deep learning needed: Nvidia's linear math swaps AI models mid-task 25x faster, cutting compute costs on long agentic sessions.
Ben Dickson
One AI module faked 86% of a pipeline's accuracy gains by feeding another the answers
A decomposer module was caught planting fake answers to inflate its pipeline's accuracy score. Most of the 'progress' wasn't real.
Ben Dickson
Brex assumes its AI agents could do anything — so it watches the network, not the code
An LLM now decides which network requests Brex's AI agents are allowed to make — and it only has to step in about 2% of the time.
Ben Dickson
Four AI agents coordinating in real time outperformed Claude Opus 4.8 on enterprise coding tasks
Enterprise codebases stump AI agents working alone. Sharing discoveries with teammates in real time, instead of waiting for a review phase, nearly doubled task accuracy in new research.
Ben Dickson
Stanford is running 37,000 AI agents as a virtual biotech — and one of its drug designs got independently confirmed by Merck
The agents that designed this drug also argue with each other. Researchers found multi-agent debate produces more robust results than a single model working alone.
Ben Dickson
Asana's AI agents share memory across your company — but not your secrets
How do you let an AI agent learn from a secret M&A deal without ever letting that memory reach an employee who isn't supposed to know? Asana's CPO explains the fix.
Ben Dickson
Structured AI data pipelines score 10.9 points below free-form code — DataFlow-Harness closes the gap
Disposable scripts are easy for AI to write and hard for engineers to trust. A new open-source framework aims to change that math for enterprise data teams.
Ben Dickson
Runway couldn't fix a bug in its AI video model, so it turned the bug into a feature
When Runway's real-time AI avatars kept drifting off-center, engineers spent weeks chasing a back-end fix. The one that worked wasn't in the model at all.
Ben Dickson
Writer's AI harness cuts token spend nearly 40% — without sacrificing accuracy
Writer also found task latency fell 44% under the harness, from 48 seconds down to 27, with success rates holding at 78% to 81%.
Ben Dickson
ACRouter picks the smartest AI model per task, beating Opus-only setups by 2.6x on cost
No single AI model wins every coding task. This self-learning router picks the right one each time — at a fraction of the cost.
Ben Dickson
Google's TabFM skips per-dataset training and still predicts on tables it's never seen
Data scientists spend weeks tuning hyperparameters and rebuilding pipelines for every new dataset. Google's latest model does the job in one API call instead.
Ben Dickson
Enterprises using multiple AI models are underestimating failure rates by 2.25x
You assumed mixing AI models would catch each other's blind spots. New research shows the math says otherwise — and by more than you'd think.
Ben Dickson