AI agents often struggle to assess their own progress and decide what to do next. They can waste compute resources following failed approaches, or they might overlook promising results or stop before completing a task. These problems become more costly as agents take on longer and more complex workflows.

Researchers at Meta Superintelligence Labs have developed a framework called agentic meta-reasoning that addresses this challenge by enabling the agent to have a dedicated reasoning stream on what to do next. The framework enables agents to evaluate their progress, explore alternative approaches, and direct resources toward the most promising next steps.

On ProgramBench, a software reconstruction benchmark, the framework achieved a 71.5% average hidden-test pass rate with GPT-5.5, compared with 58.0% for Codex in the researchers' evaluation.

Unlike self-improving harnesses that optimize the software and configurations powering agents across runs, meta-reasoning improves how agents direct their work during a run. This gives enterprise developers another way to improve agent performance without modifying or fine-tuning the underlying models.

The challenge of directing agentic work

Consider a coding agent implementing a software component. It has written the core functionality and passed several tests, but some edge cases remain unresolved. Should it investigate the failures, rewrite part of the implementation, run additional tests, try another approach, or stop? Each option consumes resources, and a poor decision could waste the remaining budget or damage an otherwise working implementation.

The researchers refer to this challenge as metacognitive control, which they describe as "assessing one’s own progress and using that assessment to decide what to do next."

For an AI agent, metacognitive control means deciding which intermediate results to trust, what previous work to build on, and when to allocate more compute resources. A bad decision can continue wasting compute on a failed approach or cause the agent to abandon a correct solution it has already found.

Existing agents often interleave these decisions with task execution. This means they choose their next action in a single step based on an accumulating history of actions and interactions with their environment. As the task progresses, relevant information can become buried in previous attempts, tool outputs, and errors.

Traditional inference-time scaling techniques have their own constraints. Approaches such as generating multiple candidates, critique-and-revision loops, and predefined search trees typically determine their computational structure in advance. They determine the number of attempts, verification steps, or refinement rounds before the problem reveals which computations would be most valuable.

The Meta researchers argue that decisions about what to compute next “deserve an agentic reasoning process of their own.” 

“That deliberation has a structure of its own, consolidating what the run has established, exploring what could be done next, and assessing what each option is worth,” they write.

How agentic meta-reasoning works

The agentic meta-reasoning framework proposed by Meta researchers enables an agent to reason about and act on its own inference process. 

The framework creates a recurring cycle: perform work, evaluate the results, identify promising next steps, and assign additional computation. Previous findings remain available throughout the process, allowing the agent to build on earlier attempts without repeatedly processing its entire execution history.

Agentic Meta-Reasoning

Agentic Meta-Reasoning (source: arXiv)

The researchers implement this approach in an inference-time harness called the Meta-Reasoning Agent. It separates two components: “workers” that perform the actual tasks, and a controller that decides what work should happen next.

Workers can be individual large language model (LLM) calls or coding agents that use tools to inspect files, modify code, and run tests. The controller can inspect their results, record its own findings, launch new workers with targeted instructions, and decide when to stop.

The controller follows a four-stage reasoning cycle:

  • Assess: The controller examines new worker results and updates its understanding of the task. For the coding example, it might record that the main functionality works, some tests fail, and error handling remains unverified.

  • Propose: Based on this assessment, it generates possible next actions, such as investigating failing tests, implementing missing functionality, or independently verifying existing code. At this stage, it explores options without restricting them based on the remaining budget.

  • Evaluate: The controller weighs the proposed actions against the available computational budget. It might decide that fixing a known failure would bring more value than rewriting a working component. 

  • Dispatch: The controller gives instructions to workers, and it uses relevant findings from previous attempts to provide context to the workers. Alternatively, it can stop and submit an existing result.

Each stage can involve its own model calls and memory operations, allowing the controller to investigate before making a decision.

A key part of this architecture is persistent artifact memory. Worker outputs and controller notes are stored as artifacts with unique identifiers. The controller maintains a compact assessment of the task's status and retrieves detailed artifacts only when needed.

These artifacts also form a graph of computation. Whenever the controller provides an earlier artifact as context for new work, the system records a connection between them. For example, a test report might lead to a targeted repair, which another worker subsequently verifies.

The graph reveals how the agent explores alternative approaches and builds on previous results. It also provides a way to analyze whether the agent is making productive use of its computation rather than generating disconnected attempts.

Meta-reasoning in action

The researchers evaluated the framework on four benchmarks: IMO ProofBench-Advanced for mathematical proofs, ARC-AGI-2 for abstract visual reasoning, LongCoT-mini for long-horizon reasoning across several domains, and ProgramBench for reconstructing programs from documentation and executable references.

They tested three frontier models: Gemini 3.1 Pro, GPT-5.5, and Opus 4.8. Their main baseline was a Direct Control Agent, a variant of the Meta-Reasoning Agent that makes control decisions in a single step over its accumulated history instead of using a separate controller. They also compared the framework against existing research harnesses and coding agents, including Codex, Claude Code, and recursive language models (RLM).

At the largest tested budgets, meta-reasoning achieved higher scores in all 12 matched comparisons. 

Agentic meta-reasoning results

Agentic meta-reasoning results (source: arXiv)

One of the most interesting findings was how agents responded to larger computational budgets.

When the ProgramBench allowance increased from 400 to 1,200 model calls, meta-reasoning with GPT-5.5 improved from 64.1% to 71.5%, while direct control remained near 64% and used only about 18% of its available calls at the highest setting.

The artifact graphs component plays a key role in enabling this behavior. Meta-reasoning generated more intermediate results and more connections between them. In one ARC-AGI-2 example, direct control produced six independent attempts and four shallow follow-ups, while meta-reasoning explored more branches and connected later work to previous results.

Meta-reasoning graph building

Meta-reasoning uses advanced graph-building techniques (source: arXiv)

Implementing agentic meta-reasoning

The researchers have not provided an official ready-to-run implementation in the materials accompanying the study, but the paper includes detailed specifications of the controller prompts, worker instructions, memory interfaces, and tools.

Developers can use these specifications to experiment by adding a separate control loop to existing agent systems, maintaining persistent records of intermediate work, and explicitly evaluating the cost and benefit of proposed actions.

The ProgramBench implementation offers practical lessons for coding agents. Workers commit their changes to Git, and the controller can inspect the repository through a read-only interface to verify their claims. Its running assessment distinguishes implemented, broken, unverified, and unexplored functionality. The framework also restricts parallel code modifications in the shared workspace to prevent workers from overwriting each other's changes.

However, meta-reasoning introduces additional overhead and LLM calls during the control cycle. At smaller budgets, the framework sometimes performed worse than direct control before overtaking it as more computation became available.

For enterprise AI teams, the findings suggest that the ability to manage computation should become an explicit part of agent evaluation and optimization. As agents tackle longer workflows, they should be able to pause and determine how to best allocate their resources.