AI agents now do real work. They interact with software, call tools, retrieve information, update records, and carry out multi-step tasks with limited human involvement. As companies give these systems more autonomy, software reliability is becoming about more than whether an application stays online.

An application can run normally, its APIs returning successful responses, while an AI agent quietly does the wrong thing. It might misread a request, loop, call the wrong tool, or fabricate an identifier. A conventional monitoring system may have nothing to report, making these failures harder to catch than a system crash.

This creates a growing gap in the infrastructure used to monitor software. Lemma is targeting that problem. The startup is building what it describes as a reliability layer for AI agents, meant to help engineering teams find failures traditional monitoring can miss, understand why they happened, and act on what is actually happening in production.

Working versus working right

Traditional observability gives engineering teams visibility into an application's technical health, including response times, error rates, and API responses. AI agents introduce a different kind of risk. A system can execute every technical step correctly and still make a semantic mistake along the way.

Consider a support agent that returns a successful 200 status and shows a green dashboard, while misunderstanding the customer's question and supplying the wrong answer. There is no software error to investigate. The system has succeeded technically, but failed at the task it was meant to perform.

This blind spot matters more as agents move deeper into business operations. An incorrect action by an agent with access to customer records can cost far more than a bad chatbot reply. When agents can take actions across business systems, teams need to evaluate not only whether those actions executed, but whether they were appropriate in the first place.

That is the problem Lemma is designed to address. According to the company, its software analyzes production traces for patterns such as loops, broken tool calls, and misinterpreted requests, then connects recurring problems across traces to their underlying causes. The company says its software now analyzes more than one million agent traces each day.

From observability to reliability

Detecting an agent's failure is only part of the challenge. Engineering teams also need to understand what happened, identify the source of the problem, and determine whether a fix prevents it from recurring.

Lemma focuses on this post-deployment stage of AI reliability, monitoring how agents behave in live production. When it flags a problem, engineers can examine the affected traces and pinpoint whether the issue comes from a prompt, a tool interaction, or another part of the workflow. According to the company, the tool surfaces the issue in Slack with the production evidence attached, and suggests a fix connected to the coding tools developers already use, so teams can confirm whether the change actually worked.

This creates a feedback loop between an agent's real-world behavior and the engineering work needed to improve it. The goal is to make that process a routine part of development, ahead of any customer complaint.

Why AI agents create a different reliability challenge

Generative AI initially entered many businesses as an assistant, in which a person asked a question and remained responsible for what happened next. Agents take on more of that responsibility themselves, deciding which tools to use and adapting based on results. This independence is what makes agents useful, but it also creates failures that are hard to define in advance.

Unlike conventional software, where engineers can often anticipate specific failure conditions and write rules to detect them, agents operate across changing language, context, tools, and instructions. This makes it harder to anticipate every possible failure in advance.

The challenge is not simply detecting a crash. It is interpreting behavior across a workflow and determining whether the agent's actions matched the user's intent.

Lemma's founders,Jerry Zhang andCole Gawin, met while studying at the University of Southern California, where their first product, Clinicode, automated medical insurance paperwork for hospital staff in the Los Angeles area. The idea behind Lemma took shape later, during separate internships. Gawin was rewriting AI prompts by hand at a healthcare AI company, and eventually built a tool to automate that work. Zhang ran into the identical problem at his own internship. The two had independently hit the same wall, and it became the starting point for Lemma.

A growing infrastructure market around AI reliability

The challenge Lemma is addressing is part of a broader shift, as companies move beyond standalone models and chat interfaces toward systems that access business software with far less human oversight. Gartner predicts more than 40 percent of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value, and weak risk controls. Showing that an agent can work is no longer enough. Companies need to know how it behaves at scale and how quickly problems are found and fixed.

Lemma is betting that the next generation of AI tooling will need to judge agents on whether they did the right thing, a question that conventional observability often leaves unanswered, even when every technical signal looks fine.

Backing the reliability layer

The company was also part of Y Combinator's Fall 2025 batch and has been recognized among notable startups from the cohort. That recognition comes as developers and enterprises experiment with agents across customer support, software development, and research. Lemma's opportunity lies in building the infrastructure that surrounds the agents companies are already deploying.

As AI systems take on more work, reliability must account for more than whether software stays online. The infrastructure challenge is to make agent failures visible, explainable, and easier to address. Lemma is building its software around that distinction, aiming to help engineering teams understand and improve agent behavior in production.


VentureBeat newsroom and editorial staff were not involved in the creation of this content.