The 'Two Poisons' Problem in AI Testing

VB Transform 2026 Highlights LLM-as-a-judge vs. human-in-the-loop oversight: The great debate on model validation As agents move from prototype to production, the evaluation bottleneck has become one of the biggest hurdles. Can we truly trust an LLM to grade another LLM, or are we simply building a hall of mirrors? In this session, panelists weigh all sides of model validation, testing and putting trusted agents into production. They’ll get into the necessity of LLM-as-a-judge for real-time, automated observability at scale vs. the real world challenges of poor agent responses and the indispensable role of golden-set validation and human-in-the-loop oversight.