
Multi-agent systems shift the evaluation challenge from individual model outputs to the integrity of the coordination layer. When a supervisor agent delegates a task with flawed context, the error propagates and amplifies through the chain, leading to distributed hallucinations that bypass traditional end-to-end testing. Debugging these systems requires treating them as distributed networks rather than isolated large language model (LLM) calls.
In this session, Oleksandra Bovkun covers tracing, evaluation, and governance for multi-agent systems and how to ensure that agentic workflows remain reliable, transparent, and secure at scale. This session provides a technical deep dive into solving observability challenges in complex agentic workflows. She explores how to move from black-box testing to a transparent architecture using MLflow tracing.