
The talk begins with a company where generative AI (GenAI) is already in production, but nothing is tested. Every change is risky, and nobody can tell whether the system has improved or simply broken something.
Marcin Szymaniuk explains how to bring structure from three perspectives:
He then covers how to evaluate non-deterministic systems:
Marcin also discusses the trade-offs: not everything should be automated, and not everything needs perfect scoring. Testing GenAI is not about achieving perfect correctness – it is about building enough confidence to ship, iterate, and avoid breaking what already works.