Marcin Szymaniuk

CEO | Senior Data Engineer | International Conference Speaker
TantusData
Poland

About

Specialising in helping clients monetise big data since the early 2000s, Marcin Szymaniuk leads a team of seasoned data engineers with expertise in data engineering, machine learning (ML), machine learning operations (MLOps), and cloud technologies.Marcin is adept at solving both non-standard challenges and everyday problems that require fast, practical solutions. His experience spans a wide range of industries and project sizes, with a strong focus on artificial intelligence (AI), ML, and deployment strategies.He has presented at numerous industry events, including Infoshare, J On The Beach, Devoxx, Huawei Eco-Connect Poland 2023, Berlin Buzzwords, Codestar, GeeCON, and Java Day Istanbul.
Talk

Marcin Szymaniuk | From No Tests to Trust: Evaluating End-to-End GenAI Systems

AI Evaluation, AI Testing, RAG, LLM-as-a-Judge, Synthetic Testing, AI Reliability

The talk begins with a company where generative AI (GenAI) is already in production, but nothing is tested. Every change is risky, and nobody can tell whether the system has improved or simply broken something.

Marcin Szymaniuk explains how to bring structure from three perspectives:

  • Developer: how to build evaluation loops, datasets, and basic tests for retrieval-augmented generation (RAG), retrieval, and document pipelines
  • Team lead/manager: what is realistic to test, what is expensive, and how to plan work around uncertainty
  • Business: how to move fast without guessing and how real user feedback is more valuable than any single metric

He then covers how to evaluate non-deterministic systems:

  • RAG and document-quality metrics, including relevance, faithfulness, and completeness
  • Conversation testing using synthetic scenarios
  • Using large language models (LLMs) as judges when rule-based checks fail

Marcin also discusses the trade-offs: not everything should be automated, and not everything needs perfect scoring. Testing GenAI is not about achieving perfect correctness – it is about building enough confidence to ship, iterate, and avoid breaking what already works.

2026-11-26
10:10
10:55
Data Meets AI
4 TICKETS FOR A PRICE OF 3