
Everyone is shipping tools and skills for agents, but almost nobody can tell you whether those tools and skills actually help. “It felt better in the demo” is not a metric, and with agents, that impression is often wrong: something can look great in one transcript and quietly fail much of the time in aggregate. This deep dive shows how to put real numbers on the effectiveness and efficiency of agent-facing tools and skills. The order matters: effectiveness first, efficiency second, because reducing costs is worthless if the agent never reaches the goal. Using real evaluation infrastructure as a worked example, Michael Hablich covers how to run controlled experiments with and without the tools or skills, scored against explicit checks; how to read the results to decide where to invest; an honest case in which adding guidance made things worse; and the efficiency levers and their trade-offs that reduce costs once quality is maintained. Attendees will leave able to design an evaluation for their own agent tools instead of trusting a good-looking demo.