Project-simulator — The Flight Simulator for Project Managers

Kirkpatrick's four levels: is your training measuring the right thing?

Published 2026-07-02 · The research behind Project-simulator

The course scored 4.6 out of 5 in the evaluation. Everyone was happy, the coffee was good, the presenter was funny. Three months later, everyone is working exactly as before. Was the training a success? Most organizations never measure deeply enough to be able to answer.

The American Donald Kirkpatrick formulated as early as the 1950s – and summarized in 1994 in the book Evaluating Training Programs – what has become the standard model for evaluating training: the four levels. The model is easy to understand and uncomfortably revealing to use.

What does the research say?

Kirkpatrick's model says that training should be assessed on four levels, each one deeper than the last. Level one is reaction: did the participants like the training? Level two is learning: do they actually know more afterwards – knowledge, skills, attitudes? Level three is behavior: are they doing anything differently in their work? And level four is results: is there any visible effect on the business – fewer incidents, better deliveries, happier customers?

The sting of the model lies in the fact that the levels don't automatically follow from one another. Participants can love a course without learning anything. They can pass a knowledge test without ever changing their behavior at work – often because daily routines, their manager, or time pressure leave no room for it. The most common criticism of the training industry is precisely that it measures level one, maybe two, and then hopes for the rest. The reason is blunt: reactions are cheap to measure with a survey in the last minute of the course, while behavior change requires someone to follow up weeks or months later, out in the organization.

For skills training – such as leading difficult conversations – level three is the decisive one: does the person do something differently the next time things heat up for real? This is often called transfer, meaning the carry-over from the training environment to reality, and it's where most training efforts fail.

What does this mean for you as a project manager?

When you commission or choose training for yourself or your team: ask how the effect will be measured beyond the satisfaction survey. What should participants be able to do afterwards that they can't do now, and how will you notice it in your projects? If the provider can't answer, that's a warning sign.

Apply the model to your own development too. After a course or a book, write down one concrete behavior you intend to change – "in the next steering committee meeting, I'll raise the risk before anyone asks" – and follow up with yourself after a month. That moves you from level two to level three. Without that step, there's a real risk the course becomes a pleasant footnote – you agreed with everything, remembered it for a week, and then carried on as usual.

And if you evaluate your projects' own competence development: don't count training hours. Count changed behaviors and, where possible, effects on delivery. It's stricter, but it's the only thing that counts – and it forces better choices of interventions from the very start.

How to practice this

Project-simulator is built with Kirkpatrick's levels in mind: the tool doesn't just measure whether you liked an exercise, but whether your assessed skills improve between sessions – and the feedback points to concrete behaviors you can bring to your next real conversation. Try it free for a week at project-simulator.com.

Want to practice this for real?

In Project-simulator you practice these conversations against an AI counterpart and get feedback grounded in research like this.

Try free for 7 days
Source: Kirkpatrick, D. L. (1994). Evaluating Training Programs: The Four Levels. Berrett-Koehler.