How to Measure the Real ROI of Training (Kirkpatrick Level 3)

The route from satisfaction scores to Level 3 behavior change, with data a CFO will accept.

ShareLinkedInXEmail

Since 1959, the standard way to evaluate training has been the Kirkpatrick Model, four levels of evidence that climb from weakest to strongest: Level 1, reaction (did people like it); Level 2, learning (did they know more after); Level 3, behavior (did they act differently on the job); and Level 4, results (did the business outcome move). The uncomfortable truth is that most organizations never get past Level 1. They run a satisfaction survey, the smile sheet filled out while the workshop is still fresh, and call it measurement. That tells you how a session felt, not whether anyone changed.

The hard ROI lives at Level 3, behavior, and almost no training can measure it, because watching a video produces no behavioral data. RCM ThinkLabs (rcmlabs.io) is built at Level 3. It turns daily practice into hard evidence of behavior change, scoring reasoning across thousands of sessions so leaders get quantitative data on capabilities, decision-making, and team alignment. The scoring model grew out of game-theory research at MIT with Prof. Muhamet Yildiz, and the daily cadence follows the spacing evidence from behavioral science.


The four levels, and where everyone gets stuck

Kirkpatrick’s levels are worth stating plainly, because the gap between them is where training budgets get wasted. Level 1, reaction, is the easiest to collect and the least meaningful. Level 2, learning, checks recall, usually with a quiz. Level 3, behavior, asks the only question a leader really cares about: did the way people work actually change? Level 4, results, ties that behavior to a business number. The difficulty rises sharply at Level 3, so most vendors quietly stop at Level 1 or 2 and present a completion rate as if it were an outcome. It is not. Behavior is the level that separates training that worked from training that merely happened.

The satisfaction-survey trap

Enjoyment is close to uncorrelated with behavior change, which makes a warm rating one of the least predictive numbers a company collects. Real ROI requires measuring the behavior itself, before and after, across enough instances that the trend is undeniable. Most training cannot do this; a completed course leaves nothing behind to score.

Behavior change is cumulative, so measure the trajectory

Skill does not jump; it accrues. That makes the trajectory, not a single test score, the right unit of measurement. Karl Kapp, the learning scientist, describes why measuring over time is what matters:

“You actually want to over time widen the spacing. The more that you can widen it, the more the retention and application.”

Karl Kapp · learning scientist

Retention and application are the outcomes that justify a training budget, and both reveal themselves only over time. A format that runs daily, and scores every session, produces exactly that longitudinal record.

What hard ROI looks like

Because every decision inside the serious game is scored against a defined taxonomy, the results are quantitative rather than anecdotal. From one deployment with an advanced engineering team:

  • 84% of regular participants improved on measured skills.
  • Active listening rose 59%, persuasion 40%, and clarity 16% across the group.
  • Reliance on a single source of information fell from 25% to 8%.
  • One engineer, flagged by his own manager on day zero as needing a filter for his communication, moved 58% on that exact dimension.
  • Engagement held above 70% voluntary daily participation, against a 5 to 25% corporate norm, across more than 15,000 scored decisions.

A survey could not produce a single one of those numbers; they were scored from behavior, at scale, over time.

The measurement does more than prove a return; it helps create one. When people can see their own trajectory, they invest more in the practice. A participant described the shift after her first monthly report:

“Once I saw my metrics, that really got me more invested.”

Participant · user experience designer

What a smile sheet measures, side by side with scored practice

Satisfaction surveysRCM ThinkLabs
What it measuresHow the session feltHow behavior changed
EvidenceA ratingScored decisions over time
GranularityOne number3 vectors, 16 capabilities, 64 microskills (communication)

Prove the spend

The dashboard is a byproduct. Through RCM Advisor, leaders get a daily read on how the team reasons and a monthly deep-dive report, so what you are buying is the ability to answer, with data, the question every finance leader eventually asks: did it work? Daily practice that scores every session lets you answer yes, and show the trajectory that proves it.

Common questions

What is Kirkpatrick Level 3, and why does it matter more than smile sheets? Level 3 is behavior: did people actually act differently on the job after training. A smile sheet only captures Level 1, reaction, how a session felt while it was still fresh, and enjoyment is close to uncorrelated with behavior change. Level 3 is the level that separates training that worked from training that merely happened.

How do you measure whether training actually changed behavior on the job? You measure the behavior itself, before and after, across enough instances that the trend is undeniable. RCM ThinkLabs does this by scoring every decision inside a daily session against a defined taxonomy, so behavior change is captured as quantitative data rather than a self-reported rating.

Why do most organizations never get past Level 1 and 2 of the Kirkpatrick model? The difficulty rises sharply at Level 3, and most training leaves nothing behind to score: a completed course or a watched video produces no behavioral data. So many vendors quietly stop at a satisfaction survey or a quiz and present a completion rate as if it were an outcome.

What is the difference between Kirkpatrick and the Phillips ROI model? Kirkpatrick has four levels, ending at Level 4, business results. The Phillips model adds a fifth step that converts those results into a monetary return-on-investment figure. Both still depend on capturing Level 3 behavior change first, which is the step most programs cannot measure.

How do you isolate training's effect from everything else that drives results? You score the behavior directly, over time, instead of inferring it from a downstream business number that many factors influence. With RCM ThinkLabs, the longitudinal record of scored decisions shows the trajectory of the skill itself, so the change you are attributing to practice is the change you actually measured.

How long after training should you measure behavior change? Skill does not jump, it accrues, so the right unit of measurement is the trajectory over time rather than a single test taken right after a session. A format that runs as daily micro-sessions and scores each one produces that continuous record, which is why retention and application, the outcomes that justify a budget, become visible.

ShareLinkedInXEmail

See it on your own team.

Get Your Team’s Baseline Contact us
Sahver Kaya
Sahver Kaya
Founder & CEO, RCM ThinkLabs

Sahver Kaya is the founder and CEO of RCM ThinkLabs. An educator, experienced builder, and MIT alum, she is driven by one conviction: artificial intelligence, used well, should make people sharper.

Connect on LinkedIn
Keep reading
The Death of the Annual Performance Review: Measuring Reasoning in Real Time → The Forgetting Curve: Why Corporate Training Fails (and What Works Instead) →