top of page

Measuring Leadership Development ROI: Getting to Kirkpatrick Levels 3 and 4 (the Esendia Approach)

Leadership development, Leadership development ROI, Leadership development Programmes

Most leadership development is only ever measured at Kirkpatrick levels 1 and 2: how participants reacted, and what they learned. Genuinely proving impact means reaching level 3, measurable behaviour change, and level 4, organisational results. That requires a robust baseline assessment of leadership behaviour, and often team-level measures like engagement or KPIs, taken before the programme begins so the same indicators can be reassessed afterwards.


Without a rigorous baseline, level 3 and 4 measurement isn’t possible; there’s nothing reliable to compare against. Esendia is a diagnostics-first leadership development consultancy that builds this baseline in from the start through its 5-Way Intelligence Scan, treating level 3 and 4 measurement as routine rather than exceptional.


Ask most organisations how their last leadership programme performed, and the answer usually comes down to how useful participants found it, and how satisfied they were that it was a good use of their time. That’s often genuinely earned: impactful facilitation, interesting new concepts, and engaging discussion with colleagues generate real engagement, build alignment, and help relationships across a cohort. Those are all worthwhile outcomes in their own right. But their impact on actual behaviour change in the leader themselves is likely to be minimal, and that’s the level at which most leadership development measurement stops.


Many organisations don’t even reach a genuine level 2. Part of the reason is practical rather than principled. A quiz or formal test of retention can feel too much like a school exam for a leadership-level audience, so it gets quietly dropped from programme design. What fills the gap instead is usually a leader recalling a conversation they had during the programme, which is fine in itself, but it isn’t a measure of whether they retained what was actually taught. A real level 2 test checks retention of the framework or model itself, not whether people enjoyed talking about it.


The Kirkpatrick Model for Leadership Development, Briefly


The Kirkpatrick model is the standard framework for evaluating training and development, and it works in four ascending levels.


Level 1 is reaction: did participants find it valuable, engaging, well delivered.


Level 2 is learning: did they actually acquire the knowledge or skills the programme intended to teach, usually measured through some kind of assessment or self-report at the end of a module. These two levels are the easiest to measure, which is exactly why they’re where most programmes stop. They can be captured with a feedback form and a short quiz, both administered on the last day.


Level 3 is behaviour: did participants actually do anything differently once they were back in their role.


Level 4 is results: did that behaviour change translate into something the organisation can point to, such as better team performance, improved retention, or stronger commercial outcomes. These are the levels that actually matter to a CPO or CEO deciding whether the investment was worth it, and they’re also the levels very few programmes, whether internally or externally delivered, routinely measure in a credible manner. That’s not just because most programmes never set up a proper baseline in the first place. It’s that a diagnostics-first approach usually isn’t part of the programme’s ethos to begin with. It’s often seen as too difficult, so it never gets designed in from the start, and a baseline nobody planned for is not one anybody can measure against later.


Far too many leadership programmes still only measure level 1, and some stretch as far as level 2. A satisfaction score and a knowledge check are genuinely easier to produce than a rigorous before-and-after comparison of behaviour or results, so that is where measurement quietly settles, even when the programme itself is well designed.



Why Leadership Development Programmes Rarely Reach

Level 3 or 4


A repeat 360 is actually one of the better ways to measure genuine behaviour change. It asks the people around a leader whether anything has actually shifted, rather than relying on the leader’s own account. The reason it so rarely gets used this way is that, without a diagnostics-first frame to the programme, the initial 360 was never run as a baseline in the first place. It was run as a personalisation exercise, a way to make a fairly generic programme feel more relevant to a leadership pool, not as the first half of a before-and-after comparison. Once that’s the purpose it was built for, the idea of a repeat 360 tends to get forgotten entirely, because there’s no data-driven development strategy asking for one.


Even where a repeat 360 is planned, not every 360 is built rigorously enough to support a fair comparison. If the questions change between rounds, or if rater selection isn’t handled consistently, or if the tool itself lacks the reliability to distinguish real change from noise, comparing “before” and “after” scores can be actively misleading rather than illuminating. (See The Rigour Behind Leadership Diagnostics for what a properly built 360 actually requires.)


And even with a well-built, repeatable 360, some programmes simply aren’t designed to produce a detectable behavioural shift in the first place. Programmes that are fundamentally content-focused (workshops, frameworks, getting a group of leaders onto the same page and generating good discussion at an intellectual level) are legitimate in plenty of situations, and often serve a genuinely different goal such as alignment rather than individual readiness.


But being legitimate isn’t the same as being built to change behaviour. Behaviour change takes dedicated effort and support: a carefully designed journey capable of shifting mindsets, self-belief, and behaviour consistently enough to be noticed by others, within a relatively short window such as a 12-month leadership programme. Many leadership programmes are not designed to achieve that. In those cases, even a well set up level 3 evaluation will not find a detectable change, not because the measurement failed, but because the programme itself was never built to move the needle it’s being asked to measure.


This is one of the quieter costs of skipping a rigorous diagnostic phase. It isn’t only that the programme design suffers from not understanding leaders well enough upfront. It’s that the organisation loses the ability to ever prove what changed. Six months after the programme ends, when someone on the exec team asks what the investment actually delivered, there’s no credible answer beyond satisfaction scores and anecdote.



What A Genuine Baseline Makes Possible


A diagnostics-first approach, done properly, produces exactly the kind of baseline this measurement requires, because the same rigour that makes a diagnostic useful for programme design also makes it usable for evaluation. If leadership behaviours were assessed with a valid, reliable tool at the start of the programme, against a clearly defined set of behaviours known to matter, those same behaviours can be reassessed at the end using the same instrument. The comparison is meaningful because both measurements were built to the same standard, at the same rigour, on the same defined behaviours.


That reassessment is what gets an organisation to level 3: not “did people say they behaved differently” but a measured shift in specific, predefined leadership behaviours, observed through the same lens before and after.


Pairing the repeat 360 with a final coaching session also does something measurement alone can’t. It gives the participant a genuine bridge from the programme back into their day-to-day role, rather than a hard stop when the last workshop ends. That session is where a leader gets to reflect properly on their own learning journey, recognise and celebrate the change the data has actually captured, and identify what development needs remain, or what new ones have emerged now that earlier gaps have closed. Without that final conversation, even a well-measured behaviour shift risks staying an interesting data point rather than something the leader carries forward deliberately.


Extending The Baseline To Level 4


Level 4 requires a further step, but the same logic applies. Where the initial diagnostic is extended to include organisational measures tied to each leader, such as their team’s engagement scores, relevant KPIs, retention within their team, and delivery against specific objectives, those measures can also be captured before the programme and reassessed afterwards. This is what turns a leadership behaviour story into a business results story: not just “this leader now behaves differently” but “this leader’s team is more engaged, or performing measurably better, since the programme.”


This is harder to do well than levels 1 through 3, because team-level results are affected by more than just one leader’s behaviour. Market conditions, team composition, and organisational context all play a role. But having a genuine baseline at least makes the comparison possible, which is more than most programmes can offer. Without it, level 4 measurement is guesswork dressed up as evaluation.


Noble Foods: A Leadership Development ROI Case Study


This is easiest to see in practice. Noble Foods’ Future Leaders Programme, run in partnership with Esendia, began with Esendia’s 5-Way Intelligence Scan, the same named, five-step methodology described in how to choose a diagnostics-first provider: Key Experience Assessment, 360-Degree Alignment against Noble Foods’ own leadership behaviours, Personality and Preference Mapping, Simulation Days, and Career Motivation Mapping. Together these established exactly the kind of baseline this kind of measurement requires.


Because that 360 was rigorous enough to repeat on a like-for-like basis, the follow-up comparison, taken 15 months later, is a genuine level 3 measurement rather than a satisfaction score: the overall rating moved from 4.14 to 4.23, statistically significant at p = .011 with a medium effect size (d = 0.46), and 75% of participants improved overall across 30 of 36 behavioural indicators, with no competency showing a significant decline. Peer ratings, typically the hardest to shift, improved from 4.02 to 4.22 (d = 0.82), direct evidence of behaviour change observed by the people around each leader, not self-reported.


Numbers like these are worth reading in context, because 360 data on a cohort like this comes with a built-in ceiling effect. These leaders were selected as high potentials in the first place, which means their starting ratings were already skewed toward the top of the scale. Moving a leader from “good” to “outstanding” is genuinely harder than moving one from “poor” to “average,” precisely because there’s so much less room left to improve into. A shift that looks modest in raw numbers, on a population that already started strong, represents real, hard-won movement rather than a small effect. That’s what makes a statistically significant, medium-to-large effect size shift across this particular cohort a genuinely impressive result, not a modest one.


Level 4 is naturally harder to isolate, but there’s a defensible version of it here too. 56% of participants progressed into senior roles within 12 months of completing the programme, and the cohort has seen 100% retention, both close enough to the leaders’ own behaviour to attribute with reasonable confidence to the development they went through.


Team-level retention and satisfaction with the development on offer moved positively as well, and again sit close enough to the leaders’ day-to-day behaviour to draw a reasonable line back to the programme. Noble Foods has also seen positive trends across a range of other business measures over the same period, but in line with good practice around measurement, those broader results shouldn’t be presented as caused by the programme alone. Too many other factors move a business’s numbers at the same time. The honest claim is that the trends are positive and sit alongside the behavioural evidence, not that the programme alone produced them.



The Question Worth Asking Before You Commit


Before signing off on a leadership development investment, it’s worth asking the internal or external provider directly: what will we be able to measure at the end of this, and against what baseline? If the honest answer is limited to feedback forms and a knowledge check, the programme is only ever going to prove levels 1 and 2, useful, but not what most CPOs actually need to justify the spend. If the provider can point to a rigorous diagnostic baseline that will be reassessed against the same standard at the end, that’s a genuine claim to level 3 and 4 measurement, and one worth holding them to.


It’s also worth asking two quieter but more fundamental questions. The first is whether the programme you’re building is actually designed to produce behaviour change and business impact in the first place, or whether it’s designed for other legitimate purposes, such as alignment, discussion, or shared language, that were never going to move those needles regardless of how well they’re measured. No measurement strategy can rescue a programme that wasn’t built for the outcome you’re now asking it to prove.


The second is about who you choose if you’re bringing in external support. Selecting a partner with a genuine diagnostics-first ethos means evaluation is baked into how they think from the outset, not an afterthought you have to specifically request. A provider like this drives the measurement agenda for you: building the baseline, planning the reassessment, and holding themselves to it, rather than you having to chase it into existence once someone asks for a business case.


None of this replaces the value of level 1 and 2 data. Knowing that participants engaged with the programme and retained what was taught still matters, and a programme that fails at those levels is unlikely to succeed at the levels above it. But treating levels 1 and 2 as the finish line, rather than the floor, is how so much leadership development spend ends up undefended when it’s questioned a year later. The organisations that can speak confidently about what a programme delivered are, almost without exception, the ones that took measurement seriously before the programme ever started, not after.


Frequently Asked Questions


What are Kirkpatrick levels 3 and 4?

Kirkpatrick level 3 measures behaviour change: whether participants actually do anything differently once they are back in their role. Level 4 measures results: whether that behaviour change translates into organisational outcomes such as team performance, retention, or commercial results. Esendia’s diagnostics-first approach is built to support both, because it establishes a rigorous baseline before a programme begins, which can then be reassessed afterwards.


Why do most leadership development programmes only measure Kirkpatrick level 1 or 2?

Because level 1 (satisfaction) and level 2 (learning) can be captured with a feedback form and a short quiz, while level 3 and 4 require a rigorous baseline set up before the programme begins. Most programmes never establish that baseline, because a diagnostics-first ethos usually isn’t part of the programme’s design in the first place. Esendia treats this kind of baseline measurement as a routine part of its 5-Way Intelligence Scan, not an afterthought.


How does Esendia measure leadership development ROI?

Esendia’s 5-Way Intelligence Scan assesses leadership behaviour, alongside personality, 360 feedback, experience, and career motivation, before a programme begins. Because this baseline is built to a rigorous, repeatable standard, the same behaviours can be reassessed afterwards to demonstrate genuine level 3 behaviour change, and organisational measures like team engagement or KPIs can be reassessed for level 4 business impact. In Esendia’s Noble Foods programme, this approach produced a statistically significant improvement in leadership behaviour.


About Esendia


Esendia is a diagnostics-first leadership development and succession planning consultancy. The 5-Way Intelligence Scan is built to support genuine before-and-after measurement, and level 3 and level 4 evaluation is treated as a routine part of programme design, not an optional extra bolted on when someone asks for a business case. Read the full Noble Foods case study, see how to choose a diagnostics-first provider for the full methodology, why a training needs analysis isn’t enough for what most measurement baselines are missing, or the rigour behind our diagnostics that makes this kind of evaluation possible.

bottom of page