A contractor, a rail project owner and a delivery authority walk into a (we)b(in)ar: Part I
Rail Projects
Project Assurance
AI-SRA

A contractor, a rail project owner and a delivery authority walk into a (we)b(in)ar: Part I

We brought together three different parties to the rail overrun problem - owner, contractor and delivery authority- for one webinar on AI-led risk. Read on for the highlights!

A contractor, a rail project owner and a delivery authority walk into a (we)b(in)ar: Part I
Written by
Colin Myer
nPlan evangelist and content creator. Passionate about major projects and the role they play in driving economic growth and raising standards of living. Ambitious infrastructure projects are awesome!

Major rail programmes overrun on schedule and cost by an average of 45% (source: Bent Flyvbjerg). nPlan has been talking and writing about the causes of this execution failure for a while now: optimism bias inflates forecasts, availability bias leads teams to prioritise risks they've encountered in the past, salience bias drags attention away from critical-but-dull risks and towards 'flashy' risks, and strategic misrepresentation means that plans and reports are designed to keep stakeholders happy rather than to reflect reality.

Earlier this summer we took a unique opportunity to bring together three people who are actively trying to counter the rail project overrun problem - using nPlan's AI - for a webinar. And there was an appealing twist to the line-up as well: each of our speakers represented a different party to the problem: Wes Cadby leads risk on the £14bn Transpennine Route Upgrade, from the owner's side. Paul Bradley runs planning at John Holland, a Tier 1 contractor. And Paul Reichl works at Victoria's Infrastructure Delivery Authority, the government body overseeing a A$120bn 'Big Build' of more than 200 road and rail projects.

We were lucky enough to get three different but equally important perspectives on the same multi-billion-pound problem. Our CEO, Dev Amratia, put the same questions to all three - and in this blog I'm going to look at some of the highlights from that webinar and discuss what was said, where our speakers views converged - and where they diverged. Let's go!

Life before AI: three problems, one root cause

Dev started with the obvious question: before AI, what was the hardest part of staying on top of delivery risk?

Paul Bradley's frustration was with the standard tool of the trade, Monte Carlo risk analysis:

The most difficult thing would've been to answer the question about traditional Monte Carlo risk analysis - if it really works, why does it never work? Because, you know, so often the outcomes were outside any predicted range.

The Monte Carlo method runs thousands of simulations on the probabilities and impacts your team gives it, but those inputs are essentially educated guesses, shaped by bias and often by whoever is most forceful in the room. Garbage in, garbage out, as the saying goes.

For Wes Cadby, the issue was speed:

The Transpennine Route Upgrade program moves at such a speed. We are literally building and commissioning assets on a weekly basis.

On a programme moving that fast, a traditional risk assessment is out of date by the time it's finished. It can take months to pull a programme-wide picture together, by which point the programme has already moved on.

These two problems compound each other. A method you can't fully trust is bad enough; a method you can't fully trust and can't produce quickly enough to be current is worse.

Paul Reichl, coming at it as the client rather than the delivery team, raised something different:

The challenge isn't often the absence of information, it's rather obtaining a system-level view across all of the things that are going on.

When a delivery authority oversees a programme delivered by many contractors and partners, each with their own priorities and reporting, crucial information filters up late and fragmented. There's no shortage of data; the difficulty is assembling it into one coherent, current picture.

The aha! moments: when AI proved its worth

Next, Dev asked our speakers about the moment each of them saw AI do something that convinced them the AI hype was (at least partially) justified.

For Paul Bradley, the change was in how his team talked about risk:

It stopped being a you and an I conversation, a risk. It started being a we conversation with risk.

In a traditional risk workshop, someone puts a number on a risk and then has to defend it, and people rarely like backing down from a position once they've taken it. What Paul found was that when the analysis came from nPlan rather than from a colleague, the room stopped debating individuals and started discussing the risk together. The conversation became collaborative rather than confrontational.

Paul Reichl's moment was about sheer scale:

AI could connect thousands of pieces of information across a program…in a way that no individual or team realistically could.

A single person can hold a handful of dependencies in their head; a large programme has hundreds of thousands of them, and the signals that matter are often buried somewhere in that mass. What struck Paul was that AI can hold the whole picture at once, and then let anyone interrogate it - putting insight that used to require a room full of specialists within reach of the wider team.

For Wes Cadby, the realisation was about what the tool was fundamentally doing:

This is new capability, not faster existing capability.

For the first few months, his team used AI to check their existing work - running its output alongside their traditional assessments to see if the two agreed. The shift came when they started asking it a different kind of question: not "is our plan right?" but "why might this programme fail?" Using AI to probe the future, rather than to double-check the past, was something they simply hadn't been able to do before.

Winning hearts and minds: where they agree

Then the hard part. Choosing a tool is the easy bit; getting people to genuinely use it is not. Dev asked what it actually took to get AI adopted (and trusted).

Wes Cadby was honest about how uncomfortable the conversation can be:

If I say to a very competent, very experienced project delivery professional, "My AI says this," it's not an easy conversation to have.

Telling an expert with decades of experience that a machine has reached a different conclusion is a delicate thing. Wes's approach was to run the AI alongside the established, trusted process rather than in place of it, so the team could see the two side-by-side and build confidence in the AI's output over time, rather than being asked to take it on faith.

Paul Bradley found that adoption came down to the tool proving its worth:

It had to prove itself. You know, it's part of the reason we trialed it. Like, if it wasn't useful, it would've been rejected.

People at project level are busy and pragmatic, and won't persist with something that slows them down. What won them over was being able to ask a plain question - like what had materially changed between one version of a programme and the last - and get a useful answer back in moments rather than days.

Paul Reichl, from the delivery authority side, stressed that adoption rarely happens on its own:

You need sometimes a push from multiple directions to get things to happen.

New technology seldom fits neatly into existing budgets or processes, so it needs deliberate backing to get going: a pilot to prove the concept, an innovation-funding mechanism to carry the early risk, and champions inside both the client and the contractor pushing from their respective sides. Get those aligned and a trial can turn into embedded practice.

Watch the full webinar

As you can hopefully tell from reading this blog (and watching the embedded clips), this webinar was full of incredibly rich insights into the ups and downs of embracing AI-led assurance. We'll shortly be publishing part II of our highlights recap from the webinar on this blog - but if you can't wait until then, the whole webinar is available to view on-demand right here. It really is recommended viewing.