AI value · an interactive exploration

Built by Aidan Webb w/ Claude Code

Spend more time on the things that matter.

What AI-enabled engineering could be worth to an organisation like ours.

Scroll to begin

Seven ideas explain what is happening to software delivery.

Tap any point on the curve.

These seven studies explore the impacts of AI on software development in organisations like ours.

Section 3 · Illustrative model

Start with one engineer
and one task.

Every organisational claim about AI value is built from a unit this small. Explore the benefit per engineer, or scaled across a team.

£400
illustrative costs only
5.0
£25
illustrative costs only
Cost without AI £2,000
Cost with AI £1,725
Gross saving £300
− Tool cost £25
Net saving after tooling £275

Positive: this task pays for its tools and leaves headroom.

The honest range across credible studies is wide. That range is the argument for measuring our own.

Team view applies a fixed +20% productivity effect — a deliberately conservative point inside the study range. The variable you set is how many engineers benefit.

A measured baseline, repeated across hundreds of tasks, proves everything.

Section 4 · Illustrative model

Value and Adoption Exist on Different Curves

Every strategy deck shows adoption climbing — but research shows realised value tracking it only while the surrounding system keeps pace. Set the three levers and watch the space between them to explore how adoption without guidance and investment creates a value gap.

Index 0 6 12 18 24 Months
Unproven spend
Lever 1 · Adoption pace
Lever 2 · Enablement Engagement and Education
Lever 3 · Rollout structure leapfrog waves, 4 months apart
Token spend — cost of active usage Activity — tools rolled out and used Realised value — measured business outcome Unproven spend — the gap between them

Steady adoption with structured enablement keeps the gap modest.

Education and communication are a key variable in the value equation, and an enabler of organisational success.

Section 5 · Culture

Doubters are not the problem. We need the doubters.

Every organisation adopting AI hears objections. Most programmes treat them as blockers to be managed, but we need to treat them as critical signals to be understood and engaged with fairly.

01"Reviewing machine output all day is not a career."
What it gets right

If engineers become approval gates for a stream of generated code, the work hollows out. Job satisfaction and critical engagement are genuine delivery risks, and disengaged reviewers make worse decisions.

What it changes

Fulfilment gets measured, not assumed. Developer experience and engagement signals sit inside the value framework alongside cost and speed. The goal is engineers spending more time on design, architecture and judgement — and the measurement proves whether that is actually happening.

02"We would never design a team this way with humans."
What it gets right

Picture a team where one person briefs the work, everyone in between executes without engaging their judgement, and one person reviews the output. Nobody would call that a healthy culture. The thought experiment deserves an answer, not a shrug.

What it changes

Team design stays a first-class question. The operating model keeps humans in the roles where judgement, creativity and initiative live — and treats "how does this feel to work in" as evidence, gathered through the same check-ins and surveys used in the current trials.

03"AI doesn't interrupt."
What it gets right

People are proactive. They phone you with an idea. They push back mid-meeting. They have inspired days that change an architecture. None of that appears in a model that treats a team as units of code output — which means the model undercounts what humans contribute.

What it changes

The value framework measures outcomes, never lines of code. What a team delivers, what it prevents, and what it improves all count. What a tool typed does not.

04"Identical tools produce identical weaknesses."
What it gets right

If every system generates code the same way, it introduces similar patterns, similar mistakes, and similar attack surfaces. Human diversity has always been a quiet form of resilience — in codebases and in decisions.

What it changes

Diversity of review is a control, not a nicety. Human oversight, varied perspectives, and deliberate guardrails stay in the loop as autonomy expands. Constraint earns trust.

05"Where do the next mid-level engineers come from?"
What it gets right

If nobody writes code by hand, nobody learns the fundamentals, and in five years the market has prompt operators who cannot explain what the machine built. This is a succession problem the whole industry is quietly creating.

What it changes

Early-career engineers earn their stripes. The evidence shows juniors gain the most from these tools — and gain the most from learning fundamentals first. A deliberate talent pipeline is part of the adoption plan, because a responsible organisation invests in the skills of the people who will run it next.

Sourced · Cui et al., 2025
06"The enthusiasts need keeping in check."
What it gets right

Every transformation has people who are certain, early, and loud. Unchecked enthusiasm produces the unfilled dashboards and unproven claims that sink programmes.

What it changes

Sceptics get a formal seat. The measurement framework exists to hold the enthusiasts to the same standard of evidence the doubters already demand. A person who doubts every step of the plan but wants it to succeed is the most valuable reviewer a programme can have.

Culture change and technology change move at different speeds. Adoption succeeds when the plan respects both clocks.

Section 6 · Enablement

The shift does not stop at the codebase.

The same tools that generate code also read documents, draft analysis, test assumptions and build prototypes. Everyone who shapes software is affected — most of them before a line of code exists.

AR

Architects

System context generated and validated faster. Effort moved from documenting what exists to evaluating what should change. Architecture as a living, evidence-informed discipline rather than a periodic review.

PO

Product owners

A backlog conversation with the evidence already gathered. Prototypes that exist by the afternoon. Faster answers to "what would it take?" — and more time spent with customers instead of documents.

BA

Business analysts

Requirements traced against real system behaviour. Edge cases surfaced before development, not during testing. Analysis that starts from a draft instead of a blank page.

UX

Designers

Working prototypes in hours instead of handoffs. Ideas tested with real interactions before a sprint is committed. Design decisions made with evidence rather than argument.

Colleagues empowered with AI tooling not only move faster — they can participate more deeply across the organisation, giving new reach to their knowledge, experience and relationships.

Section 7 · Illustrative model

No new money.
Different choices.

Investment in AI capability does not require a bigger budget. It requires a decision about where existing delivery spend goes. This model explores one of those decisions: the balance between directly employed engineers and external delivery capacity.

illustrative — set your own assumptions
illustrative costs only
Current cost (index)
Rebalanced cost
Headroom created
Reinvestment capacity

This is a decision-support model, not a proposal. It shows the shape of a choice. The real version needs real baselines.

Industry context: most executives plan to maintain or grow third-party spend even as insourcing accelerates. The emerging model is a rebalanced blend, not a purge. Sourced · Deloitte, 2025

Every point of headroom is a choice: bank it, or reinvest it in the people who created it.

Section 8 · Strategy

Two years, four stages.

These stages are not locked gates. Some will run in parallel, some will emerge organically, but the key activities within each will roughly follow this timeline. Early-moving parts of the organisation go first, paving the way for the more traditional, entrenched areas to follow with confidence.

The organisations that reach stage four are the ones that took stage one seriously.

Measured

Section 9 · Framework

The rules we follow when proving the value.

01

Baselines before scale.

Before teams are enabled with AI tooling, measurement models and instrumentation must be in place. Building measurement into the enablement process, not after it, is what allows us to demonstrate value, learn and optimise.

02

Multiple classes of evidence, soft and hard.

We use hard metrics to demonstrate change; lead time, cost per outcome, quality signals. Soft measures; developer experience, qualitative feedback, engagement and colleague retention, tell us whether the change is sustainable.

03

Outcomes, never activity.

Lines of code, commit counts and acceptance rates measure motion, not value. Only metrics directly related to outcomes belong in a value framework — lead time, quality, cost per unit of delivery, engagement. Code is not an asset. Code is a liability. Working software is the asset.

04

Scrutiny at every point.

We want to make every claim open to challenge, and every metric published with its method. The people closest to the work — including the sceptics — are invited to stress-test the evidence. By open-sourcing our approach we build credibility and trust.

05

Communicate in human terms.

Value doesn't just have to be proved — it has to be heard. We communicate what we learn in language that builds trust, helps colleagues join us on the journey, and cuts through the noise. Education, engagement and clear storytelling are just as critical to measurement and modelling — they ensure our organisation's purpose stays aligned, and we move together towards our AI-enabled future.

Section 10 · Build ledger

How this page was made.

This page was designed, written and built by Aidan Webb, working with AI tools — the same class of tools the page describes.

Design and copyShaped across several working sessions. The structure came first, then the copy, then the interaction layer — each pass reviewed before moving on. Headlines and figures went through plenty of rewrites as the underlying evidence was checked and rechecked.
Build toolClaude Code (agentic). The build ran from a single HTML file rather than a framework, with the tool generating and refining markup, styling and interaction logic within the same session. It handled the code; the decisions about what the page needed to say stayed with me.
Human judgementThe thinking behind this page was shaped by conversations with colleagues across AI and engineering — their experience, challenges and perspectives informed the position and logic throughout. This is a starting point, not a conclusion. The arguments here are open to critique, and the intent is to keep learning and collaborating with the wider business as this work moves forward.

Baseline > Enable > Measure > Scale