← Back to the BlueprintThe proof layer

Turn estimated opportunity into verified impact

The proof layer combines adoption telemetry with delivery, quality, and cost data against a pre-rollout baseline, turning estimated opportunity into evidence leaders can act on.

See the impact console ↓
01 · The evidence model

A defensible claim connects adoption to outcomes

Leading signals show whether adoption is taking hold, at what capability level, and at what cost. Operational outcomes show what changes in delivery, quality, and flow. A pre-rollout baseline connects the two.

leading signals

adoption + capability + cost

×

operational outcomes

delivery + quality + flow

=

result

decision-ready evidence

Decision-ready example

Adoption reached 78% of eligible engineers and cycle time p85 fell 21% against the Q1 baseline. Quality stayed visible: change failure rate is one point above baseline and under review.

Every clause maps to a console metric and a named source, making both verified impact and open questions visible.

Trustworthy comparisons start with a pre-rollout baseline: six to twelve months of git history, incident records, and ticket flow, measured and handed to you immutably on day one. It anchors every later comparison, costs almost nothing to capture, and cannot be recreated afterwards.
02 · The console

One page a CxO can read in thirty seconds

Every figure below is computed from the controls above — team, period and baseline. Change any of them and the whole console recomputes, quality guardrails included, even when they move against us.

Delivery Impact Console

ILLUSTRATIVE CLIENT · AI AUGMENTATION PROGRAMME · Q3

SAMPLE DATA

Adoption

78%

▲ 71 pts62 of 79 eng

Spend / eng / mo

$186

▲ $24Q1 $162

Cycle time p85

4.1d

▼ 21%Q1 5.2

Change failure rate

⚠ WATCH

11%

▲ 1 ptQ1 10

Rollout & response

ADOPTION vs CYCLE TIME · BY COHORT
Wave 1 — 4 teamsWave 2 — 3 teams

Adoption %

048.396.6Workshop · Wave 1Wave 2 rollout

Cycle time p85 (days)

3.94.85.8Q1 baseline 5.2dWorkshop · Wave 1Wave 2 rolloutJanFebMarAprMayJunJulAugSep

Wave 2 holds flat through Wave 1's improvement and bends only after its own June rollout — the staggered schedule acts as a control group.

Convergent evidence

4 / 5 AGREE

Speed

cycle time p85 · time to merge · review latency

IMPROVING

Volume

throughput · items completed

IMPROVING

Quality

failure rate · defects · reopen · MTTR · complexity

MIXED

Adoption

active engineers · capability usage

IMPROVING

Perceived

developer survey · +0.8 pts

IMPROVING

Recomputed live from the panels below: a family counts as agreeing when its member metrics move in the improving direction against the selected baseline.

Delivery

GITHUB · 90d

Cycle time p85

4.1d

▼ 21%

Throughput

3.4PR/eng/wk

▲ 17%

Review latency

5.2h

▼ 32%

PR size

312LOC

▲ 18%⚠

Rising. Bigger changes normally slow review — cycle time improved in spite of this.

Quality guardrails

CI · INCIDENTS · JIRA

Change failure rate

11%

▲ 1 pt⚠

Above baseline for 4 weeks. Under review.

Escaped defects

3.1/release

▼ 18%

Reopen rate

6.4%

▼ 2 pts

MTTR

47min

▼ 11%

Cyclomatic complexity

6.2% vs base

▲ 6.2 pts⚠

Rising since adoption began. The main long-term risk signal.

Revert rate

2.1%

▼ 0.5 pts

What we built

CAPABILITY PORTFOLIO · 30d
CapabilityUsesSuccessCost/useVerdict
code-review-agent1,24094%$0.11★ KEEP
migration-helper38091%$0.34★ KEEP
test-generator4461%$0.92✕ KILL
doc-writer3——✕ KILL

A capability nobody invokes is not a neutral outcome — it is a failed deliverable we shipped. Reporting the kill is the point.

Where the spend goes

BY CAPABILITY · 30d
code-review-agent$2,140
migration-helper$710
test-generator$392
unattributed / ad-hoc$281

How to read this

Design

Staggered rollout. Wave 2 serves as a control until June.

Baseline

Q1, captured from git history before the first workshop.

Confounders

PR size rose 18% over the same window; team composition changed on two teams.

Coverage

7 of 9 teams. Platform and Data are not yet instrumented.

03 · The claim ladder

What we can prove, and what we refuse to claim

Publishing this is the point. Being visibly disciplined about the limits is what makes everything above the line believable — and it is the fastest way to tell a measurement practice from a marketing one.

A

Directly measured

High confidence · straight from telemetry

  • How many people use it, how often, and who never adopted
  • What it costs — per person, per team, per month
  • Which capabilities we built are used, and which are ignored
  • Which tools fail, how often, and with what error
  • Which MCP servers are silently broken
  • Where the budget actually goes, by skill and by model
B

Measured, with real caveats

Directional · never quoted as precision

  • AI-assisted code share — deliberately conservative, a floor and not a measurement
  • Suggestion accept rate — a good friction signal, a poor quality signal
  • Lines of code — a real number, near-worthless as value
  • Model duration — model time, not engineer time. Never presented as the latter
C

Requires your systems

Where the value claim is actually made

  • Cycle time and lead time for change — from your git history
  • Throughput per engineer per week
  • Change failure rate, escaped defects, MTTR
  • Rework: reverts, hotfixes, reopen rate
  • Review latency and review load
  • Whether engineers feel more effective — survey, at baseline and at day 90
D

We will not claim this

Not defensible · not offered

  • "We saved you N FTEs" — requires assumptions nobody can evidence
  • "AI wrote N% of your codebase" — the attribution is a deliberate undercount and does not support the reading
  • "Quality improved because of AI" — correlation at best without a controlled rollout
  • "The team is N% faster" — unless measured on time, with a control cohort, paired with quality, over 90+ days

Confounders we name before you do

Novelty effect

A 14-day read is marketing, not measurement. We report at 90 days.

Self-selection

Enthusiasts adopt first and were already fast. Cohort comparison alone cannot separate the two.

Team composition

People join and leave mid-window. Metrics normalise by active contributors, never by headcount.

Seasonality

Q4 and holiday windows distort throughput in both directions.

Concurrent initiatives

If a platform migration lands in the same quarter, it owns part of the result.

Observation effect

Measured teams behave differently while they are being measured.

04 · The cadence

Ninety days, with the honest checkpoint at day 14

The schedule exists to stop the two failure modes that destroy credibility: declaring victory too early, and never checking whether the capabilities we shipped are used at all.

2 weeks before

Baseline capture

Git, incident and ticket history measured. Metric set and success criteria agreed in writing, before anyone is trained.

Day of the workshop

Workshop and rollout

Telemetry live on day one, with the legal clearance already signed off.

Day 14

Adoption check only

Diagnostic, never a result. Anyone reporting an outcome at two weeks is measuring novelty.

Day 30

Capability portfolio review

Keep, fix, promote or kill each thing we built. First honest cost read.

Day 90

The real measurement

Outcomes against baseline, survey re-run, families counted. This is the number that gets quoted.

Quarterly

Ongoing review

Capabilities decay, usage drifts, new teams onboard. Retire what stopped earning its keep.

The day-90 report

Two pages plus an appendix. Five sections, always the same five.

  1. 01Adoption
  2. 02Capability portfolio
  3. 03Economics
  4. 04Outcomes vs baseline
  5. 05The next 90 days

It reports at least one negative finding. A report with no bad news reads as marketing, and gets read as marketing.

05 · Field evidence

The most useful public dataset says the bottleneck may not be the code

Public research, not a client of ours.

Bartosz Ocytko, "Agentic Engineering at Zalando: a snapshot", Zalando engineering blog, 14 August 2026.

Read the original ↗

Reported across more than 250 engineering teams, with adoption measured through a gateway proxy serving roughly 2,000 monthly active users on six small pods — this is modest infrastructure, not a platform programme.

33%

of PRs auto-approved

A risk-based approval bot handles the low-risk third automatically.

20–40%

lead-time reduction

On those PRs — their largest measured win, and it came from process automation rather than coding speed.

▲

PR size rose

Consistently, in the 100–500 line bucket and above. The opposite of what most AI-productivity models assume.

The consulting implication is the part we take seriously: the highest-ROI intervention in an engagement may not be a coding agent at all. Instrument first, find the real bottleneck, and be willing to recommend an approval-automation bot over another skill. Review latency is usually the biggest single block of cycle time, and removing a queue beats making the typing faster.
06 · Where to start

The baseline audit is the cheap part

It is short, it happens before any training, and it is the thing that makes every later number mean something. Without it there is no measurement — only assertion.

Baseline audit

Two weeks before anything else. We measure your current state and hand you an immutable copy of it.

Talk to us →

See the blueprint

The full picture of a company that works augmented — five layers, every role, the governance that keeps it safe.

The Blueprint →

Check where you stand

An eight-item self-check against the blueprint, with the honest scoring.

Self-check →
Book a free diagnostic call

Get your team orchestrating AI systems while others are still vibe-coding.

After a free 30-minute skills map, we take your developers to multi-level orchestration and secure, parallel delivery, at the quality standard you already expect. No pitch.

Why Contact Us?

Quick Response
We respond to all inquiries within 24 hours
Free Consultation
No commitment required to get expert advice
Expert Advice
From senior AI specialists with 40+ years combined experience