← Back to the BlueprintThe proof layer

How you would know it actually worked

The blueprint gives you an estimate. Estimates are cheap — anyone can build a calculator that returns a flattering number. This is the apparatus that would confirm that estimate or kill it, and the part where we say out loud what it cannot prove.

See the console
01 · The shape of a defensible claim

Telemetry alone cannot prove productivity

The telemetry proves adoption, capability quality and cost — precisely. It does not, on its own, prove productivity.

telemetry (leading)

adoption + capability

×

your systems (lagging)

outcome trend over time

=

the claim

defensible

What a vendor says

"Your engineers are 30% faster."

Off tool telemetry alone. The first competent CFO takes this apart, and should.

What we would say

"Adoption reached 78% of eligible engineers, the six capabilities we built carry 61% of agent spend at a 91% success rate, and over the same period PR cycle time fell 22% against a baseline we captured before we started."

Every clause traces to a named source. The last one is only possible because of the next section.

Which requires the one thing most engagements skip: a baseline captured before we begin. Six to twelve months of your git history, incident record and ticket flow, measured and handed to you immutably on day one. It is the single highest-value hour of the whole engagement, it costs almost nothing, and it cannot be recreated afterwards.
02 · The console

One page a CxO can read in thirty seconds

Every panel below carries its own provenance. Turn on Explain all metrics — or click any ⓘ — and each number tells you where it comes from, the exact formula behind it, what it shows and why it holds up under challenge.

Delivery Impact Console

ILLUSTRATIVE CLIENT · AI AUGMENTATION PROGRAMME · Q3

SAMPLE DATA

Adoption

78%

▲ 71 pts62 of 79 eng

Spend / eng / mo

$186

▲ $24Q1 $162

Cycle time p85

4.1d

▼ 21%Q1 5.2

Change failure rate

⚠ WATCH

11%

▲ 1 ptQ1 10

Rollout & response

ADOPTION vs CYCLE TIME · BY COHORT
Wave 1 — 4 teamsWave 2 — 3 teams

Adoption %

048.396.6Workshop · Wave 1Wave 2 rollout

Cycle time p85 (days)

3.94.85.8Q1 baseline 5.2dWorkshop · Wave 1Wave 2 rolloutJanFebMarAprMayJunJulAugSep

Wave 2 holds flat through Wave 1's improvement and bends only after its own June rollout — the staggered schedule acts as a control group.

Convergent evidence

4 / 5 AGREE

Speed

cycle time p85 · time to merge · review latency

IMPROVING

Volume

throughput · items completed

IMPROVING

Quality

failure rate · defects · reopen · MTTR · complexity

MIXED

Adoption

active engineers · capability usage

IMPROVING

Perceived

developer survey · +0.8 pts

IMPROVING

Recomputed live from the panels below: a family counts as agreeing when its member metrics move in the improving direction against the selected baseline.

Delivery

GITHUB · 90d

Cycle time p85

4.1d

▼ 21%

Throughput

3.4PR/eng/wk

▲ 17%

Review latency

5.2h

▼ 32%

PR size

312LOC

▲ 18%

Rising. Bigger changes normally slow review — cycle time improved in spite of this.

Quality guardrails

CI · INCIDENTS · JIRA

Change failure rate

11%

▲ 1 pt

Above baseline for 4 weeks. Under review.

Escaped defects

3.1/release

▼ 18%

Reopen rate

6.4%

▼ 2 pts

MTTR

47min

▼ 11%

Cyclomatic complexity

6.2% vs base

▲ 6.2 pts

Rising since adoption began. The main long-term risk signal.

Revert rate

2.1%

▼ 0.5 pts

What we built

CAPABILITY PORTFOLIO · 30d
CapabilityUsesSuccessCost/useVerdict
code-review-agent1,24094%$0.11 KEEP
migration-helper38091%$0.34 KEEP
test-generator4461%$0.92 KILL
doc-writer3 KILL

A capability nobody invokes is not a neutral outcome — it is a failed deliverable we shipped. Reporting the kill is the point.

Where the spend goes

BY CAPABILITY · 30d
code-review-agent$2,140
migration-helper$710
test-generator$392
unattributed / ad-hoc$281

How to read this

Design

Staggered rollout. Wave 2 serves as a control until June.

Baseline

Q1, captured from git history before the first workshop.

Confounders

PR size rose 18% over the same window; team composition changed on two teams.

Coverage

7 of 9 teams. Platform and Data are not yet instrumented.

03 · The claim ladder

What we can prove, and what we refuse to claim

Publishing this is the point. Being visibly disciplined about the limits is what makes everything above the line believable — and it is the fastest way to tell a measurement practice from a marketing one.

A

Directly measured

High confidence · straight from telemetry

  • How many people use it, how often, and who never adopted
  • What it costs — per person, per team, per month
  • Which capabilities we built are used, and which are ignored
  • Which tools fail, how often, and with what error
  • Which MCP servers are silently broken
  • Where the budget actually goes, by skill and by model
B

Measured, with real caveats

Directional · never quoted as precision

  • AI-assisted code share — deliberately conservative, a floor and not a measurement
  • Suggestion accept rate — a good friction signal, a poor quality signal
  • Lines of code — a real number, near-worthless as value
  • Model duration — model time, not engineer time. Never presented as the latter
C

Requires your systems

Where the value claim is actually made

  • Cycle time and lead time for change — from your git history
  • Throughput per engineer per week
  • Change failure rate, escaped defects, MTTR
  • Rework: reverts, hotfixes, reopen rate
  • Review latency and review load
  • Whether engineers feel more effective — survey, at baseline and at day 90
D

We will not claim this

Not defensible · not offered

  • "We saved you N FTEs" — requires assumptions nobody can evidence
  • "AI wrote N% of your codebase" — the attribution is a deliberate undercount and does not support the reading
  • "Quality improved because of AI" — correlation at best without a controlled rollout
  • "The team is N% faster" — unless measured on time, with a control cohort, paired with quality, over 90+ days

Confounders we name before you do

Novelty effect

A 14-day read is marketing, not measurement. We report at 90 days.

Self-selection

Enthusiasts adopt first and were already fast. Cohort comparison alone cannot separate the two.

Team composition

People join and leave mid-window. Metrics normalise by active contributors, never by headcount.

Seasonality

Q4 and holiday windows distort throughput in both directions.

Concurrent initiatives

If a platform migration lands in the same quarter, it owns part of the result.

Observation effect

Measured teams behave differently while they are being measured.

04 · The cadence

Ninety days, with the honest checkpoint at day 14

The schedule exists to stop the two failure modes that destroy credibility: declaring victory too early, and never checking whether the capabilities we shipped are used at all.

T−2 weeks

Baseline capture

Git, incident and ticket history measured. Metric set and success criteria agreed in writing, before anyone is trained.

T−0

Workshop and rollout

Telemetry live on day one, with the legal clearance already signed off.

Day 14

Adoption check only

Diagnostic, never a result. Anyone reporting an outcome at two weeks is measuring novelty.

Day 30

Capability portfolio review

Keep, fix, promote or kill each thing we built. First honest cost read.

Day 90

The real measurement

Outcomes against baseline, survey re-run, families counted. This is the number that gets quoted.

Quarterly

Ongoing review

Capabilities decay, usage drifts, new teams onboard. Retire what stopped earning its keep.

The day-90 report

Two pages plus an appendix. Five sections, always the same five.

  1. 01Adoption
  2. 02Capability portfolio
  3. 03Economics
  4. 04Outcomes vs baseline
  5. 05The next 90 days

It reports at least one negative finding. A report with no bad news reads as marketing, and gets read as marketing.

05 · Field evidence

The most useful public dataset says the bottleneck may not be the code

Public research, not a client of ours.

Bartosz Ocytko, "Agentic Engineering at Zalando: a snapshot", Zalando engineering blog, 14 August 2026.

Read the original

Reported across more than 250 engineering teams, with adoption measured through a gateway proxy serving roughly 2,000 monthly active users on six small pods — this is modest infrastructure, not a platform programme.

33%

of PRs auto-approved

A risk-based approval bot handles the low-risk third automatically.

20–40%

lead-time reduction

On those PRs — their largest measured win, and it came from process automation rather than coding speed.

PR size rose

Consistently, in the 100–500 line bucket and above. The opposite of what most AI-productivity models assume.

The consulting implication is the part we take seriously: the highest-ROI intervention in an engagement may not be a coding agent at all. Instrument first, find the real bottleneck, and be willing to recommend an approval-automation bot over another skill. Review latency is usually the biggest single block of cycle time, and removing a queue beats making the typing faster.
06 · Where to start

The baseline audit is the cheap part

It is short, it happens before any training, and it is the thing that makes every later number mean something. Without it there is no measurement — only assertion.

Baseline audit

Two weeks before anything else. We measure your current state and hand you an immutable copy of it.

Talk to us

See the blueprint

The full picture of a company that works augmented — five layers, every role, the governance that keeps it safe.

The Blueprint

Check where you stand

An eight-item self-check against the blueprint, with the honest scoring.

Self-check
Book a free diagnostic call

Get your team orchestrating AI systems while others are still vibe-coding.

After a free 30-minute skills map, we take your developers to multi-level orchestration and secure, parallel delivery, at the quality standard you already expect. No pitch.

Why Contact Us?

Quick Response
We respond to all inquiries within 24 hours
Free Consultation
No commitment required to get expert advice
Expert Advice
From senior AI specialists with 40+ years combined experience