Adoption
78%
The blueprint gives you an estimate. Estimates are cheap — anyone can build a calculator that returns a flattering number. This is the apparatus that would confirm that estimate or kill it, and the part where we say out loud what it cannot prove.
See the console ↓The telemetry proves adoption, capability quality and cost — precisely. It does not, on its own, prove productivity.
telemetry (leading)
adoption + capability
your systems (lagging)
outcome trend over time
the claim
defensible
What a vendor says
"Your engineers are 30% faster."
Off tool telemetry alone. The first competent CFO takes this apart, and should.
What we would say
"Adoption reached 78% of eligible engineers, the six capabilities we built carry 61% of agent spend at a 91% success rate, and over the same period PR cycle time fell 22% against a baseline we captured before we started."
Every clause traces to a named source. The last one is only possible because of the next section.
Every panel below carries its own provenance. Turn on Explain all metrics — or click any ⓘ — and each number tells you where it comes from, the exact formula behind it, what it shows and why it holds up under challenge.
ILLUSTRATIVE CLIENT · AI AUGMENTATION PROGRAMME · Q3
Adoption
78%
Spend / eng / mo
$186
Cycle time p85
4.1d
Change failure rate
⚠ WATCH11%
Adoption %
Cycle time p85 (days)
Wave 2 holds flat through Wave 1's improvement and bends only after its own June rollout — the staggered schedule acts as a control group.
Speed
cycle time p85 · time to merge · review latency
IMPROVING
Volume
throughput · items completed
IMPROVING
Quality
failure rate · defects · reopen · MTTR · complexity
MIXED
Adoption
active engineers · capability usage
IMPROVING
Perceived
developer survey · +0.8 pts
IMPROVING
Recomputed live from the panels below: a family counts as agreeing when its member metrics move in the improving direction against the selected baseline.
Cycle time p85
4.1d
▼ 21%
Throughput
3.4PR/eng/wk
▲ 17%
Review latency
5.2h
▼ 32%
PR size
312LOC
▲ 18%⚠
Rising. Bigger changes normally slow review — cycle time improved in spite of this.
Change failure rate
11%
▲ 1 pt⚠
Above baseline for 4 weeks. Under review.
Escaped defects
3.1/release
▼ 18%
Reopen rate
6.4%
▼ 2 pts
MTTR
47min
▼ 11%
Cyclomatic complexity
6.2% vs base
▲ 6.2 pts⚠
Rising since adoption began. The main long-term risk signal.
Revert rate
2.1%
▼ 0.5 pts
| Capability | Uses | Success | Cost/use | Verdict |
|---|---|---|---|---|
| code-review-agent | 1,240 | 94% | $0.11 | ★ KEEP |
| migration-helper | 380 | 91% | $0.34 | ★ KEEP |
| test-generator | 44 | 61% | $0.92 | ✕ KILL |
| doc-writer | 3 | — | — | ✕ KILL |
A capability nobody invokes is not a neutral outcome — it is a failed deliverable we shipped. Reporting the kill is the point.
Design
Staggered rollout. Wave 2 serves as a control until June.
Baseline
Q1, captured from git history before the first workshop.
Confounders
PR size rose 18% over the same window; team composition changed on two teams.
Coverage
7 of 9 teams. Platform and Data are not yet instrumented.
Publishing this is the point. Being visibly disciplined about the limits is what makes everything above the line believable — and it is the fastest way to tell a measurement practice from a marketing one.
High confidence · straight from telemetry
Directional · never quoted as precision
Where the value claim is actually made
Not defensible · not offered
Novelty effect
A 14-day read is marketing, not measurement. We report at 90 days.
Self-selection
Enthusiasts adopt first and were already fast. Cohort comparison alone cannot separate the two.
Team composition
People join and leave mid-window. Metrics normalise by active contributors, never by headcount.
Seasonality
Q4 and holiday windows distort throughput in both directions.
Concurrent initiatives
If a platform migration lands in the same quarter, it owns part of the result.
Observation effect
Measured teams behave differently while they are being measured.
The schedule exists to stop the two failure modes that destroy credibility: declaring victory too early, and never checking whether the capabilities we shipped are used at all.
T−2 weeks
Git, incident and ticket history measured. Metric set and success criteria agreed in writing, before anyone is trained.
T−0
Telemetry live on day one, with the legal clearance already signed off.
Day 14
Diagnostic, never a result. Anyone reporting an outcome at two weeks is measuring novelty.
Day 30
Keep, fix, promote or kill each thing we built. First honest cost read.
Day 90
Outcomes against baseline, survey re-run, families counted. This is the number that gets quoted.
Quarterly
Capabilities decay, usage drifts, new teams onboard. Retire what stopped earning its keep.
Two pages plus an appendix. Five sections, always the same five.
It reports at least one negative finding. A report with no bad news reads as marketing, and gets read as marketing.
Public research, not a client of ours.
Bartosz Ocytko, "Agentic Engineering at Zalando: a snapshot", Zalando engineering blog, 14 August 2026.
Read the original ↗Reported across more than 250 engineering teams, with adoption measured through a gateway proxy serving roughly 2,000 monthly active users on six small pods — this is modest infrastructure, not a platform programme.
33%
of PRs auto-approved
A risk-based approval bot handles the low-risk third automatically.
20–40%
lead-time reduction
On those PRs — their largest measured win, and it came from process automation rather than coding speed.
▲
PR size rose
Consistently, in the 100–500 line bucket and above. The opposite of what most AI-productivity models assume.
It is short, it happens before any training, and it is the thing that makes every later number mean something. Without it there is no measurement — only assertion.
Two weeks before anything else. We measure your current state and hand you an immutable copy of it.
Talk to us →The full picture of a company that works augmented — five layers, every role, the governance that keeps it safe.
The Blueprint →An eight-item self-check against the blueprint, with the honest scoring.
Self-check →After a free 30-minute skills map, we take your developers to multi-level orchestration and secure, parallel delivery, at the quality standard you already expect. No pitch.