Adoption
78%
The proof layer combines adoption telemetry with delivery, quality, and cost data against a pre-rollout baseline, turning estimated opportunity into evidence leaders can act on.
See the impact console ↓Leading signals show whether adoption is taking hold, at what capability level, and at what cost. Operational outcomes show what changes in delivery, quality, and flow. A pre-rollout baseline connects the two.
leading signals
adoption + capability + cost
operational outcomes
delivery + quality + flow
result
decision-ready evidence
Decision-ready example
Adoption reached 78% of eligible engineers and cycle time p85 fell 21% against the Q1 baseline. Quality stayed visible: change failure rate is one point above baseline and under review.
Every clause maps to a console metric and a named source, making both verified impact and open questions visible.
Every figure below is computed from the controls above — team, period and baseline. Change any of them and the whole console recomputes, quality guardrails included, even when they move against us.
ILLUSTRATIVE CLIENT · AI AUGMENTATION PROGRAMME · Q3
Adoption
78%
Spend / eng / mo
$186
Cycle time p85
4.1d
Change failure rate
⚠ WATCH11%
Adoption %
Cycle time p85 (days)
Wave 2 holds flat through Wave 1's improvement and bends only after its own June rollout — the staggered schedule acts as a control group.
Speed
cycle time p85 · time to merge · review latency
IMPROVING
Volume
throughput · items completed
IMPROVING
Quality
failure rate · defects · reopen · MTTR · complexity
MIXED
Adoption
active engineers · capability usage
IMPROVING
Perceived
developer survey · +0.8 pts
IMPROVING
Recomputed live from the panels below: a family counts as agreeing when its member metrics move in the improving direction against the selected baseline.
Cycle time p85
4.1d
▼ 21%
Throughput
3.4PR/eng/wk
▲ 17%
Review latency
5.2h
▼ 32%
PR size
312LOC
▲ 18%⚠
Rising. Bigger changes normally slow review — cycle time improved in spite of this.
Change failure rate
11%
▲ 1 pt⚠
Above baseline for 4 weeks. Under review.
Escaped defects
3.1/release
▼ 18%
Reopen rate
6.4%
▼ 2 pts
MTTR
47min
▼ 11%
Cyclomatic complexity
6.2% vs base
▲ 6.2 pts⚠
Rising since adoption began. The main long-term risk signal.
Revert rate
2.1%
▼ 0.5 pts
| Capability | Uses | Success | Cost/use | Verdict |
|---|---|---|---|---|
| code-review-agent | 1,240 | 94% | $0.11 | ★ KEEP |
| migration-helper | 380 | 91% | $0.34 | ★ KEEP |
| test-generator | 44 | 61% | $0.92 | ✕ KILL |
| doc-writer | 3 | — | — | ✕ KILL |
A capability nobody invokes is not a neutral outcome — it is a failed deliverable we shipped. Reporting the kill is the point.
Design
Staggered rollout. Wave 2 serves as a control until June.
Baseline
Q1, captured from git history before the first workshop.
Confounders
PR size rose 18% over the same window; team composition changed on two teams.
Coverage
7 of 9 teams. Platform and Data are not yet instrumented.
Publishing this is the point. Being visibly disciplined about the limits is what makes everything above the line believable — and it is the fastest way to tell a measurement practice from a marketing one.
High confidence · straight from telemetry
Directional · never quoted as precision
Where the value claim is actually made
Not defensible · not offered
Novelty effect
A 14-day read is marketing, not measurement. We report at 90 days.
Self-selection
Enthusiasts adopt first and were already fast. Cohort comparison alone cannot separate the two.
Team composition
People join and leave mid-window. Metrics normalise by active contributors, never by headcount.
Seasonality
Q4 and holiday windows distort throughput in both directions.
Concurrent initiatives
If a platform migration lands in the same quarter, it owns part of the result.
Observation effect
Measured teams behave differently while they are being measured.
The schedule exists to stop the two failure modes that destroy credibility: declaring victory too early, and never checking whether the capabilities we shipped are used at all.
2 weeks before
Git, incident and ticket history measured. Metric set and success criteria agreed in writing, before anyone is trained.
Day of the workshop
Telemetry live on day one, with the legal clearance already signed off.
Day 14
Diagnostic, never a result. Anyone reporting an outcome at two weeks is measuring novelty.
Day 30
Keep, fix, promote or kill each thing we built. First honest cost read.
Day 90
Outcomes against baseline, survey re-run, families counted. This is the number that gets quoted.
Quarterly
Capabilities decay, usage drifts, new teams onboard. Retire what stopped earning its keep.
Two pages plus an appendix. Five sections, always the same five.
It reports at least one negative finding. A report with no bad news reads as marketing, and gets read as marketing.
Public research, not a client of ours.
Bartosz Ocytko, "Agentic Engineering at Zalando: a snapshot", Zalando engineering blog, 14 August 2026.
Read the original ↗Reported across more than 250 engineering teams, with adoption measured through a gateway proxy serving roughly 2,000 monthly active users on six small pods — this is modest infrastructure, not a platform programme.
33%
of PRs auto-approved
A risk-based approval bot handles the low-risk third automatically.
20–40%
lead-time reduction
On those PRs — their largest measured win, and it came from process automation rather than coding speed.
▲
PR size rose
Consistently, in the 100–500 line bucket and above. The opposite of what most AI-productivity models assume.
It is short, it happens before any training, and it is the thing that makes every later number mean something. Without it there is no measurement — only assertion.
Two weeks before anything else. We measure your current state and hand you an immutable copy of it.
Talk to us →The full picture of a company that works augmented — five layers, every role, the governance that keeps it safe.
The Blueprint →An eight-item self-check against the blueprint, with the honest scoring.
Self-check →After a free 30-minute skills map, we take your developers to multi-level orchestration and secure, parallel delivery, at the quality standard you already expect. No pitch.