Artificial intelligence case studies
What changes when AI becomes part of the delivery system?
Three evidence-led studies from five months operating an independent engineering environment. The goal was not to maximize agent activity. It was to learn how architecture, decision rights, verification, and human attention must change.
This work sits alongside more than 20 years leading technology organizations, including a 50+ person organization and $4.3M budget at PrePass. Explore executive experience →
01
Operating model
Designing an AI-native engineering operating model
Parallel agents initially made the human the scheduler, message carrier, reviewer, and cleanup crew. A 30-day study of 4,419 prompts, including 240 manually reviewed samples, reframed the goal around successful objectives and reduced human coordination.
The target is minimum human coordination per successful objective—not maximum agent activity.
Read case study →
02
Governance
Governing autonomous work without eliminating autonomy
The review system became so demanding that its controls created more operating burden than some of the risks they addressed. Removing it too quickly exposed a second failure. The answer was proportional review and authority tied to consequence.
Code should enforce the few boundaries that genuinely matter. Governance must also prove its own value.
Read case study →
03
Trust
Making AI-generated work trustworthy
A first review system processed about 1,700 packets, but 64% of blocked outcomes came from the machinery around the reviewers. A smaller second-generation system reached a substantive verdict in 98.2% of 329 measured attempts.
AI can verify AI, but the verifier also needs verification.
Read case study →
Evidence boundary
Substantial hands-on work, carefully scoped claims.
This was an independent, single-operator environment—not a customer-facing enterprise deployment. I defined objectives, architecture, constraints, and acceptance criteria; directed AI agents to implement; and remained accountable for the evidence and decisions. The work supports informed executive judgment. It does not establish commercial impact or enterprise-scale adoption.
Primary source
Examine the full evidence on GitHub.
The repository contains the complete case studies, method, measurements, limitations, diagrams, and selected evidence exhibits.