Most of the metrics engineering leaders have relied on, like lines of code, commits, story points and utilization, were built for a time when engineering set the pace and writing the code was the hard part.
AI changed that. Engineers can now produce more than product can spec, QA can verify or customers can absorb, and the old metrics don’t show where the work is getting stuck.
An AI-native software development lifecycle, or AI SDLC, connects planning, design, development, QA, deployment and monitoring around shared standards and measurable outcomes. Building one means looking past the coding tools your engineers use to the process around them.
Faster Engineering Moves the Bottleneck
When engineering capacity jumps, work piles up on either side of it. We’ve seen teams where engineers sit waiting for requirements defined well enough to build, and with others where code arrives faster than review and testing can handle.
QA’s job is changing too. The question used to be “does it work?” Now it’s “is it right?”
In one engagement, an AI-generated account settings screen matched the design exactly but included a delete-account button with no confirmation dialog. It would have passed every functional test.
Teams now need to check user intent, business logic and consequences, not just whether a feature runs. For agentic products, that means running representative requests and inspecting the resulting actions whenever the system changes. Security has to run from requirements through production monitoring, and NIST’s Secure Software Development Framework is a solid baseline.
Documentation is another constraint. We’ve seen AI handle release notes well, but help content still lags behind. When features change weekly, customers need guidance that changes with them.
DORA’s 2025 research describes AI as an amplifier of an organization’s existing strengths and weaknesses, and across two consecutive years of DORA reports, AI adoption has been associated with rising delivery instability. Getting value from the tools depends on fixing the system around them.
Measure the Full Lifecycle Across Five Pillars
Lines of code, completed tickets and prompt counts show activity. Leaders need to see how work reaches customers and what happens once it gets there.
York IE’s measurement framework brings together five pillars:
| Pillar | Core question | Example measures |
| AI leverage | How effectively is AI being used? | Adoption by task and phase, reusable skill usage, token efficiency |
| Quality | Is the output consistently good? | Review effectiveness, rework, escaped defects, release-checklist compliance |
| Stability | Does the software hold up in production? | Change failure rate, incident frequency, recovery time, rollbacks |
| Speed | How quickly does an idea reach a customer? | Queue time, review turnaround, deployment frequency, end-to-end lead time |
| Business value | Is delivery improving business outcomes? | 90-day feature adoption, time to first use, KPI-linked releases, customer feedback |
DORA’s software delivery performance metrics cover throughput and instability, and they fit inside this framework. More releases look good until incidents climb alongside them, and faster delivery means little if customers barely use what ships.
We roll the pillars into a weekly score for each team in Pulse, the operating dashboard we built for our own R&D organization. A composite score tells you where to look, but only if the measures, sources and weighting behind it are visible. Otherwise a strong speed number can hide a quality problem.
On one of our teams, median recovery time fell from 71 minutes to 42 while delivery volume grew and post-release checklist compliance held at 100%. Of the last 18 features that the team shipped, 12 saw meaningful use within 90 days. The other six went back to planning.
The same dashboard caught a release that never went out. Our process asks someone to file a release intent, a declared ship date the team commits to. The person who normally filed it was unexpectedly out, and nobody else thought they were allowed to do it. We named a backup. The dashboard also showed that on-call coverage rested on just two engineers, a risk we’d rather find on a chart than during an outage.
AI Usage Needs Context
High token consumption doesn’t reliably identify your most productive engineers, and a low prompt count can signal maturity rather than weak adoption. As skill libraries mature, strong engineers send fewer, simpler prompts because the skills do the heavy lifting.
We learned this firsthand. One of our engineers scored low on the first version of our prompt-quality scoring because their prompts were short. The score was wrong: the prompts were short because they leaned on reusable skills. Looking closer turned up a different problem. Those skills were poorly configured and burning far more tokens than the work required. We rebuilt them into a company-approved skill library designed for token efficiency, and our scoring now accounts for skills applied, execution efficiency and token effectiveness.
Categorize the work AI supports (exploration, creation, correction, documentation and infrastructure) instead of just counting usage. Two engineers with the same prompt count can be doing very different work, one building new features and the other mostly fixing what AI got wrong.
Use what you find to improve training and the shared skill library. We built Pulse’s people view for coaching, not policing. A below-average score routes someone to enablement, not a performance conversation, because the data stops being honest once people expect it to be used against them. When token spend rises, look at the task and workflow first.
Give Every Phase a Maturity Score
A company can have sophisticated AI workflows in development while product planning is still ad hoc. That mismatch is often why engineers run out of ready work.
Score each of the six lifecycle phases on four levels:
- Adoption: AI is part of daily work, not just licensed. Push this first and optimize later.
- Depth: Teams have repeatable skills for specific tasks, such as creation, correction, QA, documentation and infrastructure.
- Instrumentation: Automation, pipelines and quality gates carry the work, so results don’t depend on individual heroics.
- Outcomes: Quality, stability and speed are strong at the same time, and delivery moves business KPIs. This is the only level that shows up in business results.
Expect the levels to differ. Development might sit at level three while planning is still at level one, and that spread is your roadmap. For each phase, record the current level, the output the next phase needs and the next step. Assign an owner and review progress regularly.
We usually start with development, where the leverage compounds, and follow right away with QA, because generated code needs generated verification. Deployment and monitoring come next so the new speed is reversible and observable. Planning and design close the loop. If engineers are already stalled waiting for work, though, planning moves up the list. AI can help product managers draft clearer requirements, but people should keep ownership of prioritization and acceptance criteria. More requirements only help if they describe the right work.
Build Shared Context and Governance
Each pillar runs on the same foundation: a documented process, a library of reusable skills, integrations with the tools you already use (mostly configuration, not new software) and governance that confirms the process ran and scores the output. You don’t need all of it on day one. Start with the phase that needs it most.
Shared context matters more as product, design and engineering work overlap. In one prototype, an AI model that couldn’t see the design library built its own table component, and the team ended up maintaining two.
Give AI access to approved components, architecture decisions and customer context, and name the people accountable for standards and exceptions. For more on the data foundation behind agents, see our article on building reliable AI agents from document data.
The same caution applies to delivery cadence. Faster coding tempts teams to drop sprints for continuous flow. Before making that move, make sure planning, verification and release controls can keep up. DORA’s continuous delivery guidance covers the testing, security and process work that makes frequent releases reliable.
Bring Outcomes to the Board
Boards are asking three questions right now: How much has delivery throughput increased with AI? Are quality and stability holding? How do you know the team is working on what the business needs?
Bring evidence for each. Show changes per engineer-month trended over four quarters, with quality alongside so the gain is credible. Show change failure rate staying flat as volume grew, with incidents and recovery time trending down. Show 90-day feature adoption and KPI-linked delivery. A phase-by-phase scorecard turns each question into an evidence conversation instead of a vibes conversation.
For leaders deciding where to start:
- Find the constraint across the full lifecycle, including documentation and customer adoption.
- Assess all five pillars together, with the underlying measures visible.
- Score each phase, choose the next improvement and assign an owner.
- Tie AI activity to delivery quality and business outcomes.
York IE’s R&D advisory and execution teams work across product, design, engineering, QA and DevOps. We help companies assess their development processes, build repeatable AI workflows and set up the measurement needed to improve them.
For investors, our technical due diligence services offer a focused way to assess technology risk and improvement opportunities.
Start with the phase that is limiting your team’s ability to turn customer needs into reliable, useful software. That is where your next AI investment should earn its value.
