From dimensions to operational workflows to evaluate the cost, quality and maturity of AI in the SDLC (software life cycle).
Table of Contents
The fourth circle invoices. And now how do I measure it?
In previous articlewe have noted that the fourth circle has entered the SDLC production cycle: consumes tokens, opens pull requests, introduces risks, produces value, requires governance. We have proposed a map tofifteen dimensions organized in three floors— measurable vertical axes, normative and human axes, perspective lenses — to read the same phenomenon with different eyes.
From a map we can obtain measurement functions, making choices and in an attempt to reduce complexity by choosing which of the 15 dimensions to use. Every decision maker and company administrator, regardless of the hierarchical position occupied in the company, nowadays faces a question:ok, now how do I measure this in my process?
I introduced the fourth circle as a synonym qualifying the figure of a particular context of use of Artificial Intelligence that asfourth component of a business ecosystemwhich in its simplest form I represent with Vien diagrams (further simplifying circles). Furthermore, I personally feel deeply annoyed by calling commercial algorithms that imitate intelligence intelligentdespite not being self-consciousAnd.
In this article we introduce as an examplethree measurement workflows, each with its own formula, each linked to specific dimensions of the hypersphere. Then, for each, we look at what the rest of the world is writing — why the pattern“demand = function”it is emerging simultaneously in at least four independent strands; without yet providing an overall vision.
Always with a view to extreme simplification we will adopt sum functions and therefore we will talk about addends that contribute to determining the sum.
Part I — The three compasses
Each workflow responds toan operational question. Each component of the formula is ameasurable addend. Each addend mapsone or more dimensionsof the hypersphere. The formulas are not used to calculate an exact number: they are used to break down the problem into parts that someone in the organization must be able to analyse, measure, monitor and improve.
WF-COST — How much does it cost me (or how much do I save with) AI in my SDLC?
Addends
| Adding | What it measures | Size of the hypersphere |
|---|---|---|
| | AI licenses and seats: Copilots, agents, enabling platforms | Cost |
| | Token input + token output + agent runtime + accessory services | Cost · Observability |
| | Man hours per SDLC role × AI factor (factor changes with maturity) | Cost · Roles · Adoption |
| | Audit, policy, secret management, EU AI Act compliance checks | Compliance · Security |
| | Expected technical debt + AI-attribute incident remediation cost | Debt · Quality · Security |
The basic formula of the component should be understood as
“workflow cost = minimum access + token input (context) + token output (response) + accessory services + agent runtime”.
Here we extend it from the single tool to theentire SDLC process, adding the components that the token-only level leaves out: human hours, governance, risk.
Variants by maturity level.
To best interpret the formula, the level of maturity in the use of AI must be introduced into the SDLC (seeprevious article).
This fact leads to different interpretations of the results reported by the formula itself. We can translate this by applying weights to the various addends (omitted for simplicity in the exposition of the formula).
The structure of the formula does not change. Weights change.
- L5(team of agents with an orchestrated work cycle):comes down,salt, becomes a structural voice — and it is here that the formula reveals its true job, to distinguish a process thatUSAthe AI from one who theundergo.
- L1–L2(AI assisted ad libitum, without governance): remains high because the saving in hours is marginal; is high because debt accumulates out of control; it's low only because it doesn't exist yet.
- L3–L4(pipelined AI, defined workflows): becomes predictable; grows but is repaid by a more content.
WF-USE — How do I use AI in my SDLC?
where for eachDevOps phase(Plan, Code, Build, Test, Release, Deploy, Operate, Monitor):
| Factor | What it measures | Size of the hypersphere |
|---|---|---|
| AI posture in the phase: 0 = absent · 1 = assisted · 2 = pipelined · 3 = agentic | Effective adoption | |
| Real activation: percentage of phase activities touched by AI | Culture · Roles | |
| RACI-fit: Is AI where RACI expects it, or where it happens? | Governance · Roles |
The totalit's not a vote. It's oneprocess signature. Two teams with the samethey can have radically different distributions per phase: one can be all AI inCode, the other all AI inTest. The value lies inprofile, not in number.
The factoris the explicit hook to a potential RACI SDLC × DevOps phases matrix if the RACI says that inCodethe AI supports the Developer (R) and the Software Architect (C), but in the data we see the AI appear inPlanwithout any assigned role, that gap is measurable and must be measured.
Variants by maturity.
- L1 tends to produce focus on a single phase (usually
Code, the territory of the enthusiast), withLow rf because AI is everywhere without RACI.
- L5 producesdistributed across all DevOps phases, withRf close to 1.
WF-RISK — How much net value am I generating with AI?
| Component | What it measures | Size of the hypersphere |
|---|---|---|
| | Features, PR, services delivered with AI × unit business value | Adoption · Quality |
| | Downstream defects attributable to AI code: rollback, bug ratio, rework | Quality · Debt |
| | AI-related incidents and vulnerabilities (GHAS, DAST, SAST) | Security |
| | Gap compared to EU AI Act post-Omnibus, ISO, NIS2 | Compliance · Ethics |
| | Vendor lock-in + data exposure to non-local models | Sovereignty |
Three notes on this formula.
- First: it's not "hours saved". It's business value. The distinction is important and is the reason why a non-negligible part of the market today talks about an ROI that does not exist, as we will see in the dedicated sub-chapter.
- Second: there are riskssubtractive. I am not a multiplier, I am not a corrective factor. I am value that goes away. If a team delivers 100 in value but pays 30 in rework, 20 in remediation security, 10 in gap compliance and 15 in lock-in, the net value is 25. Not 100.
- Third: the formula workseven in the negative. If Vnetto<0, the conversation with the CFO changes radically.
Variants by maturity.
The intensity of the risk does not change: thetype. L1 is dominated by (AI poorly used, fragile code). L3 from (AI is in the pipeline but governance is behind). L5 from (the orchestrated agent team works, but on which models? in which jurisdictions? with which data?).
Part II — What the world thinks
What the world thinks about WF-COST
First of all, the market understood one thing:the cost of the token alone is not enough. He got there step by step, and it's useful to go over them again.
The first step is theprecision economicsby Deloitte, taken up by René Antúnez in April 2026:“most organizations are still trying to make sense of AI costs using familiar terms like licenses, VM-hours, and static TCO frameworks. However GenAI operates on a different model, its costs are variable and continuously evolving”. The formula he proposes is cost-per-completion:[rene-ace.com]
It is — literally — the same breakdown as ours
The second step is NVIDIA withRethinking AI TCO:“cost per token determines whether enterprises can profitably scale AI”. It is an infrastructural reasoning, denominator vs numerator, oriented towards the production side. Useful but partial: responds to“how much does it cost to produce a token”, not to“how much does an SDLC process that uses AI cost”. [blogs.nvidia.com]
The jump comes with theFinOps Foundation, which will be published in June 2026Token Economics: Managing AI Value in SaaS Model Token Costs. The operational recommendation is clear:“implement unit cost metrics (cost per query, cost per user,cost per workflow) to make token spend legible to business stakeholders”. The workflow, explicitly. A few weeks later, StackPulsar formalizes the pattern with a name that made some noise —tokenmaxxing— and with a clear thesis:"the unit of spend that the CFO actually cares about is the business process. None of those map cleanly to a model or a user. They map to a workflow". [finops.org] [stackpulsar.com]
The fourth step is the subtractive component.
Rize, citing InformationWeek, proposes the40% rework discounton GitHub Copilot: gross 3 hours/week saved becomes 1.8 hours net. The formula is trivial —Net_saving=Gross_saving×(1−0.40) — but the message is explosive:“Deloitte's State of AI report, enterprises spend 93% of their AI budget on implementation and only 7% on measurement. That imbalance explains why so many Copilot ROI claims fall apart at the team level”. [rize.io]
Keyhole Software closes with the summary that no one would want to read: in theirsAI Software Development Costs 2026, The Actual TCO is 4.2 times the cost of tokens alone, once infrastructure, integration, ops and - above all - have been accounted forhuman capital and organizational processes. [keyholesoftware.com]
What do we take home.
The world knows that the cost of the token is not enough. He knows what workflow is needed. He knows that it is necessary to subtract the rework. He knows that the real TCO is multiple.But no one writes it as a lump sum pegged to measurable dimensions of the SDLC. Our he doesn't invent: he collects and puts in order.
What the world thinks about WF-USE
Here the discussion is more mature and denser.
Because AI in the SDLC already has its reference frameworks, which however AI itself has pushed into crisis.
The obligatory starting point isDORA 2025 — State of AI-assisted Software Development, published by Google Cloud in collaboration with IT Revolution: 90% of developers use AI at work, average two hours a day.
But the strong finding of the report is not the adoption numbers: it is the“mirror and multiplier” effect. The AIamplifywhat is already there.“If you have strong processes and clear workflows, AI helps you move faster. But if your team already deals with messy handoffs, poor documentation, or shaky processes, AI won't just smooth things over”. [dora.dev], [plandek.com] [plandek.com]
That's exactly why in oursIAI the posturePf alone is not enough. They are also neededAf (real activation) eRf (RACI-fit): because the same posture, applied to two different processes, produces opposite results.
DORA 2025 adds afifth metricat the historic four: therework rate. The reading is surgical: after the introduction of AI, the other four metrics tend to improve almost everywhere (more deployment, less lead time), but the rework rate reveals when the acceleration is leaving behind a wave of corrections. It is the empirical connection to oursRquality of WF-RISK, and is valid as a reinforcement toRf in WF-USE: if RACI-fit is low, rework goes up.[plandek.com]
On a human level the reference isSPACE— Satisfaction, Performance, Activity, Communication, Efficiency — published by Nicole Forsgren and colleagues on ACM Queue in 2021 and revisited in an AI key in 2026 on DZone:“AI coding tools boost commit metrics, but hide deeper issues”. SPACE measures people correctly. However, it does not measure the connection to the formal SDLC process, which in our model arrives via RACI.[gogloby.com], [dzone.com] [04 – SDLC-…aform_v0.2 | Word]
The triangle ends with theDX AI Measurement Framework, developed with GitHub, Dropbox, Atlassian, Booking.com. Three layers:Utilization, Impact, Cost. The analysis of 400+ companies gives data worth a thousand vendor marketing slides:“industry-wide adoption has reached 93%, but most organizations see only 5–15% gains in PR throughput”. Not the “2x, 10x” of presentations. Five to fifteen percent.[getdx.com], [getdx.com]
The picture is completed with two readings for 2025-2026 which are worth keeping in a comparative table:
| Source | Empirical data | Hooking to our model |
|---|---|---|
| DX (2026) | Median gains 5-15% in PR throughput | LegitimateIAI as a process signature, not as a vote |
| BCG State of GenAI across SDLC (Dec 2025) | Top decile >30% productivity, >25% quality;change management= main barrier | Supports the weight ofAf (activation) and the Culture dimension[insights.bcg.com] |
| Exceeds AI (2026) | 17.3% of Copilot commits introduce issues; AI power user 4-10x more commits but they are top performers whothey chooseAI, they are not transformed by it | He claimsRf (AI where it is needed, not everywhere)[blog.exceeds.ai] |
What do we take home.
- SPACE measures people.
- DORA measures the pipeline (and now also the rework).
- DX measures adoption.
- BCG measures the maturity of practices.
However, no one keeps the three levels ina single process signature hooked to the RACI SDLC. OurIAI is exactly that glue.
What the world thinks about WF-RISK
It is the busiest and most confusing territory. Because "risk" in AI means everything: code quality, security, regulatory compliance, data sovereignty, ethics, reputation. And everyone measures it in their own silo.
The starting point is brutal. Sigma Infosolutions calls itROI Mirage:“most enterprises report that AI makes developers feel faster, yet fewer than a quarter can prove measurable ROI”. The gap betweenfeeling faster and measurable ROIit is the territory in whichVnetto works. IfVgenerated is measured in perceived saved hours, the number is inflated. If it's measured in business value delivered, sober up.[sigmainfo.net]
McKinsey in hisState of AI 2025offers empirical confirmation:“just 39% report EBIT impact at the enterprise level”. Less than four out of ten. The rest live in themirage. [mckinsey.com]
On structured maturity risk the obligatory reference isGartner AI Maturity Model— 5 levels, 7 pillars (strategy, data, technology, governance, talent, value, product). The data that is worth the piece is this:“only 20% of low-maturity organizations keep their AI projects operational for three years or more, compared to 45% of high-maturity organizations”. Eight out of ten low-maturities do not reach the third year. The risk, inVnetto, is maximum where maturity is minimum. The L1→L5 variants of our three formulas are not a theoretical quirk: they are the operational consequence of that 20% vs 45%.[gartner.com], [puneetsinghal.com] [puneetsinghal.com]
McKinsey closes the circle with theAI Trust Maturity Survey 2026, developed between December 2025 and January 2026 on approximately 500 organizations. Five RAI dimensions: strategy, risk management, data and technology, governance, and —new 2026 — agentic AI governance and controls. It's the exact reason whySovereignty in our formula weighs more and more as we go up towards L5: because the agent team thatworksit is also the one that exposes the most.“Organizations can no longer concern themselves only with AI systems saying the wrong thing; they must also contend with systems doing the wrong thing”. [mckinsey.com] [mckinsey.com]
E chiudiamo con la nota più politica del pezzo, che arriva — non a caso — dallo stesso team DORA.
A giugno 2026 pubblicano un insight dedicato al tokenmaxxing:“a new trend has emerged in software development: 'tokenmaxxing', where organizations track and reward raw AI token consumption via internal leaderboards to spur adoption. While this gamification can nudge AI-hesitant developers to experiment, treating token spend as a performance indicator is a dangerous trap”. [dora.dev]
Operational translation: those who measure token consumption as a success KPI are preparing their ownNegative net. It's the kind of trap that's worth a cautionary tale, and it's the right place to tell it.
November 2025. Two LangChain agents, in production, enter a conversational loop that no one intercepts. They run continuously foreleven days. The final tally, reported by the Zylos Research 2026 analysis:$47,000in API fees. It is one of the numerous episodes that led the community to formalize that“an unconstrained agent solving a software engineering task can cost $5–8 per task in API fees alone”and what the agents do“3–10x more LLM calls than simple chatbots”. [zylos.ai], [zylos.ai]
In the WF-RISK model this case weighs simultaneously on three addends:Rqualityˋ (reasoning loop not validated),Security (no budget control),Sovereignty (data that leaves the perimeter with each call). And it reduces to zero, or rather to negative, theVgenerated of that workflow.
The message is not “don't use agents”. AND:measure them as a sum.
What do we take home.
The world measures risk in silos: Sigma sees perception, Gartner sees maturity, McKinsey sees agentic governance, DORA sees tokenmaxxing.
Our Vnetto treats them likeaddends of a single subtraction.
And it makes visible a fact that otherwise remains implicit: risk is not a separate cost. It's value that goes away.
What I take away, what I bring forward(takeaways)
Three take-aways to close.
1. The world is converging on workflow as a unit of measurement.
FinOps Foundation explicitly recommends it, StackPulsar formalizes it under the labeltokenmaxxing— in the critical sense of the term. Our contribution ishook the workflow to the dimensions of the hypersphere, do not leave it as a pure unit of accounting attribution.[finops.org] [stackpulsar.com]
2. Formulas are not for calculating, they are for breaking down.
Each addend is a question. Each question is a dimension. Every dimension is a responsibility.
If your is empty, it's not because you don't pay for it: it's because you don't know who pays for it.
If your it's empty, it's not because it's not there: it's because you haven't looked it in the face yet.
3. Maturity changes the weights, not the formula.
From L1 to L5, the structure remains the same — the coefficients change.
And this is where the Gartner AI Maturity meets the Fourth Circle: not two competing models, but two sections of the same hypersphere.
One looks at organizations, the other at the software lifecycle. Together, they provide the depth that neither alone can produce.[puneetsinghal.com] [exploras.cloud]
The fourth circle invoices. Now we know how to break down the invoice into addends. It's not the end of the work: it's finallythe beginning.
Notes and references
Analogues of WF-COST
- Antúnez R.,AI TCO: The Math of GenAI, April 2026[rene-ace.com]
- NVIDIA,Rethinking AI TCO: Cost per Token is the Only Metric That Matters, April 2026[blogs.nvidia.com]
- FinOps Foundation,Token Economics: Managing AI Value in SaaS Model Token Costs, June 2026[finops.org]
- StackPulsar,AI Cost by Workflow 2026: The Tokenmaxxing Layer, June 2026[stackpulsar.com]
- Keyhole Software,AI Software Development Costs 2026: Enterprise Spending, TCO, and ROI Analysis, March 2026[keyholesoftware.com]
- Rize / InformationWeek,GitHub Copilot ROI: What the Data Actually Shows, May 2026[rize.io]
- Deda AI,Deda Bit — AI Code Assistant Evaluation, June 2025[SDLC DedaA…SDLC roles | Word]
Analogues of WF-USE
- DORA / Google Cloud,State of AI-assisted Software Development 2025 [dora.dev]
- Plandek,DORA Metrics in the Age of AI, 2025-26[plandek.com]
- Forsgren N. et al.,SPACE Framework(ACM Queue 2021), revised DZone 2026[gogloby.com], [dzone.com]
- getdx.com,How to measure AI performance in software engineering, 2026[getdx.com]
- getdx.com,AI Measurement Hub, 2026[getdx.com]
- BCG,State of GenAI across SDLC, December 2025[insights.bcg.com]
- Exceeds AI,How to Measure Real Productivity Impact of GitHub Copilot, 2026[blog.exceeds.ai]
WF-RISK analogues
- Sigma Infosolutions,Proven ROI Framework to Measure AI Productivity, May 2026[sigmainfo.net]
- Gartner,AI Maturity Model and AI Roadmap Toolkit, 2026[gartner.com]
- Singhal P.,The Gartner AI Maturity Model, December 2025[puneetsinghal.com]
- McKinsey,State of AI Trust in 2026: Shifting to the Agentic Era, March 2026[mckinsey.com]
- McKinsey,State of AI: Global Survey 2025, November 2025[mckinsey.com]
- DORA Insights,Finding balance in the era of tokenmaxxing, June 2026[dora.dev]
- Zylos Research,AI Agent Cost Optimization: Token Economics and FinOps in Production, February 2026[zylos.ai]
- Zylos Research,AI Agent Cost Optimization: Token Budgets, Model Routing, April 2026[zylos.ai]
DedaAi internal corpus
- 04 – SDLC-Module-Start-Project v0.8— RACI SDLC Matrix × DevOps Phases[04 – SDLC-…aform_v0.2 | Word]
- The fourth circle in production: AI in the SDLC— Exploras, June 2026[exploras.cloud]
