
Your organization can see more of its delivery system than it ever could, and the people running that system say the work is getting harder rather than easier. Both of those can be true at once. The reason is not that anyone is ignoring the data. The data never had anywhere to go.
The visibility is in place. The decisions are not.
Digital.ai's 18th State of Agile report is a vendor survey, published by a company that sells enterprise agile planning and delivery tooling, and it is worth grading as one. It asked 349 self-selected practitioners about their delivery experience between July and August of 2025. On the infrastructure question the picture is strong: 55 percent report complete visibility into what is being developed and delivered across the software lifecycle, and 64 percent say their agile teams can see into the DevOps pipeline. The charts exist. The pipelines are wired.
The same block of questions also found 63 percent agreeing that their organization struggles to deliver reliable, high-quality software, which the report calls a twelve-point increase over the prior edition. And in that same block, 68 percent said at least half their applications ship on time and with high quality. Digital.ai states plainly that both cannot be fully true, and reads the gap as organizations either trading quality for deadlines or defining quality inconsistently.
Digital.ai, 18th State of Agile Report (vendor survey, n = 349): the 18th State of Agile report
Sit with that, because it is this week's argument in miniature. A measurement that can point in two directions at once cannot change a decision. Nobody acts on a number whose meaning is still unsettled, so the number quietly becomes something you look at rather than something you use. That is not the data's fault. It is a definition problem that nobody had time to settle.
Why adding dashboards can make things worse
A dashboard used to cost something. Somebody had to build it, wire the data, and keep it alive when the schema underneath it moved. That cost was a filter, and weak metrics died of neglect because nobody wanted to maintain them. A rollup is now one prompt away, the marginal cost of another chart is close to zero, and the filter is gone.
So the overhead grows with nothing stopping it. We audit spend. We audit headcount. We audit licenses nobody opened. The dashboard wall gets a pass, because it looks like the opposite of waste, and it can sit there for years without anyone asking what any of it changed. It does not help that the proxies teams reach for break in predictable ways once AI is in the loop, which means the wall can grow and get less useful at the same time.
Trust is not the missing piece either. In the same survey, 63 percent agree they trust the metrics their organization uses to evaluate team performance. That is a different question in the same block, and coincidentally the same number as the struggle figure above. It leaves more than a third running dashboards they do not trust, and that inversion is ours rather than the report's. But notice what even the trusting side does not tell you. A metric can be accurate, trusted, and beautifully rendered, and still change nothing.
The Scrum Guide is blunt about the order these things go in. Transparency enables inspection, and inspection without transparency is, in the Guide's words, "misleading and wasteful." Inspection then enables adaptation, and an inspection that never leads to a change has no point. Schwaber and Sutherland scope those pillars to Scrum's own artifacts and events; applying the same chain to your dashboard wall is our extension of it, not theirs. The chain still has to end in a change, or it was theater.
Mark Graban's Measures of Success adds the other half. It is foundational work from 2019, a model rather than a 2026 finding, and its argument is that movement in a metric is not automatically a signal. Leaders who react to every up and down spend time they could have spent improving the system, and that is Graban's own point rather than an inference drawn from him. His method is built for metrics tracked over time, so it is not a tool for comparing two surveys. The discipline underneath it still travels, and it is the same one worth applying to that twelve-point swing: interrogate a move before you react to it. Whether a number moved is a different question from whether anything changed.
Graban's point is about how to read a metric worth charting. This one sits upstream of that. Before you ask what a number did this week, ask whether that metric has ever changed what you do.
The Dashboard-to-Decision Audit
Stop auditing dashboards by their accuracy and start auditing them by the one thing that justifies keeping them, which is the decisions they change. One record per active metric, five lines each. One metric or dashboard per record, no bundling.
DASHBOARD-TO-DECISION AUDIT
One record per active metric. Run it with the people
who read the metric, not the people who built it.
METRIC ____________________________________
READ BY ____________________________________
LAST CHANGED ____________________________________
IF GONE ____________________________________
VERDICT keep / cut / merge into ____________
DECISION RULE
A metric that has not changed a decision in 30
days is a candidate for removal. Not removed.
A candidate, with a case to answer.
RENT
Obligated metrics (compliance, board rollups,
contractual reporting) go on a separate list.
They are not audited here and they are not a
fourth verdict. Keep, cut, and merge remain the
only three.
GUARDRAIL
A metric that exists to fire on an event you hope
never happens changes no decisions by design.
IF GONE is the line that saves it.
DONE WHEN: every active metric carries a verdict,
and every cut or merge has an owner and a date.Line one. METRIC.
Name it. One metric or one dashboard per record. The moment you bundle four charts into a single line, the record can no longer produce a verdict, because the answer will always be that some part of it is useful.
Line two. READ BY.
Who actually looks at it, not who receives it. Those are different lists, and the gap between them is usually the whole story. Ask the named readers directly rather than inferring from a distribution group. You will trim this list fast.
Line three. LAST CHANGED.
The last decision this metric changed, and when. Not informed. Not supported. Changed. The question that settles it is whether the metric had said something different, would you have done something different. If the honest answer is no, write that down instead of a softer version.
Line four. IF GONE.
What would be different next month if this metric disappeared tonight? Be specific, because "people would miss it" is not a difference. A decision made later, a risk caught slower, a conversation that stops happening: those are differences.
This line also carries the one exception the audit needs. A metric that exists to fire on an event you hope never happens will fail lines two and three by design, the way a smoke detector nobody has ever heard fails them. It has changed no decisions and almost nobody watches it. If it were gone, the first thing to tell you would be the fire it existed to catch. Guardrail metrics are kept on line four, not on line three, and the record should say so in as many words.
Line five. VERDICT.
Keep it, cut it, or merge it into something that earns its place. Three verdicts, no fourth. The rule that keeps everyone honest is that a metric that has not changed a decision in thirty days is a candidate for removal, and a candidate is not a removal. It is a case to answer, and only the people who run that system can answer it.
Some metrics are rent. Compliance reporting, board rollups, anything a contract requires: those are obligations, not instruments, and they exist whether or not they change a decision. Put them on their own list so they stop polluting the audit. That is a scope carve-out and not a fourth verdict, and the difference matters, because "rent" is otherwise the place every awkward metric goes to survive.
Three signs your visibility has turned into surveillance
The first sign is the direction of the audience. When the people a metric is actually read by all sit above the team producing it, and nobody on the team uses it to decide anything, the chart has stopped being an instrument and started being a report card. Teams respond rationally to report cards by managing the number.
The second sign is that the metric has no losing move. If every possible reading of it produces "keep going" or "try harder" and none produces a specific change, the metric is not measuring the system, it is measuring compliance with the act of reporting. The gut check for any productivity number is what decision it changed, and a metric that cannot fail that check is not neutral, it is expensive.
The third sign is that removal is unthinkable but nobody can say why. Ask what would break, and you get discomfort rather than a consequence. That is the signature of a metric being kept for what it signals about the people watching it, not for what it tells them.
Try this next week
Pick one metric and fill in one record. Not the wall. One metric. If you already have one in mind, start there. If you do not, take the rollup that summarizes the other reports, because it is the one most likely to fail the test and the cheapest to stop producing. Cutting it does not remove any information, since the underlying dashboards still exist. It removes the assembly work, and it hands the remaining dashboards the same five questions.
What the team relearns from that is worth more than the hour it takes. A metric on the wall is a claim that it changes decisions, and claims get tested. If you want the version of this argument that sits one level down, in the review capacity that all this reporting is supposed to be protecting, that was last week's piece.
Three completed records, so the shape is unambiguous. All three are illustrative and fictional.
ILLUSTRATIVE EXAMPLE 1 OF 3. FICTIONAL.
METRIC Weekly Delivery Executive Summary
READ BY Presented monthly to the VP staff.
Asked the named readers directly:
none read it before the meeting.
LAST CHANGED No decision named this quarter.
IF GONE The four source dashboards still
exist. Only the assembly work is
lost.
VERDICT Cut. Owner: delivery ops lead.
Date: first week of next month.ILLUSTRATIVE EXAMPLE 2 OF 3. FICTIONAL.
METRIC Escaped Defects by Team
READ BY Two engineering managers, weekly.
LAST CHANGED Moved a regression suite earlier in
the pipeline, six weeks ago.
IF GONE The same signal reaches us through
change failure rate about a week
later.
VERDICT Merge into the delivery quality
review. Owner: QA lead. Date: end
of the quarter.ILLUSTRATIVE EXAMPLE 3 OF 3. FICTIONAL.
METRIC Payment Reconciliation Mismatch
Alert
READ BY On-call engineer, on trigger only.
LAST CHANGED Has not fired in eight months.
IF GONE The first thing to tell us would be
the mismatch itself, after customers
had already felt it.
VERDICT Keep. Guardrail metric: it exists to
fire on an event we hope never
happens.Visibility is not insight, and insight is not a decision.
If the hard part is not building the audit but getting a room of leaders to agree to cut something in public, that is the work we do together. Big Agile coaching is built around exactly that kind of decision, with the people who have to live with it afterward.
