Your team shipped more code last quarter than in any previous quarter. Pull requests are up. You look at the heat map of the repo and the changes look good. One of the cheerleader developers on the team even dropped the rocket emoji in the channel. You know that guy.
And delivery is slower than it was last spring.
You are doing all these new things and it is still happening. When you start digging into it, you can plainly deduce that nobody stopped working. The work stopped moving. That is your cue.
So let's try this. Open your board and look at the review column, or whatever you call it. I bet that is where a lot of the work is sitting. And the team is desensitized to it, by the way. It has been sitting there long enough that nobody remarks on it anymore. It is unremarkable.
The callback: you sped up one station on the line
I have been stuck on this thread for a few weeks now, which is fine. I am not complaining. As I keep pulling on it, there is more and more to do on it.
Last week I said something I glazed over, and I want to pick it back up. You sped up one station on your line, but the work did not leave the building any faster. Go-to-market was not faster. It piled up somewhere else, in a queue that few people are actually measuring or are even aware of.
That conversation was about delegation. This one is about the queue itself, because the queue is what is causing the problem.
Generation got cheap. Verification didn't.
What appears to be happening to all of us is simple to say and hard to see. Generation is getting really cheap. Verification is not, yet. For most teams that has moved the bottleneck rather than removed it, and I almost never meet a team that has checked or explored where it went. They are too busy, and rightfully so. I am not condemning anyone for that. I am just trying to call it out.
Let me start on the cheap side, because it is easy to underrate. I might be the best example of that in practice, hence the scar tissue.
How much of your code now carries AI assistance
The first source I want to bring in is GitClear, which measures code repositories that its customers opt into. That makes it a slice of the industry rather than the industry as a whole. In their 2026 research, roughly a quarter of the commits they detected carried AI assistance.
Two caveats, and I would rather give them to you up front so you can make your own deductions. It is year-to-date from a partial sample, so read it as a direction and not a closed figure. And it counts only what their tooling can detect, so treat it as the floor. I still think it is worth talking about.
What developers say about "almost right"
Now the other side of the ledger, from a source we reference a lot: the Stack Overflow Developer Survey, run in the middle of 2025. It is self-reported, and respondents were found through Stack Overflow, so read that as engaged developers rather than a census. That is a real caveat.
Of the developers who answered the frustration question, 66% picked an output that is almost right, but not quite. It was a select-all question, so the options add up past 100.
Pause there for a second and breathe, because that is what this whole week is about.
Wrong code costs you thirty seconds. Almost right costs you an afternoon.
Obviously wrong code costs you thirty seconds. You read it, you reject it, you move on. Plain as day.
Almost right code costs a careful person the time to read every line and ask what it was thinking ahead of that, and then make a decision about it. It looks finished. It compiles. The tests pass. Nothing is visibly wrong. Somebody still has to work out whether it does the thing you actually meant.
In my experience that is the slow work, and it does not speed up with practice the way writing code does. How do you automate it without a dictionary of your own rules that then has to be maintained? That might be more work than you had before. We have not gotten around that yet.
In the same survey, 45% said debugging AI-generated code is more time-consuming. Both of those answers sit at review time, and neither of them is waiting time. That is somebody's attention.
The people doing this work are not sold on the output either. In that same data, 33% trust the accuracy of AI output and 46% distrust it. Only three points of that 33 land on "highly trust." Note that those two shares do not cover everyone who answered, because there is an unlabeled group in the middle, so the complement of 33 is not 67 the way most math people would want it to be.
I would not call that distrust a problem. That is people pricing the work correctly. Somebody still has to check.
Churn is rework, and rework lands on the reviewer's desk
Then there is what happens after the code gets pushed. GitClear also tracks churn, which means code that gets revised or reverted within two weeks of being written.
Their 2026 research puts two-week churn about 15% higher than its 2023 level, and from what I can tell the increase is largely one step between 2024 and 2025 rather than a smooth climb. The current-year figure is partial, so we still need to see where it lands at the end of the year. Their method also covers specific repositories in specific languages, so do not read it as a true statement for every team or every stack. That is impossible. But it is interesting enough that maybe you should go check your own.
That churn is rework. And rework sits on the same desk review is already sitting on. In my opinion they are the same category of work.
What DORA found in 2024, and what changed in 2025
One of my most cited sources is DORA, Google's long-running research program on software delivery. Some of the best data in the industry, in my opinion, is data that is hard to get in your own teams without a pile of analytics.
Their 2025 report reverses part of the previous year's finding, so if you only carry the 2024 headline around, you have the wrong version of this. That is how fast things are moving.
In 2024, DORA found AI adoption associated with lower throughput and lower delivery stability, and that result traveled a long way. In 2025, throughput flipped positive. The speed problem resolved itself as people learned the tools, the tools improved, and the work started moving faster.
The stability problem did not resolve. They still find AI adoption associated with rising delivery instability, and that covers two things they asked about: how often a change fails on deploy, and how often you ship an unplanned fix afterward. Both are self-reported, so take that for what it is worth. They also say plainly that their own data cannot explain why that one did not improve, which is still further than most teams have gotten with their own numbers.
Instability on its own may not automatically be bad. It means something is absorbing more risk. The question is whether you chose that when you designed the process.
Where their finding ends and my explanation starts
I want to be clear about the seam. The instability finding is DORA's. The explanation I am about to give you is mine, because I love process. I am a process guy. I believe in minimal viable process and minimal viable bureaucracy, and I think that is what gives us the guardrails to visualize and improve a system.
That same DORA report has a chapter on our favorite words, value stream mapping, which just means tracing the whole path work takes from ideating all the way to production. That is a meaningful agile concept. The chapter puts it better than I would, so I am going to borrow it. It says a team may find code review to be a significant bottleneck, and that the real win is using AI to improve review, because generating more code just makes the bottleneck worse. More code, more to read, more things to check.
That is guidance in their report rather than something they measured. But it points in exactly the direction I want to go.
Review is not a step. It's a queue.
What most teams are missing here is smaller and more structural than any strategy. Review is not a step. It is a queue, because it halts work. Until we find a better way to define it, that is the definition I am working from.
I love helping teams minimize friction in work. Bottlenecks and queues. Those are the words. And the best book I have found on this is still Donald Reinertsen's 2009 book, The Principles of Product Development Flow. It is a foundational model, not new research. It predates all of this by more than a decade, and it has been the backbone of process engineering ever since. Three of his points matter this week.
1. You cannot see this inventory
Product development inventory is invisible physically, and probably more hurtful, invisible financially. A plant manager can walk the floor and see a pile of inventory. Your CFO opens the balance sheet and finds no line item for half-reviewed pull requests. So the queue keeps growing and nothing reports on it.
Reinertsen argues that queues are the largest source of economic waste in product development, and I still believe that. Largely because they have no natural predators. They are just invisible.
2. Queue size does not grow linearly with load
This one keeps me up at night. Queue size does not increase linearly with load. It rises sharply as you approach full capacity, which is dangerous.
So think about a review step. Running it near full utilization backs up badly even when average capacity looks fine on paper, or fine on whatever a manager is tracking.
3. Rate matching
This is what my whole week runs on. Reinertsen calls it rate matching. You cap the work in process between two steps, and you force the arrival rate to match the departure rate. Sounds good in theory, right?
He takes the idea from data networks. You limit how many packets can be in flight before an acknowledgement comes back, which forces a fast sender down to the speed the receiver can absorb. That sounds dangerous the first time you hear it. It is also how two machines communicate reliably when one of them is a hundred million times faster than the other.
Generation is the fast sender. Review is the slow receiver. Right now nothing is capping the work in flight. We just keep shoving material downstream, like shoving more paper into a printer to make it print faster.
Size it: the four numbers
So what is the guidance? Size it. Don't roll your eyes. You can, but I have hardly ever seen anyone do this.
One definition first. An item is one thing that needs one review decision. A pull request, a merge request, a change list, a ticket, whatever it is you are doing. Not a commit. GitClear's quarter was measured in commits, and a pull request typically holds many commits. It is like a car with ten people in it.
Four numbers:
- How many AI-assisted items you put up for review in a week
- How many review hours a week you actually have
- How long a careful review of one item actually takes
- What share of reviews send the item back for another pass
Working the example
I am going to invent numbers to show you the shape. Use your numbers, not mine.
Say three qualified reviewers. Not everyone with merge rights. The people whose approval you would actually want on a change that touches something of value or money. Say each of them has six hours a week that genuinely goes to review, not a full week, six hours. They have their own work too. That is 18 review hours.
Say a careful review of an AI-generated change takes 45 minutes. Reading it, not skimming it, deciding whether almost right is right enough. Eighteen hours divided by 45 minutes is 24 reviews a week.
But one review in four sends the item back for another 45 minutes, and some go back more than once. Three quarters of those 24 reviews end in an approval, and each approval clears one item. That leaves you 18 items. That is what the team can verify in a week without cutting corners.
Now go count what you produced last week. Go look at your pull requests. If the answer is 60, you are running at more than three times your ceiling. That difference does not evaporate. It accumulates in a column.
Checking my math
Some of you are doing your own math in your head right now and checking me on it, and I love that, and you are getting a different number.
If you assume every rejected item clears on its second pass, you get 19. That is a fair reading and an optimistic one, because it assumes the rework never recurs. In my experience, rework recurs. And if you would rather think of it as each item costing you four-thirds of a review, that also lands on 18. Two ways to the same answer. Don't you love math?
I am going to start calling this arithmetic the verification capacity planner, and I am publishing it on Wednesday. It gives you a ceiling, shows you where your current volume sits, and recommends a limit on AI-generated work in flight.
It will not predict how long your queue will be. That is not what we are looking at here. Classic queuing math assumes a single server, and your review queue has several reviewers, I hope. So I am not handing you a number that is wrong in a way you could actually go check.
Where the limit actually goes
One more principle, because the obvious move here is the wrong one, and that gets us in trouble quite a bit.
Reinertsen talks about the critical queue. You put the limit where the queue is most expensive, not where it is easiest to measure, which is typically what we want to do. The expensive queue is finished work still sitting on a reviewer. So the limit does not go on review. It goes upstream, on how much you let into the production pipe.
You cap generation to protect verification. That feels backwards the first time you say it out loud, especially for those of us in leadership and management positions.
And the limit costs you something. Reinertsen is honest about that, so I will be too, for the utilization folks out there. A cap turns away work that might have been valuable and leaves some capacity idle. He argues a light limit is worth that trade-off and a tight one is not. Set it too aggressively and you will pay more than you save, which is counter to everything we are trying to do about friction.
The mindset problem is older than AI
The last piece is a mindset problem, and it is much older than AI.
Lean manufacturing taught a whole generation of managers to file testing under necessary waste, and necessary waste is something you minimize. Reinertsen argues that is a category error, because an activity adds value when it increases the economic value of what you are building. His example is testing. I am extending it to review, which could be argued is a form of testing, and I think it holds. You can of course disagree.
Verification is not overhead added to the real work. It is the work. And that is well worth saying in a year when generation is close to free. Most of us built our habits around getting in there, writing the code, producing as much as possible. The real work is whether it is the right work and whether it works.
So the scarce person on your team is no longer the one who can produce a plausible change. It is the one who can look at a plausible change and tell you whether it is right. Sometimes almost right is worse than wrong.
Stack Overflow put a hypothetical to developers: in a future where AI does most of the coding, when would you still want to ask a person? The top answer, at 75%, was when they do not trust the AI's answer. That is people imagining a future rather than reporting what they do today. But when they picture AI at its strongest, the human job they reach for is verification.
So embrace that. Cheap generation does not give you free verification yet. That is the lesson today. Maybe one day we fix that too, and we likely will. That is why I use the word "yet" in all of my change efforts. We can't do that yet.
Go count two numbers this week
Count what you generated. Count what you can actually review. If those two numbers are not close, you do not have a productivity story. You have a queue.
If you want help mapping your team's verification queue, I run short workshops that do exactly that. Bring your four numbers and let's see what happens. In the meantime, noodle on this a bit. If you are on our weekly newsletter we will recap it and talk through next steps, and the verification capacity planner goes up on the guidance blog on Wednesday. There is a deeper cut from the leadership angle on our LinkedIn newsletter.
And if you want to go further than a worksheet, this is the kind of system-level thinking we work through in our upcoming classes. Flow, queues, and the difference between a team that is busy and a team that is delivering.