I am pretty sure most of you have a list of things you let AI do, and a second list of things you would never let it near. I am also pretty sure neither list is written down anywhere.
You know exactly what is on them. I hope everyone on your team does too. But those are two very different lists, and the space between them is where the trouble lives.
Here is the pair of numbers that started this for me. Stack Overflow's 2025 developer survey found 84% of developers using or planning to use AI tools, while 46% say they do not trust the accuracy of what comes out of them. Which tells you something plainly: using AI is not the same act as delegating to it, and most teams have never said out loud which one they are doing.
The question is no longer whether your team uses AI
Everybody does in some form or fashion, or they had better be. So that question is settled. The one worth your time now is what you let AI decide, and what it costs you when you get that wrong.
The meeting readout that went out with names in it
This showed up at one of my clients a couple of weeks ago. They had started using an AI tool to summarize meeting notes, then push a readout in their format to whoever needed it. Reasonable. Fast. Everybody is doing some version of that, and the tool does it far better than we do.
One day the meeting ended and a few people stayed behind. They started talking about some of the other people who had been on the call, and not in a generous way. The transcription was still running. It captured all of it, with names and timestamps.
My client fed those notes into the tool, generated the summary and the readout, and skimmed it. He checked whether the summary matched what he wanted to say. He did not read the details closely. The last section was action items, and the tool had pulled open-ended action items out of the private conversation. The one that was no longer private.
Notice what did not go wrong here. The tool did not hallucinate. It did not malfunction. It did exactly what it was built to do. We are the ones who malfunctioned. Nobody had ever decided what that tool was allowed to act on.
We run a short review checklist in our classes for exactly this situation. Is what came out actually ours, actually relevant, actually actionable? I got to this person too late.
Human in the loop is the right instinct. But a loop with nobody's hand on it is just a diagram.
Four roles a machine can hold on your team
Years ago I took a certification class on AI and business transformation with Thomas Malone at MIT's Center for Collective Intelligence. In Superminds, he lays out four roles a computer can play relative to the people around it, and I still reach for this framing constantly.
A tool does what you tell it while you stay in control. An assistant works without you watching it. A peer handles part of the work and hands the rest back for you to look at. A manager assigns work and evaluates what comes back.
He published that in 2018 and it has held up better than most things written about AI since. Our processes are what have not caught up.
What 49,000 developers actually said
The gap I just described is not one client having a bad week. Stack Overflow runs the largest developer survey in the industry, and the 2025 edition pulled about 49,000 responses, fielded from late May into June. Good sample size. One caveat before you run with the numbers: they recruit mostly through their own channels, so it skews toward already-engaged users. Read it as a strong signal rather than a census.
84% use or plan to use AI tools, up from 76% the year before. About half of professional developers use it every single day.
Adoption went up. Trust went down.
46% distrust the accuracy of what comes out. 33% trust it. 3% highly trust it. And the most experienced developers are the most cautious of anybody in the sample.
Sit with that for a second, because it matters. The people carrying the most production scars trust it the least. I have plenty of production scars. I am in that group, and I suspect a lot of you are too.
The expensive problem is "almost right"
The survey asked about frustrations with AI, and the top answer surprised me. It had nothing to do with AI being wrong. 66% said their biggest frustration is AI solutions that are almost right, but not quite. Another 45% said debugging AI-generated code takes more time than writing it themselves.
I love that phrasing, because I could never quite put my finger on those words. Almost right is the bane of our existence.
That first number is really this whole episode. Something that is wrong is cheap. You catch it, you fix it, you move on, and the output cost you less than having anyone else produce it. Something that is almost right is expensive. It looks finished. It passes the glance. And those of you who are lazy, you know who you are, will say that is good enough. Worse, it is wrong in a different place every single time, so you cannot build a habit around it or code a check for it.
You can use a tool like that constantly. What you cannot do is hand it work and walk away. Those are two different relationships, and the whole point of this episode is knowing which one you are in.
Your developers already delegate by risk
Deeper in that same survey is a table that changed how I think about this. They asked which parts of the workflow developers are actually handing over. Roughly 54% already use AI mostly to search for answers, the way we all used to hunt for code examples. But 76% say they do not plan to use it for deployment and monitoring, and somewhere between half and 60% say the same about project planning, commits, and code review.
Read that again and you will see what I saw. That is a ladder. Developers are already delegating by risk. Low stakes handed over, things that are hard to reverse kept close.
The problem is they built it privately, in their own heads. Nobody wrote it down. So every person on your team may be running a different ladder, and that is the dangerous part.
One aside while we are in the data. Agents are not the flood you keep hearing about. In that same survey, only 31% use AI agents at work at all, and 38% have no plans to. That can change quickly based on what gets built next, but plan for where you actually are.
Personal speed, no system gain
Put the delivery data next to all that. DORA surveyed nearly 5,000 technology professionals in roughly the same window and found 90% of them using AI in their daily work, at a median of about two hours a day.
Throughput went up this year, which is a reversal from 2024. Delivery instability stayed elevated right alongside it. Friction and burnout did not move at all.
And here is the one that should stop leaders cold. Back in the Stack Overflow data, among developers who are actually using agents, around 70% say agents cut the time they spend on specific tasks and made them more productive. Only 17% say agents improved collaboration on their team. That was the lowest-rated impact in the entire set.
Individual work got faster. The system did not.
Anyone who works in flow will recognize that instantly, and I hope you do. You sped up one station on the line and the work did not leave the building any faster. You built inventory. It piled up somewhere else, in a queue nobody is measuring. Faster upstream, less stable downstream, and no shared agreement about who checks what.
The delegation ladder
So fix the decision mechanism rather than the tool. Be careful with anyone who tells you here is exactly how you do it, because this is new to all of us and there are no experts yet, only people who have been at it longer. What I can offer is the scar tissue and a mechanism worth trying. It has five rungs and it is progressive.
Rung 1. Draft. AI writes the first version and you own everything that happens after that.
Rung 2. Assist. AI suggests while you work. You spar with it and you are in the chair the entire time.
Rung 3. Recommend. AI proposes two or three or five options and you make the call.
Rung 4. Act with review. AI does the work, but nothing moves until a person looks at it.
Rung 5. Act with monitoring. AI does the work and keeps going. You watch for signals and exceptions.
Rung three is my favorite for innovation work. We say it on our teams all the time: collective intelligence is the key, and none of us is smarter than all of us together. AI happens to be one of those voices now. So spar with it. Who cares if it is wrong? You say that is terrible and you move on. We do that with people every day.
The one question that sets the rung
Your tool quality has almost nothing to do with the rung. You will always be tweaking that, or replacing it altogether. Two things actually set the rung: what it costs you when the output is wrong, and whether you would ever find out.
So ask one question about the task. If this is wrong and nobody catches it, who finds out first, and how?
If the answer is the test suite in four minutes, you can sit high on that ladder with a clear conscience. If the answer is a customer in three weeks, you belong lower. If the answer is a colleague reading their own name in a document, you belong a lot lower.
The ladder moves in both directions
Down first. That meeting readout was rung five behavior. The tool did the work and sent it out. Nothing in that system would ever have told anyone it went wrong. No signal, no exception report, and the person monitoring was skimming. That task belongs at rung four. Nothing leaves until a human reads it, including the section at the bottom that nobody seems to read.
Now up. Take dependency version bumps. Most teams sit at rung three, where AI recommends and a human approves each one. If you have a real test suite and a fast rollback, move that to rung four. The signal is genuine and it arrives within minutes, so a human reviewing the drift before the merge is enough. You just bought back an hour a week without buying any new risk.
Most conversations about AI are stuck arguing over a switch with two positions. This is a ladder, and it is supposed to move.
A rung is a team decision
This is the part teams skip. The rung has to be set out loud and written where people can see it. Right now, three people on your team might be holding three different rungs for the same task, and none of them know it, including them.
Your 15 minutes this week
Pick the single AI task your team does most often. Not the most interesting one. The most frequent one.
Write down which rung it is actually on today. Honestly, not aspirationally. Then ask the question out loud with your team in the room: if this went wrong and nobody caught it, who finds out first, and how? If the honest answer is that nobody would find out, move it a rung this week. Then tell your team you moved it and explain why.
That last part matters more than it sounds. I am guilty of skipping it myself, because I like to speed decisions up by making the call and moving on. Coaching and facilitating taught me over time that I need to run it by the team.
So the whole thing looks like this. One task, one rung, written somewhere your team can see it, one sentence explaining the change. Fifteen minutes. Bring your Scrum Master or agile coach into the conversation.
People ask me where to put it, and my answer is boring on purpose. Put it wherever your team already looks without being told. A pinned line in the team channel. A row in your working agreements. A column on the board. What matters is that a new person could find it in their first week, and that nobody has to ask you what the rule is.
And product owners, run this on your own work too. You are almost certainly using AI for backlog items, requirements, and discovery, and in my experience that is the one nobody is checking at all. More on that another time, because we do need to focus on it.
Decide it on purpose
Using AI is not the same as delegating to it. So decide what AI is allowed to decide, do it on purpose, and put it somewhere people can read it. That is where trust and confidence actually start to build, and it costs you a quarter of an hour.
If you want to work through this with your team and get the ladder into your real workflow, take a look at our upcoming classes. That is where we run the review checklists and the delegation conversations with the people who have to make these calls.