A team I worked with recently had just finished an AI training course I delivered personally. Certificates all around. Everybody in the room could write a clean prompt with self-consistency checks, which is the part of modern prompting that actually matters. They were proud of it, and they should have been. Six weeks earlier, half of them were afraid to open the tool. Now they were using it every day, experimenting responsibly, and growing.
About two weeks after the course, a model handed one of them an answer that was confident, specific, and completely wrong. It cited a number that didn't exist. But it sounded exactly right, so it went into a document, and the document went to a client before anyone caught it.
When I asked the team how it slipped through, the answer was simple: nobody questioned it. It looked fine. And honestly, we all do that, not just with AI. They had been trained to operate the tool and check the results. What they hadn't practiced was doubt. That gap, the space between using the tool and knowing when not to trust it, is the most expensive thing in AI right now. So here's the question I'd put to any leader spending money on this: who's measuring that gap, and what are you doing about it?
What most AI training actually teaches
Once you see this, it's hard to unsee. Look at what almost every AI course teaches: prompts, features, shortcuts. Which phrasing gets a better answer, which model for which job, how to structure a request so the output comes back usable. All of that is real, and for people who are new to the tools it's valuable. You don't want a team running automations before they can understand and validate what's happening underneath.
But all of it is one thing. It's tool fluency, and it answers a single question: can this person operate the tool? For a while, that was the whole game. The interfaces were awkward and strange, and knowing how to drive them was a genuine edge. That's what all of us have been trying to learn.
That edge is gone. The tools got easier every month. The interfaces got friendlier. Fluency got cheap, and it's getting cheaper. So if fluency is the thing your training produces, you're investing hard in a skill the market is busy making worthless.
The skill that's actually scarce
The skill that's still scarce is judgment under uncertainty: knowing when to trust an output, when to verify it, and when to refuse it outright. That one isn't getting cheaper. As the tools get more capable and more convincing, it's getting more valuable and more rare.
Fluency and judgment are different skills, and teaching the first does not give you the second. You can be completely fluent and still have no idea the machine is confidently lying to you. Being good at asking is not the same as being good at doubting. I don't have a better word for it than that. You want people who will doubt the tool.
Why we keep training fluency and calling it capability
Why does almost everyone keep training fluency and then call it capability? Me included, by the way. Because fluency is easy to teach. You can build a curriculum, run a workshop, give a test, hand out a certificate. Everyone can see the progress, and it feels like a result. It is a good one.
Judgment doesn't behave that way. It doesn't show up on a test. It shows up in a real moment, when someone looks at a plausible answer and decides not to trust it. That's hard to certify in a two- or five-hour class, because it takes time and it takes coaching. So most programs quietly stop trying and optimize for what they can measure. That's reasonable, right up until you remember what you were actually trying to build. You end up with a program that gets very good at producing fluent people but never finds out whether they can tell a good answer from a convincing one.
What the Deloitte data shows
Deloitte's State of AI in the Enterprise report has real data on this. Leaders name insufficient worker skills as the single biggest barrier to getting value from AI. That tracks. Skills are a real barrier and I wouldn't argue otherwise.
But look at what organizations do about it. The same research finds that the most common response, by a wide margin, is to educate the broader workforce to raise AI fluency. It's the top talent move companies are making, and it's not a bad one. Here's the problem: the barrier they name is skills, and the fix they reach for is fluency. That only works if fluency was the skill they were missing. At first, maybe it is. But fluency was never the real barrier. The barrier is judgment.
And here's the trap. Pouring more fluency training onto a judgment gap doesn't close it. It makes it worse. (Deloitte's own numbers hint at why: 84% of companies say they have not redesigned jobs around AI. Everyone is teaching the tool; almost no one is rebuilding the work around it.)
Confident, unsafe users
Think about what you produce when you train fluency hard and never train judgment. You get people who score high on operating the tool and low on questioning it. That combination has a name: confident, unsafe users. They reach for the tool constantly, and they trust it when they should have stopped.
These are not the people who avoid AI. They're the people who use it fastest and question it least. And a certificate on the wall makes everyone feel the risk was handled when it wasn't. That's a coaching problem, not a training problem.
I want to be clear: I'm all for certifications. Prove you went and learned something. I hand them out myself. What matters is how you use the thing you got certified on, and that's what we try to build into the workshop. I'm a bleeding-edge technology guy at heart, so I've been burned at least as often as the people I'm describing. That scar tissue is a big reason Big Agile exists, and it's what I'm sharing on the NextGen Agility podcast.
What training judgment actually looks like
If fluency training is the wrong shape, what does judgment training look like? Almost nothing like a prompt drill. A prompt drill hands you a clean task, a mostly known answer, and tidy data, and it rewards you for phrasing. You're practicing form on a problem that was never going to go wrong, because I built it to be safe. Judgment practice hands you a mess. The output looks reasonable on the surface, the situation is ambiguous, and your actual job is to decide whether the thing in front of you can be trusted at all.
Exercise 1: catch the answer that's confidently wrong
Take a real question from your own work and run it so the answer comes back fluent, specific, and quietly wrong. Not obviously broken. Confidently wrong. (That's hard to engineer, by the way. I try to build it into exercises all the time.) The task is not to fix the prompt, because the prompts are usually fine. The task is to catch the error before it ships, then say out loud how you caught it.
Slow the person down enough to actually see it, because you are human and you cannot keep up with everything the tool is doing. What tipped you off? What did you check? The goal isn't a lucky catch. It's a habit the team can repeat when the stakes are real. And it only works if the example is genuinely wrong. Clean, correct examples teach people to trust the output, which is the opposite of the skill you're trying to build.
Exercise 2: the question the model can't answer
This one uses a different muscle. Give the model a question it genuinely cannot know, something that needs a person, a source, or a decision it was never entitled to make. A fluent user takes the confident answer and moves on. Someone with judgment recognizes that the answer shouldn't exist and escalates instead of accepting it. Knowing when to verify, when to escalate, and when to refuse: that's the muscle.
Psychological safety is not a soft add-on
One condition makes all of this work, and a lot of teams skip it. People only build the doubting muscle when it's safe to use. Psychological safety, in Amy Edmondson's sense, is the belief that your environment is safe for interpersonal risk: safe to raise a concern or admit doubt without looking like you don't get it.
Put that against AI. Someone on your team spots that the model is wrong. If they can't say so without seeming like they don't understand it, they'll stay quiet and accept the bad output. It happens constantly. The judgment gets trained, and then it never gets used, because using it feels dangerous. The environment isn't a nice-to-have here. It decides whether the skill you paid to build ever shows up.
The layer training programs skip most
There's a third layer, and it's the one programs skip most often. I've skipped it too, early on. None of this matters unless it changes what people do next week. So hold three layers together:
Tool fluency. You do need it. But fluency without judgment gives you confident, unsafe users.
Judgment practice. The muscle that catches the confident, wrong answer. But judgment without transfer is a classroom trick that evaporates the moment people get back to their desks.
Workflow transfer. The test. If people leave the session and their actual work is unchanged, the training didn't happen, no matter how good the room felt.
You need all three, and the value lives in the last two. Most programs spend everything on the first.
One swap to try in your next session
Here's something concrete. Take one prompt drill out of your next session and replace it with a "should we trust this output?" exercise. Use a real output from your own work, and make it a wrong one. Ask the team to catch the error, then ask what they'd do next in the actual workflow, not in theory.
That one swap moves the whole session from "can they operate the tool" to "can they decide when not to." Same hour, same people, a completely different skill walking out the door.
That's the trap in one sentence: most AI training is solving the wrong problem. It teaches people to use the tool and skips whether they should. The scarce skill isn't using AI. It's knowing when not to trust it. If you want to build that into your team on purpose, judgment as a practice and not an accident, that's exactly the work I do.
Start building judgment on purpose
If your team is fluent but you've never tested whether they can spot a confident, wrong answer, that's the gap to close first. Our public workshops and classes at big-agile.com are built to teach the tool and the judgment to question it, and to help you carry both back into the actual workflow. Come earn a credential to get started, and let's build judgment into how your team works, not just what they know.