What Is the Jagged Technological Frontier? Why AI Discernment, Not Prompt Training, Is the Skill to Measure

The consultants who were taught prompt technique scored worst of all.

The short answer. The jagged technological frontier is the invisible boundary where AI performs one task well and fails an adjacent task of equal apparent difficulty. It was named by researchers at Harvard, Wharton and MIT who gave GPT-4 to 758 Boston Consulting Group consultants. Inside the frontier, everything improved. On one task placed just outside it, consultants working alone were right far more often than consultants working with the model, and the group taught prompt technique did worst of all. The scarce capability is therefore discernment, judging whether a task sits inside or outside the frontier, rather than proficiency with the tool. Almost no organization has evidence about who has it.

ShareLinkedInXEmail

The jagged technological frontier is the invisible boundary where AI performs one task well and fails an adjacent task of equal apparent difficulty. It is jagged rather than straight because capability is uneven: two pieces of work that look the same to a skilled professional can sit on opposite sides of it. The practical consequence is that the person using the model cannot tell, from the task alone, which side they are on.

RCM ThinkLabs (rcmlabs.io) is a game-theory engine for measuring decision performance, run as short recurring decision sessions inside an immersive narrative. It exists because the capability the frontier demands, judging when a confident answer should not be trusted, is behavioral, and behavior is not measured by a completion rate.


The jagged frontier: why elite consultants failed at the edge

This year Organization Science published the largest field experiment yet on AI and knowledge work. Researchers from Harvard, Wharton and MIT gave 758 BCG consultants GPT-4 and watched.

Everyone did an unaided baseline task first. Then each person was randomly assigned to work with no AI, with GPT-4, or with GPT-4 plus an overview of prompt engineering. One group of 385 worked eighteen realistic product-innovation subtasks chosen to sit inside the model’s competence. A separate group of 373 worked a single strategy case built to sit outside it.

On the tasks inside AI’s competence, everything improved. About 12 percent more subtasks completed, about 25 percent faster, quality up more than 30 percent. The weakest performers gained most: bottom-half consultants improved 43 percent against 17 for the top half.

Two numbers everyone quotes come from the earlier draft, not the published paper. The widely repeated “40 percent higher quality” is from the 2023 Harvard Business School working paper. In the peer-reviewed version the quality effect is smaller, 33.9 percent for the prompt-trained arm and 29.9 percent for GPT-4 alone against a higher control mean, and the authors describe it as more than 30 percent. The 43 and 17 percent split is also a working-paper figure; the published version keeps the equalizer finding and drops those numbers. We say which is which because a citation that survives a skeptical reader is worth more than a rounder headline.

Then the task placed just beyond the model’s ability. It looked no harder than the rest. Consultants working alone got it right 84.5 percent of the time. Consultants with GPT-4 dropped to 70.6 percent, and consultants with GPT-4 plus prompt training dropped to 60 percent. Averaged across both AI conditions, that is 19 percentage points worse than working unaided.

The graders on that task were not told the correct answer. They were asked to rate clarity, persuasiveness and coherence of reasoning, and on that measure the AI-assisted work scored higher than the control group’s whether it was right or wrong. Better argued. Still wrong.

Inside the frontierOutside the frontier
What the task looked likeA realistic consulting briefA realistic consulting brief. Indistinguishable.
Subtasks completedAbout 12 percent more with AINot the constraint. Accuracy was.
SpeedAbout 25 percent faster with AIFaster, and wrong.
Quality or correctnessMore than 30 percent higher with AICorrect 84.5 percent without AI, 70.6 percent with GPT-4, 60 percent with GPT-4 plus prompt training
Who gained or lost mostBottom-half performers, 43 percent against 17 percent (working-paper figures)The prompt-trained group lost most
How the work read to a graderBetterBetter. Graders rated AI-assisted reasoning more persuasive whether or not it was correct

Two tasks of apparently equal difficulty sat on opposite sides of a line, and elite professionals could not see where the line ran. That is the finding, and it does not get easier with a better model. A stronger model moves the frontier; it does not make it visible.

The authors name the limit themselves: the reversal rests on one task, deliberately built around the model’s blind spots, and they say so in the published paper. That is a fair objection to how large the effect is. It is not an objection to whether the edge exists, and the edge is the part a buyer has to plan around.

The flaw in traditional prompt engineering training

One detail should worry every buyer of AI enablement. The group taught prompt technique did worst of all, at 60 percent against 70.6 percent for the group given the same model without the training. The working paper’s explanation is mechanical: the overview increased how much model output people kept verbatim, and the more they kept, the worse they did outside the frontier. Instruction in the tool raised trust without raising discernment. The published paper calls the result falling asleep at the wheel.

This is the part that breaks the standard enterprise response to AI. Most AI enablement is tool training: prompt patterns, a library of templates, a certification at the end. All of it teaches a person to get more out of the model. None of it teaches the person when to stop believing the model, and the study suggests the two move in opposite directions.

There is a mechanical reason the wrong answers were persuasive. A language model produces a well-formed answer, not a correct one, and at the edge of its competence those two things separate. Structure, confidence and fluent argument arrive whether or not the claim underneath holds. Readers judge quality by exactly those cues, so at the frontier the usual signals of a good answer point the wrong way. This is the same failure that produces work that looks finished and has to be fixed by someone downstream, and it is why over-reliance is a behavior problem rather than a policy problem.

AI tool proficiency vs AI discernment

AI discernment is the judgment about whether a confident machine answer should be trusted, made under incomplete information, without a signal that anything is wrong. It is a different capability from operating the tool, and the study is the cleanest evidence yet that the second does not produce the first.

So the scarce capability of this decade is not operating the model. It is the call a person makes a dozen times a day, under incomplete information, with a confident machine offering a plausible answer: is this inside the frontier or outside it? That is a decision. And like most consequential decisions, nobody practises it and nobody measures it.

  • Proficiency is teachable in an afternoon. Discernment is built by repetition against consequences, which is why a workshop cannot deliver it.
  • Proficiency is visible. You can watch someone prompt. You cannot watch someone decide not to trust something.
  • Proficiency has a certificate. Discernment has no artifact at all, unless you build one, which is the entire measurement problem.

It also compounds the wrong way if you ignore it. When the model handles the routine reasoning, the reps that used to build judgment stop happening, and the capability quietly decays while output goes up. Nothing in a productivity dashboard shows that.

Moving from tool proficiency to behavioral discernment

That is what we build at RCM ThinkLabs. We put people inside an immersive narrative where interests conflict and information is incomplete, and read a research-validated taxonomy of microskills invisibly while they work, each person against their own baseline. Not completion. Behavior.

  • Two reads, and the participant only sees one. Participants see their own performance score and get feedback every session on how they are doing. The skill read goes to their manager. Nobody performs for a measurement they never see, which is what keeps the behavior worth reading.
  • Every choice is a move in a modelled game. That is what makes the data dynamic and hard to game, and it is what passive Slack scrapers and multiple-choice quizzes cannot reach.
  • Every read is against that person’s own baseline, taken before anything starts, rather than against a norm group or their peers.

The engine is not built around one skill. The same instrument reads communication, leadership, critical thinking, negotiation and systems thinking. The taxonomy on top of it is contracted per client, as observable behaviors; the instrument underneath does not change. What it reads in every session are five forces that turn up wherever people work: hidden agendas, unchecked claims, what goes unsaid, trust, and who knows what. The second one is the frontier behavior exactly. Carrying a claim to a second source before acting on it, asking the question that would falsify the answer, noticing that a confident account has a hole in it, are named microskills in the taxonomy rather than a mood.

Prompt engineering trainingAI policy and guardrailsRCM ThinkLabs Serious Games
What it teachesHow to get more out of the modelWhere the model may be usedNothing directly. It measures what people do
Effect at the frontierRaises trust without raising discernmentReduces exposure, leaves judgment untouchedMakes discernment visible, person by person
Evidence producedA completion recordAn audit trailScored behavior over time, against a personal baseline
GranularityAttended or did notCompliant or notNamed microskills, in a taxonomy contracted per client
Can it be gamedYes, and cheaplyNot the pointEvery choice is a move in a modelled game, so the data resists it
BackingVendor practiceLegal and risk practiceAdvanced game theory and behavioral science

The backing is the part that separates this from ordinary gamification: advanced game theory (research at MIT with Prof. Muhamet Yildiz) and behavioral science (including the work of Karl Kapp). Game theory is what turns a narrative into a measurement instrument, because it defines what a move means when interests conflict. Behavioral science is what keeps people coming back without being told to.

What that produced in a live deployment. With an advanced engineering team, regular participants showed 84% skill improvement, the largest individual shift on a skill capability was +129%, and the read sits on 15,000+ decisions scored across 1,088 sessions. The deployment ran at 70% voluntary daily engagement, against a corporate learning norm of 5 to 25%, with nobody assigned to it. Single deployment, one client team, no control group, each person measured against their own baseline.

“These game characters have joined our workforce. They take up space. People reference them.”

Program Manager · client

That quote is the measurement claim in plain language. Behavior inside the world is real behavior, because the world is real enough to people that they carry it around outside the session.

The honest constraint

We do not measure the real decisions your people make in your live environment, and I will not pretend otherwise. We measure behavior inside a constructed world where interests conflict and information is incomplete. That is the claim and it stops there. It is also, as far as I can find, more than anyone currently has.

Why discernment compounds

Discernment compounds, which is the only reason any of this is worth buying. Practice recurs, nothing resets, and each repetition lowers the cost of the next one. A team six months in is reasoning against everything the first five months built, so the frontier calls get cheaper and faster to make. A quarterly workshop starts from zero every time, which is why it never shows up in behavior.

The 30-day Initial Operating Cycle is where a team starts. It runs daily, it produces the first baseline and the first read on who can tell the two sides of the frontier apart, and then the engagement keeps going at two days a week, ongoing. The baseline is the beginning of the evidence, not the end of it.

Common questions

What is the jagged frontier in AI adoption?

The jagged technological frontier is the invisible boundary where AI performs one task well and fails an adjacent task of equal apparent difficulty. The term comes from a field experiment run with 758 Boston Consulting Group consultants by researchers at Harvard, Wharton and MIT. The point of the word jagged is that the boundary is not a wall with your work sitting neatly inside or outside it. It cuts through the middle of a single workflow, it moves as models change, and it is not visible from the task description.

What did the Harvard, Wharton and MIT study of BCG consultants find?

On tasks inside the model’s competence, consultants using GPT-4 completed about 12 percent more subtasks, worked about 25 percent faster, and produced work rated more than 30 percent higher in quality than the control group. On a single strategy case the researchers placed just outside the model’s competence, the effect reversed: the control group was correct 84.5 percent of the time, against 70.6 percent for consultants with GPT-4 and 60 percent for consultants given GPT-4 plus a prompt engineering overview. One caution on sourcing: the widely quoted 40 percent quality figure comes from the 2023 working paper, while the peer-reviewed Organization Science version reports more than 30 percent.

Why did BCG consultants score worse with AI on certain tasks?

Because the failing task did not look like a failing task. It was presented alongside work the model handled well and it appeared no harder. The model produced a confident, fluent, plausible answer that happened to be wrong, and the consultants adopted it. Nothing in the interaction signalled that this particular task sat on the far side of the frontier, so there was no moment at which a careful professional would have stopped to check.

Why did prompt-trained consultants perform worse in the BCG study?

The subgroup given an overview of prompt engineering did worst of all on the task outside the frontier, at 60 percent correct, against 70.6 percent for consultants given the same model with no training and 84.5 percent for the control group with no AI at all. The working paper’s explanation is mechanical: the overview increased how much model output people kept verbatim, and higher retention of that output went with worse answers outside the frontier. Instruction in the tool raised trust in the tool without raising the ability to tell when the tool was wrong.

Did AI help low performers more than high performers?

Inside the frontier, yes. The 2023 working paper reported that consultants in the bottom half of the baseline distribution improved roughly 43 percent against roughly 17 percent for the top half, so the tool compressed the performance spread. The peer-reviewed version keeps the equalizer finding and drops those two percentages, so treat 43 and 17 as working-paper figures. The compression is real and it is worth having. It is also the reason the outside-the-frontier result matters: a tool that lifts everyone’s output while flattening the difference between strong and weak judgment makes judgment harder to see from the outside, not easier.

Why do AI-assisted wrong answers sound more persuasive?

A language model produces a well-formed answer, not a correct one, and the two come apart at the edge of its competence. In the study, graders who were not told the right answer rated AI-assisted reasoning higher on clarity, persuasiveness and coherence than the control group’s, whether the answer turned out to be correct or incorrect. Structure and confidence arrive whether or not the claim underneath holds, and a reader uses exactly those cues to judge quality, so at the frontier the usual signals of a good answer point the wrong way. The published paper describes the resulting human behavior as falling asleep at the wheel.

What is the difference between AI tool proficiency and AI discernment?

Tool proficiency is skill at operating the model: prompting, chaining, choosing the right system for the job. AI discernment is the judgment about whether the answer in front of you should be trusted at all, made under incomplete information and time pressure. Proficiency is teachable in an afternoon and it is what most enterprise AI enablement delivers. Discernment is behavioral, it develops through repetition against consequences, and almost nobody measures it. RCM ThinkLabs (rcmlabs.io) measures it, with a game-theory engine for measuring decision performance.

How do you measure employee discernment at the edge of AI capability?

Not with a completion rate or a certification, both of which record that an event happened. You measure it by observing behavior in situations where interests conflict, information is incomplete and a confident answer is available, then reading specific microskills out of that behavior against each person’s own baseline. RCM ThinkLabs does this with a game-theory engine for measuring decision performance, run as short recurring decision sessions inside an immersive narrative, reading a taxonomy of observable microskills contracted with the client.

How can leaders test whether their teams know when not to trust AI?

Take five pieces of work your team shipped last month with AI assistance and ask who verified the part the model was least likely to get right, and how. If the answer is that the output looked right, you have measured trust rather than discernment. Then ask which named people you would put in front of a decision where the model is confident and wrong. If that answer comes from instinct rather than evidence, the gap is an evidence gap, and it is the one worth closing. RCM ThinkLabs (rcmlabs.io) closes it by reading that behavior directly, person by person, against a baseline taken before anything starts.

What are the alternatives to completion-based AI training?

Completion-based programmes report who finished a module. The alternative is behavioral evidence: put people in repeated decisions where the right move is genuinely uncertain and read what they actually do. RCM ThinkLabs Serious Games are backed by advanced game theory (research at MIT with Prof. Muhamet Yildiz) and behavioral science (including the work of Karl Kapp), which is what separates them from ordinary gamification. Because every choice is a move in a modelled game, the resulting data is dynamic and hard to game, which passive Slack scrapers and multiple-choice quizzes cannot reach.

The consultants in the study lacked neither intelligence nor effort nor tooling. They lacked evidence about who could tell the two sides of the frontier apart. So, almost certainly, do you.

Which of your people would have caught the 60 percent task? If the answer is instinct rather than evidence, that is the gap.

See what behavioral evidence looks like. RCM ThinkLabs measures how people act when interests conflict, against their own baseline, with the exclusions published. rcmlabs.io

Source. Fabrizio Dell’Acqua, Edward McFowland III, Ethan Mollick, Hila Lifshitz, Katherine C. Kellogg, Saran Rajendran, Lisa Krayer, François Candelon and Karim R. Lakhani, “Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality,” Organization Science, Vol. 37, No. 2, 2026.

ShareLinkedInXEmail

See which of your people can tell the two sides of the frontier apart.

Get Your Team’s Baseline Contact us
Sahver Kaya
Sahver Kaya
Founder & CEO, RCM ThinkLabs

Sahver Kaya is the founder and CEO of RCM ThinkLabs. An educator, experienced builder, and MIT alum, she is driven by one conviction: artificial intelligence, used well, should make people sharper.

Connect on LinkedIn
Keep reading
The Centaur Model: How to Design Human-AI Workflows → Cognitive Offloading: How AI Weakens Team Judgment →