What AI is good and bad at
The same AI gave opposite results on different tasks
In an experiment with 758 BCG consultants, using AI on tasks it is good at raised the quality of the work by more than 40%. On a task it is bad at, accuracy fell by 19 percentage points.
- Tasks completed (AI’s strengths)
- +12.2%
- Quality (AI’s strengths)
- +40% or more
- Accuracy (AI’s weak spot)
- −19 pt
The same people using the same AI did better on one kind of task and worse on another. What made the difference was not the tool, but whether the task suited AI.
What the study found
The experiment was run by researchers from Harvard Business School and elsewhere, together with Boston Consulting Group (BCG). It involved 758 BCG consultants, randomly split into a group that could use AI (GPT-4) and a group that could not.
The researchers call the boundary of what AI can do the “jagged technological frontier”. The idea is that the boundary is not a straight line but uneven. Tasks that look equally hard can sit side by side, one easy for AI and the other still beyond it.
In this study, on tasks inside AI’s strengths, consultants using AI completed 12.2% more tasks on average than those without it. They worked 25.1% faster, and the quality of their work was more than 40% higher. The tasks were close to real work, such as planning a new shoe, analyzing a market and writing a press release.
On these tasks, AI raised the performance of both below-average and above-average consultants. The biggest gains went to those below average.
On a task placed outside AI’s strengths, however, consultants using AI were 19 percentage points less likely to reach the right answer. The task looked as if it could be solved from spreadsheet data alone, but the key clues were actually in interview notes with people inside the company.
AI can give plausible answers even when it is wrong. According to the researchers, consultants who did worse with AI tended to accept its answers as they were and question them little.
Those who used AI well showed two patterns. Some drew a clear line between what they gave the AI and what they did themselves; others worked with the AI closely throughout the whole task. The researchers call the first “centaurs” and the second “cyborgs”.
For small teams and sole proprietors
The lesson for small companies: AI is not a switch that works equally well on every job. On some jobs it works very well; on others, plausible mistakes come easily. What matters is telling them apart.
General-purpose AI tools are built for average use. Drafting a quote from your own price list and answering a question without being given the source material are very different in how reliable the result is.
You can’t tell which work suits AI by how hard it looks. In this study, too, results split on tasks of similar difficulty. The safe way is to try small, check the results, and only then widen what you hand over.
In this experiment, the clues to the right answer were not in the tables of numbers but in the details of the interview notes. Small companies’ work also rests on things that are written down nowhere, like conversations with customers and past history. Giving that information to the AI, and deciding where a person checks, is what we believe makes the difference.
For mid-sized and larger companies
At scale, invisible variations in quality become a problem. Some teams will achieve great results, while others may trust plausible answers and let quality slip without noticing.
The study also points out that where AI’s strengths end is hard for users to see. What’s needed is design more than enthusiasm.
- Before launch, sort the work into what suits AI and what doesn’t
- Have AI agents answer from your company’s data, not general knowledge
- Always require human approval for actions where a mistake would be costly
- Log every action, so mistakes can be found and fixed
- Help staff recognize when an AI answer needs checking
AI’s strengths shift as the technology advances. Any line you draw needs regular review.
How we use this
First, we sort the work into three groups. In the free AI audit (30 min) and the interviews after it, we learn your work, tools and data. Then we sort the work into jobs for AI agents, jobs where AI drafts and a person approves, and jobs that stay with people.
We make the agents work from your material. With agentic RAG, CAG and knowledge graphs, the AI agents answer from your company’s documents and data. When no source can be found, they don’t fill the gap with a guess; they pass it to a person.
Judgement stays with people. Important actions pass through an approval gate (a person always approves important actions), and every action is logged. Data is handled in a private environment (data is never used for training).
We keep watching the frontier. Our AI agent Polaris (support and monitoring) watches over how things run, and the results dashboard (time saved, cost per task and accuracy, checked weekly) keeps an eye on the numbers. As your work and AI change, we review what the agents handle.
We teach your team how to check. In team training, we show concrete ways to check AI answers, so your team can tell when one needs questioning.