AI Has Made it Much Easier to be Wrong at Scale


I used to have to work pretty hard to make a bad idea look convincing.

There were spreadsheets to manipulate, diagrams to overcomplicate, and enough meetings to slowly exhaust everyone’s desire to keep questioning the assumptions. I had to choose a PowerPoint template, align all the boxes, find suspiciously specific statistics, and invent a name that could be shortened into an acronym. If the acronym spelled an actual word, the idea was basically funded. That process required commitment. Bad ideas had to be earned.

Now there’s a prompt for that. Give AI ten minutes and your questionable assumption has a business case, an executive presentation, an implementation roadmap, and five years of projected financial returns. Give it another ten minutes and it will write the announcement congratulating everyone on the success of the initiative that has yet to begin. 😂

Honestly, I’m a little jealous of its confidence. I have spent weeks researching an idea and still walked into the meeting wondering if I missed something painfully obvious. AI can receive three paragraphs of context, pause for approximately half a second, and recommend a global rollout beginning Monday. It will include governance, stakeholder engagement, training requirements, and a tasteful three-stage maturity model. Phase One is always “Foundation,” by the way. Apparently every corporate transformation starts by pouring concrete.

When thinking suddenly gets production value

Generative AI does something strange to ideas: it makes them look finished. An incomplete thought can suddenly appear fully formed, a guess can arrive with supporting evidence, and a weak argument can be delivered using the exact vocabulary preferred by the executive team. Strategic. Scalable. Transformational. Possibly even “future-ready,” if things get out of hand.

Presentation quality influences how we interpret the thinking behind it. A huge slide deck suggests someone did a tremendous amount of research. We see structure and assume rigor. We see detail and assume depth. And the productivity gains are very real. The study Generative AI at Work, involving 5,179 customer-support agents, found that access to an AI assistant increased productivity by an average of 14%. The least experienced and lowest-skilled employees improved by 34%, suggesting that AI can help people apply patterns and practices that previously required years to learn.

A separate study, Navigating the Jagged Technological Frontier, followed 758 Boston Consulting Group consultants using GPT-4. On tasks within the model’s capabilities, the consultants completed 12.2% more work, finished 25.1% faster, and produced outputs rated more than 40% higher in quality. Those improvements represent a substantial expansion of what one knowledgeable person can accomplish. Someone can explore more possibilities, test more variations, analyze more information, and complete work that previously required an entire team.

Then the study gets interesting when you look deeper. When participants used AI for a business problem outside the model’s reliable capabilities, they were 23% less likely to reach the correct answer. The same technology that dramatically improved performance on one set of tasks reduced it on another. More output. More confidence. Less friction. Suddenly the organization has an enormous pile of beautifully formatted nonsense arriving ahead of schedule.

AI usually pursues the task it receives. When the question is poorly framed, the objective is wrong, the data is incomplete, or the original assumption is flawed, the system can still produce an extraordinary answer. Every section may connect logically to the previous one. Every recommendation may appear practical. The entire thing can travel coherently, professionally, and confidently in the wrong direction. This is what makes the obsession with doing MORE with AI worth examining. More reports, analysis, code, content, presentations, and decisions can generate real value. They can also overwhelm an organization’s ability to determine which work matters, which claims are reliable, and which ideas should have been politely escorted from the building several prompts ago.

Historically, weak ideas encountered natural resistance. Someone had to research them, explain them, defend them, convince other people, secure resources, and eventually translate them into something operational. That process was slow and frequently annoying. It also created multiple opportunities for another person to stare at the proposal for a while and ask, “Wait…why are we doing this?”

AI removes a great deal of that resistance. Some useful skepticism was apparently hiding inside it.

Judgment is becoming the job

As producing polished work becomes dramatically easier, evaluating that work becomes more valuable. Expertise increasingly includes knowing what to question, where context is missing, and when the impressive answer is very impressively wrong. The strongest AI users seem to do five things especially well:

  • They frame the problem carefully. AI can pursue a poorly chosen objective with breathtaking efficiency. Experienced people understand the larger system and recognize when a locally sensible recommendation could create a much larger problem elsewhere.

  • They inspect the empty spaces. Missing stakeholders, operational constraints, contradictory evidence, and inconvenient dependencies rarely announce themselves. Domain expertise helps someone notice what never appeared in the answer.

  • They separate polish from proof. Attractive charts and confident language influence how credible an idea feels. Validation still requires data, customers, operating experience, and occasionally a grumpy expert who has seen this movie before.

  • They understand where supervision matters. AI creates the most value when someone can evaluate what it produces. Risk rises quickly when nobody involved has enough knowledge to recognize a plausible mistake.

  • They retain accountability for the decision. AI can recommend, simulate, summarize, compare, and challenge. A human still owns what happens when the presentation closes and the idea collides with reality.

I suspect many organizations will initially measure AI by counting activity because activity is wonderfully easy to count. Prompts submitted. Hours saved. Content created. Agents deployed. One department will inevitably announce that it generated 400% more reports, which raises the reasonable question of whether anyone was suffering from a shortage of reports. The larger opportunity is improving the speed and quality of decisions. That requires technical capability, domain expertise, critical thinking, and clarity about where human judgment adds value. It also requires enough psychological safety for someone to interrupt all that AI-powered momentum and question the original premise without being labeled “resistant to innovation.”

Imagine giving your worst idea unlimited energy, perfect grammar, and access to your entire organization. Six minutes later, it returns with an ROI calculation, a steering committee structure, and a cheerful note saying it has already drafted the communication plan. At least the bad idea used to have the decency to look like one.

Anyway…how’s your AI rollout going? 😇


References:

  • Brynjolfsson, E., Li, D., & Raymond, L. R. (2023). Generative AI at work (NBER Working Paper No. 31161). National Bureau of Economic Research. https://www.nber.org/papers/w31161

    Dell’Acqua, F., McFowland, E., III, Mollick, E. R., Lifshitz-Assaf, H., Kellogg, K. C., Rajendran, S., Krayer, L., Candelon, F., & Lakhani, K. R. (2023). Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge worker productivity and quality (Harvard Business School Working Paper No. 24-013). Harvard Business School. https://www.hbs.edu/faculty/Pages/item.aspx?num=64700

Previous
Previous

What Are We Carrying Into Alignment?

Next
Next

The Law of Conserved Indecision