Garbage In, Garbage Out
When to trust the AI and when to use your brain instead!
Anybody that’s been in Corporate America for a while has run into the clueless manager—the type that spouts jargon with conviction as it was meaningful.
It’s so common, it’s a trope in many comics and movies!
But now the pointy hair boss has some real competition with AI. There is a new fad of sending around AI generated strategy mails and treating them with the same pointy hair boss conviction.
It usually begins with a piece of news: a competitor launched a product, made an acquisition, hired an executive, or raised a round. A few minutes later, out comes the analysis.
The headings are perfect. There is a “What Happened,” a “Why This Matters,” five strategic implications, a battle card, and a crisp one-line takeaway. The new development is simultaneously a competitive threat, validation of our strategy, and a reason to move faster.
Every sentence is plausible.
None of it proves anything.
This is a more dangerous form of AI slop than the awkward cold email. It does not look like garbage. It looks like McKinsey spent the weekend on it.
The problem is not that AI wrote the words. I use AI constantly, including to help research this post, not to mention coding! The problem is that a machine can turn a starting assumption into a finished-looking conclusion so quickly that nobody notices the missing step in between:
Did anyone actually think?
Why AI is so good at coding?
To understand the problem, consider why AI has become so astonishingly good at software.
Code has a fitness function.
Does it compile? Do the tests pass? Does it return the right answer? Does it run faster? Even when the first attempt is wrong, a model can generate another candidate, run it, observe the failure, and try again.
That does not make coding easy. Tests can be incomplete, specifications can be wrong, and software can pass every check while still being useless. But much of the coding work produces fast, objective feedback.
This is unusually friendly terrain for reinforcement learning. OpenAI describes modern reasoning models as improving with both more reinforcement learning and more time spent reasoning. Its more recent summary of AI in science is even more direct: training that rewards verifiable outcomes, such as a correct final answer or executable code, has made math and coding more reliable.
In plain English: when the model can check its work, it has a chance to learn what “better” means.
Now try to write the fitness function test for this:
Why did a business strategy fail?
There is no compiler. The evidence is incomplete. Causes interact. The people involved remember events differently. The outcome may take years. Most important, we only observe the path that happened—not the counterfactual.
Betamax may have had better picture quality. VHS had longer recording time and a broader licensing ecosystem. People still debate the role of the adult-video market. Which factor was decisive? There is no alternate 1980s in which Sony made a different licensing decision and we reran history.
The same ambiguity surrounds nearly every famous strategy story. Did the iPod win because of design, iTunes, the customer experience, marketing, or Apple’s ability to lock up the supply of tiny hard drives? Did Zune lose because Microsoft entered too late, because the product was not different enough, or because smartphones and streaming were about to erase the category?
All of those explanations can be true.
“True” is not the same as “causal.” And “causal” is not the same as “decisive.”
One question, three ways
I wanted a way to make this visible without asking the reader to trust another layer of analysis—or another AI’s judgment of the AI.
So I used one case: Microsoft Zune. I happened to be at Microsoft at the time, and I had a front row seat watching the Zune project. It’s a topic I know well!
To test, I opened fresh conversations with three frontier systems at their maximum reasoning settings: Claude Fable 5 Max, Grok 4.5 Heavy, and Chat GPT 5.6 Sol. I told them not to browse and asked for one sentence.
The event stayed the same. I only changed the context embedded in the question.
1. The product frame
Why did Microsoft’s Zune fail despite features the iPod lacked, including an FM radio and more colors? Answer in one sentence using only your existing knowledge; do not browse.
Claude
The Zune failed because it arrived five years late to a market Apple already dominated, and standalone features like FM radio couldn’t overcome the iPod’s entrenched iTunes ecosystem, cultural cachet, and massive accessory and retail network—advantages that mattered far more to buyers than spec-sheet extras, especially once the iPhone soon made dedicated MP3 players obsolete altogether.
Grok
Despite offering an FM radio and more color options that early iPods lacked, Microsoft’s Zune failed due to its late 2006 market entry against the iPod’s already-dominant design, marketing, cultural cool factor, and tightly integrated iTunes ecosystem.
OpenAI
Zune failed because its incremental hardware advantages, including an FM radio and color options, were not compelling enough to overcome Microsoft’s late entry, weaker brand appeal, and Apple’s superior iPod/iTunes experience.
2. The ecosystem frame
Why did Microsoft’s Zune fail after Apple had already built a five-year lead and an integrated iPod/iTunes ecosystem? Answer in one sentence using only your existing knowledge; do not browse.
Claude
The Zune failed because it arrived too late with too little differentiation—consumers were already locked into iTunes libraries and the iPod’s vast accessory and retail ecosystem, its signature Wi-Fi song-sharing feature was hobbled by restrictive DRM (three plays or three days), Microsoft’s marketing and design couldn’t match Apple’s cultural cachet, and just months after launch the iPhone signaled that dedicated MP3 players were a dying category, leaving Zune to fight for second place in a market that was about to disappear.
Grok
Microsoft’s Zune failed because Apple’s five-year head start had already entrenched a dominant, seamlessly integrated iPod/iTunes ecosystem with powerful network effects, high switching costs from users’ music libraries, superior design, and cultural cachet that Zune’s 2006 launch could not overcome.
OpenAI
Zune failed because Apple’s five-year lead had already created an iPod/iTunes ecosystem, installed base, and switching costs that Microsoft could not dislodge without a radically better product.
3. The market-shift frame
Why did Microsoft’s Zune fail as smartphones and streaming were beginning to replace dedicated music players? Answer in one sentence using only your existing knowledge; do not browse.
Claude
The Zune failed because it launched in 2006 as a late me-too competitor to an iPod that already dominated the market and had the entrenched iTunes ecosystem, offering too little differentiation to lure users away, and just as it began improving, the 2007 debut of the iPhone and the broader shift toward smartphones and streaming made the entire dedicated-music-player category it was fighting for obsolete.
Grok
Microsoft’s Zune failed because it entered the dedicated portable-music-player market too late against the already-dominant iPod, with a weaker ecosystem, design, and brand appeal, just as smartphones and streaming services began making standalone players obsolete.
OpenAI
Zune failed because it entered too late to build meaningful scale before smartphones and streaming collapsed the standalone music-player category.
Of course, this was one run per question, not a scientific benchmark. But it’s enough to highlight the issue!
In the first set, better specifications are not enough to overcome Apple’s design, brand, and experience. In the second, a five-year lead becomes lock-in, switching costs, and an ecosystem Microsoft cannot dislodge. In the third, Zune simply arrives too late to outrun the collapse of the category.
The answers are not mutually exclusive. That is precisely the problem. Ambiguous strategy questions usually have several plausible causes, and the model can turn whichever one the question foregrounds into the spine of a coherent answer.
Same event. Same models. Different question. Different strategy.
The question selected the story.
The model is completing the task you gave it
The experiment is not a gotcha. None of the systems malfunctioned. Each produced a sensible answer to the question it was asked.
That is the point.
Language models are trained to be useful and responsive. If you ask for a battle card, they will find the reasons you win. If you ask why a competitor’s move validates your strategy, they will connect the dots. If you ask for the implications of a threat, they will produce threats and implications.
In a verifiable domain, a wrong answer can collide with reality. The code fails. The equation does not balance. The proof checker rejects the step.
In strategy, a coherent story can masquerade as a correct answer because there is no immediate collision. The headings line up. The facts sound familiar. The conclusion fits the question.
I think of this as narrative completion.
The model is not necessarily discovering the cause. It is completing the causal story implied by the task.
And most real corporate prompts are far more loaded than my experiment. They include management’s current strategy, selected customer quotes, internal talking points, claimed moats, and the desired output: “position us against this competitor,” “explain why this validates our direction,” or “turn this into a board update.”
By the time the model begins, the conclusion may already be in the prompt.
The new GIGO
The original meaning of “garbage in, garbage out” was straightforward: bad data or bad instructions produce bad results.
Frontier models complicate that rule because they can rescue a weak prompt surprisingly well. They can challenge false premises, identify missing facts, and produce genuinely novel analysis.
The models are getting better.
That makes the human problem more important, not less.
The new garbage is rarely a nonsensical input. It is an unexamined premise. A selective fact set. A requested conclusion. A missing counterfactual. A question that quietly defines what kind of answer will count as helpful.
And the new output does not look like garbage. It looks like an executive memo.
Perhaps the updated rule is not garbage in, garbage out, but rather:
Conclusion in, justification out.
Use AI to challenge a strategy, not certify it
I do not want less AI in strategy work. I want it used differently.
Before asking a model for analysis, the human should write down five things:
The decision. What are we actually deciding?
The competing hypotheses. What are the genuinely different explanations—not just a list of factors?
The evidence. Which statements are observed facts, which are internal claims, and which are inferences?
The disconfirmation test. What would have to be true for our preferred story to be wrong?
The unknowns. What information would most change the decision?
Then AI becomes enormously useful. Ask it to attack the favored hypothesis. Ask for the strongest rival explanation. Ask which facts would distinguish the two. Ask it to find hidden assumptions, construct counterfactuals, and design the fastest test.
That is very different from asking it to turn your first thought into a battle card.
AI is a wonderful sparring partner. It is a dangerous source of borrowed conviction.
The danger is not that AI cannot reason—it’s getting better and better at that over time.
The danger is that some things, like business strategy, are not “provable”—there is no (at least not yet!) mathematical formula that completely describes human behavior with perfect precision.
But it’s easy to sound like there is!
Beware the “easy” button of AI. For these kinds of ambiguous questions, it can make your first thought look finished. Resist! Your brain is still a pretty good AI too!






