AI engineering

Chain-of-thought prompting, and when it backfires

Asking for step-by-step reasoning only helps above a model-size threshold. Below it, results often get worse. Why that matters more on cheap models now.

Listen to the audio explanation

8 min

Its own script, written for listening.

Most people think that if you want an AI to think harder, you just need to give it better instructions. But there's a hard physical limit to that logic.

Editor's note

Why this matters now

Chain-of-thought prompting, asking a model to work through the steps before it answers, is so ordinary now that reasoning models do it without being asked. The 2022 paper by Wei and colleagues that named it carries a caveat that rarely gets repeated.

The technique only works past a certain model size. Below about 10 billion parameters, step-by-step prompting often made answers worse than asking directly. That is what the paper means by an emergent ability: it is absent in small models, and past a threshold of scale it appears.

The source

What it says

Distilled from the original. The notes above and below are the editor's own.

The emergence of reasoning through intermediate steps

Large language models (LLMs) have long struggled with tasks that require more than just pattern matching, such as math, logic, and symbolic manipulation. This paper introduces Chain-of-Thought (CoT) prompting, a simple but powerful method to unlock these abilities. Instead of asking a model to go directly from a question to an answer, CoT prompting provides a few examples (exemplars) that show the model how to generate a series of intermediate reasoning steps.

The core thesis is that CoT is an emergent ability. This means the benefit isn't a universal fix that works on every model; rather, the ability to successfully use CoT only appears once a model reaches a specific threshold of scale. For builders, the takeaway is clear: attempting to use CoT on small models is often a waste of resources, as they may actually perform worse than standard prompting.

When used with sufficiently large models, CoT delivers notable improvements. For instance, the researchers found that prompting the PaLM 540B model with just eight CoT exemplars allowed it to achieve state-of-the-art accuracy on complex math benchmarks like GSM8K. This method moves beyond simple question-answering, allowing models to decompose multi-step problems into manageable pieces, effectively allocating more "thought" (in the form of intermediate tokens) to the harder parts of a problem.

The researchers demonstrate that CoT is not just for math; it also improves performance in commonsense reasoning and symbolic tasks. Crucially, this is an "off-the-shelf" method. It requires no expensive fine-tuning of the model; it simply works by changing how you present the task through prompting.

The scale threshold: why size matters for CoT

One of the most critical findings in this research is that CoT prompting is an emergent property of model scale. It is not a linear improvement that gets slightly better as models get bigger. Instead, there is a significant leap in effectiveness once a model hits a certain size—typically around the 100B parameter mark.

For smaller models, CoT can actually be a hindrance. The researchers observed that models with fewer than 10B parameters often performed worse with CoT than they did with standard prompting. These smaller models tend to produce reasoning chains that look correct on the surface—they are fluent and follow the right "style"—but are logically hollow or outright incorrect. They mimic the sound of reasoning without actually performing the underlying logic.

Want the technical picture?

The researchers analyzed error patterns in the 62B parameter PaLM model and compared them to the 540B version. Scaling the model significantly fixed two major error types: "semantic understanding" (the model failing to grasp the meaning of the question) and "one-step missing" (the model skipping a necessary logical leap). This suggests that the "reasoning" capacity is tied to the model's ability to build a robust internal representation of the world and the rules governing it, which only stabilizes at high scales.

The difference in performance is stark. While a 137B LaMDA model shows significant gains from CoT, the jump in capability becomes most profound in the massive PaLM 540B. This scale threshold creates a clear divide:

Model ScaleImpact of CoT Prompting
Small (<10B)Often negative; produces fluent but illogical chains.
Medium (~60B-100B)Shows signs of emergence; starts to fix semantic errors.
Large (~100B - 540B+)Dramatic performance gains; achieves state-of-the-art on complex tasks.

For a developer building agentic workflows or complex reasoning pipelines, this means your choice of model is the primary driver of whether CoT will work. If you are constrained to smaller, more efficient models, CoT might actually introduce more noise and errors than it solves.

How CoT unlocks complex reasoning

The mechanism of CoT is deceptively simple: you provide the model with a few-shot prompt containing triples of <input, chain of thought, output>. By seeing these intermediate steps in the prompt, the model learns to follow a similar pattern of decomposition for new, unseen problems.

This decomposition is what allows the model to tackle "multi-hop" problems. In a single-step prompt, the model must jump from the question to the final answer in one go, which is a massive cognitive leap. With CoT, the model can allocate more computation—expressed as a sequence of intermediate tokens—to each part of the problem.

One of the most impressive side effects of this method is length generalization. In symbolic reasoning tasks (like concatenating the last letters of names), the researchers tested whether models could handle inputs longer than those seen in the examples. For example, if the model only saw examples with two words, could it handle names with four words?

The results showed that large models using CoT achieved "upward scaling curves" in these out-of-domain (OOD) settings. While performance is naturally lower on tasks the model hasn't seen before, CoT allows the model to apply its learned logic to longer, more complex sequences that would otherwise break a standard prompt.

The effectiveness of this method was validated across several high-stakes benchmarks:

  • GSM8K: A benchmark for grade-school math word problems.
  • SVAMP: A dataset featuring math problems with varying structures.
  • MAWPS: A repository of diverse math word problems.

In these tests, the PaLM 540B model didn't just improve; it reached new state-of-the-art levels. This was achieved using only eight exemplars, proving that you don't need a massive dataset of reasoning steps to trigger this ability—you just need the right model scale and a handful of high-quality examples.

When to use (and when to skip) CoT

While CoT is a powerful tool, it is not a "silver bullet." The researchers identified a clear "sweet spot" for when this technique provides value and a "dead zone" where it is unnecessary or even detrimental.

The Sweet Spot: Multi-step, Complicated Problems CoT thrives on complexity. The performance gains are most pronounced when a task requires several logical leaps or semantic interpretations. For example, on the GSM8K benchmark—which has a low baseline performance because the problems are hard—the use of CoT more than doubled the solve rate for the largest GPT and PaLM models. If the problem is hard and the "scaling curve" (the model's baseline performance) is relatively flat, CoT is your best lever.

The Dead Zone: Simple, Single-Step Tasks For tasks that are relatively easy or only require a single operation, CoT provides little to no benefit. The researchers tested a subset of the MAWPS benchmark called "SingleOp," which only requires a single reasoning step. For these easy tasks, performance improvements were either negative or very small. If the model can already solve a problem with 90% accuracy using standard prompting, adding CoT is essentially just adding extra tokens and latency for no real gain.

Natural Language vs. Equation-Only Prompting The researchers also explored whether a model could perform reasoning by simply outputting a mathematical equation instead of a full natural language explanation. They found that "equation-only" prompting works well for simple one- or two-step problems where the math is easy to derive.

However, for semantically challenging problems like GSM8K, equation-only prompting failed. The model struggled to translate the nuances of a word problem directly into a math formula. In these cases, the natural language "chain" is essential because it allows the model to reason through the meaning of the words before attempting the math.

Task TypeRecommended StrategyWhy?
Simple / Single-stepStandard PromptingCoT adds unnecessary latency and complexity.
Complex / Multi-stepChain-of-ThoughtAllows decomposition of the problem into parts.
Purely MathematicalEquation-only (may work)Efficient if the semantics are trivial.
Semantic / Wordy MathChain-of-ThoughtNatural language is needed to bridge meaning to math.

The 'Interpretability Illusion' and its risks

One of the most attractive features of CoT is that it provides an interpretable window into the model's "thought" process. By reading the generated chain, a human can see how the model arrived at an answer, which makes it much easier to debug where a reasoning path went wrong.

However, the researchers issue a stern warning: CoT is not a guarantee of correctness.

The existence of a reasoning chain does not mean the reasoning is sound. There are two primary ways this can fail:

  1. Correct by Chance: A model can follow a completely illogical or incorrect reasoning path but somehow land on the correct final answer. This is particularly common in multiple-choice or binary classification tasks (like "yes/no" questions), where the model has a statistical chance of guessing correctly even if its "logic" is broken.
  2. Incorrect Logic, Correct Answer: Even in free-response math, the model might perform a correct step but fail to connect it to the next, or arrive at a correct answer through a series of "almost right" steps.

The researchers noted that while most correct answers in their GSM8K analysis were backed by correct logic, a small percentage were "correct by chance."

Beyond the risk of error, there is a deeper, more fundamental uncertainty: Does the model actually "reason"?

The paper notes that while CoT emulates the thought process of a human, it does not definitively prove that the underlying neural network is performing true reasoning. It is currently an open question whether the model is genuinely executing a logical process or simply predicting the next most likely "reasoning-sounding" token based on linguistic patterns seen during training.

For builders, this means you should treat the CoT output as a diagnostic tool, not a proof of truth. Use the chain to understand how the model is failing, but never assume that a well-written explanation means the answer is right.

Building with emergent reasoning

For developers and product managers building with LLMs, this research provides a tactical roadmap for when and how to deploy reasoning capabilities.

1. Prioritize Model Scale for Agentic Workflows If you are building an autonomous agent or a complex reasoning engine, do not attempt to "prompt your way out" of the limitations of a small model. If your task requires multi-step logic, you must use high-parameter models (typically 100B+). Using CoT on a 7B or 8B model may result in a "fluent hallucination"—an agent that sounds incredibly confident and logical while being fundamentally broken.

2. Augment CoT with External Tools The research shows that even with CoT, models can make "calculator errors" (simple arithmetic mistakes). One of the most effective ways to boost performance in math-heavy tasks is to augment the CoT process with an external calculator. Instead of relying on the model to do 45 * 12 in its "head," your system can be designed to recognize when the model has generated an equation and then pass that equation to a reliable Python eval() function or a calculator tool. This combines the semantic reasoning of the LLM with the precision of symbolic computation.

3. Leverage CoT as a Low-Cost Upgrade One of the biggest advantages of CoT is that it is an "off-the-shelf" method. You do not need to collect a dataset of reasoning steps to fine-tune a model, which is often prohibitively expensive and difficult. You can unlock higher performance in your existing applications simply by updating your prompt templates to include a few high-quality, hand-written CoT exemplars.

Summary of Tactical Moves:

  • If the task is simple: Use standard prompting to save latency and cost.
  • If the task is complex: Use CoT, but ensure you are using a high-parameter model.
  • If the task is math-heavy: Use CoT + an external tool (like a Python interpreter).
  • If you need to debug: Use the CoT output to identify if the model is failing due to semantic misunderstanding (meaning) or computational error (math).

Further Reading

  • Cobbe et al. (2021): Training verifiers to solve math word problems.
  • Brown et al. (2020): Language models are few-shot learners.

Editor's note

What to do with this

If you are running a prompt with step-by-step instructions on a small or cheap model, test it against the same prompt with those instructions stripped out. That is a 10-minute experiment, and the result can surprise you.

The broader habit is worth more than the technique. Something that works on a frontier model is a technique plus a model size. When you move work down to something cheaper to save money, the prompt does not automatically come with it.

The original

  • AI engineering

    The GPT-3 paper: where prompting came from

    Language Models are Few-Shot Learners is why you put examples in your prompts. What in-context learning claimed, and the caveats it admitted.

    9 min read7 min listen

  • AI engineering

    The RAG paper: what retrieval actually buys you

    Human judges called the retrieval-backed model more factual 42.7% of the time against 7.1% for the plain one. The original is more careful than its descendants.

    7 min read7 min listen

All posts