Chain-of-Thought Prompting

Language models generate one token at a time, and each token is computed from what came before. When a model jumps straight to an answer for a multi-step problem (a math word problem, a logic puzzle, a policy decision with several conditions), it has no room to work things out. Chain-of-thought (CoT) prompting asks the model to reason step by step before answering, and the intermediate reasoning measurably improves accuracy on arithmetic, logic, planning, and complex analysis.

The technique evolved from a prompt trick ("Let's think step by step") into a core model capability. Reasoning models and extended thinking modes are trained to produce long internal reasoning before answering, and they can be controlled with thinking budgets or effort settings. Knowing when reasoning helps, and when it just adds latency and cost, is part of modern prompt engineering.

TL;DR

Quick Example

Prompted chain-of-thought with a separated answer:

Using extended thinking via the API instead of prompt-level CoT:

Core Concepts

Why Reasoning Helps

Each generated token gives the model more computation to spend on the problem. Writing out intermediate results (sub-totals, which rule applies, what's known versus unknown) lets later tokens condition on them, much like a person using scratch paper. Gains are largest for multi-step math, logical deduction, applying several rules, planning, code debugging, and analysis requiring synthesis. They're smallest for simple factual recall or classification.

Prompting Techniques

Separating Reasoning From Output

Applications usually need only the final answer. Put reasoning and answer in distinct tags, parse the answer tag, and either discard the reasoning or keep it for logging and debugging. With structured outputs, you can include a reasoning field before the answer fields, so the model reasons first. Field order matters, because generation proceeds sequentially.

Reasoning Models and Extended Thinking

Reasoning models (and extended thinking modes in general-purpose models) are trained with reinforcement learning to produce long internal chains of thought, exploring, checking, and backtracking before answering. Practical differences:

See reasoning models.

When Not to Use CoT

Evaluate: compare accuracy, latency, and cost with and without reasoning on your task. See LLM evaluation.

Faithfulness Caveat

Written reasoning isn't guaranteed to reflect how the model actually reached its answer: models can produce plausible-looking rationales that don't match their underlying computation. Treat CoT as a performance aid and a debugging signal, not as a verified explanation. Verify important conclusions independently.

Best Practices

Guide What to Think About

Tell the model which aspects matter (the rules to check, the constraints to verify, the edge cases to consider). Guided reasoning is more reliable than a generic "think step by step".

Ask the Model to Check Its Work

Adding "Before answering, verify the result against the requirements" or "check your calculation" catches many errors, especially for arithmetic and code. Tests and executable checks are even better.

Keep Final Answers Machine-Parseable

Put final answers in tags or structured fields, and validate them. Never parse answers out of free-form reasoning text.

Budget Reasoning by Difficulty

Route easy requests to fast, direct answers, and hard ones to reasoning models or larger thinking budgets. Classification-based routing keeps costs in check. See prompt chaining.

Common Mistakes

Answer Before Reasoning

Showing Raw Reasoning to End Users

Internal reasoning can be long, tentative, or include considerations not meant for users. Display polished answers, and keep reasoning for logs, or offer it as optional detail when appropriate.

Over-Prescribing Steps to Reasoning Models

Forcing a thinking model through a rigid 12-step template can reduce quality compared with giving the goal and key considerations. Start with high-level guidance, and add structure only where evaluation shows it helps.

FAQ

What is chain-of-thought prompting?

A technique that has a language model write out intermediate reasoning steps before its final answer. The extra reasoning tokens let the model break down multi-step problems, which improves accuracy on math, logic, planning, and complex analysis.

Do I still need chain-of-thought with reasoning models?

Not in the same form. Reasoning models think internally before responding, so prompt-level "think step by step" instructions are mostly unnecessary. Guide them with high-level considerations, and control depth with thinking budgets or effort settings.

Does chain-of-thought always improve results?

No. It helps most on problems requiring multiple steps. For simple tasks it adds latency and cost with little benefit, and occasionally it can lead models to overthink. Measure on your task.

What is self-consistency?

Generating several independent reasoning paths (with sampling) for the same question, then selecting the most common final answer. It improves accuracy on problems with a single correct answer, at the cost of multiple generations.

Related Topics

References