Few-Shot Prompting
Few-shot prompting means including a handful of input-output examples in the prompt so the model learns the task's pattern from demonstrations. The ability to learn "in context" was one of the headline discoveries of large language models (the GPT-3 paper's title was "Language Models are Few-Shot Learners"), and it remains one of the most reliable ways to control format, tone, classification boundaries, and edge-case behavior.
Modern instruction-tuned models often perform well zero-shot, from instructions alone, so few-shot examples are now used more surgically: to pin down output structure, demonstrate judgment on ambiguous cases, and match a house style that's hard to describe in words.
TL;DR
- Zero-shot: instructions only. Few-shot: instructions plus a few worked examples (typically 2–8).
- Examples are most valuable for format, style, labeling boundaries, and tricky edge cases.
- Make examples diverse and representative, including hard and negative cases, and consistent with your instructions.
- Wrap examples in clear tags, and say they illustrate a pattern, so the model doesn't copy them literally.
- Dynamic few-shot: retrieve the most similar examples per input from an example bank using embeddings.
- Watch for bias from label imbalance, ordering, and overly similar examples. Evaluate with and without examples.
Quick Example
Few-shot classification of support tickets with edge cases:
Dynamic example selection by similarity (Python sketch):
Core Concepts
Zero-Shot, One-Shot, Few-Shot
With long-context models, many-shot prompting (hundreds of examples) can approach fine-tuning quality for some classification and extraction tasks, and prompt caching keeps it affordable.
What Makes Good Examples
- Representative: reflect real inputs (length, messiness, language), not idealized ones.
- Diverse: cover different categories, phrasings, and difficulty, so the model doesn't latch onto superficial features.
- Include edge cases: ambiguous inputs, "none of the above", refusals, and empty results, the cases where instructions alone are unclear.
- Correct and consistent: examples must follow your own instructions exactly. Contradictions between instructions and examples confuse models, and examples usually win.
- Balanced: roughly even label distribution, unless you intentionally reflect real priors.
Formatting Examples
Use clear delimiters (<example>, <input>, <output> tags, or consistent headings) so the model distinguishes examples from instructions and from the real input. State that examples illustrate the pattern and aren't exhaustive, otherwise the model may overfit to their content, length, or phrasing.
For chat APIs, examples can also be provided as prior user/assistant message pairs. That's effective for conversational style, but it's less clear when mixed with real conversation history.
Dynamic Few-Shot Selection
Rather than a fixed set, keep an example bank (curated input-output pairs, including corrected production mistakes) and retrieve the most relevant ones for each input using embeddings and similarity search, optionally diversified by label. It's RAG applied to demonstrations, and it scales to large, varied task spaces.
Biases and Failure Modes
- Majority label bias: the model over-predicts labels that dominate the examples.
- Recency bias: the last examples influence outputs most, so shuffle or balance ordering.
- Surface copying: outputs mimic example length, phrases, or specific values.
- Format over substance: few-shot fixes format quickly, but may not fix reasoning errors. Combine it with chain-of-thought or better instructions.
Few-Shot vs Fine-Tuning
See fine-tuning.
Best Practices
Start Zero-Shot, Add Examples for Specific Failures
Write clear instructions first, evaluate, and add examples targeting the cases the model gets wrong. Examples chosen to fix observed errors are the most effective.
Grow an Example Bank From Production
Save misclassified or poorly handled inputs with corrected outputs. They become high-value examples for dynamic selection, and for evaluation sets.
Evaluate Example Sets
Measure accuracy with and without examples, with different example sets, and with shuffled orderings. Keep examples that help and remove ones that don't. See LLM evaluation.
Cache Static Examples
Fixed example sets belong in the cached prefix of the prompt (see system prompts), and dynamic examples after it.
Common Mistakes
Examples That Contradict Instructions
If instructions say "return only the category" but examples include explanations, models follow the examples. Keep them perfectly aligned.
Too-Similar Examples
Four examples of short, polite billing tickets teach the model little about long, angry, multi-issue tickets. Vary the examples.
Leaking Test Data Into Examples
Using evaluation cases as few-shot examples inflates measured accuracy. Keep evaluation sets separate from example banks.
FAQ
How many examples should I include?
Typically 3–5 for format and style, and more (10–50, or even hundreds with long context) for nuanced classification or extraction. Returns diminish, and each example costs tokens, so measure accuracy as you add them, and stop when gains flatten.
Does few-shot prompting still matter with modern models?
Yes, though more selectively. Strong instruction-following models often handle tasks zero-shot, but examples remain the most reliable way to lock in exact formats, match a specific style, and handle ambiguous edge cases consistently.
Should examples go in the system prompt or user message?
Either works. Static examples that apply to every request often live in the system prompt (and benefit from caching). Dynamically selected examples are usually inserted in the user turn near the input. Clear tags matter more than placement.
What is dynamic few-shot prompting?
Selecting examples per request from a larger bank based on similarity to the current input, usually via embeddings and a vector search. It provides the most relevant demonstrations for each input without bloating every prompt with all examples.
Related Topics
- Prompt Engineering — Techniques overview
- System Prompts — Where static examples often live
- Chain-of-Thought Prompting — Combining examples with reasoning
- Fine-Tuning — Learning from many examples permanently
- Embeddings — Selecting similar examples
- Structured Outputs — Guaranteeing output formats