Few-Shot Prompting

Few-shot prompting means including a handful of input-output examples in the prompt so the model learns the task's pattern from demonstrations. The ability to learn "in context" was one of the headline discoveries of large language models (the GPT-3 paper's title was "Language Models are Few-Shot Learners"), and it remains one of the most reliable ways to control format, tone, classification boundaries, and edge-case behavior.

Modern instruction-tuned models often perform well zero-shot, from instructions alone, so few-shot examples are now used more surgically: to pin down output structure, demonstrate judgment on ambiguous cases, and match a house style that's hard to describe in words.

TL;DR

Quick Example

Few-shot classification of support tickets with edge cases:

Dynamic example selection by similarity (Python sketch):

Core Concepts

Zero-Shot, One-Shot, Few-Shot

With long-context models, many-shot prompting (hundreds of examples) can approach fine-tuning quality for some classification and extraction tasks, and prompt caching keeps it affordable.

What Makes Good Examples

Formatting Examples

Use clear delimiters (<example>, <input>, <output> tags, or consistent headings) so the model distinguishes examples from instructions and from the real input. State that examples illustrate the pattern and aren't exhaustive, otherwise the model may overfit to their content, length, or phrasing.

For chat APIs, examples can also be provided as prior user/assistant message pairs. That's effective for conversational style, but it's less clear when mixed with real conversation history.

Dynamic Few-Shot Selection

Rather than a fixed set, keep an example bank (curated input-output pairs, including corrected production mistakes) and retrieve the most relevant ones for each input using embeddings and similarity search, optionally diversified by label. It's RAG applied to demonstrations, and it scales to large, varied task spaces.

Biases and Failure Modes

Few-Shot vs Fine-Tuning

See fine-tuning.

Best Practices

Start Zero-Shot, Add Examples for Specific Failures

Write clear instructions first, evaluate, and add examples targeting the cases the model gets wrong. Examples chosen to fix observed errors are the most effective.

Grow an Example Bank From Production

Save misclassified or poorly handled inputs with corrected outputs. They become high-value examples for dynamic selection, and for evaluation sets.

Evaluate Example Sets

Measure accuracy with and without examples, with different example sets, and with shuffled orderings. Keep examples that help and remove ones that don't. See LLM evaluation.

Cache Static Examples

Fixed example sets belong in the cached prefix of the prompt (see system prompts), and dynamic examples after it.

Common Mistakes

Examples That Contradict Instructions

If instructions say "return only the category" but examples include explanations, models follow the examples. Keep them perfectly aligned.

Too-Similar Examples

Four examples of short, polite billing tickets teach the model little about long, angry, multi-issue tickets. Vary the examples.

Leaking Test Data Into Examples

Using evaluation cases as few-shot examples inflates measured accuracy. Keep evaluation sets separate from example banks.

FAQ

How many examples should I include?

Typically 3–5 for format and style, and more (10–50, or even hundreds with long context) for nuanced classification or extraction. Returns diminish, and each example costs tokens, so measure accuracy as you add them, and stop when gains flatten.

Does few-shot prompting still matter with modern models?

Yes, though more selectively. Strong instruction-following models often handle tasks zero-shot, but examples remain the most reliable way to lock in exact formats, match a specific style, and handle ambiguous edge cases consistently.

Should examples go in the system prompt or user message?

Either works. Static examples that apply to every request often live in the system prompt (and benefit from caching). Dynamically selected examples are usually inserted in the user turn near the input. Clear tags matter more than placement.

What is dynamic few-shot prompting?

Selecting examples per request from a larger bank based on similarity to the current input, usually via embeddings and a vector search. It provides the most relevant demonstrations for each input without bloating every prompt with all examples.

Related Topics

References