Prompt Optimization & Iteration
Most prompts start as a reasonable first draft and then get "tweaked" based on a few examples someone tried by hand. That feels productive, but often isn't: a change that fixes one case silently breaks three others, and nobody notices until users do. Prompt optimization replaces anecdotal tweaking with an engineering loop: define success criteria, build an evaluation set, analyze errors, change one thing at a time, and measure.
The loop applies whether you edit prompts by hand, use a model to suggest improvements, or run automated optimizers like DSPy. It's also how you migrate safely to new models, which can respond differently to the same prompt.
TL;DR
- Define success criteria first: accuracy, format compliance, tone, latency, and cost targets.
- Build an evaluation set from real inputs, covering common cases, edge cases, and past failures.
- Do error analysis: read failures, categorize them, and fix the biggest category first.
- Change one variable at a time, re-run evals, and keep changes that help overall.
- Use LLM-as-judge for subjective criteria (calibrated against humans) and code checks for objective ones.
- Consider automated optimization (DSPy, model-generated prompt revisions) and A/B tests in production. Re-evaluate on model upgrades.
Quick Example
A minimal prompt evaluation loop:
Core Concepts
Success Criteria
Before optimizing, define what "better" means, specifically and measurably:
- Task quality: accuracy, F1, faithfulness, coverage of key points, correctness against references.
- Format: schema validity, length limits, required sections.
- Behavior: tone, refusal appropriateness, escalation correctness, tool usage.
- Operational: latency, token cost, and failure and retry rates.
Vague goals ("make it better") lead to endless, unmeasurable tweaking.
Evaluation Sets
A good eval set:
- Uses real inputs (sampled production data, anonymized as needed), not only synthetic ones.
- Covers common cases, edge cases, adversarial inputs, and known past failures.
- Includes references or rubrics: expected outputs, key facts, or grading criteria.
- Is versioned and grows over time, as production failures are added.
50–200 well-chosen cases catch most regressions. Larger sets improve statistical confidence. See LLM evaluation.
Grading Methods
Error Analysis
The highest-leverage step: read the failures. Categorize them (missed instruction, wrong format, hallucinated detail, overly long, wrong tool, misunderstood domain term), count them, and target the largest category. Fixes then address root causes: missing context, ambiguous instructions, lack of examples, or a task that should be split (see prompt chaining).
Iteration Discipline
- Change one thing per experiment: an instruction, an example, the output format, or the model.
- Compare against the baseline on the full eval set, and inspect per-case regressions, not just averages.
- Keep a changelog of prompt versions with eval results.
- Watch for overfitting to the eval set: hold out a test split, and refresh cases periodically.
Levers to Try
- Clarify the task and context: who the audience is, and what good looks like. See system prompts.
- Add examples for failing patterns. See few-shot prompting.
- Structure input and output: tags, schemas, structured outputs.
- Allow reasoning for multi-step tasks. See chain-of-thought.
- Improve context: better retrieval, fewer irrelevant documents. See context engineering.
- Split the task into a chain.
- Change the model or thinking budget, weighing cost and latency.
Automated Prompt Optimization
- Model-assisted revision: ask a strong model to critique a prompt against failing cases and propose improvements, then evaluate the proposals like any change. Prompt generator and improver tools in provider consoles do this.
- DSPy: declare the task as modules with signatures and a metric, and optimizers (such as MIPROv2 and bootstrapped few-shot) search over instructions and example selections to maximize the metric on your data.
- Search-based methods (OPRO-style, genetic variants) iterate prompts using eval scores as feedback.
Automated methods need good metrics and eval data, and they can overfit. Keep a held-out test set, and read the resulting prompts.
A/B Testing in Production
Offline evals don't capture everything. For user-facing changes, run A/B tests: route a percentage of traffic to the new prompt version, and compare user outcomes (resolution rate, thumbs-up rate, escalations, conversions), plus cost and latency. Log prompt versions with every request. See A/B testing.
Model Migrations
New model versions follow instructions differently: often more literally, sometimes with different default verbosity or tool behavior. Treat a model change like a prompt change. Run the eval suite, compare, adjust prompts (often removing old workarounds), and roll out gradually.
Best Practices
Invest in Evals Before Tweaking
A day spent building an eval set saves weeks of guesswork, and makes every future change, including model upgrades, safer and faster.
Look at Outputs, Not Only Scores
Aggregate metrics hide important shifts. Always read a sample of outputs, especially regressions and borderline cases.
Manage Prompts Like Code
Version prompts in source control or a prompt registry, review changes, tie deployments to eval results, and record the prompt version in logs and traces. See LLMOps.
Optimize for Cost and Latency Too
After quality meets targets, try smaller models, shorter prompts, prompt caching, and fewer chain steps, verifying quality holds. The best prompt is often the cheapest one that meets the bar.
Common Mistakes
Tweaking Based on One Example
Fixing the case in front of you while unknowingly breaking others is the most common prompt-engineering failure. Always re-run the full eval set.
Adding Rules for Every Failure
Piling on special-case instructions creates long, contradictory prompts. Look for the underlying cause (missing context, unclear goal) and fix that.
Trusting Uncalibrated LLM Judges
A judge prompt that disagrees with human raters often produces misleading "improvements". Validate judges on a labeled sample, and monitor their agreement.
FAQ
How do I know if a prompt change is an improvement?
Run the old and new prompts on the same evaluation set and compare metrics tied to your success criteria. Inspect per-case regressions, then confirm with production A/B tests for user-facing impact. Without an eval set, you're guessing.
What is DSPy?
A framework for programming LLM pipelines declaratively: you define modules with input and output signatures, and a metric, and DSPy's optimizers automatically tune prompts and few-shot examples to maximize that metric on your data. It turns prompt engineering into an optimization problem.
How big should my evaluation set be?
Start with 50–100 diverse, realistic cases, including known hard ones, which is enough to catch major regressions. Grow it as you find new failure modes, and to a few hundred or more when you need statistical confidence for small differences.
Should I re-tune prompts when upgrading models?
Yes. Run your evaluation suite against the new model with your current prompts, look at differences, and adjust. Newer models often need less scaffolding, and may interpret instructions more literally, so some old workarounds become unnecessary or counterproductive.
Related Topics
- Prompt Engineering — Techniques overview
- LLM Evaluation — Building and running evals
- LLMOps — Prompt management and monitoring in production
- System Prompts — The prompt you're usually optimizing
- A/B Testing — Validating changes with real users
- RAG Evaluation — Evaluating retrieval-heavy prompts