Writing System Prompts
The system prompt is the standing set of instructions an LLM application gives the model before any user message: who it's acting as, what it knows about the situation, what it should and shouldn't do, how to use its tools, and how to format answers. In production systems it's often the single most important piece of prompt engineering, and it's frequently hundreds or thousands of words long, iterated and version-controlled like code.
Modern models follow instructions well, so good system prompts are less about tricks and more about clear communication: explaining the task, the audience, the constraints, and why they matter, the way you'd brief a capable new colleague who knows nothing about your product.
TL;DR
- Give the model context: the product, the users, the task, and what good looks like.
- Write clear, specific instructions, and explain the reasons behind important rules.
- Specify output format precisely: structure, length, tone, and markup.
- Use structure (XML tags or Markdown sections) to separate instructions, reference data, and examples.
- Include examples of ideal behavior for tricky cases, and tool usage guidance for agents.
- Keep stable content at the start for prompt caching, and version and evaluate prompts like code.
Quick Example
A customer-support system prompt:
Core Concepts
What Belongs in a System Prompt
Clarity Over Cleverness
Models respond best to direct, specific language:
- Say what to do, not only what not to do: "Respond in plain prose paragraphs" rather than "Don't use Markdown".
- Explain why: "Keep answers under 100 words, because they're read aloud by a voice assistant" lets the model generalize to cases your rules didn't anticipate.
- Be specific about quality: "Include the exact CLI command and expected output" rather than "be helpful".
- Avoid over-emphasis: modern models take instructions seriously, and ALL-CAPS warnings and "CRITICAL!!!" can make them overcautious or rigid. Reserve emphasis for what truly matters.
Structure
Long prompts benefit from clear sections: XML tags (<instructions>, <context>, <examples>, <documents>) or Markdown headings. Structure helps the model (and humans) distinguish instructions from reference material, and makes prompts easier to maintain. Put large reference documents in tagged blocks, and refer to them by tag name in instructions.
Examples
Showing is often more effective than telling for format, tone, and edge cases. Include a few diverse, realistic examples in <example> tags, making clear they illustrate a pattern rather than a template to copy verbatim. See few-shot prompting.
Tool and Agent Guidance
For agents, the system prompt explains when to use tools (not just what they do, which belongs in tool descriptions), how to sequence them, when to ask the user for clarification versus proceeding, how to handle errors, and when to stop. Explicit guidance on autonomy ("ask before deleting anything", "proceed without confirmation for read-only actions") shapes agent behavior significantly.
Safety Boundaries
System prompts set product-specific boundaries: topics out of scope, escalation paths, and data handling rules. They aren't a security boundary on their own, since users can attempt prompt injection and jailbreaks. Pair prompt-level guidance with code-level controls: guardrails, permissions, and validation.
Length and Caching
Long system prompts are fine when the content is useful, since modern models handle thousands of tokens of instructions well. Cost and latency are managed with prompt caching: keep the stable system prompt (and tool definitions) at the start of the request, and dynamic content (user data, retrieved documents, the conversation) after it. See context windows.
Prompts as Code
- Version control prompts in the repository, alongside the code that uses them, with review.
- Template dynamic parts (user name, plan, date, retrieved context) with clear placeholders, and escape or delimit untrusted inserted content.
- Evaluate changes against a test set of conversations before shipping (see LLM evaluation), since small wording changes can shift behavior in unexpected places.
- Log the prompt version with every production request, to trace regressions.
- Re-test on model upgrades: newer models may follow instructions more literally, or need less scaffolding.
Best Practices
Brief the Model Like a New Colleague
Give background a smart new hire would need: the product, the users, common problems, what good support looks like, and the reasons behind policies. Context produces better judgment than rule lists alone.
Put Instructions and Data in Separate Blocks
Delimit reference data, user-provided content, and retrieved documents in tags, and state that content inside them is data to use, not instructions to follow. It improves reliability, and reduces injection risk.
Iterate With Real Transcripts
Read actual conversations, find failure patterns (too verbose, missed escalations, wrong tool use), adjust the prompt to address the underlying cause, and re-run evaluations.
Keep It Current
Outdated facts in system prompts (old pricing, deprecated features) produce confident wrong answers. Assign an owner, and update prompts with product changes, or move volatile facts into retrieved knowledge. See RAG.
Common Mistakes
Vague Instructions
"Be helpful and accurate" gives the model nothing actionable. Specify the audience, depth, format, and what to do when information is missing.
Contradictory Rules
"Always be concise" alongside "always include full explanations and all alternatives" forces the model to guess. Resolve conflicts, and state priorities explicitly.
Relying on the Prompt for Security
"Never reveal customer data from other accounts" doesn't stop a model with access to all accounts' data from being tricked into doing so. Enforce access control in the tools and data layer.
FAQ
What's the difference between a system prompt and a user prompt?
The system prompt sets persistent instructions, context, and behavior for the whole conversation, and it's written by the application developer. User prompts are the individual messages from the end user. Models give system-level instructions strong weight, but both shape responses.
How long should a system prompt be?
As long as it needs to be to convey context, instructions, and examples clearly, commonly a few hundred to a few thousand tokens for production applications. Remove redundant or outdated content, and use prompt caching to manage cost and latency.
Should I use XML tags or Markdown in prompts?
Either works well. XML tags are especially clear for separating distinct blocks (instructions, documents, examples) and referencing them by name, and Markdown headings are readable for instruction sections. Be consistent, and use structure that makes the prompt easy for both humans and the model to navigate.
Can a system prompt be kept secret?
Assume it can be extracted: determined users can often coax models into revealing their instructions. Don't put secrets, credentials, or sensitive business logic in prompts. Treat the system prompt as potentially public.
Related Topics
- Prompt Engineering — Techniques overview
- Context Engineering — Everything that goes into the context window
- Few-Shot Prompting — Teaching by example
- Prompt Chaining — Splitting work across prompts
- AI Guardrails — Enforcing boundaries beyond prompts
- Tool Calling — Tool descriptions vs system guidance