LLM Tool Calling
Tool calling (also called function calling or tool use) lets a language model do more than generate text: it can request that your application run a function, such as searching a database, calling an API, reading a file, or sending a message, and then use the result to continue. The model doesn't execute anything itself. It emits a structured request (tool name plus JSON arguments), your code runs the tool, and you send the result back. That loop is the foundation of AI agents.
Every major model API supports it, and protocols like the Model Context Protocol standardize how tools are exposed. The quality of an agent depends heavily on tool design: clear names and descriptions, well-typed inputs, concise and informative outputs, and safe handling of anything with side effects.
TL;DR
- Define tools with a name, a description, and an input JSON Schema; the model decides when and how to call them.
- The loop: model returns a tool call → your code executes it → you return a tool result → the model continues, until it gives a final answer.
- Models can make parallel tool calls. Execute independent calls concurrently and return all results.
- Return errors as tool results with actionable messages, so the model can recover.
- Control usage with tool choice (auto, any, a specific tool, or none).
- Treat tool arguments as untrusted input: validate, authorize, and require confirmation for destructive actions.
Quick Example
A tool-use loop with the Anthropic Python SDK:
SDK helpers (tool runners) and agent frameworks implement this loop for you, but it's worth understanding the raw version.
Core Concepts
Tool Definitions
A tool definition is a contract the model reads:
- Name: short, specific, verb-first (
search_tickets,create_refund), consistent across tools. - Description: what it does, when to use it (and when not to), what it returns, and important constraints. This is the single biggest factor in correct tool selection.
- Input schema: JSON Schema with types, enums, formats, required fields, and per-parameter descriptions. Some APIs support strict schema validation, which guarantees the arguments match.
The Tool-Use Loop
- Send the conversation plus the tool definitions.
- The model responds with text and/or one or more tool calls (the stop reason indicates tool use).
- Your application validates and executes each call.
- Append the assistant turn and a message containing tool results (matched by tool call ID).
- Repeat until the model returns a normal answer, or a step limit is reached.
Tools can also be server-side (executed by the provider, such as web search, code execution, or computer use) or exposed via MCP servers that your client connects to.
Parallel Tool Calls
Models can emit several independent calls in one turn ("get weather in Paris and Tokyo"). Execute them concurrently and return all results together, in one user message. That cuts latency substantially in agent workflows.
Tool Choice
Results and Errors
Tool results are ordinary content the model reads, so design them for the model:
- Concise and relevant: return the fields needed, not a 50 KB raw API response. Large results consume the context window and distract the model.
- Informative errors:
"No customer with email x; try searching by name with search_customers"lets the model recover, while a stack trace doesn't. Mark errors as errors (is_error: true) where the API supports it. - Stable identifiers: include IDs the model can pass to follow-up tools.
Designing Good Tools
Test tools the way you'd test an API for a new colleague: run realistic tasks, read the transcripts, and refine descriptions where the model hesitates or misuses them. See context engineering.
Security
Tool calling connects model output to real systems, so treat the model as an untrusted caller:
- Validate arguments against schemas, and check business rules server-side.
- Authorize as the end user: tools must enforce the user's permissions, not the agent's superuser credentials.
- Require human confirmation for destructive or irreversible actions (payments, deletions, sending messages).
- Beware prompt injection: content returned by tools (web pages, emails, documents) can contain instructions aimed at the model. See prompt injection and AI guardrails.
- Limit blast radius: least-privilege credentials, rate limits, sandboxes for code execution, and audit logs of every tool call.
Best Practices
Write Descriptions Like Documentation
Spend more effort on tool descriptions than on system prompts. Include when to use the tool, parameter semantics, units and formats, and examples of good inputs.
Keep the Toolset Focused
Dozens of overlapping tools confuse selection and bloat every request. Offer the minimal set for the task, load tools dynamically (tool search), or split work across specialized agents. See multi-agent systems.
Bound the Loop
Set maximum iterations, timeouts, and token budgets, and detect repeated identical calls, so a confused model can't loop forever.
Log Every Call
Record tool name, arguments, results, latency, and errors for debugging, evaluation, and auditing. Traces of agent runs are the primary way to improve tool design. See LLM evaluation.
Common Mistakes
Dumping Raw API Responses
Map responses to the handful of fields the model needs, and offer a detail tool or pagination for more.
Vague Tool Descriptions
"description": "Gets data" leads to wrong tool selection and invented arguments. Be specific about purpose, inputs, and outputs.
Executing Side Effects Without Confirmation
An agent that can send_email or issue_refund on its own interpretation of an ambiguous request will eventually do something unintended. Gate high-impact tools behind explicit approval.
FAQ
Does the model execute the tools itself?
No. For client-side tools, the model only produces a structured request; your application decides whether and how to execute it and returns the result. Provider-hosted tools (web search, code execution) run on the provider's infrastructure, but the model still only requests them.
What's the difference between tool calling and structured outputs?
Structured outputs constrain the model's final response to a schema. Tool calling lets the model request actions mid-conversation and continue with the results. Forcing a single tool call with a schema is also a common way to get structured data.
How many tools can I give a model?
Models handle dozens of well-described tools, but accuracy and cost degrade as the toolset grows and overlaps. For large toolsets, use dynamic tool loading or search, group tools by domain, or route tasks to specialized agents with smaller toolsets.
How is MCP related to tool calling?
The Model Context Protocol is a standard for exposing tools (and resources and prompts) from servers to AI applications. An MCP client discovers a server's tools and presents them to the model as ordinary tool definitions, so any MCP-compatible app can use any MCP server's tools.
Related Topics
- AI Agents — Agents built on tool-use loops
- Model Context Protocol — Standard tool interfaces
- Structured Outputs — Schema-constrained responses
- Agent Design Patterns — Loops, planning, and workflows
- Prompt Injection — Securing tools against malicious content
- AI APIs — Model APIs that support tools