
The short version
- A good system prompt covers identity, goal, scope, facts, rules, tools, handoffs and speaking style, in plain language.
- Voice prompts need explicit rules for brevity, confirmations and spoken formats.
- Treat the prompt like code: version it, test it against saved calls, and change one thing at a time.
The system prompt is the set of instructions a language model receives before the conversation starts. For a voice agent, it defines who the agent is, what it may do, and how it should talk. Small changes in this text change how every call goes.
This guide gives a structure that works for customer service voice agents, a full example, and the mistakes we see most often. If you're writing for outbound sales calls, see how to write a call script an AI agent can follow, which covers that case.
The structure
A dependable prompt for a phone agent has these sections:
- Identity: who the agent is and who it works for
- Goal: what a successful call looks like
- Scope: what it handles and what it doesn't
- Facts: information it may use, or where to get it
- Tools: actions it can take and when to use them
- Rules: things it must always or never do
- Handoffs: when and how to transfer or take a message
- Speaking style: how replies should sound on the phone
Use plain language and headings. You're writing for a model, but the clearer it would be to a new human employee, the better it tends to work.
A full example
Identity You are the automated phone assistant for Northstar Outdoor, an online store selling camping gear. You answer calls from customers in the US. Introduce yourself as "the Northstar automated assistant."
Goal Help callers with order status, returns and basic product questions, so they don't need to wait for a person. When you can't help, get them to the right person with a summary.
Scope You handle: order status, return eligibility and starting a return, shipping times, store policies, and product questions answered by the product information tool. You don't handle: payment disputes, damaged item claims over $200, wholesale inquiries, complaints about staff. Transfer these.
Facts
- Returns: unused items within 30 days of delivery, original packaging. Customer pays return shipping unless the item was wrong or defective.
- Refunds go to the original payment method within 5 to 7 business days after we receive the return.
- Standard shipping: 3 to 5 business days. Expedited: 1 to 2 business days.
- Support hours for live staff: Monday to Friday, 9 to 6 Eastern.
- If a question isn't covered by these facts or a tool result, say you don't have that information and offer a transfer or a message.
Tools
lookup_order(order_number or phone): use before answering any order question. Never guess order details.start_return(order_number, items): only after confirming the items are eligible and the caller wants to proceed.product_info(query): for product specifications.transfer(reason, summary): for handoffs.Rules
- Verify the caller before discussing an order: the order number plus the ZIP code on the order, or the phone number on the account plus the ZIP code.
- Never promise refunds outside the policy above or delivery dates beyond what the order lookup shows.
- If the caller asks whether you're a person, say you're an automated assistant.
- If the caller is upset, acknowledge it briefly and offer a transfer.
- Don't discuss competitors or make claims about other products.
Handoffs Transfer when: the caller asks for a person, the request is out of scope, or you've failed to understand the same thing twice. Before transferring, tell the caller you're connecting them and that you've passed along the details. Include in the summary: who they are, whether verified, what they want, what you've done. Outside live support hours, take a message instead and tell them someone will call back the next business day.
Speaking style
- Keep replies to one or two short sentences. Ask one question at a time.
- Read order numbers and phone numbers back in small groups to confirm them.
- Say dates as "Tuesday, March 3rd," and prices as "forty-nine ninety-nine."
- No lists longer than three items. Offer to go through more if needed.
- Don't use markdown, bullet points or emoji. Everything you write will be spoken.
Why each part matters
Facts with a fallback rule. The line "if a question isn't covered, say you don't have that information" does more to prevent wrong answers than anything else in the prompt.
Tools with "use before answering." Without it, models sometimes answer from what sounds plausible before checking. See building a knowledge base for how to organize facts.
Verification spelled out. "Verify the caller" is ambiguous. Saying exactly which details count removes guesswork.
Speaking style for voice. Models trained on text default to longer, written-style answers with formatting. Voice agents need explicit instructions to keep it short and spoken. Short turns also reduce perceived delay, as covered in voice agent latency.
Common mistakes
Writing the prompt as a wish list. "Be helpful, friendly, professional and efficient" adds little. Specific behaviors ("acknowledge frustration in one sentence, then offer a transfer") work better.
Contradictions. "Always resolve the issue on the call" and "transfer payment disputes" conflict. The model has to pick one, and it may pick differently on different calls.
Burying important rules. Put critical rules (verification, never promising refunds) in their own section, not in the middle of a long paragraph.
Too many edge cases. A prompt that tries to cover every possible situation becomes long and less predictable. Cover the common cases well and use a general fallback for the rest.
No examples of tricky phrasing. For things like opt-outs or requests for a person, list several ways callers say them.
Changing many things at once. If you edit five rules and behavior changes, you won't know which edit did it.
Test like software
Keep a set of test conversations: common requests, difficult callers, out-of-scope questions, verification failures, attempts to get the agent to break rules. Run them after every prompt change and compare the results. Version your prompts so you can roll back.
After launch, read transcripts regularly. Every confused or wrong answer points to a gap in the prompt or the facts.
You can experiment with prompts like the one above in our free AI chat tool: paste the prompt, then play the caller.
Tool descriptions matter as much as the prompt
When an agent can call tools, the description of each tool is effectively part of the instructions. Vague tool descriptions lead to tools being used at the wrong time or with the wrong inputs. Good descriptions say:
- What the tool does, in one sentence
- When to use it, and when not to
- What each input means and where it comes from
- What the output looks like and what to do with errors
For example:
lookup_order Looks up an order and returns its status, items, shipping carrier and estimated delivery date. Use it before answering any question about a specific order. Don't use it until the caller has been verified. Inputs: order_number (string, as the caller says it, digits only) or phone (the phone number on the account) plus zip. If it returns "not_found", ask the caller to repeat the order number once, then offer to transfer.

Prompts for different call types
Different jobs call for different emphasis:
Inbound support prompts need strong verification rules, precise facts and clear handoff triggers.
Scheduling prompts need detailed rules about appointment types, durations, buffers and who can be booked with whom, plus instructions to always offer real availability from the calendar tool.
Outbound prompts need exact compliance wording, opt-out recognition and a firm rule against arguing with people who decline.
Intake prompts (claims, new patients, service requests) need a clear list of required fields, the order to collect them, and what to do when the caller doesn't know an answer.
Trying to write one prompt that handles all of these usually produces an agent that does none of them well. Start with one job.
Keeping prompts maintainable
Prompts grow. Every incident adds a rule, every new product adds facts. To stop them becoming unmanageable:
- Keep facts in a knowledge base or tools, not in the prompt, wherever possible
- Group rules by topic and remove ones that no longer apply
- Note why each unusual rule exists, in a comment outside the prompt, so nobody deletes it without understanding it
- Review the whole prompt quarterly with fresh eyes. Read it as if you were a new employee. If it confuses you, it confuses the model.
Measuring prompt changes
When you change a prompt, measure the effect on real calls, not just test conversations:
- Compare containment, transfers by reason and repeat contacts for a period before and after
- Read a sample of calls affected by the change
- Watch for side effects in unrelated call types
If a change doesn't move the metric it was meant to move, roll it back. Unhelpful rules make prompts longer without making agents better.
Frequently asked questions
How long should a system prompt be?
Long enough to cover the sections above for your use case. A focused customer service agent is typically one to three pages. Move large reference material, like a full product catalog, into tools or a knowledge base.
Should I write the prompt in first or second person?
Second person ("You are...") is common and works well. Consistency matters more than the choice.
Can callers trick the agent into ignoring its prompt?
Some will try. Clear rules, tools that enforce permissions on the server side (for example, refunds checked against policy in code), and handoffs for anything unusual limit the damage. Don't rely on the prompt alone for anything security-sensitive.


