How to write a good prompt, with bad and good examples side by side
There is no magic word and no secret formula. There is a difference in method between asking and specifying, and it shows up in the first answer.
A good prompt has four things: the context of the situation, the exact task, the expected format of the answer and at least one example of the result you want. Most bad prompts fail by omission, not by lack of technique: the writer has all of that in their head and sends only the last sentence. Showing an example of the desired result pays off more than any other single technique.
What you get from this article
- A good prompt is a specification, not a request: what you have in your head and did not write, the model does not have.
- One example of the desired result is worth more than three paragraphs describing the desired result.
- Say what to do, not what not to do: a prohibition with no alternative tends to be ignored.
- Put the long text before the instruction, and the instruction last.
- Giving a role only works when the role changes the vocabulary, not when it is decoration.
- If the output goes into a system, ask for a structured format and validate before using it.
Why most prompts fail
The cause is almost always the same, and it is not lack of technique: omission. The writer has the customer, the history, the objective, the company tone and the expected format in their head, and sends only the last sentence of that reasoning. The model receives the sentence, not the head.
Compare "write an email to the customer" with what you actually want: an email to a customer who has not replied to a proposal for three weeks, who is already a client for another service, who you do not want to pressure, with the aim of reopening the conversation without sounding like chasing payment, short because they do not read long emails, signed by the company rather than by you.
The second version is not more technical, it is just more complete. And that is what separates a generic answer from a usable one. The model is very good at meeting a specification and terrible at guessing the one it did not receive.
There is a simple test that solves most cases: if you sent that same text to a competent intern who just arrived, could they do it? If the answer is no, the problem is not the model.
The model does not guess what was left out. And almost everything that decides the quality of the answer is precisely what was left out.
The four parts of a prompt that works
It is not a rigid formula, it is a checklist. Good prompts almost always have these four things, in some order.
Context: what the situation is, who the audience is, what happened before. It is the part most often omitted and the one that most changes the result. "For a customer who has already complained twice about the same problem" rewrites the entire answer.
Task: the specific verb. Summarize, compare, classify, rewrite, list, extract. "Help me with this text" is not a task. "Cut it to 150 words keeping the three main arguments" is.
Format: how the answer should arrive. Length, structure, whether it has a heading, whether it is a list or continuous prose, whether technical terms are allowed. Without that, you receive the average format of the internet, which is almost never yours.
Example: a sample of the result you want. It is the part that pays off most and the one almost nobody does, probably because it takes work the first time. It is worth the work: one example communicates tone, structure and level of detail all at once, without you having to describe any of it.
Bad and good, side by side
Three real business cases, with the request that does not work and the one that does.
Case 1, replying to a negative review. Bad: "reply politely to this bad review". Good: "I own a pizzeria. A customer left 2 stars saying the pizza arrived cold and the delivery rider was curt. The complaint is valid, we had a routing problem that day. Write a public reply of no more than 4 lines that acknowledges the mistake without making excuses, offers to resolve it by direct message and does not use call-centre stock phrases. Tone: a business owner talking, not a corporation. Do not promise a discount."
Case 2, message triage. Bad: "classify these messages". Good: "Classify each customer message into exactly one of these categories: QUOTE, SUPPORT, COMPLAINT, OTHER. Return one line per message in the format number|category|urgency from 1 to 3. If a message does not fit clearly, use OTHER instead of picking the closest. Do not explain the classification."
Case 3, page copy. Bad: "write some copy about our maintenance services". Good: "Write the copy for a service page about preventive air conditioning maintenance for residential buildings. Audience: a building manager, not a technician, who decides on cost and on the risk of penalties. Structure: an opening paragraph about the problem, three subheadings with two paragraphs each, and a closing with a call for a quote. Between 400 and 500 words. Do not invent savings figures or timelines."
Notice that none of the good examples uses a sophisticated technique. They only say what the person already knew and had not written down.
Say what to do, not what not to do
Negative instructions work worse than positive ones, and the reason is practical: a prohibition closes a door without opening another, and the model has to fill that space with something.
"Do not be too formal" leaves open what it should be. "Write like somebody explaining to a colleague in the corridor" solves it. "Do not use jargon" is weaker than "if you need a technical term, explain it in brackets the first time".
The exception is a concrete, verifiable prohibition, which works well precisely because it leaves no ambiguity: do not quote a price, do not promise a deadline, do not exceed 300 words. These are rules a person can check by looking at the result.
A valuable variation on that idea is authorizing the emergency exit. Instead of forbidding the model to invent, give it something to do when it does not know: "if the information is not in the material provided, write NOT FOUND and move on to the next item". That reduces invention far more than a generic prohibition, as explained in AI hallucination.
Order matters, more than it seems
When you send a long text along with the request, a contract, a report, a transcript, the position of the instruction changes the result.
The order that works best is: long material first, instruction last. The model processes the text in sequence, and the instruction at the end is fresh at the moment of answering. An instruction at the start of a twenty-page document has a real chance of being diluted.
That connects to a measured limitation of these systems: they pay more attention to the beginning and the end of what was sent than to the middle. Critical information buried in the middle of a huge block is the easiest to be ignored. Where you can, cut out the relevant passage instead of sending the whole document.
And there is a cost reason for the same order, which matters to anyone building a system: if the fixed block comes first and never changes, it can be cached and billed at a fraction of the price on subsequent calls. Putting today date at the start of the instruction breaks that, as detailed in tokens and the context window.
- Long, fixed material at the start, so it fits the cache and does not get diluted.
- Context and rules in the middle.
- The specific task at the end, immediately before the model answers.
- Examples between the rules and the task, clearly marked.
What is real technique and what is folklore
Giving a role works when it changes the vocabulary. "You are a tax lawyer" changes the repertoire used and is useful. "You are the best copywriter in the world" changes nothing, because it does not point at a specific repertoire. The test is: does this role correspond to a recognizable way of writing, or is it just flattery?
Asking it to think step by step still helps, with a caveat. In older models it was a big gain. In models with built-in reasoning, it already does that internally and the request only pads the answer. Where it still pays off is to see the reasoning when you need to check the logic.
Being polite does not improve the result. Please and thank you do not change the quality of the answer. They do no harm, and they are not technique.
Threatening or pleading does not work. Saying your job depends on it, promising a tip or insisting with urgency circulates in tutorials and produces no consistent improvement. What produces improvement is specifying better.
Complaining about the result does not fix it. "That was bad, do it again" returns something else equally random. Say what was wrong and what should be there instead: "the first two paragraphs came out generic, replace them with concrete examples from residential buildings".
When the output goes into a system
A prompt for a person to read and a prompt for a program to consume are different things, and mixing them is a silent source of error in production.
If the answer is going to be read by code, ask for an explicit structured format, define every field, and say what to do when a field does not exist. "Return JSON with the fields name, amount and due date; use null when the information is not in the document" is far more reliable than hoping the model picks a reasonable format.
And always validate before using. The model can return an extra field, an explanatory sentence before the JSON, or a value in the wrong format. Code that trusts model output without checking breaks in production the day a different input arrives.
The rule worth keeping: what is verifiable by machine should be verified by machine. If the model returned an amount, recalculate it. If it returned a date, validate the format. If it returned a product code, check whether it exists in the catalog.
If you are building a flow like that and want to define where the validations sit before it becomes code, that is the kind of decision ROO3 AI consulting resolves at the start of the project, while it is still cheap to change.
Frequently asked questions
Is there a prompt formula that always works?
There is no magic formula, but there is a checklist that solves most cases: context of the situation, a task with a specific verb, the expected format of the answer, and at least one example of the desired result. Most bad prompts fail by omitting one of those four.
Does being polite to the AI improve the answer?
No. Please and thank you do not change the quality of the result. They do no harm, but they are not technique. What improves the answer is specifying the context, the task and the format better.
Is it worth saying "you are an expert in X"?
It is worth it when the role points at a recognizable vocabulary, such as a tax lawyer or a workplace safety engineer, because that changes the repertoire used. It is not worth it when it is just flattery, such as "you are the best copywriter in the world", which corresponds to no specific way of writing.
Why does the AI ignore part of my instructions?
Usually for three reasons: too many instructions competing with each other in a very long prompt, an important instruction buried in the middle of a large text, or a negative instruction that closes a door without opening another. Cutting down, moving to the end and swapping prohibition for positive guidance solves most cases.
Do I need to ask the AI to think step by step?
In models with built-in reasoning, it already does that internally and the request only pads the answer. The request stays useful when you need to see the reasoning to check the logic, or in simpler models that do not reason by default.
How do I make the AI always return the same format?
Ask for the format explicitly, define every field, show a complete example of the output and say what to do when a field does not exist. And validate the answer in code before using it, because no format is guaranteed just by being requested.
Rodrigo Fávaro
Founder of ROO3, a marketing and technology agency in São José do Rio Preto, Brazil. Builds AI products running in production (Tobia, gerar.app, Pense Mercado) and maintains the AI Benchmark, a public ranking of AI models. See ROO3 AI consulting.
X @rodmf LinkedIn rodrigofavaroKeep reading

AI hallucination: why it invents and how to reduce it
What AI hallucination is, why it happens by design, what the research shows about the cause, and the techniques that...
10 min read
RAG, fine-tuning or prompt: which to use in each case
When to use a prompt, when to use RAG and when fine-tuning is genuinely justified. The question that decides, the cost...
10 min read
What an LLM is: how it works inside and what it does not do
What an LLM is, explained without maths: how the model predicts the next word, why that works so well, and which limits...
12 min readWant to apply this in your company?
ROO3 diagnoses what can be automated first in your business. The first conversation is free.