Skip to main content

Discover the best AI tools curated for professionals.

AIUnpacker

Search everything

Find AI tools, reviews, prompts, and more

Quick links
EducationVerified

7 AI Prompt Structures That Generate Better Content Every Time

No prompt produces perfect content every time. But seven research-backed prompt structures (chain-of-thought, ReAct, tree-of-thoughts, few-shot, XML-tagged, structured output, self-consistency) reliably beat unstructured prompts in the actual benchmark data. Includes a framework comparison table, copy-paste templates, and the honest caveats.

AIUnpacker

AIUnpacker Editorial

16 min read
AIUnpacker

AIUnpacker

16m read

16 min

Key Takeaways

No prompt produces perfect content every time. But seven research-backed prompt structures (chain-of-thought, ReAct, tree-of-thoughts, few-shot, XML-tagged, structured output, self-consistency) reliably beat unstructured prompts in the actual benchmark data. Includes a framework comparison table, copy-paste templates, and the honest caveats.

Summarize with AI

Editorial Disclosure & Affiliate Notice

This content is published for informational and educational purposes only. It is not intended as a substitute for professional, legal, financial, or medical advice. AIUnpacker is funded by sponsorships, affiliate commissions, and display advertising — nothing here is free to produce. When you buy through our links, we may earn a commission at no extra cost to you. Our editorial picks are never influenced by compensation.

  • For educational purposes only. Nothing here should be taken as a guarantee, recommendation, or professional recommendation.
  • AI-assisted editing. Drafts are produced with AI assistance and reviewed by our human editorial team.
  • Opinions are our own. Also, we are not affiliated with most tools we cover unless explicitly stated.
  • Information may be outdated. Verify pricing, features, and policies directly with the vendor.
  • Last reviewed: . Published .

Read more on our About page, Terms and Editorial Policy.

I have to start with a confession. The headline of this article is lying to you a little. No prompt produces perfect content every single time, and anyone who tells you otherwise is selling a course. What I can promise you is that seven prompt structures, all of them backed by published research papers and official documentation from OpenAI, Anthropic, Google, and IBM, dramatically raise the floor and ceiling of what GPT-5.6, Claude Opus 4.8, and Gemini 3 can produce for you. After a couple of years writing with these models almost daily, I have stopped arguing about whether prompt engineering “still works” in 2026 and started treating prompt structures the way I treat CSS frameworks: a small set of patterns I reach for depending on the job.

This article walks through the seven structures I actually use, the benchmark data behind each one, a side-by-side comparison of the popular named frameworks (RTF, RISEN, CO-STAR, CRISPE, RACE, and TAG), and the honest caveats the research papers themselves flag. You can copy the templates, swap your variables, and ship.

What “better” actually means here

Every structure below has the same basic job: reduce ambiguity between you and the model. Anthropic’s own prompt engineering guide is blunt about it: “Think of Claude as a brilliant but new employee who lacks context on your norms and workflows. The more precisely you explain what you want, the better the result” [1]. OpenAI’s guide makes the same point: “Because the content generated from a model is non-deterministic, prompting to get your desired output is a mix of art and science. However, you can apply techniques and best practices to get good results consistently” [2].

A 2024 academic survey of the field tallied more than 50 distinct text-based prompting techniques and 40 multimodal variants, and reported that “reordering examples in a prompt produced accuracy shifts of more than 40 percent” and that “some studies have shown up to 76 accuracy points across formatting changes in few-shot settings” [3]. That number is uncomfortable. It means a sloppy prompt is leaving most of the model on the table. It also means the right structure is not a nice-to-have. It is the difference between a usable draft and a re-write.

“On the Game of 24 task, while GPT-4 with chain-of-thought prompting only solved 4% of tasks, our method [Tree of Thoughts] achieved a success rate of 74%.” - Yao et al., Tree of Thoughts, NeurIPS 2023 [4]

That single number is the cleanest argument for prompt structures I have ever read. Same model, same problem, same day. Change the prompt structure from a single greedy chain-of-thought to a tree that explores and backtracks, and you go from 4% to 74%. Nobody beats that with vibes.

The 7 prompt structures

Here are the seven structures I reach for, in roughly the order I teach them. They layer on top of each other. Structures one and two are scaffolding. Three through five add reasoning and self-checking. Six adds structure to what the model returns. Seven adds reliability through voting.

1. Role + Context + Task + Constraints (the RTF family)

Every content prompt needs four things: who the model is, what world it is writing in, what it should produce, and the rails around length, format, voice, and what to avoid. OpenAI’s docs make the four-way split explicit through message roles: developer (system rules), user (the actual request), plus context and constraints inside each [2]. Anthropic recommends the same separation and adds a “Golden rule” for testing prompts: “Show your prompt to a colleague with minimal context on the task and ask them to follow it. If they’d be confused, Claude will be too” [1].

This is the structure behind almost every named framework on the internet: RTF (Role, Task, Format), RISEN (Role, Instructions, Steps, End goal, Narrowing), CO-STAR (Context, Objective, Style, Tone, Audience, Response format), and CRISPE (Capacity/Role, Request, Insight, Statement, Personality, Experiment). They are the same idea with different mnemonics.

A working template:

<role>
You are a senior B2B content strategist who has written for HubSpot,
Salesforce, and Notion. You write in clear, conversational English
with short sentences.
</role>

<context>
Our product is a SOC 2 compliance tool for startups. The reader is a
CTO at a 50-person SaaS company in the US who has been told by sales to
"get serious about compliance" before a big enterprise deal.
</context>

<task>
Write a 900-word blog post titled "What SOC 2 Actually Costs a Startup
in 2026." Open with the dollar figure. End with a checklist of 7 items.
</task>

<constraints>
- Avoid jargon: do not use "audit-ready," "control matrix," or "type II."
- No emoji. No exclamation points.
- Cite specific numbers from the context above. Do not invent stats.
- Use H2 and H3 headings, then prose paragraphs under each. No bullet lists
  except the closing checklist.
</constraints>

Anthropic specifically recommends XML tags like <role>, <context>, and <task> because they “help Claude parse complex prompts unambiguously” [1]. OpenAI agrees and adds that mixing “a combination of Markdown formatting and XML tags” is the modern best practice [2].

2. Few-shot with anchor examples

Few-shot is the most reliable trick in the book. Give the model 3 to 5 examples of the exact format, tone, and edge cases you want, and the model will generalize from them. Anthropic recommends 3 to 5 examples [1]. OpenAI’s docs are built around the same idea.

The trick is to make your examples diverse and relevant. Anthropic warns against giving only similar examples, because “Claude doesn’t pick up unintended patterns” is exactly what happens when your examples all look alike [1]. Mix one common case, one edge case, and one “do not do this” case.

<examples>
<example>
Input: "The cloud migration was hard but we got there."
Tweet: "Migrated 14 services to AWS in 90 days. Here's what I'd do
differently. [thread]"
</example>

<example>
Input: "Salesforce announced new agent features."
Tweet: "Salesforce just shipped Agentforce 3. Three things stand out:
1) low-code agent builder 2) native Slack handoff 3) pricing tied to
outcomes, not seats. Worth a closer look for RevOps teams."
</example>

<example>
Input: "Our Q3 earnings beat expectations."
Tweet (intentionally NOT made): "🎉🚀 WE CRUSHED IT!!! BIGGEST QUARTER EVER!!!"
</example>
</examples>

Now write a tweet for: "OpenAI released GPT-5.6 with a 1M context window."

That third negative example is something most people skip. It teaches the model what not to do, which is often more useful than another positive case.

3. Chain-of-thought (“think before you answer”)

Chain-of-thought (CoT) prompting was the 2022 paper that launched the modern era. Wei et al. at Google Brain showed that “prompting a 540B-parameter language model with just eight chain of thought exemplars achieves state of the art accuracy on the GSM8K benchmark of math word problems” [5]. The trick is tiny: ask the model to show its work. “Let’s think step by step” was enough to unlock zero-shot CoT, according to follow-up work from Google and the University of Tokyo [3].

Today, every frontier model has reasoning baked in. OpenAI’s reasoning models “generate an internal chain of thought to analyze the input prompt” [2]. Anthropic ships adaptive thinking on Claude Opus 4.7, Claude Opus 4.8, Claude Sonnet 4.6, Claude Fable 5, and Claude Mythos 5, where the model “dynamically decides when and how much to think” [1]. You still get a massive lift by adding a <thinking> block to your prompt when you are using a non-reasoning model, or by instructing the model to “think through this carefully before answering” when you are using one that exposes its scratchpad.

Before you write the final answer, plan your reasoning in a
<thinking> block. Then produce the post in <answer> tags.

If you skip CoT on hard problems, you are paying for a reasoning model and using it like a base model. That is the most expensive mistake in modern prompting.

4. ReAct: reason, then act, then read the result

ReAct, from Yao et al. at Princeton in 2022, interleaves reasoning with actions against external tools or APIs [6]. On the ALFWorld benchmark the paper reports ReAct outperforming imitation and reinforcement learning baselines “by an absolute success rate of 34%” [6]. Anthropic’s Building Effective Agents engineering post lists prompt chaining, routing, parallelization, orchestrator-workers, and evaluator-optimizer as the canonical workflows [7]. ReAct is the simplest of those patterns: think, do, observe, repeat.

For content work, ReAct shows up when you give the model tools: a web search, a file reader, a function call. The structure is:

<scratchpad>
I need to write a pricing page for our SOC 2 tool. Before I write, I
will:
1. Search our internal knowledge base for the current pricing tiers.
2. Search the public site for any positioning notes from the product team.
3. Compare the two, then draft the page.
For each step, state the reason, take the action, and read the result
before moving on.
</scratchpad>

Now execute step 1.

This pattern is what underpins every serious agent loop in 2026, from Claude Code to the OpenAI Agents SDK to the workflow agents Anthropic describes in its engineering blog [7].

5. Tree of thoughts: explore, self-critique, backtrack

When the problem has multiple plausible answers and the first one is often wrong (strategy, planning, creative ideation, multi-step math), Tree of Thoughts (ToT) generalizes CoT. The 2023 paper reports 74% on Game of 24, versus 4% for vanilla CoT on the same model [4].

In practice, this is two prompts. First you generate a few candidate branches:

Generate 3 different approaches to the headline for this article. For
each approach, write the headline, a one-sentence rationale, and a
predicted CTR reasoning.

Approach A: ...
Approach B: ...
Approach C: ...

Then you evaluate and pick:

Rate each approach 1-10 on: clarity, curiosity, specificity, and
whether it would survive a director's edit. Reject any approach below
6 on clarity. Pick the strongest remaining approach and explain why.

IBM lists ToT as one of its core “agentic prompting” techniques for exactly this kind of structured exploration [8].

6. Structured output + XML framing

Stop asking for prose when you want structured data. Both OpenAI and Anthropic ship first-class structured output. OpenAI’s structured_outputs feature constrains the response to a JSON schema you provide [2]. Anthropic ships a structured_outputs API and recommends wrapping different parts of your prompt in tags so the model can tell them apart [[1]](#sources]. For long-context work (documents over 20k tokens), Anthropic reports that putting “longform data at the top” and putting queries at the end “can improve response quality by up to 30% in tests” [1].

<documents>
<document index="1">
<source>Q2-board-update.pdf</source>
<document_content>
{{PASTE_BOARD_UPDATE}}
</document_content>
</document>
<document index="2">
<source>competitor-pricing.xlsx</source>
<document_content>
{{PASTE_COMPETITOR_SHEET}}
</document_content>
</document>
</documents>

<task>
Return a JSON object with this exact schema:
{
  "highlights": ["string", "string", "string"],
  "risks": ["string", "string"],
  "competitor_threats": [{"name": "string", "threat_level": "low|med|high"}]
}
</task>

This is the structure I use for every internal tool that touches the API. It kills the “the model returned prose when I wanted JSON” failure mode, and it makes downstream code trivial.

7. Self-consistency: vote on multiple drafts

The Wang et al. 2023 paper introduced self-consistency: run the prompt several times, then pick the answer that comes up most often [9]. On GSM8K it added 17.9 points on top of CoT, 11.0 on SVAMP, 12.2 on AQuA, and 6.4 on StrategyQA [9]. Anthropic’s Building Effective Agents calls this the “voting” pattern under parallelization and lists it as one of the canonical workflows [7].

For content work, I almost never literally sample five times and tally. Instead, I use the same idea in two cheaper ways:

  • Generate three drafts, then pick the strongest and explain why. The reasoning pass usually beats the strongest draft.
  • Generate three headlines, three subject lines, three opening sentences, then pick the best.
Generate 5 subject lines for this email under 60 characters each.
Then pick the strongest and explain why in one sentence.

Same idea, no API gymnastics, and the second pass forces the model to internalize a standard rather than pick a loud-but-flaky first answer.

Quick comparison: named frameworks side by side

These are not seven separate inventions. They are seven flavors of the same idea, packaged by different communities. The table below is what I actually keep in my notes app.

Framework Acronym stands for Best for Origin / popularity Caveat
RTF Role, Task, Format Quick one-shot content Sora/Qasar community, 2023 Lightweight; easy to outgrow
RISEN Role, Instructions, Steps, End goal, Narrowing Long-form content and creative work Frameworks.ai / Mar-Gregorio, 2023 Five fields force rigor
CO-STAR Context, Objective, Style, Tone, Audience, Response Marketing copy, stakeholder work Sheila Teo, GovTech Singapore, 2023 Widely adopted in Asia-Pacific
CRISPE Capacity, Request, Insight, Statement, Personality, Experiment Research-heavy or analytical prompts Matt Nigh, 2023 Insight and Personality fields can overlap
RACE Role, Action, Context, Expectation Operational prompts and SOPs practitioner community Less suited for creative briefs
TAG Task, Action, Goal API-mediated and tool-calling prompts OpenAI cookbook tradition Pair it with a JSON schema

You do not need all of them. RTF, RISEN, and CO-STAR will get you through 95% of content work. The framework matters far less than the four underlying moves: name the role, ground it in context, define the task, and lock the constraints.

If you want the strongest single sentence for a content prompt, this is what I would suggest you steal from RISEN: always state the End goal. “The reader should leave this post with a one-line answer to X” is the kind of sentence that turns generic blog output into something a human could be proud of.

How to stack the seven

You almost never use one structure in isolation. The moves layer. My default stack for a 1,500-word article:

  1. Structure 1 (Role + Context + Task + Constraints) as the wrapper.
  2. Structure 2 (Few-shot) for tone and format, especially on the first project with a new voice.
  3. Structure 3 (CoT) inside the assistant instructions: “Plan the outline in <thinking>, then write the post.”
  4. Structure 5 (ToT) on the headline, the lead, and the closing line.
  5. Structure 6 (Structured output) when I want a brief or a metadata block, not the full prose.
  6. Structure 7 (Self-consistency) on the sections that have to land (the headline, the first sentence, the call to action).

ReAct (structure 4) comes in only when the model has tools, and only on tasks where grounding matters: anything that touches live data, internal documents, or external APIs.

Honest caveats the research flags

Three things the marketing copy about prompt engineering does not tell you:

Prompts are brittle. Wikipedia’s prompt engineering entry, citing recent studies, notes that “minor surface-level changes in phrasing, punctuation, or word order can produce dramatically different outputs, even when the semantic intent remains identical” [3]. That is why you ship tests with your prompts, not vibes.

What worked on GPT-3 may not work on Claude Opus 4.8. “Effective prompting strategies are highly model specific,” the same source notes, “a technique that improves performance on one model may degrade it on another” [[3]](#sources]. When OpenAI deprecated reusable prompt objects in favor of typed code-as-prompt in 2026, it was for this reason [2]. Treat every prompt like a piece of code: version it, test it, and re-test when the model changes.

The role of “prompt engineer” is shrinking. The Wall Street Journal reported in 2025 that the prompt engineer role, the hottest of 2023, has cooled as models better intuit user intent and as companies train staff [[3]](#sources]. What has not shrunk is the need for people who can design prompt systems: chains, evaluators, structured outputs, and tool calls. The structures above are not tricks for one-off chats. They are the bricks of those systems.

Where this leaves you

If you take nothing else from this article, take this: prompts are not magic spells, they are interfaces. The seven structures above are the patterns I have watched survive every model upgrade since GPT-4. Combine them with evaluation (literal pass/fail tests of model output), version control (commit your prompts to git), and a clear sense of what the model does not know (use RAG, not vibes, for facts), and you will be well ahead of the median content team in 2026.

You will still hit bad days. The model will still hallucinate. The headline will still land flat sometimes. But the gap between a well-structured prompt and a sloppy one is not 5 percent. It is, in the right benchmarks, 70 percentage points. That gap is what this article is about.

FAQ: AI prompt structures in 2026

What is the single most impactful AI prompt structure? Give the model role, context, task, and constraints in that order. The named frameworks (RTF, RISEN, CO-STAR) are all packaging this one idea.

Do chain-of-thought and tree-of-thought still help on GPT-5.6 and Claude Opus 4.8? Yes. GPT-5.6 is a reasoning model that already runs an internal CoT, and Claude Opus 4.8 ships adaptive thinking [1] [2]. Forcing explicit CoT blocks or asking for “branch-and-prune” reasoning still measurably improves structured tasks like planning and classification.

How is ReAct different from chain-of-thought? CoT is pure reasoning. ReAct interleaves reasoning with actions, like searching a knowledge base or calling an API, then reads the result before the next step. It is the prompt pattern every modern agent loop inherits [[6]](#sources] [[7]](#sources].

When should I use structured output instead of prose? Whenever a downstream system needs to parse the response. OpenAI’s structured_outputs and Anthropic’s structured outputs API both constrain the model to a JSON schema [[1]](#sources] [2]. Save prose for prose.

Do prompt frameworks like RTF, RISEN, and CO-STAR actually beat free-form prompts? In published comparisons, mostly yes when the task has audience, tone, or format requirements. The frameworks do not invent new techniques so much as force discipline: stating the role, context, and constraints is the thing that moves the needle.

What is the difference between few-shot and chain-of-thought prompting? Few-shot teaches format and tone with examples. Chain-of-thought teaches the model to externalize its reasoning. You can stack them: few-shot CoT, where each of your 3 to 5 exemplars includes a thinking block, is one of the strongest single-shot prompting techniques in the literature.

Are these structures different for ChatGPT, Claude, and Gemini? The pattern is identical. The phrase choice varies. Anthropic recommends <example> tags and system messages; OpenAI prefers developer messages and Markdown with XML; Google Gemini emphasizes system instructions and structured output [[1]](#sources] [[2]](#sources] [10]. The bones are the same.

Is prompt engineering obsolete in 2026? As a hand-tuned “magic phrase” hobby, yes. As a system-design discipline (chains, evaluators, structures), absolutely not. Anthropic’s own engineering post is titled Building Effective Agents, not Do You Even Need a Prompt [[7]](#sources].

Sources

  1. Anthropic. Prompting best practices for Claude. docs.anthropic.com, Claude Opus 4.8, Sonnet 5, Fable 5 / Mythos 5 documentation. https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/claude-prompting-best-practices
  2. OpenAI. Prompt engineering guide. platform.openai.com, GPT-5.6 Sol / Terra / Luna documentation. https://platform.openai.com/docs/guides/prompt-engineering and https://platform.openai.com/docs/models
  3. Wikipedia contributors. Prompt engineering. Wikipedia, accessed July 2026. https://en.wikipedia.org/wiki/Prompt_engineering (cites Wei et al. 2022, Yao et al. 2023, Lewis et al. 2020, APE paper, the Wall Street Journal April 2025 piece on the prompt-engineering role, and the EMNLP 2024 paraphrase-types paper by Wahle et al.)
  4. Yao, Shunyu; Yu, Dian; Zhao, Jeffrey; Shafran, Izhak; Griffiths, Thomas L.; Cao, Yuan; Narasimhan, Karthik. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. arXiv:2305.10601, NeurIPS 2023 camera-ready. https://arxiv.org/abs/2305.10601
  5. Wei, Jason; Wang, Xuezhi; Schuurmans, Dale; Bosma, Maarten; Ichter, Brian; Xia, Fei; Chi, Ed; Le, Quoc; Zhou, Denny. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. arXiv:2201.11903. https://arxiv.org/abs/2201.11903
  6. Yao, Shunyu; Zhao, Jeffrey; Yu, Dian; Du, Nan; Shafran, Izhak; Narasimhan, Karthik; Cao, Yuan. ReAct: Synergizing Reasoning and Acting in Language Models. arXiv:2210.03629, ICLR 2023. https://arxiv.org/abs/2210.03629
  7. Eriksson, Erik; Zhang, Barry. Building effective agents. Anthropic Engineering blog, December 19, 2024. https://www.anthropic.com/engineering/building-effective-agents
  8. IBM. What is prompt engineering? IBM Think topic guide, with sub-pages on ReAct, tree of thoughts, meta-prompting, iterative prompting, DSPy, prompt caching, and role prompting. https://www.ibm.com/think/topics/prompt-engineering
  9. Wang, Xuezhi; Wei, Jason; Schuurmans, Dale; Le, Quoc; Chi, Ed; Narang, Sharan; Chowdhery, Aakanksha; Zhou, Denny. Self-Consistency Improves Chain of Thought Reasoning in Language Models. arXiv:2203.11171, ICLR 2023. https://arxiv.org/abs/2203.11171
  10. Google. Prompting strategies and Introduction to prompt design in the Gemini Enterprise Agent Platform documentation (Vertex AI / Gemini 3). https://cloud.google.com/vertex-ai/generative-ai/docs/learn/prompts/prompt-design-strategies
  11. Bai, Yuntao; Kadavath, Saurav; Kundu, Sandipan; Askell, Amanda; et al. Constitutional AI: Harmlessness from AI Feedback. arXiv:2212.08073; Anthropic, December 15, 2022. https://www.anthropic.com/research/constitutional-ai-harmlessness-from-ai-feedback
  12. Lewis, Patrick; Perez, Ethan; Piktus, Aleksandra; Petroni, Fabio; Karpukhin, Vladimir; Goyal, Naman; Küttler, Heinrich; Lewis, Mike; Yih, Wen-tau; Rocktäschel, Tim; Riedel, Sebastian; Kiela, Douwe. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. arXiv:2005.11401, NeurIPS 2020. https://arxiv.org/abs/2005.11401
  13. Khattab, Omar; Singhvi, Arnav; Maheshwari, Paridhi; et al. DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines. arXiv:2310.03714. https://arxiv.org/abs/2310.03714
  14. Zhou, Yongchao; Muresanu, Andrei Ioan; Han, Ziwen; Paster, Keiran; Pitis, Silviu; Chan, Harris; Ba, Jimmy. Large Language Models Are Human-Level Prompt Engineers (APE). arXiv:2211.01910. https://arxiv.org/abs/2211.01910
  15. Anthropic. Claude Opus 4.8 model page (1M context window, adaptive thinking, May 28 2026 release). https://www.anthropic.com/claude/opus

Weekly digest

Get our weekly AI digest

The latest AI tools, prompts, and insights — delivered every Tuesday.

No spam. Unsubscribe anytime.

AIUnpacker

AIUnpacker Editorial Team

Verified

A collective of engineers, journalists, and AI practitioners dedicated to providing hands-on, transparently disclosed analysis of the AI tools shaping tomorrow.