ImageFirm Intelligence · 2026 Edition

One task.
Nine AIs.
Nine prompt dialects.

A research-grounded field guide to the transferable science—and model-specific craft—of prompting major generative AI systems.

University level9 model ecosystems30+ worked patternsFirst-party sourcesReviewed 28 July 2026
Central thesis. Models do not possess fixed “characters” in the human sense. Yet they do exhibit repeatable interaction profiles created by post-training, instruction hierarchy, reasoning style, tool architecture, context treatment, and product design. Treat “character” as a useful interface metaphor—not a psychological claim.
01 · The common grammar

A robust prompt is a specification.

Across providers, strong prompts answer seven questions. Persona is optional; evidence, constraints and an output contract usually matter more.

01

Outcome

What decision, artifact or transformation must exist when the work is done?

02

Context

What background, audience, definitions and source material change the answer?

03

Evidence

Which sources may be used, how current must they be, and what must be cited?

04

Constraints

Scope, exclusions, length, language, risk boundaries, tools and budget.

05

Process control

What may the model decide autonomously? When should it ask, verify or stop?

06

Output contract

Required sections, schema, units, ordering, tone and acceptance criteria.

07

Evaluation

How will correctness, completeness, groundedness and usefulness be tested?

02 · Comparative matrix

Same fundamentals, different leverage points.

These are interaction tendencies and documented affordances, not immutable rankings. Product and API behavior can differ even when they use related models.

EcosystemHigh-leverage moveUseful structureReasoning guidanceLong context / retrievalFrequent mistake
OpenAI GPTState the goal, constraints, evidence and “done” conditionDeveloper/system policy + concise task; Markdown/XML delimitersFor reasoning models, ask for outcome and verification—not hidden stepsStable material first; changing task later; use tools for freshnessOver-engineering with repeated instructions and “think step by step”
Anthropic ClaudeMake implicit expectations explicit; supply 3–5 representative examplesSemantic XML tags are especially naturalAsk for criteria and checks; configure thinking separately when availableClearly fence documents; put queries after large evidence blocksNegative-only rules, vague “be concise,” or examples that contradict prose
Google GeminiDirect instruction + multimodal context + explicit output shapeHeadings/tags; instructions after very large contextCurrent reasoning models prefer concise, direct promptsAnchor: “Based on the preceding information…”Placing the real question far before a huge document set
xAI GrokSpecify whether current web/X evidence or pure reasoning is requiredGoal, search scope, evidence standard, formatFocus on result and source criteriaDistinguish model knowledge from live search and attached-file toolsAssuming “real-time” without enabling or requesting the relevant tool
Meta LlamaUse the exact chat template for the deployed checkpointCorrect system/user/assistant tokens via the tokenizerBenchmark per checkpoint and quantizationContext limits and behavior depend on host/configurationHand-writing prompt tokens or assuming all Llama hosts behave alike
MistralPut durable behavior in system; show the desired formatClear system role + instruction + demonstrationsUse direct decomposition where measured usefulChoose model and endpoint for task; validate language adherenceExpecting persona text alone to enforce production guardrails
DeepSeekMatch prompt and API mode: thinking, tool use or JSONExplicit system task; JSON example when using JSON modeDo not mix hidden-reasoning transport with ordinary text handlingPreserve required reasoning/tool state across API turnsRequesting JSON without the API flag, the word “json,” example and token headroom
Microsoft CopilotGoal + context + source + expectationsReference the active file, meeting, mail or app artifactAsk for assumptions and verification where decisions matterGrounding depends on tenant permissions and selected Microsoft 365 contextWriting an abstract prompt while failing to name the available work source
PerplexityWrite a retrieval query, not merely a conversational requestQuestion + definitions + date/domain/source constraintsAsk for synthesis, disagreements and uncertaintyUse search parameters for filters; request primary evidenceFew-shot prose that pollutes search, or asking it to fabricate URLs in text

Interpretation rule: a platform feature such as browsing, connectors or memory is not a property of the base model. Always identify the surface you are using.

03 · Model-specific field notes

Learn each system’s prompt dialect.

Filter the profiles, then copy and adapt the worked prompt. The “archetypes” are mnemonic shorthand only.

OpenAI GPT

The outcome-driven operator
GPT / ChatGPT

Do

  • Use lean, non-duplicative instructions.
  • Define autonomy, evidence, completion and output format.
  • Separate durable behavior from the current task.

Avoid

  • Demanding private chain-of-thought.
  • Maxing reasoning effort without an evaluation.
  • Conflicting prose and examples.

Worked prompt · strategic review

Goal: Evaluate whether our AI music catalogue should prioritize hotel groups or independent venues in Latin America. Use: the attached performance data and current, cited market evidence. Evaluate: addressable reach, sales-cycle length, integration effort, margin, retention and reputational risk. Deliver: 1. Executive recommendation in ≤120 words. 2. Weighted decision matrix; show weights and scores. 3. Three strongest counterarguments. 4. A 90-day validation plan with owners and measurable gates. Rules: - Distinguish evidence, inference and assumption. - Do not invent missing figures. Mark them “unknown.” - Consider the task complete only when every criterion is addressed.

Anthropic Claude

The careful constitutional editor
Claude

Do

  • Use descriptive XML tags for role, context, documents and criteria.
  • Tell Claude what to do, not only what to avoid.
  • Use representative examples when style or classification boundaries matter.

Avoid

  • Assuming it will infer an unspoken executive standard.
  • Overloaded role-play that adds no domain criteria.
  • Vague requests for “professional” prose.

Worked prompt · document analysis

<role>You are a senior research editor. Preserve nuance and challenge unsupported claims.</role> <task>Turn the source report into a board-level analytical brief.</task> <source>{{REPORT}}</source> <method> - Extract the central claim and causal chain. - Test each major claim against evidence inside the source. - Separate fact, author interpretation and your inference. - Identify omissions that could reverse the recommendation. </method> <output> Return: thesis; evidence map; five material weaknesses; repaired recommendation; questions for management. Use precise prose. No generic conclusion. </output>

Google Gemini

The multimodal synthesizer
Gemini

Do

  • Use direct, concise instructions for current reasoning models.
  • Put the question after very large context blocks.
  • Name which image, audio, video or document evidence supports each finding.

Avoid

  • Verbose legacy “prompt magic.”
  • Unclear references such as “this” across many media items.
  • Expecting current facts without Search grounding.

Worked prompt · multimodal audit

The preceding files contain: A. mobile homepage screenshots, B. desktop homepage screenshots, C. Lighthouse results, D. the brand guide. Based only on that evidence, audit the page for: - visual hierarchy, - brand adherence, - mobile conversion friction, - accessibility, - Core Web Vitals risk. For every finding, cite the file label and visible or measured evidence. Rank the top 10 issues by user impact × confidence. End with a mobile-first revised page outline. Do not infer implementation details that are not observable.

xAI Grok

The live-signal investigator
Grok

Do

  • Specify whether live web/X signals are required.
  • Define the observation window and source-quality hierarchy.
  • Request corroboration outside social posts for factual claims.

Avoid

  • Confusing popularity with reliability.
  • Leaving “recent” undefined.
  • Accepting social sentiment as market size.

Worked prompt · real-time narrative scan

Investigate how the launch is being discussed in the last 72 hours. Search scope: official company statements, credible technology press, domain experts, and relevant public X discussion. Return: 1. Verified event timeline. 2. Five dominant narratives with representative evidence. 3. Claims that are viral but unverified or contradicted. 4. Differences between expert and general-user reactions. 5. What evidence in the next 30 days would confirm or falsify each narrative. Use exact dates. Label source type. Do not treat repost volume as independent corroboration.

Meta Llama

The configurable instrument
Open weights

Do

  • Use the official tokenizer/chat template for the exact checkpoint.
  • Test system-message adherence on the chosen host and quantization.
  • Keep demonstrations close to the production distribution.

Avoid

  • Treating “Llama” as one stable product.
  • Hand-inserting special tokens.
  • Importing a prompt tuned on a proprietary model without evaluation.

Worked prompt · local classifier

SYSTEM You classify support requests. Follow the label definitions exactly. Output one JSON object and no commentary. LABELS billing: charges, invoices, refunds access: login, credentials, permissions bug: reproducible product malfunction how_to: request for operating instructions other: none of the above SCHEMA {"label":"billing|access|bug|how_to|other","confidence":0.0,"evidence":"short quote"} USER {{TICKET}} [Apply this content through the deployment’s official chat template; do not manually copy control tokens.]

Mistral

The compact technical specialist
Mistral / Le Chat

Do

  • Put stable role, values and format guidance in the system prompt.
  • Use demonstrations for language, label and format adherence.
  • Pair prompts with sampling and model-selection tests.

Avoid

  • Using prompt text as the only safety layer.
  • Ignoring model-size differences.
  • Changing prompt and sampling parameters simultaneously during tests.

Worked prompt · multilingual extraction

System: You are an information-extraction engine. Preserve the source language. Never translate names or quoted wording. Return valid JSON only. Task: Extract commitments from the message. JSON schema: { "language": "ISO-639-1", "commitments": [{ "actor": "string", "action": "string", "deadline": "ISO date or null", "certainty": "explicit|inferred", "evidence": "verbatim short excerpt" }] } If no commitment exists, return an empty commitments array. Message: {{MESSAGE}}

DeepSeek

The mode-sensitive reasoner
DeepSeek

Do

  • Select thinking/tool/JSON mode deliberately.
  • For JSON mode, request JSON, show the shape and allow enough tokens.
  • Preserve required reasoning state in tool-call continuations.

Avoid

  • Parsing internal reasoning as the final answer.
  • Requesting JSON only through prose when an API mode exists.
  • Dropping required state between tool turns.

Worked prompt · strict JSON analysis

System: Analyze the contract excerpt and return a JSON object matching the example. Do not add Markdown. Example JSON: {"risk":"high","clause":"7.2","reason":"Uncapped liability","missing_information":[]} Definitions: - high: could create uncapped loss, illegality or loss of core IP - medium: material cost or operational constraint - low: limited, reversible impact Analyze: {{CLAUSE}} Return keys: risk, clause, reason, missing_information. [API: enable JSON output mode and set sufficient max_tokens.]

Microsoft Copilot

The context-bound colleague
Microsoft 365

Do

  • Use Microsoft’s goal–context–source–expectations pattern.
  • Name the document, meeting, email thread or table.
  • Specify the business action the result must support.

Avoid

  • Assuming it can access content you cannot access.
  • Unbounded “summarize everything” prompts.
  • Failing to name the target audience and decision.

Worked prompt · meeting follow-through

Goal: Create the CEO follow-up from today’s “LATAM Venue Distribution” Teams meeting. Context: We must decide whether to run a 90-day pilot in Asunción, Montevideo or both. Sources: Use the meeting transcript, the attached pilot budget and the email thread titled “Venue partner shortlist.” Do not use unrelated tenant content. Expectations: - 150-word decision brief; - decision points and unresolved disagreements; - action table with owner, deadline and dependency; - flag any action not explicitly agreed; - draft a concise follow-up email in my direct executive tone.

Perplexity

The retrieval-first research analyst
Perplexity / Sonar

Do

  • Use discriminating search vocabulary and define ambiguous terms.
  • Constrain dates, geography, domains and source type with parameters where possible.
  • Ask for primary sources, disagreements and uncertainty.

Avoid

  • Long fictional examples that distort retrieval.
  • Asking the model to emit invented URL lists.
  • Equating a citation with proof that the cited page supports the claim.

Worked prompt · evidence review

Research whether AI-generated background music measurably affects dwell time or spending in hotels, restaurants and retail venues. Scope: 2018–2026; hospitality and physical retail; exclude purely online advertising. Prioritize: peer-reviewed studies, field experiments, systematic reviews and primary industry datasets. Use vendor claims only as clearly labeled secondary evidence. For each relevant study report: sample, setting, intervention, outcome, effect size if available, limitations and DOI/source. Synthesize where findings converge or conflict. Conclude with what is known, what remains uncertain, and a defensible pilot design. Do not claim causation from correlational evidence.
04 · Prompt surgery

Upgrade vague requests into testable briefs.

The improvement does not come from ornamental wording. It comes from reducing ambiguity and making quality observable.

Weak
“Analyze our market and make a professional report.”
Stronger, cross-model
“Estimate the serviceable market for AI-curated music in four-star and five-star hotels in Paraguay and Uruguay, 2026–2029. Use current primary sources; state definitions and assumptions; provide low/base/high scenarios; separate verified facts from inference; end with five validation interviews and a go/no-go threshold.”
Weak
“Think step by step and tell me the right strategy.”
Stronger for reasoning models
“Recommend one strategy under a US$50,000 test budget. Evaluate it against reach, speed to evidence, gross-margin potential and reversibility. Show the decision matrix and the evidence supporting each score; then identify the fact most likely to change your recommendation.”
Weak
“You are the world’s greatest lawyer. Review this.”
Stronger domain specification
“Review the attached agreement as counsel to the supplier. Identify clauses affecting liability, IP ownership, termination, payment, exclusivity and governing law. Quote clause numbers, explain commercial impact, propose replacement language and flag questions requiring locally qualified counsel. Do not imply this is a formal legal opinion.”
Weak
“Search the web deeply and give reliable links.”
Stronger for search-grounded systems
“Find primary evidence published since January 2024. Prefer regulator filings, official datasets and peer-reviewed work. For each claim, cite the source that directly supports it, publication date and evidence type. Exclude SEO summaries unless they lead to a primary source. Report material disagreement and failed searches.”
05 · Production method

Prompting is experimental design.

A prompt that “worked once” is an anecdote. A production prompt is a versioned hypothesis tested against representative cases.

Define success

Create a rubric and a small set of real, difficult, edge and adversarial inputs before polishing the prompt.

Establish baseline

Run the simplest direct prompt first. Record model snapshot, parameters, tools, latency, cost and failures.

Change one variable

Add context, examples, structure or reasoning effort separately so you know what produced the change.

Compare blindly

Where stakes justify it, hide model and prompt identity from human evaluators; use deterministic checks for schemas.

Stress test

Test ambiguity, missing data, long context, prompt injection, multilingual input and distribution shift.

Monitor drift

Pin model versions where possible. Re-run the evaluation set when a model, tool, retrieval layer or system prompt changes.

06 · Interactive lab

Generate a disciplined first draft.

This local builder assembles a portable prompt. It is a starting hypothesis, not proof of quality.

Prompt draft

07 · Quality rubric

Score outputs, not eloquence.

Use a five-part rubric. Fluency without groundedness can make a weak answer look deceptively strong.

20

Correctness

Claims, calculations and interpretations withstand checking.

20

Groundedness

Evidence supports the adjacent claim; uncertainty is visible.

20

Completeness

All required criteria and edge cases are addressed.

20

Compliance

Format, scope, policy and tool boundaries are followed.

20

Utility

The output enables the intended decision or action.

08 · Evidence base

Research and first-party guidance.

Provider documentation is authoritative for current product behavior; peer-reviewed and preprint research supports general techniques. Neither substitutes for evaluations on your own tasks.

01
OpenAI — Model guidance and current prompting best practicesLeaner prompts, autonomy boundaries, verbosity, reasoning and tools.
02
OpenAI — Reasoning best practicesDirect prompts, delimiters, zero-shot first; avoid chain-of-thought instructions.
03
Anthropic — Prompting best practicesClarity, examples, XML, output control, tools and agentic systems.
04
Google — Gemini prompt design strategiesClear instructions, examples, context, structure and iteration.
05
Google — Gemini 3 developer guideConcise reasoning prompts and instructions after large context.
06
xAI — Grok developer guidanceCurrent model capabilities, tools and agentic workflows.
07
Meta — Llama prompt engineeringPrompt construction and deployment-specific formatting.
08
Mistral — Prompting best practicesSystem prompting, instructions, demonstrations and output formatting.
09
DeepSeek — JSON output guideAPI mode, explicit JSON instruction, example shape and token budget.
10
DeepSeek — Thinking modeReasoning state and multi-turn tool-call handling.
11
Microsoft — Effective Copilot promptsGoal, context, source and expectations.
12
Perplexity — Agent API prompt guideRetrieval specificity, filters and evidence-oriented queries.
13
Schulhoff et al. — The Prompt ReportA systematic taxonomy of 58 text prompting techniques and 40 multimodal techniques.
14
Sahoo et al. — A Systematic Survey of Prompt Engineering in LLMsTechniques, applications, strengths and limitations.
15
Wei et al. — Chain-of-Thought PromptingFoundational evidence for intermediate reasoning exemplars on older large models.
16
Wang et al. — Self-ConsistencySampling multiple reasoning paths improved benchmark accuracy, with extra inference cost.
Scholarly caution. Prompt techniques are conditional interventions, not universal laws. Chain-of-thought results from earlier model generations do not imply that asking every modern reasoning model to expose step-by-step reasoning is beneficial. Formatting sensitivity, model updates and task distribution can reverse a result. The scientifically defensible practice is controlled, task-specific evaluation.