ImageFirm / FIELD MANUAL 01
Field Manual 01 · Operating Claude

The model is not the bottleneck. The brief is.

Most guides hand you twelve tricks and call it leverage. Tricks produce better single answers. Leverage produces assets that keep paying — and that difference is a system, not a tip.

This manual is the ladder from asking to operating: five rungs, each with the moves that matter, the failure modes that follow, and the review gate you do not remove to go faster. Read it in twenty minutes. Use the kit at the end today.

5Rungs, not 12 tips
9Named failure modes
3Templates to copy
1Test everything passes
Rewrite  / request → brief
Give me 10 content ideas for my brand.
+You are planning a month of content for a studio that
+sells calm, not urgency. Audience: venue owners who
+choose music by feel, not by genre tags.
+
+Produce 10 concepts. Each: a hook line, the tension it
+resolves, the format, and the one metric it should move.
+
+Constraints: no listicles, no fear framing, nothing that
+needs a face on camera. Two must work without sound.
+
+Bar: I should want to make 3 of them immediately. Mark
+any concept you think is filler and say why.
Four defects removed
  1. No audience → the answer was written for nobody.
  2. No output contract → shape had to be guessed.
  3. No exclusions → the obvious answer wins by default.
  4. No success bar → nothing to judge the output against.
01 — THE ESCALATION LADDER

Five rungs. Each one changes what you're doing, not just how well you do it.

The rungs are ordered because they depend on each other. You cannot systematise a brief you have not yet written well, and you cannot safely delegate a task you cannot yet verify. Autonomy rises as you climb — and so does the number of gates you owe the work. That is the whole discipline in one sentence.

R1

Ask — write a brief, not a query

You are still asking for one answer. The only variable you control is the quality of the request, and it is worth more than any other single lever on this page.

Model autonomy
Gates you owe
1.1

Six slots, every time

Role, context, task, constraints, output contract, success bar. A request missing any of the six is asking the model to invent your intent — and it will, plausibly and wrongly. The six-slot habit is the single highest-yield change most people can make.

1.2

Show three, don't rule twenty

Two or three diverse, canonical examples of what good looks like beat a long list of rules. Rules describe the boundary; examples describe the target. Pick examples that differ from each other — near-identical samples teach the shape of one case, not the shape of the job.

1.3

Name the exclusions

The most-skipped, highest-yield line in any brief: what a good answer must not be. Without it, you get the statistically obvious answer — which is exactly the answer your competitors are also getting.

1.4

Set a bar you can judge against

"Make it good" is unjudgeable. "A skeptical operations lead should be able to act on this without asking a follow-up question" is testable — by you, in ten seconds, on every draft.

R2

Direct — steer the reasoning, not just the wording

You stop rewriting the request and start shaping how the work is done: what gets thought through first, what format locks, and who argues the other side.

Model autonomy
Gates you owe
2.1

Reasoning first, conclusion last

On any task with a judgement in it, ask for the working before the verdict. A conclusion stated first tends to get defended; a conclusion reached last tends to get earned. On genuinely hard problems, turn on extended thinking rather than simulating it in the prompt.

2.2

Tag your inputs

When you paste source material, fence it and say what it is. Instructions and data that look alike get treated alike. Delimiters cost nothing and prevent the most common long-input failure: your directions read as part of the document.

Input fencing
<source_material>
  [paste the transcript / draft / data here]
</source_material>

<instructions>
  Work only from source_material. Where it is silent, say
  "not stated" rather than filling the gap.
</instructions>
2.3

Ask for the strongest counter-case — not for brutality

"Critique this brutally" optimises for the tone of criticism. You get harshness, which feels like rigour and often isn't. Name a standard and an adversary instead, and you get findings you can act on.

Adversarial review
Review this against one standard: would a skeptical buyer
who has seen four competing proposals this month act on it?

Give me:
1. The three weakest claims, and what evidence each needs.
2. The strongest argument someone would make against it.
3. What I have assumed without stating.
4. What to cut. Be specific about which lines.

Do not soften. Do not perform harshness either — I want
findings, not adjectives.
R3

Structure — the context is the product

The shift that separates power users from everyone else: you stop optimising sentences and start curating what is in the window at all. Anthropic's engineering team calls this context engineering — the natural successor to prompt engineering.

Model autonomy
Gates you owe
3.1

Attention is a budget

Every token you add spends from a finite attention budget, and recall degrades as the window fills — a gradient, not a cliff, but a real one. The goal is the smallest set of high-signal tokens that reliably produces the outcome. Minimal does not mean short; it means nothing in the window is dead weight.

3.2

Write at the right altitude

Two failure modes bracket every set of standing instructions. Too low: brittle if-this-then-that rules that break the first time reality varies. Too high: vibes that assume context you never supplied. Aim for strong heuristics with concrete signals — specific enough to steer, loose enough to survive a new case.

3.3

Chain instead of cramming

One job per call. Research → outline → draft → adversarial review → final, with the artifact of each stage passed forward. A single mega-prompt asking for all five produces five mediocre stages, because the model is optimising a compromise across incompatible goals.

3.4

Put durable context where it reloads itself

Your voice spec, house rules, non-negotiables and canonical examples do not belong in the chat you happen to be in. They belong in the layer that reloads on every session — a project's instructions and knowledge, a skill, a committed instructions file. Re-pasting is a symptom; it means the asset lives in your head instead of in the system.

3.5

Start a clean thread more often than feels natural

A long thread accumulates dead ends, discarded drafts and stale corrections — all of which still compete for attention. When quality starts sliding mid-project, the fix is usually not a better prompt. It is a fresh window seeded with a deliberate summary: decisions made, constraints still binding, current state, next step. This is compaction, done by hand, and it is one of the most reliable quality recoveries available to a non-technical user.

R4

Systematise — produce assets, not answers

Compounding starts here. An answer is consumed once; an asset — a template, a rule, an eval, a connected source — pays out on every future run. Most people never reach this rung, which is precisely why it is where the advantage is.

Model autonomy
Gates you owe
4.1

Every correction becomes a rule

The second time you give the same note, stop giving the note and write it into the standing instructions. Corrections repeated by hand are unpaid labour; corrections written down are infrastructure. This is the mechanism by which a system gets better while you sleep.

4.2

Evals before scale — the real dividing line

Before you run a prompt a hundred times, run it against twenty real cases you already know the right answer to, score them against explicit criteria, and set a pass bar. Anthropic's own guidance puts success criteria and evaluations before prompt engineering for a reason: without them you are tuning on vibes and calling drift improvement.

4.3

Connect the source instead of pasting it

Pasting is a manual, stale, error-prone copy of the truth. Tools and connectors let the model fetch what it needs, when it needs it — the current file, the live calendar, the actual repository. Keep the tool set small and unambiguous: if you can't say which tool should be used in a given situation, neither can the model.

4.4

Retrieve just in time

Don't front-load everything that might be relevant. Keep lightweight pointers — file paths, links, queries — and pull the content in at the moment it's needed. The structure itself carries signal: a folder name, a naming convention and a timestamp tell you what a file is for before you open it.

R5

Delegate — loops, with the brakes wired in

An agent is a model using tools in a loop. The loop is what creates the leverage and the exposure at the same time, so autonomy is earned by verification, never granted by enthusiasm.

Model autonomy
Gates you owe
5.1

Delegate the loop, not the judgement

Hand over the repetitive traversal — search, gather, transform, check, repeat. Keep the calls that set direction, accept risk, or commit the organisation to something. That boundary is not a limitation of the technology; it is where accountability lives.

5.2

Three techniques for long-horizon work

Compaction — summarise and restart before the window degrades. Structured notes — have the run write its state to a file it can re-read after a reset. Sub-agents — send a clean context to explore deeply and return a condensed result, so the main thread stays focused on synthesis.

5.3

Size the blast radius before the capability

Ask what the worst plausible wrong action costs, and whether it is reversible. Sandbox first. Read before write. Draft before send. Staging before production. An approval gate on irreversible actions is not friction — it is the thing that lets you raise autonomy everywhere else.

5.4

Instrument the run

A loop you cannot inspect is a loop you cannot trust. Log what it did, what it fetched, what it decided and where it stopped. When something goes wrong you want a trace, not a theory.

02 — THE VERIFICATION GATE

The poster never mentions checking. That omission is the whole risk.

Fluent output is not verified output, and fluency is exactly what makes unverified output dangerous — it reads finished. Before anything leaves your hands, run the gate. It takes under a minute and it is the difference between using a tool and being liable for one.


Brief integrity check

Tick what your brief actually contains. Nothing is stored or sent anywhere — this runs in your browser and forgets when you close the tab.

0/8
Not a brief yetA request with none of these is asking the model to invent your intent. It will — plausibly, and not the way you meant.
03 — FAILURE LEDGER

Nine symptoms, and what is actually causing each one.

Most people respond to a bad output by rephrasing. Rephrasing fixes roughly one of these nine. Diagnose first: the symptom tells you which rung you're missing.

SymptomActual causeFix
Generic output No exclusions and no audience — the statistically obvious answer is the only target available. R1.3 Name what it must not be. Name who reads it.
Confidently wrong details The brief demanded facts the model had no way to verify, and left no legal way to say so. R4.3 + gate Connect the source, or permit "unknown" explicitly.
Drifts off your voice The voice spec lives in your head. Each session starts from nothing and reverts to house-neutral. R3.4 Move the spec and three exemplars into the layer that reloads.
Quality decays in a long chat Accumulated dead ends and superseded drafts still compete for a finite attention budget. R3.5 Compact: new thread, seeded with decisions and current state.
Ignores half the instructions Too many rules at once, or instructions buried inside pasted material and read as content. R2.2 + R3.3 Fence the inputs. Split into stages, one job each.
Great draft, wrong direction Execution began before the framing was agreed. Speed spent in the wrong direction. R3.3 Approve the outline as its own stage before any drafting.
Agreeable to a fault You asked for critique inside the thread that produced the work, so it inherits every assumption. R2.3 Adversarial review, clean context, named standard.
Worked once, unreliable at volume Tuned on a handful of easy cases. No pass bar, no held-back set, so drift reads as improvement. R4.2 Twenty real cases, explicit criteria, a bar it has to clear.
Agent loops or overreaches Ambiguous or overlapping tools, no stop condition, no gate on irreversible actions. R5.3 Prune the tool set. Define done. Gate anything you can't undo.
04 — CORRECTIONS

Where the beginner poster is wrong, and why it matters.

Several of its twelve tips are sound — build a conversation, be specific, use it to think rather than to type. Four are not, and each one fails in a way that costs the user something real.

"Use Claude's memory to your advantage — it remembers your goals, brand voice, audience, preferences."

Persistence is a feature you configure, not an ambient property. Inside one conversation, "memory" is just the context you supplied. Across conversations, it depends on which product you're in and what you've set up — projects, memory settings, an instructions file.

Assuming it remembers is how briefs quietly get thinner until quality drops and nobody can say when. Put durable context in a layer that reloads by design, then verify it is actually loading.

"Act as a therapist helping me improve discipline."

Role prompting is a real technique — it selects a perspective and a vocabulary. It does not confer clinical judgement, duty of care, or the ability to notice what's going wrong with a specific person over time.

Fine: act as a strategist and pressure-test this plan. Not fine: outsourcing care to a role label. Use roles to select expertise, not to replace a relationship with someone accountable to you.

"Critique this brutally and improve it."

This optimises for the tone of criticism rather than its accuracy. You get harshness, which feels like rigour — and harshness is cheap to generate about anything, including work that is fine.

Name a standard and an adversary instead. "Would a skeptical buyer who has seen four competing proposals act on this?" produces findings you can act on. "Be brutal" produces adjectives.

"Claude won't replace you. But someone using Claude will."

Fear is excellent at producing saves and shares, and poor at producing capability. It also gets the mechanism wrong: the gap between operators is not access to the tool — nearly everyone has that now.

The gap is systems and verification. Whether corrections become rules, whether outputs become assets, whether anything is checked before it ships. Twelve disconnected tips cannot close that gap, however well they perform on a feed.

05 — GOVERNANCE AS AN OPERATING LAYER

Five principles, each with a test you can apply this afternoon.

Ethics stated as values is a poster. Ethics stated as gates is an operating system. Each principle below carries one concrete test — if the test fails, the work does not ship, regardless of how good the output looks.

Principle 01

Transparency is mandatory

Anyone affected by the output can find out how it was made. Disclosure is a design decision made before publication, not a defence prepared after.

Test: could I state how this was produced, to the person it affects, without editing the sentence?

Principle 02

Humans retain final authority

A named person signs off on anything that commits money, reputation, or someone else's time. Autonomy scales inside that boundary, never through it.

Test: who is accountable if this is wrong — by name, not by role?

Principle 03

Expand capability, don't create dependency

The system should leave the team more able, not less. If removing the tool would leave someone unable to do work they could previously do, that is not leverage.

Test: after six months of this, is the team stronger or only faster?

Principle 04

Own the consequences

"The model wrote it" is not a position. Everything that ships under your name carries your judgement, including what you chose not to check.

Test: am I prepared to defend this line in a room, without mentioning how it was drafted?

Principle 05

Protect sovereignty

Data, IP and decision rights stay under your control. Convenience is not a reason to hand any of the three to a system you cannot audit or leave.

Test: if this vendor disappeared on Friday, what of ours goes with it?

The overriding test

Does this make us more human?

Every rung on this ladder is in service of that question. A system that produces more output while producing less judgement has failed, no matter what the metrics say.

If the answer is no, the leverage isn't leverage. It's drift.

06 — STARTER KIT

Three artifacts. Copy them, fill the brackets, start compounding.

The first is for a single task. The second is the standing layer that makes every future task cheaper. The third is the one almost nobody builds, which is exactly why it's the advantage.

Kit 01 · Rung 1

The operator brief

Six slots. Use it until you stop needing it.

Task brief
ROLE
You are [expertise], working for [who], who cares about
[what they optimise for].

CONTEXT
Situation: [what is true right now]
Stakes: [what happens if this is right / wrong]
Prior: [what has already been tried or decided]

TASK
Produce [exact artifact], for [audience], to be used
[where and how].

CONSTRAINTS
Must: [format, length, medium, non-negotiables]
Must not: [the obvious answer you refuse to accept]

OUTPUT CONTRACT
[structure, section by section]

SUCCESS BAR
[a test I can apply in ten seconds]

If anything above is missing or contradictory, ask before
you start. If a fact isn't available to you, write
"not stated" rather than filling the gap.
Kit 02 · Rung 3

The standing layer

Written at the right altitude: heuristics with concrete signals, not brittle rules.

Project / system instructions
<background>
[What we do, who we serve, what we refuse to be.
 Keep to what changes a decision. Cut the rest.]
</background>

<voice>
Sounds like: [3 adjectives + 1 sentence of anti-pattern]
Never: [tics, clichés, formats we don't use]
Two exemplars of good live below — match their judgement,
not their wording.
</voice>

<standing_rules>
- [Rule earned from a correction you gave twice]
- [Rule earned from a correction you gave twice]
</standing_rules>

<output_defaults>
Lead with the answer. No preamble. Flag uncertainty
inline rather than hedging the whole response.
</output_defaults>

<exemplars>
[Two diverse samples of work that cleared the bar]
</exemplars>
Kit 03 · Rung 4

The eval harness

The rung most people skip. Run this before you scale anything.

Grade before scale
Here are 20 real cases and the outputs my prompt produced.

<criteria>
1. [criterion] — fail if [specific observable]
2. [criterion] — fail if [specific observable]
3. [criterion] — fail if [specific observable]
</criteria>

For each case: score each criterion pass/fail with a
one-line reason. No overall impressions.

Then:
- Which criterion fails most, and what pattern connects
  the failures?
- What single change to the prompt would fix the most
  failures without breaking the passes?
- Which cases are ambiguous because my criteria are
  underspecified? Name them — that's my problem, not
  the prompt's.

Pass bar: [n]/20 with zero failures on criterion [x].
Recommended next move

Don't adopt the ladder. Climb one rung.

Claude won't replace you, and neither will someone using Claude. Systems replace improvisation — including yours.

  • This week: take the task you repeat most and write it as an operator brief. Run it twice — once as you always have, once with the brief. Keep the better one as your exemplar.
  • This month: move your standing rules out of your head and into the layer that reloads. Every correction you've given twice becomes a line in it.
  • This quarter: build one eval — twenty real cases, three criteria, a pass bar. It is the only thing on this page that tells you the truth about whether any of the rest is working.

Then apply the overriding test. If the system produces more output and less judgement, take it apart and build the other one.