How to Write an Instruction Someone Can Actually Follow

Why experts are the worst judges of their own instructions — and four mechanical tests that find what they can’t see.

Most instructions fail before anyone follows them. The problem isn’t the user. It’s that experts forget what it feels like not to know.

The instruction said: Review the call detail for anything unusual.
The customer had been losing calls for nine days. The agent pulled the call detail — a grid of timestamps, durations, and disconnect codes — and looked at it. Nothing jumped out. He checked for outages in her area. There weren’t any. He told her everything looked healthy on our end, walked her through a power cycle, and closed the ticket.
She called back four more times.
The pattern had been in the record the whole time. Nearly every call was being dropped by the network rather than by either person on it, and nearly all of them died in the same place — as her phone moved between the same two sites. To someone who had read a few thousand of those records, the shape was unmistakable. To him it was a column of numbers.
He was not careless. He read it. He read every word.
He reviewed the call detail for anything unusual, and nothing was unusual to him, because unusual is not something a person can look for. It’s something you can only recognize if you already know what usual looks like — and the man who wrote that step already knew.
The failure wasn’t in the sentence. It was in one word of it.

Two people in incompatible conditions


Here is what was true when that step was written. The writer had the whole system loaded. He knew what a normal disconnect distribution looked like because he had seen thousands of them. He knew where the step was going and why it mattered. Nothing was on fire. He had time.
Here is what was true when it was read. The agent had a fragment. He had a customer who had already explained this four times and was audibly finished explaining it. He had a handle-time metric he could feel. He had never seen a call record with a known pattern in it, so he had no baseline to measure against, and he had about forty seconds of working attention before the silence on the line became a second problem.
Those two states are not different levels of skill. They are different jobs. Clarity is not a property of prose — it is a relationship between a text and the condition of the person reading it, and the second half of that relationship is invisible from where you’re sitting.
Some readers really don’t read. That’s true and it isn’t interesting. What’s interesting is how much of what looks like not reading turns out to be a document quietly requiring an ideal reader — someone unhurried, sequential, already holding the context — and then failing everyone else. The useful question was never who’s at fault. It’s what could this document have prevented?
Which brings up the uncomfortable part. You are the worst available judge of your own instructions, because you possess exactly the knowledge the instruction is supposed to supply. You can only write the step because you know the system, and knowing the system is precisely what makes the gaps in your description of it invisible to you. That’s not a character flaw and you can’t care your way out of it.
So don’t judge them. Test them.
The tests below are mechanical on purpose. They operate on the surface of the text — words, ordering, dependencies — rather than on your sense of whether something seems clear, because your sense of whether something seems clear is the compromised instrument.
It helps to stop thinking of an instruction as information. It’s a device for operating someone else’s hands at a distance. You aren’t transferring understanding; you’re issuing motor commands to a body you can’t see, belonging to a person who doesn’t share your model of the machine. You’d never tell a hand what you intended. You’d tell it where to move.
One thing to watch for as you go, because it comes back at the end: not everything you find is a writing problem. Sometimes the missing sentence belongs in the document. Sometimes its absence is telling you something upstream is broken.

Test 1: The Two-People Test


Match the package to the channels the customer actually watches.
That was the step. It ran in a television sales guide I wrote, on a page that also carried a comparison chart of four packages and what was in each. It looks like an instruction. It’s a decision with an instruction painted on it.
An agent ran it exactly as written. She mentioned her team’s games early in the call and he made a note of it. Then he built her the third tier — more sports networks, movie channels, better documentaries, and even carrying all of that it came in under what she was already paying. He wasn’t squeezing her. He was handing her more television for less money, which is the rarest thing in that job.
The regional network that carried her team wasn’t in the third tier. It wasn’t in any tier. It was an add-on, priced separately, and nothing in the guide mentioned that some channels don’t come in packages at all.
It got caught in the recap, which was my job — terms, billing, install window, and a line-by-line check that the channels the customer asked for were on the order. Hers weren’t. So we rebuilt it. Good outcome. Consider what produced it. Not the document — a person performing by hand a check the document never specified.
Now look at where the step broke. It wasn’t the adverb. Actually is doing something, but the real weight is on match, which compresses an entire procedure into one word. To match, the agent had to work out what she valued, identify which network carried it, determine whether that network lived in a package or outside all of them, and find the combination that covered everything. Five judgments, none of them written down, all of them performed silently by whoever wrote the guide before the guide existed.
That’s the pattern. Some words don’t describe an action. They stand in for a procedure the writer already knows and the reader has to reconstruct.
They’re learnable by sight. Two families:
Words that hide a standard: appropriate · correct · relevant · applicable · best · suitable · significant · necessary · unusual
Words that hide a procedure: match · determine · identify · assess · evaluate · select · choose · review
The test. Circle those words, plus any verb asking the reader to choose, judge, compare, interpret, or decide. That part is mechanical — you’re finding words, not evaluating clarity.
Then ask the question that does the real work: could two competent people, given only this step, independently arrive at the same answer?
If yes, the step is fine. If no, you’ve found a decision you made so fast you never registered making it.
Then do one of three things with it. If there’s a right answer, put the right answer in the step; most of them resolve here. If the answer varies but the rule doesn’t, write the rule — ask which teams she follows; regional sports networks are an add-on, not part of any tier. And if the judgment genuinely belongs to the reader, stop hiding that a judgment is happening. Put it on its own line, name it, and give them the lookup that settles it.
The goal was never to eliminate judgment. It’s to stop burying it inside a word that looks like an action. A visible decision costs the reader ten seconds. A buried one costs a rebuild.

Test 2: The Search-Bar Test


You numbered them. It helped less than you think, because a numbered list is not experienced as a sequence. It’s experienced as a menu.
The dropped-call guide had a step near the bottom that let the agent push the device off the site it was registered to, so the customer could power cycle and reattach somewhere else. It was the only step in the document that did anything — everything above it was checking. So agents went straight to it. Mid-call, customer waiting, scroll, there it is.
Which would be fine, except that step quietly assumed three things had happened four steps earlier: that the disconnects were confirmed network-side, that no maintenance was flagged at the sites involved, and that this was a new condition rather than a nine-month-old one. Forcing a re-registration on a device whose real problem is its own software accomplishes nothing, except that you’ve now spent the customer’s patience on a maneuver — and the maneuver has a ceremony to it. Power it off. Wait. Power it back on. Which makes the failure feel bigger when it arrives.
The assumption wasn’t in step four. It was in the white space between three and four. That’s where these live: not inside the instructions but in the unwritten membrane between them, invisible to anyone who wrote the steps in order and invisible on purpose to everyone who didn’t read them that way.
The test. Read each step as though the reader arrived at it directly from a search result. What does it assume they already did?
Every assumption you find either gets stated at the step or the step fails for everyone entering mid-document — which happens constantly, through search, bookmarks, Ctrl-F, and a colleague pasting you one paragraph out of nine. AI makes it more important, not less: a retrieval system can lift a single step out of a larger document and hand it over without the context that made it safe.
Stating the assumption costs a clause. If the call record shows network-side disconnects and no site maintenance in the last two weeks, force a re-registration. Now the step carries its own preconditions and survives being read alone, which is how it will be read.
This article is built the same way, for the same reason.

Test 3: The Hands Test


Ensure there are no outages in the customer’s area.
Try to do that with your hands.
You can’t, because as written it names a state you’d like to arrive at and leaves the reader to invent the procedure that produces it. To the writer this reads as complete — he knows the procedure, and the sentence retrieves it instantly. To the reader it’s a destination with no route.
The verb isn’t the villain. Verify that the serial number on the device matches the one under Settings > About is perfectly executable. Ensure, verify, confirm, review, and make sure are warning lights, not forbidden words. They tell you the writer has described the desired state. Your job is to check whether he also supplied the operation that produces it.
The test: could a pair of hands perform this? If carrying it out requires already knowing what the result should look like, you’ve written a description, not an instruction.
The rewrite is unglamorous. Name the tool, the field, what to enter, what to look at.
Open the site status tool. Enter the sites listed in the call record — not the customer’s ZIP, which returns the nearest sites and not necessarily the ones handling her calls. Look at the maintenance column for the last fourteen days.
Then add the half almost nobody writes: how the reader knows it worked. You should get at least one row per site. An empty result means the site IDs didn’t take. It does not mean the network is clean.
Leave that out and a reader who fumbles the input gets a blank screen and reads it as good news. He didn’t skip the step. He completed it, and it paid him in a confident wrong answer — which is worse than a failure, because a failure at least announces itself. He tells the customer everything looks healthy on our end. She calls back four more times.

Test 4: Find the Decay Points


The dangerous document isn’t the wrong one. It’s the one that went wrong eleven months ago and still looks right.
The channel chart was accurate the day it shipped. Then a carriage agreement lapsed, a block of channels came off the lineup while two companies renegotiated, and the people inside the building knew within a day or two, the way you know about weather. It arrived as an announcement, which is to say it reached whoever was paying attention that morning.
The document knew nothing. It went on saying what it had always said, in the same confident chart, to every agent who opened it cold — and to everyone working from a printout, which was most of them, because a printout doesn’t make you wait for a page to load while a customer is talking.
Instructions aren’t right or wrong. They’re right for now, and the parts that go first are always the ones leaning on something you don’t control.
A decay point is anything in a document whose truth depends on something outside your authority: a menu label someone can rename, a price, a policy, an API, a lineup someone else negotiates, a tool another team owns, a link, a law.
The test. Mark them. Then build them to break loudly instead of quietly — source, date, and a way to check. Lineup current as of [date]; confirm in the live tool before quoting. When the two disagree, the disagreement registers as a disagreement rather than as the reader’s own error. You haven’t made the document permanent. You’ve made it capable of admitting when it isn’t, which is the only durable version of the same thing.

Learning mode is not doing mode


Before any of this I was the one taking the calls, and what I had was a wall of dry text.
Part of the job was sending signals to a customer’s phone — one that forced a software update onto the device, another that pulled it off the network so it would come back somewhere else. I could tell you which buttons produced which, because that’s what the documentation said. What it never said, anywhere, was what those signals were doing or why either would fix anything.
It said: press this, it will fix it. If it doesn’t fix it, do the other thing.
The other thing was open a case.
So I’d press it. She’d hang up and go try it. And when it didn’t work she called back — not to me, to whoever picked up — and that person opened the account and found a note, and the quality of that note decided whether she was about to repeat the last twenty minutes of her life. That’s the argument for writing any of this down, and I learned it from the receiving end. Not so the next person knows the procedure. So the next person knows what’s already been ruled out.
None of which told me why the signal worked. I still don’t entirely know.
So the guides I eventually wrote had two halves. The first taught: here’s the call type, here’s what causes it, here’s what the customer is describing when she says the thing she says. That half was allowed to explain, because a person reading to learn wants causes. The second was reference: trigger, steps, expected result, exceptions, and very little else — because a person reading mid-task wants speed, and every sentence of explanation is a tax.
Learning mode wants models, causes, and why. Doing mode wants trigger, steps, expected result, and what to do when it isn’t that. Most documentation tries to make one layer do both and fails at both — too slow to use, too thin to learn from. In sales and in documentation, the translation is the whole job: converting what you know into something another person can act on. But you have to know which moment you’re writing for.

What a good document is hiding


Every excellent instruction absorbs a defect that should have been fixed upstream.
The handoff procedure existed because the disconnect data was legible only to people who’d already memorized what normal looked like. The recap existed because the ordering tool would let an agent submit a package that didn’t contain what the customer had spent ten minutes describing, and nothing in the workflow objected. The re-registration ceremony existed because there was no way to see, from where we sat, what was happening one layer down.
Write the instruction well and all of that disappears. The agent succeeds. The customer is satisfied. The metric moves. And the defect stays exactly where it was, because pain is the signal that travels upward in an organization, and you have just absorbed all of it.
Which is the question underneath the article: at what point does excellent documentation stop exposing a broken system and start making it comfortable enough to stay broken?
The useful move is to treat the writing itself as a source of evidence. Every warning you had to include is a place the system is wrong. Every buried decision you surfaced is a decision the interface should have made. Every precondition you stated is a step the workflow left ambiguous. Every dated caveat is something that changes without telling anyone. Every recap is a check the tool declined to run.
Done properly, you finish with a list — an actual one, produced as a byproduct of the work, of every place a person has to compensate for a system that should have handled it itself.
Nobody is going to ask you for that list. The documentation records the workaround. The list records why the workaround was necessary. Only one of those ever reaches anyone who could make it unnecessary.

The Instruction Audit


Before you publish anything someone else has to follow:

  1. The Two-People Test Could two competent people, given only this step, reach the same answer? If not, you’ve hidden a decision inside a word.
  2. The Search-Bar Test If someone landed here from a search result, what does this step assume they already did?
  3. The Hands Test Could a pair of hands perform this? And will the reader know it worked?
  4. Decay Points What in here depends on something you don’t control? Source it, date it, give it a way to fail loudly.
  5. Learning or doing? Are you explaining, or helping someone act right now? Pick one per layer.
  6. The log What did you have to work around — and what does that tell you about the thing upstream?
    None of this is specific to support documentation. It applies to onboarding guides, recipes, checklists, tutorials, SOPs, the handoff you write before vacation, the note you leave yourself about how you finally fixed the printer, and the prompt you give a model that doesn’t possess the context you forgot you were carrying.

The step should have said: the call record sorts disconnects by who ended the call; look at the network-side rows and check whether they cluster at the same pair of sites.
A sentence longer. Four callbacks shorter.
She would have had an answer the first time she called.

Similar Posts