⏱️ Reading time: 11 min

Terence Eden paid a handful of strangers 25 euros an hour to tear apart the README of his project, ActivityBot, in front of his camera. The experiment, published on his blog on October 11, 2026 and discussed the same day on Hacker News, is a real case of documentation testing with real users: the method that exposes exactly what an author can’t see in their own text.

📑 En este artículo
  1. TL;DR
  2. What is documentation testing with real users?
  3. Why it matters
  4. How README testing with volunteers works
  5. Practical examples: how to apply this method to your own project
  6. Real use cases
  7. Common mistakes and author biases
  8. Comparison with alternatives
  9. Going deeper
  10. Frequently Asked Questions
    1. What’s the difference between README testing with real users and review by teammates?
    2. How much does it cost to run usability tests on documentation with paid volunteers?
    3. Can simulating users with an LLM replace real README testing?
    4. How many technical documentation validation sessions are needed to find the main problems?
    5. Which author biases are hardest to detect without documentation testing?
  11. References

Each session lasted an hour, and in total he spent around 150 euros, meaning roughly six hours of conversation recorded by hand. The results weren’t trivial: he found a broken link, an entire section nobody understood, and jokes that only confused first-time readers.

TL;DR

  • Paying real users 25 euros an hour to follow a README uncovers errors the author can never catch alone.
  • The think-aloud protocol during a video call reveals the exact sentence where each reader gets stuck.
  • Updating the documentation after each session and repeating with the next person turns testing into an iterative cycle.
  • Five users are enough to find most usability problems, according to Jakob Nielsen’s 1993 research.
  • Author biases, like shared vocabulary or obvious steps, are the real reason a README fails.

What is documentation testing with real users?

Documentation testing with real users is a validation practice in which people outside the project follow written instructions (a README, an installation guide, a manual) while thinking out loud, so the author can see exactly where they get confused, give up, or misread each step.

Unlike a unit test, which checks code, these documentation tests check human comprehension of the text. The goal isn’t to fix grammar: it’s to spot the exact moment a real reader loses the thread.

Why it matters

Every author suffers from the curse of knowledge: once you understand how your own project works, it’s almost impossible to remember what it looks like from scratch. Eden admitted this on his own blog: he knew certain commands needed sudo, that -foo was actually --foo, and that, obviously, you had to restart afterward. Those certainties were never written down, because to him they were self-evident.

Skipping documentation testing isn’t a minor oversight: it’s the reason so many open source projects have a README that’s perfect for someone who already uses them and useless for someone just arriving.

At GOV.UK, the UK government’s content service where Eden worked as a technical writer, every piece of text went through a second editor before publishing. That second pair of eyes caught errors no spell checker finds and, more importantly, could hear the frustration in a reader’s voice. Still, a teammate shares most of the author’s vocabulary; an outside user doesn’t.

How README testing with volunteers works

Eden’s project, ActivityBot, receives funding from the NLnet foundation, which requires evidence that the installation experience was tested with real users. To meet that requirement, Eden posted a call on Mastodon, gathered a handful of volunteers, and agreed to pay them 25 euros per hour of their time.

The format was always the same: he asked each volunteer to share their screen and narrate every decision out loud before making it. What they understood, what they didn’t, what frustrated them, what made them laugh. Eden took notes by hand, without relying on any automatic transcription tool.

Among the findings from those sessions were, among others:

  • The link to the demo tool was broken.
  • Some people read the README directly from the terminal, not in a browser.
  • Nobody understood what “renaming a file” actually meant.
  • Renaming a hidden file had no clear instructions.
  • It wasn’t clear whether the demo tool should run on the web or on the user’s own machine.
  • Some parts of the text needed quotes and others didn’t, with no obvious reason why to the reader.
  • The author’s jokes weren’t funny: they just caused confusion.
  • Certain technical terminology needed to be explained before being used.
  • The order of the sections confused first-time readers.
  • The README never explained, in one simple sentence, what the software actually did.
  • A section the author found fascinating turned out to be incomprehensible to everyone else.

Each finding went into the README before the next call, so the next person no longer hit the same obstacle. The full cycle looks like this:

flowchart TD
    A["Recruit volunteer on social media"] --> B["1-hour session: screen share and think aloud"]
    B --> C["Author takes notes by hand on every point of confusion"]
    C --> D["Update the README with the findings"]
    D --> E{"Volunteers remaining"}
    E -->|"Yes"| B
    E -->|"No"| F["README validated by multiple real users"]
Eden paid each volunteer 25 euros an hour in 2026. Foto de CDC en Unsplash

Practical examples: how to apply this method to your own project

You don’t need an NLnet grant to replicate the format. Four decisions before the first session are enough: who to recruit (people who’ve never used the project, ideally outside your close technical circle), how much to pay or what to offer in exchange, exactly what you’ll ask them to do (read the README and narrate each step), and how you’ll record what they say.

During the session, the golden rule is not to intervene. If someone gets stuck on a command, the natural impulse is to explain it out loud; that ruins the test, because the explanation you give live will never be in the text the next thousand users read. Let them get stuck, note where, and move on.

💡 Tip: Ask the person to narrate every decision before making it (“now I’m going to copy this command because…”). That second of narration is what gives away the wrong assumption, not the error itself.

After each session, update the document immediately, while the problem is still fresh. That’s what Eden did: he didn’t wait to collect feedback from all sessions, he fixed things between calls. This matters because it prevents the same stumble from repeating five times before it gets fixed.

Real use cases

Projects funded by foundations like NLnet increasingly ask for evidence of user testing as a condition for releasing funds, not just working code. That turns documentation testing from an optional gesture into a reporting requirement.

Technical documentation teams at large companies apply a cheaper version of the same idea: using a newly hired employee as an unwitting tester of the onboarding guide, before they finish memorizing the system. Once that person masters the product, they stop being useful as a test volunteer, because they inherit the same biases as the original author.

Maintainers of smaller open source projects often apply a free variant: posting a call in their own community (Discord, a forum, a mailing list) asking someone to follow the README from scratch and report where they got stuck. Payment isn’t always necessary; sometimes a clear request and public thanks are enough.

Common mistakes and author biases

  • Curse of knowledge: the author takes for granted steps they mentally automated months ago.
  • No shared vocabulary: terms like “rename” mean something different to someone who doesn’t code every day.
  • Confirmation bias in self-review: rereading your own text only confirms what you already believe you wrote, not what it actually says.
  • Humor that doesn’t travel: an author’s inside joke rarely lands with a new reader and can distract right at the critical step.
  • Assuming a single reading channel: not everyone opens the README in a browser, some read it with cat or less in the terminal.
  • Narrative order, not usage order: the author orders sections the way they wrote them, not the way someone installing the project needs them.
GOV.UK requires a second editor to review every technical text. Foto de Trnava University en Unsplash

Comparison with alternatives

Not every form of README testing costs 25 euros an hour; there are alternatives with different costs and levels of rigor.

OptionWhen to use itAdvantageLimitation
Paid real usersBefore a major launch or a grant applicationExposes biases the author can’t see, with genuine reactionsCosts money and coordination time
Internal second editor (GOV.UK model)Ongoing review of every text before publishingFast, no extra cost, always availableShares much of the author’s technical vocabulary
LLM simulationQuick first pass for obvious writing errorsFree and available in minutesDoesn’t get frustrated or react like a real human
No testing, just self-reviewAlmost never, except for trivial one-word changesZero costThe author can’t detect their own biases

⚠️ Heads up: someone suggested Eden simulate users with an LLM and he turned it down. A model doesn’t get frustrated, doesn’t have a cat walking in front of the camera, and doesn’t convey in its voice the exact moment something stops making sense.

Going deeper

Eden’s figure isn’t arbitrary. Jakob Nielsen showed in 1993 that five users are enough to find most of the usability problems in an interface or a text; adding a sixth or seventh yields fewer and fewer new findings each time. If the 150 euros Eden spent are divided by the 25 euros per hour he paid, the result is about six sessions, right in the range where the curve of new findings starts to flatten out.

The think-aloud protocol, documented for decades in usability studies, works because it forces people to verbalize a decision before acting on it: silence is just as good a signal of a problem as a question spoken out loud.

The method has a real ceiling. It doesn’t scale to every commit, and paying users for every minor change would be absurd. Eden reserved it for first contact with the project, the moment where the most author biases survive unfiltered.

Your next step: pick the README of one of your own projects, record yourself reading it out loud for 15 minutes as if seeing it for the first time, and note every sentence where you hesitated or had to reread.

📬 Get new articles by email

We only email about big articles (1-2 a month).

Frequently Asked Questions

What’s the difference between README testing with real users and review by teammates?

A teammate shares most of the author’s technical vocabulary and so doesn’t catch the same gaps; an outside user has no prior context and gets stuck exactly where the text assumes too much.

How much does it cost to run usability tests on documentation with paid volunteers?

Eden paid 25 euros an hour and spent about 150 euros total across several sessions; the cost varies by region and call length, but the format doesn’t require expensive tools, just time and a video call.

Can simulating users with an LLM replace real README testing?

It works as a quick first pass for obvious writing errors, but it doesn’t replace a human reaction: a model doesn’t get frustrated and doesn’t let you hear in its voice the exact moment an instruction stops making sense.

How many technical documentation validation sessions are needed to find the main problems?

Jakob Nielsen’s 1993 research suggests that five users already reveal most problems; Eden did something similar with the sessions he paid 25 euros an hour for, updating the text after each one.

Which author biases are hardest to detect without documentation testing?

The curse of knowledge and shared vocabulary are the most persistent, because the author has no way of noticing them by reading their own text: they need someone from outside to get stuck out loud.

References

  • Terence Eden’s Blog: the original post where the author describes paying volunteers to test ActivityBot’s README.
  • Nielsen Norman Group: Jakob Nielsen’s research on how many users are enough for a usability test.
  • Wikipedia: definition and general methodology of usability testing, including the think-aloud protocol.
  • GOV.UK Content Design: the UK government’s content style guide, which requires review by a second writer.
  • NLnet Foundation: the foundation that funds projects like ActivityBot and requires evidence of real-user testing.

📱 Enjoying this content? Follow @programacion on Telegram for daily tech content in Spanish: quick summaries, fresh content every day. @programacion

Featured image: Foto de Vitaly Gariev en Unsplash

Did it work for you? Got a different error? Say so below: questions get answered and help the next reader.

Leave a comment
Categories: Tech NewsTutorials

Andrés Morales

Developer and AI researcher. Writes about language models, frameworks, developer tooling, and open source releases. Covers ML papers, the tech startup ecosystem, and programming trends.

0 Comments

Leave a Reply

Avatar placeholder

Your email address will not be published. Required fields are marked *

You can include code inside <code>…</code> or, for several lines, <pre><code>…</code></pre>.

This site uses Akismet to reduce spam. Learn how your comment data is processed.