A better number can hide a worse result.
Imagine a library asking an AI assistant to help attract more people to its book discussions. The initial target is simple: increase registrations. The assistant suggests dramatic headlines and promises that each meeting will reveal the one answer everyone needs.
The invitation might attract attention. It could also misrepresent a thoughtful conversation where disagreement is welcome. More registrations would not tell the librarian whether visitors found the discussion useful or felt invited to return.
The narrow target: Increase registrations.
The human purpose: Help interested readers find a conversation they value.
A better brief: Write a clear invitation describing the actual topic, format, and audience. Offer three truthful versions. Explain who each might reach and what expectations it creates.
This is a fictional worked example, not a reported experiment. It does not prove that a particular wording increases attendance. Its purpose is to expose the gap between a convenient number and the experience that number is meant to represent.
Four questions before you optimize
- What are we trying to make better? Name the human result first. “Help a reader choose a suitable book” is a purpose. “Increase clicks” is one possible measure.
- Who experiences the result? Include the people who bear costs as well as those who benefit. In the library example, consider newcomers, regular readers, and the person running the discussion.
- How could the measure improve while the purpose fails? Imagine a truthful invitation attracting fewer people but a more suitable group. Imagine a crowded meeting where nobody has room to speak. Neither attendance alone nor a single satisfaction score settles the question.
- What small test would make us reconsider? Compare two accurate invitations. Ask participants whether the meeting matched their expectations. Keep what helps; revise the brief when observations contradict it.
Use AI to challenge the brief.
Try this prompt with a fictional or non-private task:
My purpose is ____. The measure I am considering is ____. Identify two ways that measure could improve while the human result gets worse. Name someone whose experience I may have overlooked. Suggest a small reversible test, and separate what you know from what you are assuming.
Treat the response as a set of possibilities to examine. An assistant can help identify a missing perspective; it cannot speak for the people concerned. Checking facts and choosing a worthwhile goal are related jobs, but one does not replace the other.
Where the research fits
Amodei and colleagues’ Concrete Problems in AI Safety (2016) distinguishes problems involving poorly chosen objectives, side effects, reward hacking, supervision, and learning. It is a technical research paper, not evidence that this library exercise has been experimentally validated.
Take the question into a conversation.
AI: The Struggle for Perfection: A Bridge to Peace or Chaos develops the broader human question through cooperation and public choices. For a group, use the five discussion prompts for this book. For a factual answer, try the separate AI fact-checking guide.
This original companion exercise can be read independently. It contains no book excerpts. Explore all four books in Artificially Superintelligent when you want to follow the argument further.