TL;DR: Six rules. Never include your brand name unless you are deliberately testing branded recall. Write sentences, not keywords. Cover problem-aware questions, not just vendor-selection ones. Include the follow-up. Keep the set small enough to read. And write them in your buyer's words, not your category's.
Why the prompt set decides everything
A visibility measurement is only as good as its questions, and a bad prompt set fails in a specific and flattering way: it tells you that you are doing well on questions nobody asks, or that you are doing badly on questions you could never win.
The commonest failure is subtler than either. A prompt set can be made entirely of vendor-selection questions — "best X", "X vs Y", "alternatives to Z" — every one of which looks reasonable on its own. The problem only appears when you look at the citations across all of them, and find that every source is a competitor or a comparison site. We did exactly this to ourselves: twenty-two tracked questions, all vendor-selection, and the twelve most-cited domains were all rivals or roundups.
Rule 1: do not name your own brand
The single most common mistake. "Is [brand] good for X" guarantees the answer names you, because the question did. You have measured nothing.
Branded prompts have one legitimate use — testing what the engine says about you, which is a sentiment question — and they must be kept separate from presence measurement. Mixing them inflates your numbers and nobody notices until a leadership team asks how the figure is calculated.
Rule 2: write sentences, not keywords
"ai visibility software" is a search query. Nobody types that into a chatbot. They type "what should I use to find out if ChatGPT recommends my company".
The difference is not stylistic. A conversational question is longer, carries more context, and frequently triggers a different retrieval behaviour from a two-word phrase. A prompt set written as keywords is measuring a surface that is not the one you care about.
Rule 3: cover problem-aware questions
The rule that most sets break, and the one with the largest consequence.
Every vendor-selection question assumes the asker already knows your category exists and is choosing inside it. Those answers cite the people selling in that category. By definition, they cannot surface the blogs, documentation, forum threads and newsletters that a brand can actually get into.
A problem-aware question is what somebody types before they know what to buy: "why is my brand not showing up in ChatGPT", "how do I get my website cited by AI". The answers cite completely different sources, and those sources are reachable.
A healthy set is roughly 60% vendor-selection and 40% problem-aware. Most sets we see are 100% / 0%.
Rule 4: include the follow-up
Turn two is where a problem question becomes a buying question, and it is where most tools stop measuring.
Somebody asks why their site is invisible, gets an explanation, and then asks "what tools would help with that". Whether you are named in that answer is the measurement worth having, and it is a different answer from the one that comes back on turn one.
If your tool supports a second turn, use it. If it does not, ask it by hand occasionally.
Rule 5: keep the set small
Twenty to thirty questions measured weekly beats two hundred measured monthly, for a reason that has nothing to do with cost: you have to be able to read the result.
The point of measurement is to notice change and act on it. A two-hundred-row report is a thing nobody opens twice, and an unread measurement is worth exactly nothing.
Rule 6: use your buyer's words
Write the question the way a customer would say it, including the parts that make you wince. If your buyers call it "AI SEO" and you call it "answer engine optimization", measure "AI SEO".
The test: read the prompt aloud. If it sounds like something from your website rather than something a person would say, rewrite it.
A worked example
A fictional invoicing tool for construction subcontractors.
Bad set
- invoicing software (keyword, not a question)
- is [brand] good invoicing software (names the brand)
- best invoicing software (too broad to ever win)
- invoicing software pricing (keyword again)
Good set
- What is the best invoicing software for construction subcontractors?
- I run a small building firm and chase payment constantly. What should I use?
- How do I get paid faster on construction jobs?
- What is the easiest way to invoice for retention and variations?
- Is invoicing software worth it for a two-person building firm, or is a spreadsheet fine?
- [rival] vs [rival] for construction — which is better?
Six questions, four of them problem-aware, none naming the brand, all in a subcontractor's language rather than a software vendor's.
How to tell your set is working
The citations are varied. If every cited source is a competitor or a comparison site, your set is entirely vendor-selection and you are measuring the competitive landscape rather than your visibility.
You are losing some of them. A set you win entirely is a set that is too easy and will show you no movement.
You are winning some of them. A set you lose entirely is aspirational, and tracking only aspirational questions makes the measurement useless as a feedback loop.
The answers vary between engines. If all four agree on everything, the questions are probably too generic.
Frequently asked questions
How many prompts do I need?
Twenty to thirty for most businesses. Enough to cover the buying journey, few enough to read every week.
Should I track questions I know I will lose?
A few, deliberately, and flagged as such so they do not drag your headline number down. Those are your targets rather than your scoreboard.
How often should I change the set?
Rarely. The value is in comparing like with like over time, and a set you keep editing cannot show you a trend. Add rather than replace.
Where do good prompts come from?
Search Console queries, sales call recordings, support tickets, and the actual words customers use in their first email. Not from a keyword tool.