A prompt set defines the questions behind an AI visibility result. It should reflect the buyers and decisions you care about, with enough stability to compare results over time. Selecting only questions where the brand already performs well produces a reassuring number and a weak measurement.
Start with observed buyer needs
Use customer conversations, support questions, on-site search and other relevant evidence to identify real constraints: budget, compatibility, size, delivery, use case and alternatives. These sources help choose questions; they do not automatically reveal how often the whole market asks an assistant.
Write prompts in natural language and preserve their exact wording. Record where each question came from and what decision it represents.
Cover distinct intents
Intent | Purpose | Illustrative prompt |
|---|---|---|
Buyer selection | Find products meeting constraints | “A 700 ml or larger bottle for a 7 cm cup holder” |
Branded | Check a specific business or product | “What is Example Maker's return policy?” |
Educational | Understand a relevant concept | “What is the difference between washed and natural-process coffee?” |
Comparison | Choose between named alternatives | “How do Example A and Example B differ in base diameter?” |
These are sample questions, not evidence of market demand. For counting, assign each prompt to one primary intent using a documented rule. A named comparison could also mention a brand; do not double-count it in the allocation.
Use an explicit starting allocation
For an illustrative 40-prompt set, a team could start with 24 buyer-selection, six branded, six educational and four comparison prompts. That is 60%, 15%, 15% and 10%, adding to 100%.
This is a planning heuristic, not an empirically optimal mix. Adjust it when your buyer evidence supports a different balance, and document the reason. A category with difficult compatibility decisions may need more precise comparison questions than a category driven by replenishment.
Category examples can be concrete without adding leading wording:
Coffee: “Whole-bean coffee for pour-over under $20 per bag.”
Beauty: “A fragrance-free face wash under $20.”
Apparel: “Waterproof walking shoes available in wide sizes.”
Home goods: “A bookcase less than 80 cm wide with adjustable shelves.”
Record region and other relevant context where the answer depends on it.
Separate the baseline from exploration
Keep a stable baseline set for trend reporting. Use a separate exploratory set for new products, seasonal questions and emerging competitors.
When a baseline question becomes irrelevant, retain the reason and effective date of removal. Compare the overlapping set across the change, and show the new full set separately. A different set of questions is a different measurement population.
Do not delete a relevant prompt simply because the brand has not appeared in its answers. Persistent absence can identify an important discovery gap. Investigate whether the question is relevant, whether the collection works and which alternatives are being recommended.
A flat result is also not sufficient evidence that a prompt is malformed. Stable behavior can be a real finding.
Keep execution comparable
For every observation, save the prompt identifier and version, assistant product, mode, locale, model when exposed, date, collection outcome and complete answer. Keep branded and unbranded results distinct when interpreting discovery: a question that names the brand is testing a different task.
Record attempts that fail or produce no answer separately from valid answers that omit the brand. If a provider or mode changes, identify the break rather than silently treating all observations as interchangeable.
Review the set for relevance
Ask whether each question still represents a real buyer decision, whether important constraints are missing and whether the brand mix has changed. Update because the decision changed, with a recorded reason, rather than to make the metric improve.
The benchmark guide helps assess whether a comparison is relevant to your buyers. The sampling example shows why repeated answers and missing records need careful treatment.