Last chapter, you built an attribute list: the raw material of how you want AI to see your brand. Now comes the part where we turn those attributes into prompts - the questions that reveal what AI thinks you are and how it currently represents you.
The first trap in this prompt-writing endeavor is volume. You look at the endless ways someone could phrase a question to AI and think I need to catch them all. That instinct is wrong, and acting on it is how prompt programs collapse under their own weight before they produce a single useful read.
There are many frameworks for creating prompts - this is ours. This chapter is all about what we believe each type needs to accomplish, how to write them, and where to put them to work.
You can't track every prompt (and you shouldn't try)
With SEO, you had a sense of what people were searching for. Search volume gave you a number, a clean string you could drop into a tool and sort by intent. The inputs were knowable, even if the work itself was, at times, messy.
AEO doesn’t play by those rules. The same buyer, chasing the same outcome, can get there in ways that look nothing alike:
- A long, back-and-forth conversation, where everything the user said in earlier turns is influencing how AI answers their question now.
- A single dense prompt loaded with constraints.
- A personalized exchange where the platform already knows half the context before they finish typing.
Prompt volumes tell you what people are asking. They don't tell you how the answer changes once a buyer adds detail, or what happens later in the conversation. That's the limit of the data, not a flaw to fix - you don't need to chase every variant by hand. You need a way to act on the patterns volume does show you. That's what Profound Agents are for: they take a prompt cluster where competitors are winning citations and turn it into a content brief, automatically, instead of leaving you to manually sort through the data. Volume is the orientation point; coverage across variants is how you build a read you can trust.
So if you can't catalog every path a buyer might take, what are you doing when you build a prompt tracking program?
“You’re after coverage. The prompts in your tracker don’t have to match what buyers type, word for word. They just need to be representative and varied enough that you can trust the read on each attribute.“
You’re after coverage. The prompts in your tracker don’t have to match what buyers type, word for word. They just need to be representative and varied enough that you can trust the read on each attribute.
Noah Greenberg, CEO, Stacker
Treat each attribute like a diversified portfolio
Let’s make it concrete. Say you’re tracking the attribute “CRM.” If your only prompt is “What’s the best CRM?”, you have no idea how the answer changes if someone says “top” instead of “best,” asks about a vertical, frames it as a job to be done, or swaps in a related category. All those questions orbit the same need, but none of them guarantee the same answer.
We like to think of each attribute as a diversified portfolio. In finance, you spread your bets across many holdings so no single stock tanks your portfolio. In prompt development, you spread prompts across each attribute so no single phrasing tanks your read on visibility.
You still dig into individual prompts to spot where things go sideways. But it’s the rolled-up view - how you’re doing at the attribute level - that tells you if your AEO program is working.
In Profound, each attribute maps to a Topic or a Tag - your organizing unit for tracking. Each one groups the prompts you build for that attribute, so when you're measuring visibility and rank, you're reading performance at the attribute level, not prompt by prompt. That's the rollup logic in practice: the individual prompts do the measuring, the Tag/Topic holds the read you act on.

To counter that issue, one prompt framework I think more teams should using is what we call ""PSUC"" which stands for:
- Persona - who is the person doing the prompt
- Segment - what type of company do they work for
- Use Case - what use case do they have or what problem are they trying to solve
Real users tend to add context. They tell the model who they are, what kind of company they work for, what problem they are trying to solve, what constraints they have, and what kind of recommendation they need.
The reason this matters is that AI recommendations can change dramatically based on context. A vendor that shows up for a generic category prompt may not show up when the buyer adds industry, company size, geography, workflow needs, or integration requirements.
So for AEO, we don’t just want to know whether a brand appears for broad category prompts. We want to know whether a brand appears in the specific buying scenarios where it should be recommended. PSUC prompts give you a much better read on that." - John-Henry Scherck, Founder & CEO, Growth Plays
If you're wondering just how much a single word can impact your visibility, we actually have data on it. The team at OGM took one question - which tools help brands rank higher in AI search? - and asked it four different ways, swapping only the category label: GEO Tools, AI Search Optimization Tools, AI Visibility Tools, Answer Engine Optimization Tools.
Same eight AI platforms, same fourteen days, everything else held constant. Then they looked at which brands AI recommended in each variant.
Across the tools tracked, visibility swings ranged from 25 to 50 percentage points, depending on which label was in the prompt. Profound swung 25. Another tool swung 34. The platform-level picture was even more uneven: on a single prompt, on a single day, Profound's mention rate ranged from 7% on Perplexity to 100% on Google AI Mode. One word changed, and which brand AI recommended changed with it.
That's the case for coverage. If one word can swing your visibility by 50 points, trusting a single prompt is a fool’s errand. Build a portfolio of prompts for each attribute and look at the aggregate - that’s how you get a read you can trust.

| PROFOUND QUERY FANOUT Profound's Query Fanout feature does this automatically, generating prompt variants from each attribute. The principle works the same way whether you use the feature or write the variants by hand. |
|---|
Six types of prompts
Not all prompts do the same job. There are different types because they answer different strategic questions, and it’s a costly mistake to treat the prompt list as one undifferentiated bucket.
We like to use six types:
- Discovery prompts ask whether your brand gets recommended when someone is exploring solutions. "What are the best AEO tools for enterprise brands?" You're asking AI to produce a list, and you're tracking whether you're on it.
- Competitor prompts ask who AI picks when it's narrowed to a head-to-head. "I'm looking for an email marketing platform with the best AI features. Would you recommend ActiveCampaign or HubSpot?" You're seeking a verdict.
- Validation prompts ask whether AI even understands that you have a specific capability. "Does Profound offer prompt volume tracking?" It's a yes-or-no question.
- Accuracy prompts check whether what AI says about your product is factually correct. Pricing, features, integrations, specs - anything that has a right answer. You're asking whether the perception matches reality.
- Sentiment prompts ask for an evaluation. "Evaluate Profound as an AEO tool and give me the pros and cons." You want AI's read on you, in its own words.
- Market perception prompts zoom out one level further: how does AI frame the category itself? "What is a digital asset management platform and how should I think about evaluating one?" You're not chasing your brand here. You're checking whether the way AI explains the space favors capabilities you have or capabilities your competitors lead on.

What each type needs to do
Each prompt type has its own success criterion - and a prompt that doesn't match its type's criterion produces unusable data.
Discovery: Must trigger a brand list
Discovery is the most common type and the one most people start with. The whole point is to extract a list of recommended brands.
This is harder than it looks. Ask "What's the best way to improve my 3PL?" and AI will return a list of tips: optimize warehouse layout, improve shipping speeds, negotiate better carrier rates. Not a 3PL provider in sight. The fix is a trigger word - software, tool, platform, provider, solution - that tells AI you want a list of products, not a list of advice.

The difference looks like this:
- “What are the best 3PL providers for a retail D2C business?” → returns brands ✓
- “What email marketing software would you recommend for a mid-size SaaS company?” → returns brands ✓
- “How do I improve my email marketing?” → returns tips ✗
- “Tell me about 3PL trends in 2026” → returns industry analysis ✗
The validation check: after writing a discovery prompt, run it through the platforms you're tracking and confirm you're getting brand names back. If you're getting strategies, the prompt is doing the wrong job.

Competitor: Must produce a firm recommendation
Discovery tells you who's in the conversation. Competitor tells you who wins when it's narrowed to two named brands. Different data, different implications. You can show up in every discovery list and still lose every head-to-head, or vice versa.
Prompt structure matters more than you think. “Compare X and Y” is too soft; AI will hedge and move on. To get a real recommendation, you need three things:
- A specific use case
- Two named brands
- An explicit ask for a recommendation
It looks like this: "I'm looking for an email marketing platform with the best AI features. Who would you recommend, ActiveCampaign or HubSpot?" The use case forces AI to pick a side. Without it, you'll get a draw, and a draw isn't useful data.
A real buyer might never type that prompt verbatim. They might get there through a winding conversation - starting broad, narrowing down, until two names are left. We can’t mimic that journey, but we can ask the comparison directly and get a clean read on the core question.



Validation: Your diagnostic for trust vs. comprehension**
Validation prompts ask whether AI knows you have a capability, independent of whether it recommends you for it.
The prompt is a direct yes-or-no: "Does Profound offer prompt volume tracking?" That's it. You're not asking AI to compare you against alternatives; you're asking whether AI even understands the capability is yours.
Suppose Profound runs a discovery prompt - “*What are the best AI visibility tools that offer prompt volume tracking?” - *and we don’t show up. The knee-jerk reaction is to blame content: AI must not know we do this, so we need to write more. But that’s just a guess. We don’t know if AI knows; we only know it didn’t recommend us.
Now we run the validation: “Does Profound offer prompt volume tracking?” If AI says yes, you’ve learned something specific. AI knows you have the capability; it just didn’t recommend you. That’s not a comprehension problem. That’s a trust and amplification problem, and the fix is different.
Here’s a useful framework to work through potential scenarios:
- If you show up in discovery for an attribute, AI knows you have the capability and trusts you enough to recommend you for it. That's the goal state, no action needed.
- If you show up in validation but not discovery, AI knows you have the capability but doesn’t recommend you. The solution isn’t more on-page content, it’s external signals: citations, third-party authority, off-page descriptions. The work happens off your site.
- If you show up in neither, AI doesn’t know you have the capability at all. That’s a comprehension problem and the fix is upstream: clearer descriptions, structured documentation, content that makes the capability legible.

These are completely different programs of work, and without validation prompts, you can't tell which one you need.
Accuracy: Checks the facts, not the perception
Fact check prompts ask whether AI's description of your product is factually correct. Not whether buyers love you, not whether you're recommended, but whether the specs, pricing, integrations, and capabilities AI brings up about you are true.
This is its own discipline. Take a phone manufacturer launching a new model. The product comes in five colors. Marketing has the colors on the product page, in the press release, in the launch video. But when a buyer asks AI "what colors does the [model] come in?", AI lists three of them and gets one wrong.
That's a factual error. It has nothing to do with how the brand is perceived. It has everything to do with whether AI is reading the right sources, and whether those sources have the right information.
Profound has a built-in feature for accuracy testing: you input what's true about your product, e.g., the specs, features, pricing, integrations, and Profound checks AI answers against the source of truth, flagging where reality and the AI version diverge.



Sentiment: Ask for an opinion
Sentiment prompts ask AI to give you a read on how you're perceived.
Say you're tracking the attribute "CRM." A prompt like "What's the best CRM?" is a reasonable starting point, but on its own, it can't tell you how the answer shifts when someone says "top" instead of "best," asks about a vertical, frames it as a job to be done, or swaps in a related category.
The attribute-specific version is more powerful. You can tag sentiment prompts by attribute/topic and follow the perception thread by capability, instead of getting one blended “sentiment score” that hides the real signal. If your sentiment is strong on one attribute and weak on another, you can act on that. A single average tells you nothing.

Market perception: How AI frames the category
Market perception asks "how should someone think about this category?" Brands may surface in market perception answers, but that's a side effect; what you're really after is the framing AI uses to explain the category to a buyer who doesn't know it well yet.
A market perception prompt looks like this: "What is a digital asset management platform and how should I think about evaluating one?" You're watching for:
- Which features AI treats as table-stakes
- Which it treats as differentiating
- Which criteria it surfaces as the right ones for a buyer to weigh
Imagine you sell a creative operations platform. Your strength is workflow, approvals, and brand governance for in-house creative teams. You ask the market perception prompt and AI explains the category as "digital asset management for creative teams," emphasizing storage, tagging, and asset retrieval as the evaluation criteria.

You can be perfectly visible in that category and still lose at the moment of decision because AI has framed the buying decision around capabilities your competitors already own. The category framing is upstream of your visibility, and you're losing on the framing.
That's a strategic problem with brand-level consequences. Market perception prompts are how you spot it.
How to write prompts that work
Craft is actually important in prompt development. Even the smallest ambiguity can land you in a pit of bad AI visibility data, so let’s talk about how to avoid the usual traps.
Make sure AI understands what you mean
When you write a prompt, you know what you mean, but AI doesn't always. The space between what you intended to ask and what AI parsed is where bad data sneaks in.
Before you commit a prompt to your tracker, run it through the platforms you're tracking and read what comes back. Ask:
- Did AI interpret the question the way you wrote it?
- Did it read your intent as discovery when you meant validation?
- Did it interpret a vague term in a way you didn't anticipate?
Five minutes of spot-checking will catch most of this before you start building data on a flawed foundation.
One important clarification: intent alignment isn’t the same as biasing the answer. A prompt like "Tell me all of the AEO platforms named Profound that have great answer engine visibility" is a leading question, and the data it produces is worthless.
You should ask the question that reflects the intent you're trying to measure, not the answer you're hoping to see.
"Start small, keep it simple, and keep it bottom of funnel focused.
- Category prompts: “What are the best [category] platforms?”
- ICP/use-case prompts: “What’s the best [category] for [industry/company type/use case]?”
- Pain/problem prompts: “What tools help solve [specific operational problem]?”
- Comparison prompts: “How does [brand] compare to [competitor] for [use case]?”
- Brand-perception prompts: “What is [brand] best known for?”
- “What type of customer is [brand] best suited for?”" - Gaetano DiNardi, Principal Consultant @ Marketing Advice
Cut vague terms before they cut your data
If your prompt contains a term that could mean different things in different contexts, AI will pick one - or worse, mix several - and you won't know which until you read the answer. Vague terms produce unreliable data, and the most common offenders are easy to miss because you know what you mean by them.
The clearest example is acronyms. Ask "What are the best GEO platforms?" in an SEO context, and you know GEO means Generative Engine Optimization. AI might not.

Without context, it'll hedge - "depends what you mean by GEO" - and produce an answer that mixes:
-
Geographic information systems (Esri, ArcGIS)
-
Geothermal energy platforms (entirely different industry)
-
Generative Engine Optimization tools (the answer you wanted)
The solution is to spell it out the first time. "GEO (Generative Engine Optimization) Tools" gives AI the disambiguation it needs in a single phrase.
In a real conversation, that ambiguity self-resolves. In prompt tracking, there's no follow-up turn. Whatever AI parses on the first read is the data you collect. That's why you resolve ambiguity upfront: not to lead the answer, but to measure what AI thinks when it correctly understands the intent.
The same problem extends beyond acronyms, in any term that means different things across verticals or product categories. For instance:
- "What are the best CRO tools?" → Conversion Rate Optimization (Hotjar, Optimizely) or Chief Revenue Officer tools (Outreach, Gong)
- "What are the best attribution platforms?" → Marketing attribution (AppsFlyer, Branch) or security attribution (SIEM tools)
- "What's the best content management platform?" → CMS (WordPress, Webflow) or headless CMS (Contentful) or DAM (Bynder, Brandfolder)

Each of these gets you a Frankenstein answer, stitched together from different categories, because the term is ambiguous. The solution is still to name the category clearly enough that AI can’t grab the wrong one.
In a real ChatGPT session, if AI hedges, you clarify and move on. In your tracker, you don't get a follow-up turn. Whatever AI parses on the first read is the data you're collecting, and if it's parsing the wrong category half the time, your visibility numbers are far from accurate.
Check that the response format matches the type
As we’ve established, each prompt type has an expected response format. When the response format doesn't match the type, the prompt has drifted - even if the answer technically reads as reasonable.
This is a different failure mode from vague terms. Vague terms confuse AI about what you're asking; format mismatch means AI understood the question but answered it as a different type than you intended.
The mismatches to watch for are:
- Discovery returning tips instead of brands → the prompt is missing a trigger word (software, tool, platform, provider). Rewrite to ask for products explicitly.
- Competitor returning "it depends" → the prompt is missing a use case. Add the specific scenario that forces AI to pick a side.
- Validation returning a discovery-style list → the prompt isn't direct enough. Reframe as a yes-or-no question about the named brand: "Does [brand] do X?"
- Sentiment returning a list of brands → the prompt is acting as discovery. Rewrite to explicitly ask for an evaluation, with pros and cons.
The QA check is the same as the one for intent alignment: run a handful of your prompts through the platforms you're tracking and read the responses.
Where to do this work
You have the types, the requirements, and the craft principles. The remaining question is operational: where do you run this program?
There are three viable paths. They're not equivalent in effort, fidelity, or coverage, but the methodology in this chapter works on all three.
Run it in Profound
No surprises here, as Profound is purpose-built for all things AEO. Within our platform, you define your attributes, we help you generate prompts (including the variants we covered earlier in the index fund logic), and the tracking, scoring, and reporting happen automatically across the major AI engines. The query fanout feature means you don't have to write every variant by hand because we generate them from the attribute and run them at scale.
This is the natural handoff from strategy to execution. If you've worked through Chapters 1 to 3, you can step into this module of Profound University’s 101 course to learn everything about prompt setup and configuration inside the platform.
Build it yourself
If you have engineering resources (or just someone who’s really good at Claude Code) and an attribute list, you can build a system that does the core job.
The trade-off is the same you’ll encounter across every “build vs. buy” debate. Building means you own the roadmap, but you also own the maintenance, the platform integrations, the model updates, and every edge case that shows up when ChatGPT changes how it returns responses. For most teams, this isn't a feasible option.
Run it manually
You can also run prompts directly in ChatGPT, Gemini, Perplexity, or whichever platforms you're tracking, and log results in a spreadsheet. Labor-intensive, but accessible, and a way to start if nothing else. A team that runs even twenty prompts manually for two weeks will learn more about how AI describes their brand than teams who skip straight to a platform without the foundational read.
The limits are also obvious. You won't track at scale. You won't see day-over-day trends without a lot of manual aggregation. You'll miss platform-level variance unless you're disciplined about running every prompt across every platform every time. But for a team that's just trying to build intuition before committing to a tool, this is the cheapest way to see the shape of the problem.
Let there be audits
The instinct, after a chapter like this, is to want to start tracking immediately. Resist it for a beat.
What you have right now is essential setup work: a defensible identity, an attribute list shaped by your real buyers, six types of prompts each doing a distinct strategic job, and enough coverage per attribute that no single phrasing can wreck the read.
The prompts are the lens. Chapter 4 is where you finally look through it: running your first audit, reading what AI is saying about you, and figuring out the most pressing issues to address first.
“To start, teams should do research as if they are coming to an LLM cold with no context to see what an LLM spits out. It's so basic, but people skip this and it's very illuminating. We've built two skills that automate much of this.
One, an Incognito AEO Audit, which audits your site with zero context on your company. It checks whether your content is readable without JavaScript, whether your robots.txt addresses LLM crawlers like GPTBot and ClaudeBot, and whether your pages have the structure and substance (clean H1s, schema, enough body text) for an LLM to actually cite you.
Two, an Incognito Competitive Research skill, which runs the same cold-context approach across you and your competitors at once: how your positioning, messaging, and proof points read to an LLM vs. theirs, and where they have AEO advantages you don't." - Emily Kramer, Founder, MKT1 Newsletter + Dear Marketers Podcast