7 Layer Prompt Universe Framework SEO: 24-Prompt Example
Learn the 7 layer prompt universe framework SEO teams can run: what each layer decides, a 24-prompt worked example, and how many runs each prompt needs.

On this page
This 7 layer prompt universe framework SEO teams can run turns one buyer task into the set of prompts you track in AI search, through seven layers from the task itself to repeated runs. The same phrase also names a seven-part content prompt, and the two connect: the universe feeds the prompt.
You probably saw "7 layers" in a post and can't tell which one you need. Fair, because the phrase is overloaded. And there is a second problem underneath it: ask an AI tool the same question twice and you get a different list of brands.
So a prompt universe is a sampling design, not a list of prompts you hope are representative. This guide builds one layer by layer, counts a full example, and shows where the content prompt plugs in.
In this guide:
- The seven layers, each with the decision it makes and how it fails
- A worked example: 24 prompts, 6 URL clusters and the run counts behind them
- How many runs each prompt needs before you can trust a result
- The seven-part content prompt, filled from the universe
The 7 layer prompt universe framework SEO teams can run, defined
A 7 layer prompt universe framework SEO teams can run is a build order that turns one buyer task into the prompts you track in AI search, across seven layers: task, stage, persona, phrasing, fan-out, engine and repeated runs. It is our build order, not an industry standard.
Each layer makes one decision and adds one field to the record. Each also fails in its own way, which is the useful part. When a tracked number looks wrong, you can name the layer that broke.
| Layer | Decision it makes | Record it adds | Fails when |
|---|---|---|---|
| 1. Buyer task | Which job the buyer hires the product for | One-sentence job statement | The "task" is a keyword or a feature name |
| 2. Journey stage | Where in the decision the prompt sits: problem, shortlist, comparison, validation | Stage label | Every prompt sits at one stage, usually the shortlist |
| 3. Persona and constraints | Who is asking and which limit (size, stack, budget, market) changes the recommendation | Persona id, constraint fields | Constraints are missing, so the prompt cannot trigger a specific recommendation |
| 4. Phrasing | How that persona words the prompt | Three or more verbatim wordings per cell | One canonical wording stands in for the intent |
| 5. Fan-out | Which sub-queries an engine spawns and which of your URLs should answer them | Cluster and target URL | A new page is written per prompt |
| 6. Engine and market | Which surfaces, modes and countries the prompt runs on | Engine, mode, country, logged-in state | One engine stands in for all |
| 7. Runs and record | How many times each cell runs and what is stored | Run number, mention, citation, recommendation, date | A single run is read as a fact |
Layers 4, 5 and 7 fail in ways a public source measures, and the sections below cite it. Layers 1 to 3 and 6 rest on our judgment, so treat their failure tests as working rules you can revise. Gap scoring, the step after you have a universe, is covered in our AI citation gap analysis guide.
What the phrase means elsewhere
Three other things carry the same label, and none of them is the tracking set. One is a seven-part prompt for producing content: role, task schema, anti-pattern filters, intent and audience, data inputs, examples and output format. Another is a seven-stage pipeline of agents that runs from strategy to publishing. The third is the general "7 layers of SEO" post.
One vendor does publish a page under this exact name, as its own standard for an agent pipeline, and that claim is fair for its pipeline. It describes a publishing workflow, not a set of buyer prompts.
There is no shared specification behind the phrase. It joins the tracking set and the content prompt, so this guide keeps the tracking set first and returns to the content prompt in the last section.
Prompt research and the prompt universe
Prompt research is the activity of finding the full questions buyers put to AI tools. The AI search prompt universe is the structured set that activity produces, held fixed so trends stay comparable.
A prompt also differs from a keyword. It carries the decision context: company size, stack, the constraint that makes a recommendation specific. "Payout reconciliation software" is a keyword. "What do small Shopify stores use to reconcile payouts without a finance hire?" is a prompt.
Prompt research for AI SEO: layers 1 to 3, from buyer task to prompt cell
Layers 1 to 3 fix what a prompt is about: one buyer task, the stage of the decision it sits in, and the person and constraints behind it. Each stage-and-persona pair is one intent cell.
You do prompt research for AI SEO by moving from one buyer task to intent cells, then to real wordings, fan-out sub-questions and a recorded set of runs. Here is the process in five steps:
- Write the buyer task as "do X without Y".
- Cross the four stages with the personas whose constraints change the recommendation, which gives your intent cells.
- Collect three real wordings per cell from support, sales calls and community threads.
- Group the fan-out sub-questions each prompt spawns and map each cluster to its strongest URL.
- Run each prompt on each engine and record the result.
The rest of this section covers steps 1 and 2. Steps 3 to 5 are the next three sections.
The intent cell, not the single prompt, is the unit you plan and report. Here is how each of the first three layers works:
- Layer 1, buyer task. Say the task as "do X without Y". If you can't, you have a keyword or a feature name. Sales-call and support wording are the raw inputs. Example: "Match Shopify payouts to bank deposits every month without a spreadsheet."
- Layer 2, journey stage. Use four stages: problem, shortlist, comparison, validation. Every stage gets at least one cell, and no stage holds more than half the cells. Otherwise the universe measures only the shortlist, where most teams already look.
- Layer 3, persona and constraints. A persona earns a place only if its constraint would change which vendor gets recommended. A founder with no finance hire and a finance lead who already runs an accounting package get different answers, so both stay. Two personas who differ only in job title collapse into one.
Persona choice also decides how much you can lean on demand estimates. AI prompt volume is a rough number, so use it to rank cells and never to include or drop one.
In the running example, 1 task x 4 stages x 2 personas gives 8 intent cells. The example section below carries that count through to the end.
Layer 4: phrasing, why one wording per intent misleads
Layer 4 stores several real wordings for each intent cell, because people word the same intent almost entirely differently, yet AI tools keep returning overlapping sets of brands.
SparkToro's study with Gumshoe measured both halves in one purchase scenario. Here is the trail:
- Volunteers wrote 142 prompts asking for headphone advice for a family member.
- Across those prompts, average semantic similarity was 0.081, so barely any two looked alike.
- The 994 responses still put Bose, Sony, Sennheiser and Apple in front of users 55-77% of the time.
The reading is that wording is noise around an intent. A phrasing is one draw from the intent, not the intent itself. Google's own guidance points the same direction for page authors: AI systems understand synonyms and general meanings, so you don't have to capture every long-tail variation. That line is advice about retrieval, not a second measurement, so we use it as context only.
From that, three rules for the record:
- Store three wordings per cell, one per source type. Support tickets, sales calls and community threads are the usual three. Three is our working number, not a measured threshold.
- Never treat one wording as the intent. A canonical prompt you invented at your desk is a guess about how one person types.
- Split a cell when its wordings stop overlapping. If two wordings return disjoint brand sets, they are two intents. That is our judgment, and your own record will test it.
Two limits apply. The study used one scenario, and it ran on model behavior from late 2025. The 142 counts prompts, not people.
Layer 5: fan-out, map sub-queries to URLs, not to new pages
Layer 5 maps the sub-queries an engine spawns from a prompt onto the URLs that should answer them, one strongest URL per cluster, and adds a new page only for a distinct user task.
Google's guidance for generative AI search defines query fan-out as concurrent related queries the model generates. It also says that creating separate content for every variation, primarily to manipulate rankings or generative responses, violates its scaled content abuse policy. That is a line about intent to manipulate, so a genuine gap can still justify a page.
This changes how you read a universe. It is a measuring instrument, not a briefing list for 24 pages. Read sub-queries from the engine's visible search steps, the People Also Ask box and follow-up suggestions, group them, and then decide with this table:
| Cluster situation | Action |
|---|---|
| A page already answers the cluster | Strengthen that page |
| The cluster is a variant of what a page answers | Add the passage to that page |
| The cluster is a distinct task and no page exists | Write one new page |
In the example below, 24 prompts gather into 6 clusters, so the work is a few edits and perhaps one new page, not 24 documents. The wider practice of getting cited sits under generative engine optimization, and the SEO-versus-GEO question is answered there too.
Layers 6 and 7: engine and market, repeated runs and the record
Layers 6 and 7 decide where each prompt runs and how many times, and they carry the most weight because AI engines return a different list of brands on almost every run.
SparkToro's study with Gumshoe ran 12 prompts through 600 volunteers on ChatGPT, Claude and Google's AI, which was the AI Overview or AI Mode when no Overview showed. It logged 2,961 runs in November and December 2025.
The chance that two responses returned the same list of brands was under 1 in 100 for ChatGPT and Google's AI. The same order came up about 1 in 1,000 times. The author's verdict: visibility percentage across dozens to hundreds of prompts, run multiple times, is a reasonable metric, and a "ranking position in AI" is not.
Repeating runs is the easy part. The hard part is the budget: how many runs a cell needs before absence means something. The study ran each of its 12 prompts 60-100 times, and its author says to ask over and over, usually at least 60-100 times, then average.
Across 12 prompts and 3 tools, the 2,961 runs averaged about 82 per prompt and tool (2,961 divided by 12 x 3).
Now the arithmetic that sets the plan. Take a brand that truly appears in 10% of answers (an illustrative rate). The chance it is missing from n independent runs shrinks fast, because every added run is another chance to catch it and the misses multiply:
- 5 runs: 59% missing
- 30 runs: 4% missing
- 60 runs: 0.2% missing
So a prompt run 5 times can show that a brand appears, and it can never show that the brand is absent. Absence needs runs on the order of the study's own 60-100. Wordings don't help here. Three wordings of one intent are three different prompts, not three repeats of one, so they never stand in for runs.
That gives a two-tier plan, and it is our sizing built on the author's minimum, not a standard. Screen every prompt at 5 runs per engine to find where a brand appears or leads. Then measure only anchor prompts, one per cell, at 60 runs.
Layer 6 is where you fix the surface. Log engine, mode, country and logged-in state for every run, because a prompt on one engine says little about another. Running the same panel across engines, on a schedule, is covered in how to track brand mentions in AI search engines.
Keep a stable core of cells you never change, so trends stay comparable, and a small experimental slice for new wordings and stages. Retire an experiment or promote it into the core; don't let it drift.
Mission Growth's platform tracks AI citations and visibility for customers. For the metrics and the reporting stack around a universe like this, see AI search analytics.
7 layer prompt universe framework example: Shopify payout reconciliation
One seed task, reconciling Shopify payouts without a spreadsheet, becomes 8 intent cells, 24 prompts and 6 URL clusters, screened in 360 runs, once you apply the seven layers. Everything below is illustrative: the wordings are examples, not measured prompts, and no results are shown.
Treat this 7 layer prompt universe framework example as a template you can copy. As a 7 layer prompt framework example it fits any B2B or e-commerce task: swap the seed task and keep the multiplication.
The product is payout reconciliation software for Shopify brands. Here is what each layer contributes:
| Layer | Choice in the example | Count |
|---|---|---|
| 1. Buyer task | Match payouts to bank deposits and the ledger each month without a spreadsheet | 1 task |
| 2. Journey stage | Problem, shortlist, comparison, validation | 4 stages |
| 3. Persona | Founder of a 6-person brand with no finance hire; finance lead at a brand that already runs an accounting package | 2 personas, 8 intent cells |
| 4. Phrasing | 3 wordings per cell, one per source type | 24 prompts |
| 5. Fan-out | Payout mismatch causes, reconciliation how-to, tool shortlist, comparison pages, migration off spreadsheets, trust and accuracy | 6 URL clusters, 4 prompts per URL |
| 6. Engine and market | 3 engines, US, with mode and logged-in state logged | 3 engines |
| 7. Runs and record | Screen every prompt at 5 runs per engine | 360 runs a cycle |
The chain of multiplication is 1 x 4 x 2 x 3 = 24 prompts. Then 24 x 3 engines x 5 runs = 360 screening runs.
Take one cell to see Layer 4 at work. The founder at the problem stage gets three wordings, one per source type:
- Support: "my Shopify payout doesn't match my bank deposit, how do I reconcile it"
- Sales call: "how do small stores reconcile Shopify payouts against fees and refunds"
- Community: "is there something better than reconciling Shopify payouts in Excel every month"
Then the measurement tier: run the anchor prompt of each cell that screens as absent or leading 60 times per engine. At most that is 8 anchor prompts x 3 engines x 60 runs = 1,440 runs. And only 6 URLs are touched, since the 24 prompts gather into 6 clusters at about 4 prompts per URL.
Here is what the record tells you:
- Absent at 5 runs is a candidate, not a gap. A brand at 10% is missing from 5 runs 59% of the time.
- Absent at 60 runs is a finding. Now it is worth a content decision.
- One wording surfaces the brand and two don't. That is a Layer 4 finding. Check the cell before you commission a page.
The other 7 layer prompt: a content prompt fed by your universe
The seven-layer content prompt (role, task schema, anti-pattern filters, intent and audience, data inputs, examples, output format) takes its intent and data inputs from the universe. No primary evidence grades its other layers, and the role layer has no accuracy evidence.
The table maps each layer to the universe layer that fills it and to the evidence that exists. "Adjacent" means measured on something else, and "judgment" means ours.
| Content-prompt layer | Filled by | Evidence | Verdict |
|---|---|---|---|
| Role setup | Nothing in the universe | Adjacent: no accuracy gain on average in a test of open-source models on multiple-choice questions; tone untested | Keep one line for voice; do not count on it for correctness |
| Task schema | Layer 5 cluster: the sub-queries become the sections | Judgment: the cluster is the set of queries the engine fans out to, so sections answer what retrieval asks | Build sections from the cluster, not from a format habit |
| Anti-pattern filters | Nothing | None either way; editorial control | Keep a short banned list; it is house style |
| Intent and audience | Layers 1 to 3, verbatim from the cell | Judgment; it is the input the universe exists to give | Pass the cell as written |
| Data inputs | Layer 4 wordings, Layer 5 sub-queries, your own first-party data | Adjacent: statistics, quotations and cited sources added to source pages gave a 30-40% relative gain on one visibility metric; prompts were not tested | Supply real numbers with sources; expect the gain, if any, to come from the page containing them |
| Examples and references | Your best-cited page from the Layer 7 record | None either way | Use one, not five |
| Output and citation format | Layer 5 | Google: no special schema.org markup needed for generative AI search; its chunking line concerns published pages | Ask for the structure the heading's question needs, no format tricks |
The universe fills three layers: intent and audience, task schema and data inputs. Only the data layer has evidence behind it, and that evidence comes from page rewrites, not from prompts.
Role, filters and examples
The role line comes first in almost every version of this prompt. A test of 162 roles on four families of open-source models and 2,410 multiple-choice questions found no accuracy gain over no persona, and the effect of any single persona was largely random. The authors note gains are possible in certain settings. The test did not cover commercial models, open-ended writing or tone.
A later multi-model test, covered in our AI SEO mistakes guide, reaches the same place: keep the persona for tone. So the role line can set voice. Count on nothing more, including the popular claim that it locks an authoritative tone, which no published test checks.
The anti-pattern filter is the same story. A banned list of robotic phrases is editorial control, and it appears as one Constraints line in the filled prompt below. The examples slot takes the page from your Layer 7 record that was cited most. No evidence says how many examples work, so use one.
Data inputs and output format
In Google's guidance, structured data isn't required for generative AI search, and there's no special schema.org markup to add. Its line about ignoring "chunking" is about how you publish pages. Neither says anything about prompt layout. So the output layer asks for the structure the heading's question needs, such as a table for a comparison, and stops there. The passage rules for the written page live in SEO content writing.
The data layer is the one with adjacent evidence. In a KDD 2024 paper on generative engine optimization, rewriting source pages to add statistics, quotations and cited sources moved a visibility metric 30-40% and a subjective impression metric 15-30% relative to no optimization, on a benchmark of 10,000 queries. Keyword stuffing gave little to no improvement.
Two caveats. The paper rewrote pages and did not test prompts, and it measures visibility once a source is retrieved. Our inference: a prompt that supplies real, sourced numbers lets the draft contain the kinds of content that paper found helped. So "add real, sourced numbers" is the instruction to write into that layer.
The filled prompt
The content prompt for the example cell reads as follows. Everything in square brackets is yours to replace:
Role: You are a writer for [brand]; plain, direct, no hype.
Task: Answer this buyer prompt in a 400-word section with an H2 that restates it, an answer in the first sentence, then support. Cover these sub-questions in order: [Layer 5 cluster].
Intent and audience: Founder of a 6-person Shopify brand, no finance hire, asking at the problem stage: "my Shopify payout doesn't match my bank deposit, how do I reconcile it".
Data inputs: Use only these sourced figures and quotes: [paste your own numbers with source and date]. Add nothing else.
Examples: Match the structure of [your best-cited page].
Constraints: Avoid these phrases: [your banned list].
Output: Markdown, one table if a comparison helps, sources named in the sentence.
Your next step this week: write your one buyer task as "do X without Y", multiply it out on paper (stages x personas x wordings), and screen the 8 cells at 5 runs before you write a single page. The tracking set comes first. The content prompt only works once it has a cell to draw from.
Frequently asked questions
Is there an official 7 layer prompt universe framework?
No. There is no shared specification. One vendor publishes a page under this exact name as its own agent-pipeline standard, other pages use the number for content prompts, and buyer-prompt sets are described without seven layers. The stack in this guide is our build order, not a standard.
Why does the same prompt return a different brand list each time?
In SparkToro's study, the chance of the same brand list was under 1 in 100 for ChatGPT and Google's AI, with Claude slightly higher, and about 1 in 1,000 for the same order. The set of brands still overlaps, so measure visibility as a percentage across many runs.
How many prompts and runs does a prompt universe need?
Runs per prompt matter more than prompt count. The example screens 24 prompts at 5 runs per engine and measures anchors at 60. The study averaged about 82 runs per prompt and tool, and its author says at least 60-100.
Should every prompt in the universe get its own page?
No. Cluster prompts to the strongest URL and write a new page only for a distinct user task. Google says creating separate content for every fan-out variation, primarily to manipulate rankings or generative responses, violates its scaled content abuse policy.
Figures and images in this post are free to reuse under CC BY 4.0 with credit to Mission Growth.
Get Mission Growth highlighted in your Google results.

