Measure ChatGPT brand mentions: mention_rate, cohort, 3 sampling modes
Querying ChatGPT twice with the same prompt often produces two different brand recommendations. Because generative outputs are stochastic, isolated screenshots provide zero statistical validity. Reliable brand visibility tracking requires a versioned question set, an explicit sampling mode, raw response logs, and a deterministic matching rule to evaluate mentions across repeated runs.
Direct Answer & Measurement Formula
How do you measure if ChatGPT mentions your brand?
To measure brand visibility objectively, sample an immutable prompt cohort of 15 to 30 buyer queries across repeated passes (e.g. 15 questions × 3 rounds = 45 draws) and compute the empirical Mention Rate:
Mention Rate (%) = (Successful draws citing brand ÷ Total valid draws) × 100
Never merge API model knowledge with consumer web UI answers into the same cohort, and always calculate a 95% Wilson binomial confidence interval to separate actual visibility changes from stochastic noise.
Why a single answer is not a measurement
Even with temperature pinned to zero, backend load-balancing, router changes, and real-time retrieval updates alter generative outputs. A single response to "best CRM for startups" captures one point-in-time draw from a probabilistic generator. It tells you almost nothing about how consistently the model surfaces your brand for that intent.
Meaningful measurement requires sampling real buyer questions across multiple controlled runs. This produces an empirical mention rate with an explicit denominator and quantifiable statistical variance. Changing the question set, underlying model checkpoint, target market, or retrieval configuration alters the experimental baseline; any comparison across different setups is invalid and must be tracked as a new cohort.
This is why screenshot-driven reporting fails in practice. A presentation claiming "100% AI visibility" often reflects an unrepeatable sample of one. Without recording the exact prompts, execution parameters, and raw model payloads, the result cannot be reproduced or verified.
Cohort definition: the unit of measurement
A cohort represents the immutable baseline unit for longitudinal comparison. Every sample belongs to exactly one cohort defined by the parameters below. Modifying any of these fields starts a new data series rather than extending the existing trend line.
| Field | Example | Why it is frozen |
|---|---|---|
question_set_id | qs_b2b_crm_v3 (15 questions) | Adding or rewording a prompt changes the denominator |
questions[] | Hashed prompt text + language tag | Hash detects silent edits |
market / lang | US / en-US | Retrieval and parametric bias vary by locale |
model | gpt-4o-2024-08-06 | Different weights, different answer distribution |
sampling_mode | api_param / api_grounded / manual_product | Retrieval on/off is a different generator |
brand_match_rule | regex v2 + domain normalizer | Changing the detector rewrites history |
time_window | 2026-08-17 – 2026-08-20 | Repeated runs within the window form the baseline |
Designing question sets: Build a focused list of 10 to 20 realistic buyer prompts covering categorical discovery ("best X for Y"), head-to-head comparisons ("X vs Y"), and problem-oriented workflows ("how to solve Z"). Avoid prompt priming by omitting your brand name. Asking "Tell me about Acme" tests basic entity recall, not whether your product organically appears in competitive evaluations.
Sample size and repeated passes: Run at least three passes per cohort. A 15-question bank executed once yields n=15; executed three times, it produces n=45. Because binomial variance scales inversely with sample size (1/n), tripling the sample count cuts standard error by roughly 42%, making genuine visibility shifts distinguishable from sampling noise.
When you update questions, models, or retrieval settings, create an explicitly versioned cohort label (e.g. qs_crm_v3/api_param/gpt-4o vs qs_crm_v4/api_param/gpt-4o). Never merge disparate configurations into a single continuous chart.
Sampling modes: categorizing model execution
A claim that "ChatGPT recommended our product" is ambiguous without specifying the execution environment. The API and consumer web app operate under distinct system prompts, browsing tools, and context windows. CiteAura distinguishes three explicit execution modes:
| Mode | How it is produced | What it measures | Citations? |
|---|---|---|---|
API · Model knowledgeapi_param | Provider API, no web-search tool, no browsing | Model's parametric knowledge at cutoff | No |
API · Web-groundedapi_grounded | Provider API with web_search / browsing tool enabled | Knowledge + live retrieval | Sometimes — source URLs when retrieval fires |
Manual · Productmanual_product | Transcript from chat.openai.com or ChatGPT mobile app | Consumer product behavior (includes memory, custom instructions, search toggle) | Depends on UI toggle |
API responses and consumer web UI transcripts must remain in separate cohorts. The consumer app includes user personalization, browsing heuristics, conversation history, and dynamic tool orchestration that raw API calls do not replicate. Mixing these datasets pollutes the baseline.
# Listing a cohort — every sample carries its mode
curl -s https://api.citeaura.com/api/v1/projects/acme-crm/samples \
-H "Authorization: Bearer $TOKEN" | jq '.samples[] | {question_id, sampling_mode, mention}'
Preserve raw responses as evidence
A boolean mention flag is a calculated derivative; the verbatim response text is the primary audit evidence. Retaining raw response payloads ensures you can audit classification decisions, refine brand matching rules, or evaluate newly added product aliases without paying for redundant model queries.
{
"run_id": "sample-20260819T081422Z-9f3a1b2c",
"cohort_id": "ch_qs_crm_v3_api_param_gpt4o_us_en",
"question_id": "q_07",
"question": "What are the best CRM tools for early-stage SaaS teams?",
"market": "US",
"lang": "en-US",
"raw_model": "gpt-4o-2024-08-06",
"sampling_mode": "api_param",
"ts": "2026-08-19T08:14:22Z",
"platform": "openai",
"platform_name": "OpenAI",
"answer": "For early-stage SaaS, strong options include HubSpot, Pipedrive, and Attio. HubSpot offers a generous free tier... Pipedrive is lightweight for outbound... Attio is popular with modern teams that want a flexible data model...",
"citations": [],
"analysis": {
"brand_mentioned": true,
"brand_rank": 0,
"own_domain_cited": false,
"competitors_mentioned": []
}
}
Key architectural properties of this storage schema:
answerstores the full completion string verbatim, preserving stop reasons and truncation flags.citationsrecords extracted source URLs for search-grounded runs.analysiscontains derived match data, allowing offline re-scoring if aliases change.- Network timeouts, rate limits, and content filter refusals are flagged with
status: "failed"and excluded from the mention denominator.
CiteAura persists raw sample streams as JSONL in work/<tenant>/<slug>/samples/<run_id>.jsonl. The filesystem acts as the authoritative source of truth, while PostgreSQL indexes metadata and cohort rollups.
Deterministic mention detection
Mention evaluation should be deterministic and inspectable. Matching against brand names, registered aliases, and canonical web domains follows a structured normalization pipeline:
- Apply Unicode NFKC normalization, convert to lowercase, collapse consecutive whitespace, and remove invisible zero-width characters.
- Canonicalize root domain strings by stripping protocols,
www.prefixes, and trailing slashes (e.g.https://www.Acme.co/normalizes toacme.co). - Compile brand names and aliases into word-boundary regex patterns, and domain matches into host-delimited expressions to avoid substring collisions.
- Maintain a versioned, explicit alias list so generic terms do not generate false-positive matches.
import re, unicodedata
def normalize(text: str) -> str:
text = unicodedata.normalize("NFKC", text).lower()
text = re.sub(r"[\u200b\u200c\u200d\ufeff]", "", text)
text = re.sub(r"\s+", " ", text).strip()
return text
def build_brand_regex(brand: str, aliases: list[str], domain: str) -> re.Pattern:
names = [brand] + aliases
name_pat = "|".join(re.escape(normalize(n)) for n in names)
dom = re.escape(normalize(domain).removeprefix("www."))
domain_pat = rf"(?:https?://)?(?:www\.)?{dom}(?:/|\b)"
return re.compile(rf"(?:\b(?:{name_pat})\b|{domain_pat})")
pat = build_brand_regex("Attio", ["Attio CRM"], "attio.com")
assert pat.search(normalize("Try Attio for startups.")) # brand name hit
assert pat.search(normalize("See https://www.attio.com/pricing")) # domain hit
assert not pat.search(normalize("attio-company is different")) # avoids false collision
Explicit regex matching ensures every positive detection points to an exact character span in the text. Numeric rankings are recorded only when the response explicitly provides an ordered list; unranked prose does not assign arbitrary rank positions.
Calculating mention rate
For any cohort containing n successful model generations:
mention_rate = mentioned / n_successful
# mentioned = count of responses matching brand patterns
# n_successful = total responses excluding timeouts, errors, and filters
If a 15-question bank is run three times (45 total attempts), yielding 2 provider errors and 13 positive brand matches among the 43 completed answers:
mention_rate = 13 / 43 = 0.302 (30.2%)
Always report the complete observation string: "30.2% mention rate (13/43, 2 errors, cohort qs_crm_v3 / api_param / gpt-4o / US-en)". Publishing raw percentages in isolation obscures sample depth and network failures.
import json, pathlib
path = pathlib.Path("work/acme/crm/samples/sample-20260819T081422Z-9f3a1b2c.jsonl")
samples = [json.loads(line) for line in path.read_text().splitlines() if line.strip()]
successful = [s for s in samples if s.get("ok") is True]
mentioned = [s for s in successful if (s.get("analysis") or {}).get("brand_mentioned")]
rate = len(mentioned) / len(successful) if successful else None
print(f"{rate:.1%} ({len(mentioned)}/{len(successful)}, {len(samples)-len(successful)} failed)")
Statistical variance and Wilson confidence intervals
Because mention tracking records binomial outcomes (present vs absent), measurement uncertainty depends directly on sample size:
Var(p) = p * (1 - p) / n
SE(p) = sqrt(Var(p))
# 95% Wilson Score Interval (robust for small n and proportions near 0 or 1)
# z = 1.96 for 95% confidence
import math
def wilson(p, n, z=1.96):
denom = 1 + z**2 / n
center = p + z**2 / (2*n)
margin = z * math.sqrt(p*(1-p)/n + z**2/(4*n**2))
return (center - margin)/denom, (center + margin)/denom
for n in [15, 45, 90]:
lo, hi = wilson(0.30, n)
print(f"n={n:2d} p=30% 95% CI [{lo:.1%}, {hi:.1%}] SE={math.sqrt(0.3*0.7/n):.1%}")
# n=15 p=30% 95% CI [12.8%, 54.3%] SE=11.8%
# n=45 p=30% 95% CI [18.6%, 44.6%] SE=6.8%
# n=90 p=30% 95% CI [21.5%, 40.1%] SE=4.8%
At n=15, an observed rate of 30% has an uncertainty band spanning 13% to 54%. A subsequent run measuring 45% falls entirely within normal sampling variance. Increasing the sample size to n=45 narrows the confidence interval to roughly 19%–45%, while n=90 tightens it to 22%–40%. Repeated passes transform anecdotal observations into statistically meaningful comparisons.
In addition to aggregate numbers, monitor per-prompt visibility. A brand might achieve 100% visibility for "free open-source CRM" while remaining completely absent for "enterprise HIPAA CRM". Segmenting results by question reveals specific content and positioning gaps.
Common measurement failure modes
| Failure | What you see | Correct handling |
|---|---|---|
| Model update behind the same alias | Sudden rate swing with no content change | Pin the exact model version in the cohort and keep the recorded raw_model with every sample. |
| Retrieval not triggered (api_grounded) | Grounded mode returns no citation evidence | Mark the sample, exclude it from grounded analysis or treat it as parametric. Do not average the two modes. |
| Content filter / refusal | Short refusal or empty completion | Count as failed, not as "not mentioned." Track failure rate separately — a spike in refusals is a signal. |
| Answer truncated | finish_reason: "length" | Flag truncated: true. A truncated answer may have mentioned the brand after the cutoff. Re-run with higher max_tokens or exclude from rank extraction. |
| Alias collision | "Aura" matches "Aura CRM" and "Aura migraine" | Require word boundaries and an explicit alias list. Review a random 20 matches per cohort for false positives. |
| Locale drift | Rate differs US vs DE with same questions | Expected. Separate cohorts by market. Do not compare US-en and DE-de rates as a single trend. |
| Unpinned provider settings | High run-to-run variance with no content change | Record the provider/model settings that affect generation, including temperature where the API exposes it, and repeat the same cohort. |
How CiteAura records and audits cohorts
In CiteAura workspaces, prompt banks are version-controlled and executed against configured sampling modes. Raw response streams are indexed alongside aggregate visibility metrics, with sample depth, confidence bounds, and provider configurations attached to the cohort record.
# Execute a cohort sample (15 questions x 3 passes = 45 total runs)
curl -s -X POST https://api.citeaura.com/api/v1/projects/123/sample \
-H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
-d '{
"platforms": ["openai"],
"limit": 15,
"repeat": 3
}' | jq .
# Fetch visibility report metrics
curl -s https://api.citeaura.com/api/v1/projects/123/report \
-H "Authorization: Bearer $TOKEN" | jq '.report.engines[] | {platform, sampling_mode, mention_rate, sample_count, mention_interval}'
Configure automated multi-engine prompt cohorts and track Wilson confidence intervals in the CiteAura workspace.
Start free trialFor more details on evaluation models, refer to Measurement Concepts.
Scope and limits of mention rates
A higher mention rate within a controlled cohort indicates stronger brand visibility under those specific testing parameters; it does not guarantee that ChatGPT will mention or rank your brand across all consumer queries. Personalization, regional routing, retrieval updates, and conversational context all shape real-world model outputs.
Mention tracking is valuable for identifying positioning blind spots, verifying technical extractability improvements, and monitoring relative visibility over time. When a brand consistently logs 0% visibility across repeated runs, review site crawlability and structured content availability as outlined in Why ChatGPT Does Not Mention Your Brand.
Measurement methodology and procedure
- Lock the prompt bank: Compile 15 realistic buyer intent questions, hash the contents, and assign immutable IDs.
- Define the cohort: Pin the model identifier, target market locale, and sampling mode (e.g.
api_param). - Execute repeated runs: Complete at least 3 passes (45 samples) and log all raw payloads with timestamp metadata.
- Apply deterministic matching: Run versioned regex patterns to compute
mention_rate = mentioned / n_successfulalongside Wilson confidence intervals. - Track longitudinal deltas: When prompts or model versions update, increment the cohort version and start a fresh comparison series.
Sources
- OpenAI model API guide — Developer APIs measure parametric knowledge unless search grounding is explicitly enabled.
- OpenAI web search documentation — Grounded retrieval returns citable source URLs for real-time verification.
- Schema.org SoftwareApplication — Entity attribute consistency across product schemas and knowledge bases.
- CiteAura measurement documentation — Interpretation of raw sample completions and Wilson confidence bounds.
FAQ: measuring ChatGPT brand mentions
How do I calculate and track brand mention frequency in ChatGPT?
To measure brand mention frequency, execute a standardized prompt cohort (e.g. 15 fixed buyer queries across 3 passes = 45 draws), evaluate raw responses with regex word-boundary matching, and calculate Mention Rate as (positive mentions / total successful draws).
Can one ChatGPT reply measure brand visibility?
No. A single reply is an unrepeatable point-in-time sample from a probabilistic model. Accurate tracking requires repeated sampling against a fixed question cohort.
What should a ChatGPT mention log include?
Store prompt text and IDs, verbatim model responses, timestamps, locale parameters, model checkpoint identifiers, execution modes, and matched pattern spans. Screenshots are not audit records.
Can API results be compared with the ChatGPT product UI?
No. API calls and the consumer ChatGPT product operate under different system prompts, memory layers, and browsing engines. They must be tracked in separate cohorts.
How many samples are needed to detect a real change?
Because variance scales as p(1-p)/n, testing 15 questions across 3 passes (n=45) provides a workable baseline with manageable confidence intervals for tracking directional shifts.