Measurement — Engineering Guide

Measure ChatGPT brand mentions: mention_rate, cohort, 3 sampling modes

Querying ChatGPT twice with the same prompt often produces two different brand recommendations. Because generative outputs are stochastic, isolated screenshots provide zero statistical validity. Reliable brand visibility tracking requires a versioned question set, an explicit sampling mode, raw response logs, and a deterministic matching rule to evaluate mentions across repeated runs.

Updated 14 September 2026 · 12 min read · Technical

Editorial Transparency: Fact-checked against search engine specifications and empirical crawling tests. Our Editorial Standards →

Direct Answer & Measurement Formula

How do you measure if ChatGPT mentions your brand?

To measure brand visibility objectively, sample an immutable prompt cohort of 15 to 30 buyer queries across repeated passes (e.g. 15 questions × 3 rounds = 45 draws) and compute the empirical Mention Rate:

Mention Rate (%) = (Successful draws citing brand ÷ Total valid draws) × 100

Never merge API model knowledge with consumer web UI answers into the same cohort, and always calculate a 95% Wilson binomial confidence interval to separate actual visibility changes from stochastic noise.

Why a single answer is not a measurement

Even with temperature pinned to zero, backend load-balancing, router changes, and real-time retrieval updates alter generative outputs. A single response to "best CRM for startups" captures one point-in-time draw from a probabilistic generator. It tells you almost nothing about how consistently the model surfaces your brand for that intent.

Meaningful measurement requires sampling real buyer questions across multiple controlled runs. This produces an empirical mention rate with an explicit denominator and quantifiable statistical variance. Changing the question set, underlying model checkpoint, target market, or retrieval configuration alters the experimental baseline; any comparison across different setups is invalid and must be tracked as a new cohort.

This is why screenshot-driven reporting fails in practice. A presentation claiming "100% AI visibility" often reflects an unrepeatable sample of one. Without recording the exact prompts, execution parameters, and raw model payloads, the result cannot be reproduced or verified.

Cohort definition: the unit of measurement

A cohort represents the immutable baseline unit for longitudinal comparison. Every sample belongs to exactly one cohort defined by the parameters below. Modifying any of these fields starts a new data series rather than extending the existing trend line.

FieldExampleWhy it is frozen
question_set_idqs_b2b_crm_v3 (15 questions)Adding or rewording a prompt changes the denominator
questions[]Hashed prompt text + language tagHash detects silent edits
market / langUS / en-USRetrieval and parametric bias vary by locale
modelgpt-4o-2024-08-06Different weights, different answer distribution
sampling_modeapi_param / api_grounded / manual_productRetrieval on/off is a different generator
brand_match_ruleregex v2 + domain normalizerChanging the detector rewrites history
time_window2026-08-17 – 2026-08-20Repeated runs within the window form the baseline

Designing question sets: Build a focused list of 10 to 20 realistic buyer prompts covering categorical discovery ("best X for Y"), head-to-head comparisons ("X vs Y"), and problem-oriented workflows ("how to solve Z"). Avoid prompt priming by omitting your brand name. Asking "Tell me about Acme" tests basic entity recall, not whether your product organically appears in competitive evaluations.

Sample size and repeated passes: Run at least three passes per cohort. A 15-question bank executed once yields n=15; executed three times, it produces n=45. Because binomial variance scales inversely with sample size (1/n), tripling the sample count cuts standard error by roughly 42%, making genuine visibility shifts distinguishable from sampling noise.

When you update questions, models, or retrieval settings, create an explicitly versioned cohort label (e.g. qs_crm_v3/api_param/gpt-4o vs qs_crm_v4/api_param/gpt-4o). Never merge disparate configurations into a single continuous chart.

Sampling modes: categorizing model execution

A claim that "ChatGPT recommended our product" is ambiguous without specifying the execution environment. The API and consumer web app operate under distinct system prompts, browsing tools, and context windows. CiteAura distinguishes three explicit execution modes:

ModeHow it is producedWhat it measuresCitations?
API · Model knowledge
api_param
Provider API, no web-search tool, no browsingModel's parametric knowledge at cutoffNo
API · Web-grounded
api_grounded
Provider API with web_search / browsing tool enabledKnowledge + live retrievalSometimes — source URLs when retrieval fires
Manual · Product
manual_product
Transcript from chat.openai.com or ChatGPT mobile appConsumer product behavior (includes memory, custom instructions, search toggle)Depends on UI toggle

API responses and consumer web UI transcripts must remain in separate cohorts. The consumer app includes user personalization, browsing heuristics, conversation history, and dynamic tool orchestration that raw API calls do not replicate. Mixing these datasets pollutes the baseline.

# Listing a cohort — every sample carries its mode
curl -s https://api.citeaura.com/api/v1/projects/acme-crm/samples \
  -H "Authorization: Bearer $TOKEN" | jq '.samples[] | {question_id, sampling_mode, mention}'

Preserve raw responses as evidence

A boolean mention flag is a calculated derivative; the verbatim response text is the primary audit evidence. Retaining raw response payloads ensures you can audit classification decisions, refine brand matching rules, or evaluate newly added product aliases without paying for redundant model queries.

{
  "run_id": "sample-20260819T081422Z-9f3a1b2c",
  "cohort_id": "ch_qs_crm_v3_api_param_gpt4o_us_en",
  "question_id": "q_07",
  "question": "What are the best CRM tools for early-stage SaaS teams?",
  "market": "US",
  "lang": "en-US",
  "raw_model": "gpt-4o-2024-08-06",
  "sampling_mode": "api_param",
  "ts": "2026-08-19T08:14:22Z",
  "platform": "openai",
  "platform_name": "OpenAI",
  "answer": "For early-stage SaaS, strong options include HubSpot, Pipedrive, and Attio. HubSpot offers a generous free tier... Pipedrive is lightweight for outbound... Attio is popular with modern teams that want a flexible data model...",
  "citations": [],
  "analysis": {
    "brand_mentioned": true,
    "brand_rank": 0,
    "own_domain_cited": false,
    "competitors_mentioned": []
  }
}

Key architectural properties of this storage schema:

CiteAura persists raw sample streams as JSONL in work/<tenant>/<slug>/samples/<run_id>.jsonl. The filesystem acts as the authoritative source of truth, while PostgreSQL indexes metadata and cohort rollups.

Deterministic mention detection

Mention evaluation should be deterministic and inspectable. Matching against brand names, registered aliases, and canonical web domains follows a structured normalization pipeline:

  1. Apply Unicode NFKC normalization, convert to lowercase, collapse consecutive whitespace, and remove invisible zero-width characters.
  2. Canonicalize root domain strings by stripping protocols, www. prefixes, and trailing slashes (e.g. https://www.Acme.co/ normalizes to acme.co).
  3. Compile brand names and aliases into word-boundary regex patterns, and domain matches into host-delimited expressions to avoid substring collisions.
  4. Maintain a versioned, explicit alias list so generic terms do not generate false-positive matches.
import re, unicodedata

def normalize(text: str) -> str:
    text = unicodedata.normalize("NFKC", text).lower()
    text = re.sub(r"[\u200b\u200c\u200d\ufeff]", "", text)
    text = re.sub(r"\s+", " ", text).strip()
    return text

def build_brand_regex(brand: str, aliases: list[str], domain: str) -> re.Pattern:
    names = [brand] + aliases
    name_pat = "|".join(re.escape(normalize(n)) for n in names)
    dom = re.escape(normalize(domain).removeprefix("www."))
    domain_pat = rf"(?:https?://)?(?:www\.)?{dom}(?:/|\b)"
    return re.compile(rf"(?:\b(?:{name_pat})\b|{domain_pat})")

pat = build_brand_regex("Attio", ["Attio CRM"], "attio.com")
assert pat.search(normalize("Try Attio for startups."))  # brand name hit
assert pat.search(normalize("See https://www.attio.com/pricing"))  # domain hit
assert not pat.search(normalize("attio-company is different"))  # avoids false collision

Explicit regex matching ensures every positive detection points to an exact character span in the text. Numeric rankings are recorded only when the response explicitly provides an ordered list; unranked prose does not assign arbitrary rank positions.

Calculating mention rate

For any cohort containing n successful model generations:

mention_rate = mentioned / n_successful
# mentioned    = count of responses matching brand patterns
# n_successful = total responses excluding timeouts, errors, and filters

If a 15-question bank is run three times (45 total attempts), yielding 2 provider errors and 13 positive brand matches among the 43 completed answers:

mention_rate = 13 / 43 = 0.302 (30.2%)

Always report the complete observation string: "30.2% mention rate (13/43, 2 errors, cohort qs_crm_v3 / api_param / gpt-4o / US-en)". Publishing raw percentages in isolation obscures sample depth and network failures.

import json, pathlib

path = pathlib.Path("work/acme/crm/samples/sample-20260819T081422Z-9f3a1b2c.jsonl")
samples = [json.loads(line) for line in path.read_text().splitlines() if line.strip()]
successful = [s for s in samples if s.get("ok") is True]
mentioned = [s for s in successful if (s.get("analysis") or {}).get("brand_mentioned")]
rate = len(mentioned) / len(successful) if successful else None
print(f"{rate:.1%} ({len(mentioned)}/{len(successful)}, {len(samples)-len(successful)} failed)")

Statistical variance and Wilson confidence intervals

Because mention tracking records binomial outcomes (present vs absent), measurement uncertainty depends directly on sample size:

Var(p) = p * (1 - p) / n
SE(p)  = sqrt(Var(p))

# 95% Wilson Score Interval (robust for small n and proportions near 0 or 1)
# z = 1.96 for 95% confidence
import math

def wilson(p, n, z=1.96):
    denom = 1 + z**2 / n
    center = p + z**2 / (2*n)
    margin = z * math.sqrt(p*(1-p)/n + z**2/(4*n**2))
    return (center - margin)/denom, (center + margin)/denom

for n in [15, 45, 90]:
    lo, hi = wilson(0.30, n)
    print(f"n={n:2d}  p=30%  95% CI [{lo:.1%}, {hi:.1%}]  SE={math.sqrt(0.3*0.7/n):.1%}")

# n=15  p=30%  95% CI [12.8%, 54.3%]  SE=11.8%
# n=45  p=30%  95% CI [18.6%, 44.6%]  SE=6.8%
# n=90  p=30%  95% CI [21.5%, 40.1%]  SE=4.8%

At n=15, an observed rate of 30% has an uncertainty band spanning 13% to 54%. A subsequent run measuring 45% falls entirely within normal sampling variance. Increasing the sample size to n=45 narrows the confidence interval to roughly 19%–45%, while n=90 tightens it to 22%–40%. Repeated passes transform anecdotal observations into statistically meaningful comparisons.

In addition to aggregate numbers, monitor per-prompt visibility. A brand might achieve 100% visibility for "free open-source CRM" while remaining completely absent for "enterprise HIPAA CRM". Segmenting results by question reveals specific content and positioning gaps.

Common measurement failure modes

FailureWhat you seeCorrect handling
Model update behind the same aliasSudden rate swing with no content changePin the exact model version in the cohort and keep the recorded raw_model with every sample.
Retrieval not triggered (api_grounded)Grounded mode returns no citation evidenceMark the sample, exclude it from grounded analysis or treat it as parametric. Do not average the two modes.
Content filter / refusalShort refusal or empty completionCount as failed, not as "not mentioned." Track failure rate separately — a spike in refusals is a signal.
Answer truncatedfinish_reason: "length"Flag truncated: true. A truncated answer may have mentioned the brand after the cutoff. Re-run with higher max_tokens or exclude from rank extraction.
Alias collision"Aura" matches "Aura CRM" and "Aura migraine"Require word boundaries and an explicit alias list. Review a random 20 matches per cohort for false positives.
Locale driftRate differs US vs DE with same questionsExpected. Separate cohorts by market. Do not compare US-en and DE-de rates as a single trend.
Unpinned provider settingsHigh run-to-run variance with no content changeRecord the provider/model settings that affect generation, including temperature where the API exposes it, and repeat the same cohort.

How CiteAura records and audits cohorts

In CiteAura workspaces, prompt banks are version-controlled and executed against configured sampling modes. Raw response streams are indexed alongside aggregate visibility metrics, with sample depth, confidence bounds, and provider configurations attached to the cohort record.

# Execute a cohort sample (15 questions x 3 passes = 45 total runs)
curl -s -X POST https://api.citeaura.com/api/v1/projects/123/sample \
  -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{
    "platforms": ["openai"],
    "limit": 15,
    "repeat": 3
  }' | jq .

# Fetch visibility report metrics
curl -s https://api.citeaura.com/api/v1/projects/123/report \
  -H "Authorization: Bearer $TOKEN" | jq '.report.engines[] | {platform, sampling_mode, mention_rate, sample_count, mention_interval}'

Configure automated multi-engine prompt cohorts and track Wilson confidence intervals in the CiteAura workspace.

Start free trial

For more details on evaluation models, refer to Measurement Concepts.

Scope and limits of mention rates

A higher mention rate within a controlled cohort indicates stronger brand visibility under those specific testing parameters; it does not guarantee that ChatGPT will mention or rank your brand across all consumer queries. Personalization, regional routing, retrieval updates, and conversational context all shape real-world model outputs.

Mention tracking is valuable for identifying positioning blind spots, verifying technical extractability improvements, and monitoring relative visibility over time. When a brand consistently logs 0% visibility across repeated runs, review site crawlability and structured content availability as outlined in Why ChatGPT Does Not Mention Your Brand.

Measurement methodology and procedure

  1. Lock the prompt bank: Compile 15 realistic buyer intent questions, hash the contents, and assign immutable IDs.
  2. Define the cohort: Pin the model identifier, target market locale, and sampling mode (e.g. api_param).
  3. Execute repeated runs: Complete at least 3 passes (45 samples) and log all raw payloads with timestamp metadata.
  4. Apply deterministic matching: Run versioned regex patterns to compute mention_rate = mentioned / n_successful alongside Wilson confidence intervals.
  5. Track longitudinal deltas: When prompts or model versions update, increment the cohort version and start a fresh comparison series.

Sources

FAQ: measuring ChatGPT brand mentions

How do I calculate and track brand mention frequency in ChatGPT?

To measure brand mention frequency, execute a standardized prompt cohort (e.g. 15 fixed buyer queries across 3 passes = 45 draws), evaluate raw responses with regex word-boundary matching, and calculate Mention Rate as (positive mentions / total successful draws).

Can one ChatGPT reply measure brand visibility?

No. A single reply is an unrepeatable point-in-time sample from a probabilistic model. Accurate tracking requires repeated sampling against a fixed question cohort.

What should a ChatGPT mention log include?

Store prompt text and IDs, verbatim model responses, timestamps, locale parameters, model checkpoint identifiers, execution modes, and matched pattern spans. Screenshots are not audit records.

Can API results be compared with the ChatGPT product UI?

No. API calls and the consumer ChatGPT product operate under different system prompts, memory layers, and browsing engines. They must be tracked in separate cohorts.

How many samples are needed to detect a real change?

Because variance scales as p(1-p)/n, testing 15 questions across 3 passes (n=45) provides a workable baseline with manageable confidence intervals for tracking directional shifts.

About the Author: CiteAura Research Team

Search Systems Engineers & Evaluation Analysts

The CiteAura Research Team is the engineering working group behind the open-source CiteAura GEO engine, empirical crawler audits, and verifiable citation benchmarks across generative AI search systems.

Explore Team & Research → Editorial Standards →