Baseline research
Flowtrace AI visibility baseline: 10 live ChatGPT and Gemini answers
On August 21, 2026, Flowtrace sent five neutral B2B Generative Engine Optimization questions to both ChatGPT and Gemini. All 10 final responses came from live models and passed a manual check for the intended GEO meaning. Flowtrace appeared in 0% of responses, was recommended in 0%, and received zero citations. This is a dated product baseline, not an industry benchmark or a claim of future performance.

Findings
Technical readiness did not produce immediate AI visibility
The production website had already passed its rule-based site audit before this baseline was recorded. The live-model results still showed no Flowtrace mentions, recommendations, or citations. The study therefore keeps on-site readiness separate from external visibility and preserves a zero-result starting point for later comparisons.
Live model sample
ChatGPT used gpt-4o-mini-2024-07-18 and Gemini used gemini-2.5-flash. Each platform answered the same five questions, producing 10 retained response records.
Neutral prompts
The prompts did not include the Flowtrace name, URL, product details, or competitor material. Brand, recommendation, competitor, and citation checks were applied after each answer returned.
Zero-result baseline
Flowtrace mention rate was 0%, recommendation rate was 0%, and the responses contained zero cited URLs. Failed requests and historical mock data were excluded.
Meaning was manually checked
Earlier wording sometimes drifted toward geographic or local SEO. The final questions explicitly defined GEO as improving mentions, citations, and recommendations in AI answers, and all 10 final responses aligned with that meaning.
Method
From neutral questions to a reproducible baseline
- 1
Fix the question set
Cover tool comparison, visibility monitoring, combined SEO/GEO/conversion diagnosis, the SEO-versus-GEO distinction, and citation-oriented page signals.
- 2
Keep the brand out of prompts
Send only the user question and neutral response constraints so the monitored brand is not introduced by the test itself.
- 3
Retain execution evidence
Store the platform, provider, actual model, gateway, latency, response, detected mentions, recommendations, competitors, and citations.
- 4
Review semantic alignment
Check whether each answer uses the intended Generative Engine Optimization meaning before including it in the final baseline.
Open data
Download the response-level baseline
The dataset contains the neutral questions, platform, actual model, execution time, response latency, and response-level metrics. It excludes production account identifiers, internal task IDs, and long raw answers.
Use a real baseline
Audit your own website and AI visibility
FAQ
Scope and limitations
These answers separate verified product behavior from outcomes that still depend on indexing, external references, model changes, and market demand.
Can 10 answers represent the AI search market?
No. This dataset describes one brand, one date, two models, and five questions. Model versions, language, time, and wording can change the result. Its purpose is a reproducible baseline for later like-for-like tests.
Why publish a result of zero?
A zero result separates technical preparation from actual external visibility and creates an auditable starting point. Replacing it with a positive marketing claim would undermine the monitoring method.
How did the study avoid confusing GEO with local SEO?
The final prompts define Generative Engine Optimization operationally as improving brand mentions, citations, and recommendations in answers from systems such as ChatGPT and Gemini. They also state that the term does not mean local SEO or mass AI content generation.
When should this baseline be repeated?
Repeat the same questions and metrics after relevant pages are indexed, verifiable third-party references appear, or real customer evidence is published. Record model and date changes instead of overwriting the original baseline.