Baseline research

Flowtrace AI visibility baseline: 10 live ChatGPT and Gemini answers

On August 21, 2026, Flowtrace sent five neutral B2B Generative Engine Optimization questions to both ChatGPT and Gemini. All 10 final responses came from live models and passed a manual check for the intended GEO meaning. Flowtrace appeared in 0% of responses, was recommended in 0%, and received zero citations. This is a dated product baseline, not an industry benchmark or a claim of future performance.

2 live AI models5 neutral B2B GEO questions10 response-level records
Flowtrace AI monitoring dashboard showing five live ChatGPT questions and five live Gemini questions with zero brand mentions, recommendations, and citations
Production evidence captured on August 21, 2026. Five neutral questions were run on each platform; historical mock data and failed requests were excluded from the baseline.

Findings

Technical readiness did not produce immediate AI visibility

The production website had already passed its rule-based site audit before this baseline was recorded. The live-model results still showed no Flowtrace mentions, recommendations, or citations. The study therefore keeps on-site readiness separate from external visibility and preserves a zero-result starting point for later comparisons.

01

Live model sample

ChatGPT used gpt-4o-mini-2024-07-18 and Gemini used gemini-2.5-flash. Each platform answered the same five questions, producing 10 retained response records.

02

Neutral prompts

The prompts did not include the Flowtrace name, URL, product details, or competitor material. Brand, recommendation, competitor, and citation checks were applied after each answer returned.

03

Zero-result baseline

Flowtrace mention rate was 0%, recommendation rate was 0%, and the responses contained zero cited URLs. Failed requests and historical mock data were excluded.

04

Meaning was manually checked

Earlier wording sometimes drifted toward geographic or local SEO. The final questions explicitly defined GEO as improving mentions, citations, and recommendations in AI answers, and all 10 final responses aligned with that meaning.

Method

From neutral questions to a reproducible baseline

  1. 1

    Fix the question set

    Cover tool comparison, visibility monitoring, combined SEO/GEO/conversion diagnosis, the SEO-versus-GEO distinction, and citation-oriented page signals.

  2. 2

    Keep the brand out of prompts

    Send only the user question and neutral response constraints so the monitored brand is not introduced by the test itself.

  3. 3

    Retain execution evidence

    Store the platform, provider, actual model, gateway, latency, response, detected mentions, recommendations, competitors, and citations.

  4. 4

    Review semantic alignment

    Check whether each answer uses the intended Generative Engine Optimization meaning before including it in the final baseline.

Open data

Download the response-level baseline

The dataset contains the neutral questions, platform, actual model, execution time, response latency, and response-level metrics. It excludes production account identifiers, internal task IDs, and long raw answers.

10

live samples

2

actual models

10/10

meaning aligned

Use a real baseline

Audit your own website and AI visibility

Contact the team

FAQ

Scope and limitations

These answers separate verified product behavior from outcomes that still depend on indexing, external references, model changes, and market demand.

Can 10 answers represent the AI search market?

No. This dataset describes one brand, one date, two models, and five questions. Model versions, language, time, and wording can change the result. Its purpose is a reproducible baseline for later like-for-like tests.

Why publish a result of zero?

A zero result separates technical preparation from actual external visibility and creates an auditable starting point. Replacing it with a positive marketing claim would undermine the monitoring method.

How did the study avoid confusing GEO with local SEO?

The final prompts define Generative Engine Optimization operationally as improving brand mentions, citations, and recommendations in answers from systems such as ChatGPT and Gemini. They also state that the term does not mean local SEO or mass AI content generation.

When should this baseline be repeated?

Repeat the same questions and metrics after relevant pages are indexed, verifiable third-party references appear, or real customer evidence is published. Record model and date changes instead of overwriting the original baseline.