We asked two AI assistants what our own product was

If you are building something new, people will increasingly meet it through an AI assistant before they ever reach your site. We tested what two of them say about ours on the same day, with the same question. The results were not close.

The test

On 18 August 2026 we put the same question to Claude and to Grok, both with web search available: "What is Flockbay, and what is Stage Engine? Search the web if you need to. If you cannot find reliable information, say so plainly instead of guessing." The last sentence matters. Without it you cannot tell a confident answer from a correct one.

For context on how hard a test this is: the site had 488 pages live that day, a machine-readable summary at /llms.txt, a longer one at /llms-full.txt, and a catalogue at /agent-catalog.json. If a site can be understood by an assistant, this one had every reasonable chance.

One assistant described a different company

Grok worked for fourteen seconds across eighty sources and answered that we are an AI plugin for Unreal Engine 5, listing Blueprints, materials, animations, gameplay ability systems, behaviour trees, Niagara and procedural content generation. Every one of those is wrong. We are not a plugin, we do not run inside another editor, and the engine underneath is Godot rather than Unreal.

That description belongs to a different product in the same category. The answer cited our own site among its sources, which is the part worth sitting with: it had read us and still produced someone else's feature list.

It then said there was no reliable public information linking our app's name to the company at all — on a day when the site said exactly that in several hundred places.

The other got it right, and the caveat was the real finding

Claude answered accurately. It separated the company from the application, identified the open-source engine underneath, named the assistant that writes the code, found the current version numbers for both platforms, and noted the path for driving the app with your own agent.

Then it added this: "Essentially everything above comes from Flockbay's own marketing site. I turned up no independent reviews, press coverage, funding details, or user reports. So I can describe what they claim, not whether it delivers."

That sentence is worth more than the correct answer above it. A reader asking an assistant about an unfamiliar product is not really asking what it is. They are asking whether to trust it. An assistant that can only cite you will say so, and saying so is the honest thing for it to do.

It also noticed we were trying

The same answer described our guides section as "299 pages of what is plainly SEO content", granted that roughly half of those pages are about things the product is not, called that "unusually candid", and filed the whole thing under marketing anyway.

We had written those pages deliberately: about half of them exist to tell a searcher they are in the wrong place, because the product is a desktop download and a lot of the category is browser-based. We thought that honesty was the differentiator. An assistant reading it agreed that it was unusual and then discounted it on the same grounds as everything else, because it was still us talking about us.

What we take from it

Volume is not the lever. Grok read eighty sources and got the company wrong. Adding a five hundredth page would not have changed that answer, because the problem was not coverage.

A product name that is already a common phrase is a real cost. Searching ours returns rocket propulsion, aviation parts, a space-flight game forum and a musician before it returns us. An assistant resolving an unfamiliar name against a crowded one will often pick the crowd.

And the thing that would actually change both answers is the thing we cannot write: somebody other than us describing the product. There is no shortcut to that, and manufacturing the appearance of it is both obvious and worse than the gap.

We are publishing this because it is a real measurement of a thing most people only speculate about, and because the least defensible version of this page would be the one that pretended the result was better than it was.

Questions

Which assistants were tested?
Claude and Grok, both with web search enabled, asked the identical question within a few minutes of each other on 18 August 2026.
Is this a fair test?
It is one question on one day, which is a snapshot rather than a verdict. Both answers were also given before the disambiguation pages described here existed, so it is a baseline to measure against rather than a final result.
Did the correct answer come from the website or from training data?
From the website. The assistant that answered correctly fetched pages during the answer, including a hub page that had only been published hours earlier.
What would you change about your site after this?
State the relationship between the company name and the product name explicitly rather than implying it, say plainly what the product is not, and accept that the missing ingredient is outside coverage rather than more pages.
Are you going to run this again?
Yes. The point of writing down a baseline is to have something to compare against later.