I Tested AI Engines on My Own Sites—None Approved
claude open-source training
| Source: Dev.to | Original article
A developer's test of five AI engines on his own websites revealed no consensus, with the open‑source LLM visibility checker showing zero overlap among the models.
A recent experiment by a developer on the DEV Community platform shows that five leading AI search engines can’t agree on the visibility of a single website. The author ran his open‑source LLM visibility checker against his own domains and compared the results from Claude, ChatGPT and three other unnamed engines. Two of the tools produced the only citations, yet the domains they referenced did not overlap at all. Claude, running on its Sonnet 5 model, returned no results for either site, a pattern the author attributes to the domains not appearing in Claude’s training data. Had he relied on Claude alone—as he did in a July version of the tool—he would have concluded the sites were completely invisible to AI.
The findings echo a May 2026 analysis titled “One Question, Five Engines: Why AI Answers Differ About You,” which warned that AI search tools can be “confidently wrong” at widely varying rates, making any single engine an unreliable oracle for business information. A separate log from a week ago highlighted that three of four tested engines still described the recently acquired company Reforge as independent, underscoring how quickly outdated facts can persist across models.
Why it matters is clear: marketers, SEO professionals and anyone relying on AI‑driven discovery are faced with contradictory signals. A study of 62 queries published four weeks ago found that while engines agreed on recommended brands 85 % of the time, they shared a mere 2.7 % of cited source domains, suggesting that the underlying evidence for answers is highly fragmented. The broader community is already taking note—Reddit users pointed to a Tow Center for Digital Journalism study that recorded a 60 % error rate across eight AI search services.
Going forward, observers will watch for two developments. First, whether AI providers improve citation consistency and data freshness, perhaps through shared indexing standards. Second, how businesses adapt their digital strategies if AI search continues to deliver divergent, and sometimes inaccurate, portrayals of their online presence. The experiment adds fresh urgency to calls for greater transparency and accountability in AI‑powered search.
Sources
Back to AIPULSEN