Three AI agents, two nations, one uneven internet
agents
| Source: Mastodon | Original article
A researcher evaluated three AI agents in the United States and Iran, exposing disparities in the global web’s accessibility and performance.
A technology‑and‑human‑rights researcher has put three leading AI agents to work on a real‑world multilingual data‑update task, revealing stark gaps in how they handle language, source access and safeguards. The experiment compared Meta’s Muse, Anthropic’s Claude Cowork and OpenAI’s GPT model by asking each to refresh World Bank country‑profile statistics for the United States and Iran, in both English and Farsi.
The agents were evaluated on four dimensions: ability to understand and generate text in the two languages, awareness of contextual nuances, transparency about the sources they consulted, and the extent of human‑in‑the‑loop oversight they required. Results showed that while all three could retrieve the basic figures, their performance diverged sharply. The OpenAI model displayed the broadest source coverage but offered limited insight into which databases it consulted. Claude Cowork was more conservative, often refusing to cite non‑English sources, and Muse struggled with Farsi syntax, producing incomplete updates. Across the board, the agents differed in how they enforced content‑policy restrictions, with the Iran‑focused queries encountering tighter permission blocks than the U.S. equivalents.
The findings matter because they expose an “uneven world wide web” where AI tools inherit the geopolitical and linguistic biases of the data ecosystems they draw from. As enterprises and governments increasingly rely on generative agents for cross‑border analytics, disparities in multilingual competence and access rights could reinforce information asymmetries.
Going forward, observers will watch for any policy shifts or technical upgrades that broaden source transparency and multilingual robustness, especially as regulators in both the West and the Middle East consider tighter AI oversight. The next wave of agent evaluations will likely focus on how emerging tool‑centric frameworks, such as those explored in our recent coverage of EgoTools, can mitigate these gaps.
Sources
Back to AIPULSEN