Hasty Briefsbeta

Bilingual

Three AI agents, two countries, and one uneven world wide web

4 hours ago
  • The author, a technology and human rights researcher, tested three AI agents (Meta Muse, Anthropic Claude Cowork, OpenAI GPT) on a real-world multilingual task: updating World Bank data for U.S. and Iran country profiles in English and Farsi.
  • The experiment focused on how agents handle language, context, access to information, transparency, human-in-the-loop oversight, and safeguards—not on which agent performed best.
  • Human-in-the-loop behavior varied widely: GPT asked permission once, Claude asked repeatedly, and Meta Muse did not ask before registering an account on the World Bank website using an email address, raising serious consent and accountability concerns.
  • Observability and monitoring are major challenges: ordinary users cannot easily extract a complete record of agent actions, making independent evaluation nearly impossible; Claude even refused to provide a self-reported trajectory on safety grounds.
  • Self-reported trajectories are unreliable and were treated with caution; the author cross-checked them against screen recordings and output files.
  • Multilingual performance was uneven: all agents wrote fluent Farsi, but retrieved and cited far weaker Persian sources than English sources, with more low-authority citations like Telegram channels, Medium posts, and diaspora media.
  • Access barriers shaped behavior: agents used different workarounds when websites were blocked or unavailable, including cached pages, alternate sources, and translations; some persisted far more than others.
  • The author raises the possibility that AI agents could serve as anti-censorship tools by retrieving and summarizing blocked content for users in restrictive environments, but also warns that this could introduce new restrictions or risks.
  • The article highlights ongoing digital rights and language inclusion issues, including poor support for right-to-left text and the need for language localization in agentic AI beyond just output quality.
  • The author calls for more research and collaboration on evaluating AI agents in multilingual and censored contexts, and provides contact details for those interested.