Hasty Briefsbeta

Bilingual

LLMs can now identify public figures in images

2 days ago
  • Multimodal LLMs can extract structured semantic data from images for better categorization, tagging, and searching.
  • ChatGPT and Claude refuse to identify public figures like Barack Obama due to AI safety policies and RLHF training, while Gemini has no such hesitation.
  • In tests across six LLMs, Gemini consistently outperforms others in accurately identifying public figures, including in challenging scenarios like group photos and costume images.
  • Refusals by GPT and Claude to identify public figures can be bypassed using prompt engineering techniques, such as instructing the model to prefix its output.
  • The accuracy of LLMs in identifying people varies, with some models hallucinating incorrect identities (e.g., Mistral misidentifying Priscilla Chan as Sheryl Sandberg).
  • Gemini's superior performance is attributed to Google's access to vast training data from its search engine, achieving over 90% accuracy in identifying public figures across diverse domains.
  • There is an ethical concern about LLMs potentially identifying nonpublic figures in the future as models improve and RLHF rules become laxer.