LLMs can now identify public figures in images
2 days ago
- Multimodal LLMs can extract structured semantic data from images for better categorization, tagging, and searching.
- ChatGPT and Claude refuse to identify public figures like Barack Obama due to AI safety policies and RLHF training, while Gemini has no such hesitation.
- In tests across six LLMs, Gemini consistently outperforms others in accurately identifying public figures, including in challenging scenarios like group photos and costume images.
- Refusals by GPT and Claude to identify public figures can be bypassed using prompt engineering techniques, such as instructing the model to prefix its output.
- The accuracy of LLMs in identifying people varies, with some models hallucinating incorrect identities (e.g., Mistral misidentifying Priscilla Chan as Sheryl Sandberg).
- Gemini's superior performance is attributed to Google's access to vast training data from its search engine, achieving over 90% accuracy in identifying public figures across diverse domains.
- There is an ethical concern about LLMs potentially identifying nonpublic figures in the future as models improve and RLHF rules become laxer.