The famous o3 "GeoGuessr" prompt did not work
2 days ago
- The elaborate 'GeoGuessr' prompt for OpenAI's o3 model did not improve geolocation performance compared to a simple default prompt.
- A benchmark of 200 images showed the default prompt performed slightly better in median and mean distance from actual location.
- The author warns that prompt engineering can create a false sense of improvement, as models often rationalize their own mistakes.
- The benchmark was easy and cheap to run ($15, six hours), yet no one had tested the prompt's effectiveness during the initial hype.
- o3's geolocation abilities have not transferred to newer models like GPT-5.4 and GPT-5.5, which performed worse in the same benchmark.
- The author addressed concerns about images being in training data, noting that public domain images still allow useful comparisons.