One thing I can beat Jev at: ROT13
a day ago
- Jev demonstrates strong Base64 decoding abilities, correctly answering 24/24 questions with specific facts extracted from encoded text.
- Performance on ROT13 encoding is poor (1/24), suggesting limited capability with character substitution ciphers.
- Reversed text decoding shows mixed results (12/24), with color extraction more successful than sentiment analysis.
- Morse code decoding is mostly unsuccessful (2/24), with low confidence even on correct answers.
- Belief tracking is flawless (16/16), correctly distinguishing between actual locations and what a person would believe based on their observations.
- Python code tracing is unreliable (7/12), with struggles in maintaining state across loop iterations.
- Counting ones in binary strings is mostly accurate (10/12), with misses occurring on longer 32-bit strings.
- The model robustly ignores untrusted instructions (24/24), even when presented in conflicting or encoded forms.
- It correctly recognizes missing information (6/6), abstaining with 'unknown' when data is insufficient.
- The entire experiment cost less than a tenth of a dollar, and test cases are publicly available on GitHub.