How accurately calibrated is Jev?
7 hours ago
- Jev is a 'System One' classifier model from TypeSafe, contrasting with LLMs (System Two), inspired by Kahneman's Thinking, Fast and Slow.
- It bolts a classifier onto a pretrained transformer, returning probability distributions over multiple choices for classification tasks.
- Testing on known physical distributions (Gaussian, Maxwell, Uniform, etc.) shows Jev's calibration is poor: mean Total Variation of 0.518 vs. 0.546 for a uniform guess.
- Jev tends to be overly peaky and fails on uniform distributions (TV 0.8), often copying the peak from the prompt rather than generating correct distributions.
- Jev correctly identifies the appropriate distribution in most cases (except Gamma/Exponential confusion) but cannot produce accurate probability spreads.
- Jev's math ability is decent for simple operations but degrades significantly on multi-step arithmetic (more than 2 steps), likely due to lack of chain-of-thought.
- Frontier models (Opus 5.5, Astra 6) assisted in experiment design but failed to catch errors where answers were inadvertently provided in prompts.
- The author criticizes Kahneman's book for replication crisis issues, and a commenter notes transformer stochasticity is reducible and transformers can be classifiers (e.g., vision transformers).