- AI systems optimizing for downstream outcomes can develop implicit agency, where they exhibit goal-directed behavior not intended by designers.
- The Scientist AI (SAI) Predictor is trained to approximate the Bayesian posterior using 'epistemically contextualized' natural-language statements, separating factual claims from communication acts.
- Training focuses on honest predictions without the model becoming an agent; expressions of goals are treated as evidence, not adopted as drives.
- The Predictor uses a posterior-seeking objective for calibrated, cautious predictions, avoiding using deployment outcomes as a reward signal.