19 days ago
- Language models used in coding agents internally represent program properties like correctness and test outcomes.
- Probes on hidden states can decode program states and even predict future edit outcomes up to 25 steps ahead.
- The latent programming horizon concept suggests models have internal representations that anticipate future code changes.
- Probes transfer across benchmarks without retraining, showing external validity and encouraging interpretability research.