Scaling Agentic RL: 365,000 Environments for SWE, Terminal, and Search
15 hours ago
- Prime Intellect integrates 23 agentic tasksets (365,000+ tasks) across software engineering, terminal, and search domains into a unified API.
- The platform provides a single command for evals and RL training, with verifiers v1 decomposing environments into taskset, harness, and runtime layers.
- All tasksets are validated via gold-patch and no-op checks, with cleaned re-uploads and transparent exclusion lists to ensure reliable reward signals.
- Grading material is withheld from the agent until scoring time to prevent reward hacking, with future plans for isolated grading sandboxes.
- The project includes detailed breakdowns of SWE, terminal, and search tasksets, each with their own integrations and validation methodologies.