Inspect: An open-source framework for large language model evaluations
6 hours ago
- Inspect is an open-source framework for frontier AI evaluations developed by the UK AI Security Institute and Meridian Labs.
- It provides composable building blocks (datasets, agents, tools, scorers) for easy evaluation creation and reuse.
- Includes over 200 pre-built evaluations ready to run on any model.
- Offers extensive tooling such as Inspect View for monitoring and a VS Code Extension for authoring and debugging.
- Supports flexible tool calling with custom, MCP, and built-in tools (bash, python, web search, etc.).
- Enables agent evaluations with built-in agents, multi-agent primitives, and support for external agents like Claude Code.
- Features a sandboxing system for running untrusted model code in Docker, Kubernetes, and other environments.
- Provides Python API and comprehensive documentation for creating and running evaluations.