Hasty Briefsbeta

Bilingual

Inspect: An open-source framework for large language model evaluations

6 hours ago
  • Inspect is an open-source framework for frontier AI evaluations developed by the UK AI Security Institute and Meridian Labs.
  • It provides composable building blocks (datasets, agents, tools, scorers) for easy evaluation creation and reuse.
  • Includes over 200 pre-built evaluations ready to run on any model.
  • Offers extensive tooling such as Inspect View for monitoring and a VS Code Extension for authoring and debugging.
  • Supports flexible tool calling with custom, MCP, and built-in tools (bash, python, web search, etc.).
  • Enables agent evaluations with built-in agents, multi-agent primitives, and support for external agents like Claude Code.
  • Features a sandboxing system for running untrusted model code in Docker, Kubernetes, and other environments.
  • Provides Python API and comprehensive documentation for creating and running evaluations.