- AI guardrails need the same level of evaluation as the models they govern, moving from static policies to context- and language-specific evaluations.
- Evaluation shapes guardrails by identifying specific failures, such as harmful content, enabling the design of targeted guardrails.
- Open-source guardrails and policy-prompt systems allow independent evaluation, shifting from proprietary classifiers to dynamic policies.
- A community-informed evaluation of 120 refugee- and asylum-focused scenarios across five languages revealed recurring failures, leading to open-source data and concrete guardrail policies.
- Agentic guardrails with tools like web search improve reliability by verifying facts, though effectiveness depends on the underlying LLM.
- Mozilla AI's open-source tools, any-guardrail and Otari, enable configurable guardrails and seamless LLM switching for practical deployment.