Hasty Briefsbeta

Bilingual

From Evaluation to Guardrails: What We Brought to ACM FAccT 2026

3 hours ago
  • AI guardrails need the same level of evaluation as the models they govern, moving from static policies to context- and language-specific evaluations.
  • Evaluation shapes guardrails by identifying specific failures, such as harmful content, enabling the design of targeted guardrails.
  • Open-source guardrails and policy-prompt systems allow independent evaluation, shifting from proprietary classifiers to dynamic policies.
  • A community-informed evaluation of 120 refugee- and asylum-focused scenarios across five languages revealed recurring failures, leading to open-source data and concrete guardrail policies.
  • Agentic guardrails with tools like web search improve reliability by verifying facts, though effectiveness depends on the underlying LLM.
  • Mozilla AI's open-source tools, any-guardrail and Otari, enable configurable guardrails and seamless LLM switching for practical deployment.