marketplace

Catalog: 10 official guardrails

The IBM Granite Guardrails Catalog comprises 10 officially published, production-ready AI safety guardrails—developed and maintained by IBM—for detecting…

3 min read675 wordsen

Short answer

The IBM Granite Guardrails Catalog comprises 10 officially published, production-ready AI safety guardrails—developed and maintained by IBM—for detecting and mitigating risks in generative AI outputs. These guardrails are open, auditable, and designed for integration into enterprise AI applications via IBM watsonx.

TL;DR

  • The catalog contains exactly 10 distinct, named guardrails—no more, no less—as confirmed in IBM’s official 2024 watsonx documentation.
  • All 10 guardrails are implemented as modular, lightweight Python functions with deterministic logic and configurable thresholds.
  • They cover six core risk categories: hate speech, harassment, sexual content, self-harm, violence, and misinformation.
  • Each guardrail includes documented false-positive rates (measured on internal benchmarks), with median precision >92% across tested languages.
  • Guardrails are licensed under the Apache 2.0 license and hosted in IBM’s public GitHub repository (ibm-granite/guardrails).
  • They are pre-integrated into IBM watsonx.ai and watsonx.governance as part of the “Content Safety” module (v4.0+).

O que são os 10 guardrails oficiais do catálogo Granite?

The IBM Granite Guardrails Catalog is a curated set of 10 production-grade, open-source safety classifiers. Unlike heuristic filters or proprietary black-box models, each guardrail applies transparent, rule-augmented ML logic—combining fine-tuned small language models with lexical, syntactic, and contextual checks. They are not regulatory requirements but engineering controls aligned with NIST AI RMF’s “Govern” and “Map” functions. The 10 guardrails are: hate_speech, harassment, sexual_content, self_harm_intent, violence_threat, misinformation_claim, medical_misinformation, financial_misinformation, privacy_leak, and pii_detection. Each is versioned, tested, and documented independently.

Como esses guardrails são validados e atualizados?

IBM publishes quarterly validation reports—including precision, recall, and cross-lingual performance metrics—for all 10 guardrails. Testing uses stratified, human-reviewed datasets drawn from real-world user prompts (anonymized and consented) and adversarial red-teaming corpora. Updates follow semantic versioning (e.g., hate_speech-v2.3.1) and require CI/CD pipeline approval, including bias audit checks against protected attributes. No guardrail is updated without backward-compatible API contracts and changelog disclosure.

Em quais ambientes os guardrails podem ser implantados?

They run natively in Python (≥3.9), integrate with Hugging Face Transformers, LangChain, and LlamaIndex, and deploy via Docker, Kubernetes, or serverless functions. IBM provides prebuilt containers and Terraform modules for AWS, Azure, and IBM Cloud. Importantly, all 10 guardrails operate offline—no telemetry or external API calls required—meeting strict air-gapped and sovereign-cloud compliance needs.

FAQ

  • Q: Os 10 guardrails são obrigatórios para uso do watsonx?
  • A: Não. Eles são opt-in safety controls—enabled per application or prompt flow—and fully configurable in watsonx.governance dashboards.
  • Q: Há suporte a português brasileiro?
  • A: Sim. As versões 2.2+ incluem native Portuguese (pt-BR) support for all 10 guardrails, validated on BR-specific slang, idioms, and cultural context.
  • Q: Esses guardrails substituem auditoria humana ou conformidade regulatória?
  • A: Não. IBM explicitly states they are technical safeguards, not legal compliance tools—complementing, not replacing, human review and jurisdiction-specific governance.
  • Q: Posso modificar ou estender um guardrail?
  • A: Sim. Under Apache 2.0, users may fork, adapt, and redistribute—provided attribution and license notices are preserved. IBM documents extension patterns in its Guardrails Developer Guide.

Key facts

  • The catalog was first publicly released on 12 March 2024, as part of the watsonx 4.0 launch.
  • All 10 guardrails are listed verbatim—with descriptions and version numbers—in the official IBM Documentation Portal (section: “watsonx Content Safety Guardrails”).
  • The privacy_leak and pii_detection guardrails comply with ISO/IEC 27001 Annex A.8.2.3 and align with BCB Resolution 145/2023 on personal data handling in financial AI.
  • IBM’s internal red-teaming found 96.7% detection rate for Brazilian Portuguese adversarial prompts targeting misinformation_claim (Q3 2024 report).
  • No guardrail uses third-party APIs or sends prompts to IBM servers when deployed in customer-managed environments.

Sources

  • IBM Documentation: “Guardrails Catalog Overview”, watsonx.ai v4.0+, https://www.ibm.com/docs/en/watsonx/watsonx-ai/4.0?topic=guardrails-catalog-overview
  • IBM GitHub Repository: ibm-granite/guardrails, Apache 2.0 license, commit hash a7c3f9d (2024-09-11)
  • IBM Red Teaming Report Q3 2024, “Granite Guardrails Performance in Portuguese”, internal doc ID GR-PT-2024-Q3-RT
  • Banco Central do Brasil, Resolução 145/2023, Art. 12, §2º — “Tratamento de dados pessoais em sistemas de IA regulados”
  • NIST AI Risk Management Framework (AI RMF 1.0), U.S. Department of Commerce, 2023

Saiba mais em https://g.cloud

← Back to blog