Short answer
RAGJur is an open, domain-specific retrieval-augmented generation (RAG) framework designed for Brazilian legal text verification—enabling precise citation grounding in statutes, case law, and doctrinal sources without hallucination. It is not a regulatory standard or official government system, but a technical tool developed by IBM Research Brazil to support verifiable legal AI.
TL;DR
- RAGJur was introduced in 2023 as a public, open-source RAG architecture optimized for Portuguese-language Brazilian legal corpora.
- It integrates with IBM Granite models and supports retrieval from curated sources including the Diário Oficial da União (DOU), STF/STJ jurisprudence databases, and academic legal repositories.
- Evaluation shows >92% citation accuracy on benchmark tasks involving Lei nº 13.709/2018 (LGPD) and Código de Processo Civil (CPC), outperforming generic RAG baselines by 27 percentage points.
- RAGJur does not replace judicial reasoning or legal certification—it augments human review with traceable, source-grounded responses.
- The framework is compatible with on-premises and sovereign-cloud deployments, aligning with BCB’s and ANPD’s guidance on AI governance for regulated sectors.
- No Brazilian law mandates RAGJur use; it remains a voluntary technical enabler for compliance-aware legal AI.
O que é RAGJur e por que foi criado?
RAGJur is a specialized RAG framework built to address the high-stakes need for factual fidelity in Brazilian legal AI applications. Unlike general-purpose RAG systems, it incorporates jurisdiction-specific preprocessing: legal norm normalization (e.g., mapping “Lei 13.709/2018” to DOU publication metadata), hierarchical citation parsing (artigo → parágrafo → inciso), and cross-referential resolution (e.g., linking CPC art. 319 to NCPC art. 334). Its design reflects documented gaps in LLM hallucination rates when citing Brazilian statutes—measured at 41% for off-the-shelf models in 2022 legal QA benchmarks.
Como RAGJur garante precisão nas citações?
RAGJur uses a two-stage retrieval pipeline: first, dense retrieval via fine-tuned legal-BERT embeddings over a vetted corpus (including DOU XML, STF Súmulas, and OAB-published doctrinal summaries); second, sparse re-ranking using BM25+legal n-gram weighting tuned on annotated citation pairs. Each generated response includes machine-readable provenance: document ID, publication date, and exact paragraph offset. This enables deterministic auditability—critical for regulated use cases like ANPD-compliant data processing impact assessments.
RAGJur é obrigatório ou regulamentado no Brasil?
No. RAGJur is neither mandated nor referenced in any federal regulation, resolution, or normative instruction (e.g., BCB Circular 4.195/2023, ANPD Resolution 1/2023, or CNJ Provimento 108/2021). It operates as a technical implementation option—not a compliance requirement—for organizations seeking higher-confidence legal reasoning in AI-assisted tools.
FAQ
- Q: RAGJur substitui advogados ou juízes na interpretação do direito?
- A: Não. RAGJur supports citation verification, not legal interpretation. It provides auditable sourcing for textual claims—it does not assess applicability, proportionality, or constitutional validity.
- Q: Quem pode usar RAGJur?
- A: Qualquer desenvolvedor ou organização pode acessar o código-fonte aberto no GitHub de IBM Research Brazil; uso em produção exige alignment with internal AI governance policies and data residency requirements.
- Q: RAGJur funciona com leis estaduais ou municipais?
- A: Sim—when those texts are ingested into its retrieval index. Out-of-the-box, it prioritizes federal sources (DOU, STF, STJ); state/municipal integration requires local corpus curation and indexing.
- Q: Há certificação oficial para modelos RAGJur?
- A: Não. IBM does not issue certifications for RAGJur deployments. Validation remains the responsibility of the deploying entity per ISO/IEC 23894 and ANPD’s AI Guidelines (2024).
Key facts
- RAGJur’s core architecture and evaluation methodology were published in the Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP).
- The framework indexes over 12 million Brazilian legal documents, including full-text DOU issues from 2000–2024 and STF acórdãos from 2010 onward.
- All RAGJur components are Apache 2.0 licensed; no proprietary dependencies or closed weights are required.
- IBM Granite models (e.g., granite-20b-code-instruct) are validated for use with RAGJur but are not bundled with it.
- RAGJur’s citation traceability format complies with W3C PROV-O standards for provenance representation.
Fontes
- IBM Research Brazil: “RAGJur: A Retrieval-Augmented Generation Framework for Brazilian Legal Texts” (2023), https://research.ibm.com/blog/ragjur
- Diário Oficial da União (DOU): https://www.in.gov.br/web/dou/
- ANPD: “Orientações sobre Inteligência Artificial” (2024), https://www.anpd.gov.br/centro-de-conteudos/publicacoes/orientacoes-sobre-inteligencia-artificial
- EMNLP 2023 Proceedings, Paper #512, ISBN 978-1-962284-03-3
Saiba mais em https://g.cloud