Spaces:
Running
[Data Sovereignty] Open-Source Legal Instruction Dataset for African Civil Law (OHADA / AUDCG)
Hello @PleIAs team,
I am Ivan Bertrand NGOMEN BISIL, an attorney at the Cameroon Bar. I have been following your incredible work on PleIAs and the Common Corpus initiative with great admiration. Your dedication to building truly open, high-quality, and legally compliant multilingual training data is a reference for the entire AI community.
I am currently developing a supervised fine-tuning (SFT) dataset dedicated to OHADA Business Law (specifically the Acte Uniforme relatif au Droit Commercial Général - AUDCG). OHADA law unifies corporate and commercial regulations across 17 French-speaking African countries. Because regional civil law frameworks outside of Western Europe are heavily neglected or hallucinated by mainstream foundational models, this project is a step toward data sovereignty and cultural diversity in open-source AI.
I have just released a 10-row open-source production sample on Hugging Face to validate the architecture before expanding it to the dataset, and I would deeply value your data engineering and curation critique:
👉 Bisilivan/dataset-ohada-droit-commercial-general-echantillon
Technical Architecture & Data Ethics:
- Format: Clean JSONL mapping
instruction,input, andoutput. - Curation Strategy: It features a highly granular multi-tier taxonomy metadata mapping (
category,subcategory,difficulty, andentry_type) to segment basic conceptual definitions from advanced real-world business case studies (cas pratiques). - Zero-Hallucination Constraints: To ensure highest reliability, every generated target output strictly enforces step-by-step reasoning and precise citations of Articles and Paragraphs of the official public domain AUDCG text.
- Copyright & Compliance: This dataset is built completely from scratch, authored by practicing legal counsel, and fully open-source (permissible use).
Questions for PleIAs:
Given your unmatched expertise in building vast, ethnically and jurisdictionally diverse open text corpora:
- Does this multi-tier metadata structure look solid for fine-tuning open weights models to adapt to specific regional legal reasoning without losing generalization capabilities?
- Are there specific open-source tools, licenses, or documentation standards your lab recommends to ensure this dataset can easily integrate into broader global open-source training pipelines in the future?
Thank you so much for your inspiring work, your commitment to open-source, and any feedback you might share!
Best regards,
Ivan Bertrand NGOMEN BISIL
Attorney at Law, Cameroon Bar