[Data Sovereignty] Open-Source Legal Instruction Dataset for African Civil Law (OHADA / AUDCG)

#2
by Bisilivan - opened

Hello @PleIAs team,

I am Ivan Bertrand NGOMEN BISIL, an attorney at the Cameroon Bar. I have been following your incredible work on PleIAs and the Common Corpus initiative with great admiration. Your dedication to building truly open, high-quality, and legally compliant multilingual training data is a reference for the entire AI community.

I am currently developing a supervised fine-tuning (SFT) dataset dedicated to OHADA Business Law (specifically the Acte Uniforme relatif au Droit Commercial Général - AUDCG). OHADA law unifies corporate and commercial regulations across 17 French-speaking African countries. Because regional civil law frameworks outside of Western Europe are heavily neglected or hallucinated by mainstream foundational models, this project is a step toward data sovereignty and cultural diversity in open-source AI.

I have just released a 10-row open-source production sample on Hugging Face to validate the architecture before expanding it to the dataset, and I would deeply value your data engineering and curation critique:
👉 Bisilivan/dataset-ohada-droit-commercial-general-echantillon

Technical Architecture & Data Ethics:

  • Format: Clean JSONL mapping instruction, input, and output.
  • Curation Strategy: It features a highly granular multi-tier taxonomy metadata mapping (category, subcategory, difficulty, and entry_type) to segment basic conceptual definitions from advanced real-world business case studies (cas pratiques).
  • Zero-Hallucination Constraints: To ensure highest reliability, every generated target output strictly enforces step-by-step reasoning and precise citations of Articles and Paragraphs of the official public domain AUDCG text.
  • Copyright & Compliance: This dataset is built completely from scratch, authored by practicing legal counsel, and fully open-source (permissible use).

Questions for PleIAs:

Given your unmatched expertise in building vast, ethnically and jurisdictionally diverse open text corpora:

  1. Does this multi-tier metadata structure look solid for fine-tuning open weights models to adapt to specific regional legal reasoning without losing generalization capabilities?
  2. Are there specific open-source tools, licenses, or documentation standards your lab recommends to ensure this dataset can easily integrate into broader global open-source training pipelines in the future?

Thank you so much for your inspiring work, your commitment to open-source, and any feedback you might share!

Best regards,
Ivan Bertrand NGOMEN BISIL
Attorney at Law, Cameroon Bar

Sign up or log in to comment