datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
civil-code-phil
Civilex — Philippine Legal RAG & SFT Dataset
Retrieval corpus and supervised fine-tuning (SFT) data for a retrieval-augmented generation (RAG) pipeline over Philippine law: the Civil Code (Republic Act No. 386) and Supreme Court jurisprudence. Produced by the civilex-thesis research pipeline.
Contents: 11k+ Supreme Court jurisprudence cases spanning 1949–2025, and 2,270 articles from the Civil Code (Republic Act No. 386).
Dataset structure
.
├── README.md
├──… See the full description on the dataset page: https://huggingface.co/datasets/renzzyyy1028/civil-code-phil.code-civil
Code civil, non-instruct (2025-09-20)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language models based… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-civil.Italian_Civil_Code
Abstract
The Italian Civil Code, hereinafter referred to as ICC, is the legislation source containing norms that regulate private law in Italy and it consists of 2969 articles. This number actually corresponds to 3225 articles considering all variants and subsequent insertions, which are designated by using Latin-term suffixes (e.g., “bis”, “ter”, “quater”). Then, we obtain a total of 3039 if we remove the articles no longer in force (i.e., articles which are replaced by other… See the full description on the dataset page: https://huggingface.co/datasets/AndreaSimeri/Italian_Civil_Code.nepal_civilcode_en
Nepal Civil Code Instruct and Response Dataset (English)
Overview
The chhatramani/nepal_civilcode_en dataset contains instruction-response pairs in English, derived from the Nepal Civil Code (Muluki Devani Samhita, 2074 BS), designed for fine-tuning instruction-based large language models (LLMs) such as LLaMA, GPT, or similar models. Each entry includes a prompt (instruction), an optional input field (currently empty), and a detailed response based on the legal provisions… See the full description on the dataset page: https://huggingface.co/datasets/chhatramani/nepal_civilcode_en.
