datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
US_Regulation_ECFR_20260101
US Regulation eCFR 2026-01-01 Dataset
This repository contains a structured, machine-readable version of the Electronic Code of Federal Regulations (eCFR), captured as of January 1, 2026. Unlike the annual CFR snapshots, this dataset reflects the editorialized, near real-time version of federal regulations.
Dataset Description
The eCFR is a daily updated editorial compilation of CFR material and Federal Register amendments. This dataset captures a specific point-in-time… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/US_Regulation_ECFR_20260101.US_Regulation_CFR_20260101VDR_Renewable_Regulation
VDR_Renewable_Regulation - Overview
Dataset Summary
VDR_Renewable_Regulation is a curated multimodal dataset focused on renewable energy technical documents, regulations, and legal frameworks. It combines text and image data extracted from real scientific and regulatory PDFs to support tasks such as RAG DSE, question answering, document search, and vision-language model training.
Dataset Creation
This dataset was created using our open-source tool… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_Renewable_Regulation.china-effective-laws-regulations
全国现行法律法规合集
现行有效的中华人民共和国法律、行政法规、监察法规、地方性法规、司法解释结构化文本。一部法规一行,一条法条一行,供查阅、检索、RAG 和法律 NLP 使用。
数据来自全国人大常委会办公厅 国家法律法规数据库,下载口径为官网的 「有效及尚未生效」。正文由 Word 原文用脚本抽取,未经大模型改写。
这不是官方汇编,不能替代公报或标准文本,也不能作为法律意见。 电子文本与标准文本不一致时,以法律规定的标准文本为准。
快照日期:2026-08-26
效力说明
本数据集 以现行有效法律法规为主体:
效力 status
法规份数
说明
有效
17,649
现行有效,默认应使用这一部分
尚未生效
7
已公布、施行日晚于快照日
失效
45
文件名含「失效」,多为已到期的全国人大常委会试点授权决定
使用时请筛选 status == "有效",即可得到现行有效文本。同一部法若有修正前后多个版本,均予保留,用 filename_date 区分,采用最新日期即可。… See the full description on the dataset page: https://huggingface.co/datasets/senry5433/china-effective-laws-regulations.us-regulations
US Federal Regulations — the CFR, held word for word
Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them.
222,767 regulations — 219,061 CFR sections and 3,706 appendices — across all 49 titles
of the Code of Federal Regulations, each one as the agency publishes it.
Statutes say… See the full description on the dataset page: https://huggingface.co/datasets/docketx/us-regulations.regulation-retrieval
Turkish Legal Özelge Corpus Dataset
📊 Dataset Summary
Turkish Legal Özelge Corpus is a comprehensive Information Retrieval dataset consisting of özelge (tax ruling) decisions published by the Turkish Revenue Administration (Gelir İdaresi Başkanlığı - GİB).
Key Features
Format: BEIR (Benchmarking IR) format with corpus-queries-qrels structure
Language: Turkish 🇹🇷
Domain: Tax Law, Administrative Law, Turkish Law
Source: GİB Özelge Decisions
Use Cases:… See the full description on the dataset page: https://huggingface.co/datasets/newmindai/regulation-retrieval.tdtu-student-regulations-qa
TDTU Vietnamese University Regulations QA Dataset
Dataset Description
Tập dữ liệu hỏi-đáp tiếng Việt về quy chế, quy định sinh viên của Trường Đại học Tôn Đức Thắng (TDTU), được xây dựng cho bài toán Retrieval-Augmented Generation (RAG) và fine-tuning LLM tư vấn sinh viên.
Ngôn ngữ: Tiếng Việt
Domain: Quy chế đại học, chính sách sinh viên
Mục đích: Huấn luyện chatbot tư vấn sinh viên TDTU
Dataset Details
Dataset Sources… See the full description on the dataset page: https://huggingface.co/datasets/hungminhss/tdtu-student-regulations-qa.us-regulations-16k-64k
Dataset Card for Long US Regulations (16k–64k)
Dataset Summary
Long US Regulations (16k–64k) is a curated, high-integrity long-context corpus derived from the primary text of the United States Code of Federal Regulations (CFR). It consists of 793 extensive regulatory documents totaling 41,227,048 tokens, with an average document length of 51,988 tokens (measured using OpenAI's cl100k_base tokenizer).
Every document strictly spans between 16,384 and 65,536 tokens.… See the full description on the dataset page: https://huggingface.co/datasets/beyond369/us-regulations-16k-64k.Regulations-of-the-Free-State-Militia
Dataset Card: Regulations of the Free State Militia – Binding Constitutional Law
⚖️ STATUS: BINDING CONSTITUTIONAL LAW ⚖️
Links
Google Docs (Original): The Regulations of the Free State Militia
Hugging Face Dataset: https://huggingface.co/datasets/RaddicalSilly/Regulations-of-the-Free-State-Militia
Status Declaration
These Regulations are binding constitutional law.
The Regulations of the Free State Militia fulfill the Second Amendment's… See the full description on the dataset page: https://huggingface.co/datasets/RaddicalSilly/Regulations-of-the-Free-State-Militia.thai-gov-procurement_regulation-17-amend-21
🇹🇭 Dataset Card for Thai Government Procurement Dataset
ℹ️ This dataset is optimized for procurement-related NLP tasks in Thai.
This dataset contains a collection of procurement regulations, instructions, and responses focused on public sector purchasing, contract management, and compliance with Thai government standards. It aims to support natural language processing tasks involving procurement assistance, such as chatbot development, procurement dialogue generation… See the full description on the dataset page: https://huggingface.co/datasets/amornpan/thai-gov-procurement_regulation-17-amend-21.CANADA_ACT_REGULATION_QA
Canadian Acts and Regulation QA
source- https://laws-lois.justice.gc.ca/eng/XML/Legis.xml
model_name="gemini-1.5-flash-latest" with 1 million context length,
First summarize the text scrapped text from xml tree of urls using gemini.
then generate QA from sumarised text.
Performance of Gemini was way way better than GPT-4.
Fitering was done based on Heuristics after rigrous analysis because llms were not always accurate.
summary_prompt_template= """
You'r legal expert… See the full description on the dataset page: https://huggingface.co/datasets/Guggu/CANADA_ACT_REGULATION_QA.regulation-retrieval
Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından Türkçeye çevrilmiştir. Orijinal veri seti newmindai tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir.
🔗 Orijinal Kaynak: newmindai/regulation-retrieval
🔗 Derleyen Platform: VeriPazarı
Turkish Legal Özelge Corpus Veri Seti
📊 Veri Seti Özeti
Turkish Legal Özelge Corpus, Gelir İdaresi Başkanlığı (GİB) tarafından… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/regulation-retrieval.CFR-Title-41-Federal-Travel-Regulation
Federal Travel Regulation
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains document-grounded question-and-answer records based on the Federal Travel Regulation, as reproduced in the source volume of title 41 of the Code of Federal Regulations.
The Federal Travel Regulation establishes government-wide policies governing official civilian travel and relocation at Federal expense. It addresses temporary duty travel… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/CFR-Title-41-Federal-Travel-Regulation.regulationsiraq_student_regulationsSft_dataset_hust_regulation
