datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
regulations_filtered
Regulations.Gov
Description
Regulations.gov is an online platform operated by the U.S. General Services Administration that collates newly proposed rules and regulations from federal agencies along with comments and feedback from the general public.
This dataset includes all plain-text regulatory documents published by a variety of U.S. federal agencies on this platform, acquired via the bulk download interface provided by Regulations.gov. These agencies include the… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/regulations_filtered.china-effective-laws-regulations
全国现行法律法规合集
现行有效的中华人民共和国法律、行政法规、监察法规、地方性法规、司法解释结构化文本。一部法规一行,一条法条一行,供查阅、检索、RAG 和法律 NLP 使用。
数据来自全国人大常委会办公厅 国家法律法规数据库,下载口径为官网的 「有效及尚未生效」。正文由 Word 原文用脚本抽取,未经大模型改写。
这不是官方汇编,不能替代公报或标准文本,也不能作为法律意见。 电子文本与标准文本不一致时,以法律规定的标准文本为准。
快照日期:2026-08-26
效力说明
本数据集 以现行有效法律法规为主体:
效力 status
法规份数
说明
有效
17,649
现行有效,默认应使用这一部分
尚未生效
7
已公布、施行日晚于快照日
失效
45
文件名含「失效」,多为已到期的全国人大常委会试点授权决定
使用时请筛选 status == "有效",即可得到现行有效文本。同一部法若有修正前后多个版本,均予保留,用 filename_date 区分,采用最新日期即可。… See the full description on the dataset page: https://huggingface.co/datasets/senry5433/china-effective-laws-regulations.regulations
Description
Regulations.gov is an online platform operated by the U.S. General Services Administration that collates newly proposed rules and regulations from federal agencies along with comments and feedback from the general public.
This dataset includes all plain-text regulatory documents published by a variety of U.S. federal agencies on this platform, acquired via the bulk download
interface provided by Regulations.gov. These agencies include the
Bureau of Industry and… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/regulations.us-regulations-16k-64k
Dataset Card for Long US Regulations (16k–64k)
Dataset Summary
Long US Regulations (16k–64k) is a curated, high-integrity long-context corpus derived from the primary text of the United States Code of Federal Regulations (CFR). It consists of 793 extensive regulatory documents totaling 41,227,048 tokens, with an average document length of 51,988 tokens (measured using OpenAI's cl100k_base tokenizer).
Every document strictly spans between 16,384 and 65,536 tokens.… See the full description on the dataset page: https://huggingface.co/datasets/beyond369/us-regulations-16k-64k.Regulations-of-the-Free-State-Militia
Dataset Card: Regulations of the Free State Militia – Binding Constitutional Law
⚖️ STATUS: BINDING CONSTITUTIONAL LAW ⚖️
Links
Google Docs (Original): The Regulations of the Free State Militia
Hugging Face Dataset: https://huggingface.co/datasets/RaddicalSilly/Regulations-of-the-Free-State-Militia
Status Declaration
These Regulations are binding constitutional law.
The Regulations of the Free State Militia fulfill the Second Amendment's… See the full description on the dataset page: https://huggingface.co/datasets/RaddicalSilly/Regulations-of-the-Free-State-Militia.nepal_legal_regulations_corpus_np
nepal_legal_regulations_corpus_np
Nepali (Devanagari) text of Nepal Regulations.
This dataset is a cleaned, chunked text corpus (approx. 400–500 words per chunk) constructed from official legal documents of Nepal.
It is designed for training a tokenizer and for continued pre‑training of language models in the legal domain.
Source
The original documents were collected from:
नेपाल सरकार, कानून, न्याय तथा संसदीय मामिला मन्त्रालय
(Government of Nepal, Ministry of Law… See the full description on the dataset page: https://huggingface.co/datasets/chhatramani/nepal_legal_regulations_corpus_np.
