CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-law-data-thailand /soc-ratchakitcha Royal Gazette Thailand (Ratchakitcha) Dataset ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable) โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย Dataset Description ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.tabulartext-retrieval1M<n<10M14 likes28k downloads9h agoHugging Face02senry5433 /china-effective-laws-regulations 全国现行法律法规合集 现行有效的中华人民共和国法律、行政法规、监察法规、地方性法规、司法解释结构化文本。一部法规一行,一条法条一行,供查阅、检索、RAG 和法律 NLP 使用。 数据来自全国人大常委会办公厅 国家法律法规数据库,下载口径为官网的 「有效及尚未生效」。正文由 Word 原文用脚本抽取,未经大模型改写。 这不是官方汇编,不能替代公报或标准文本,也不能作为法律意见。 电子文本与标准文本不一致时,以法律规定的标准文本为准。 快照日期:2026-08-26 效力说明 本数据集 以现行有效法律法规为主体: 效力 status 法规份数 说明 有效 17,649 现行有效,默认应使用这一部分 尚未生效 7 已公布、施行日晚于快照日 失效 45 文件名含「失效」,多为已到期的全国人大常委会试点授权决定 使用时请筛选 status == "有效",即可得到现行有效文本。同一部法若有修正前后多个版本,均予保留,用 filename_date 区分,采用最新日期即可。… See the full description on the dataset page: https://huggingface.co/datasets/senry5433/china-effective-laws-regulations.tabularquestion-answering100K<n<1M0 likes471 downloads29d agoHugging Face03gkour /israeli_law Open Israeli Law (Hebrew Wikisource) Israel's entire body of law — every statute, regulation, and order that Hebrew Wikisource volunteers have transcribed for the "Open Book of Laws" project (ספר החוקים הפתוח) — in one file you can actually load and query. Nearly 6,000 pages. Almost 100 million characters. One jsonl. This isn't a scrape of a summary or a curated subset — it's the raw material: full MediaWiki wikitext, straight from the source, with the metadata you need to trace… See the full description on the dataset page: https://huggingface.co/datasets/gkour/israeli_law.tabulartext-generation1K<n<10K0 likes300 downloads4d agoHugging Face04guychuk /case-law-israel Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/guychuk/case-law-israel.tabulartable-question-answering10K<n<100K2 likes125 downloads1y agoHugging Face05minnesotanlp /lawflow-reasoning-simulation LawFlow: Collecting and Simulating Lawyers' Thought Processes Debarati Das, Khanh Chi Le*, Ritik Parkar*, Karin De Langis, Brendan Madson, Chad Berryman, Robin Willis, Daniel Moses, Brett McDonnell†, Daniel Schwarcz†, Dongyeop Kang† Minnesota NLP, University of Minnesota Twin Cities *equal contribution, †senior advisors Arxiv Project Page Dataset Summary and Purpose LawFlow: Collecting and Simulating Lawyers' Thought Processes The purpose of this dataset is aim… See the full description on the dataset page: https://huggingface.co/datasets/minnesotanlp/lawflow-reasoning-simulation.tabulartext-generationn<1K2 likes97 downloads1y agoHugging Face06free-law /Caselaw_Access_Project_FAISS_index The Caselaw Access Project In collaboration with Ravel Law, Harvard Law Library digitized over 40 million U.S. court decisions consisting of 6.7 million cases from the last 360 years into a dataset that is widely accessible to use. Access a bulk download of the data through the Caselaw Access Project API (CAPAPI): https://case.law/caselaw/ Find more information about accessing state and federal written court decisions of common law through the bulk data service documentation here:… See the full description on the dataset page: https://huggingface.co/datasets/free-law/Caselaw_Access_Project_FAISS_index.tabulartext-generationn<1K10 likes50 downloads3y agoHugging Face07free-law /Nevada_Caselaw_Access_Projectgated The Caselaw Access Project In collaboration with Ravel Law, Harvard Law Library digitized over 40 million U.S. court decisions consisting of 6.7 million cases from the last 360 years into a dataset that is widely accessible to use. Access a bulk download of the data through the Caselaw Access Project API (CAPAPI): https://case.law/caselaw/ Find more information about accessing state and federal written court decisions of common law through the bulk data service documentation here:… See the full description on the dataset page: https://huggingface.co/datasets/free-law/Nevada_Caselaw_Access_Project.tabulartext-generation10K<n<100K0 likes12 downloads3y agoHugging Face08lilgoose777 /nepal-law-commission-nepaligated ⚖️ Nepal Law Commission — Nepali Legal Corpus Dataset Summary A cleaned Nepali-language text corpus extracted from official annual reports published by the Nepal Law Commission (lawcommission.gov.np). The corpus spans fiscal years 2067/68 – 2081/82 (approximately 2010–2025), covering legal research, legislative drafting, law reform activities, and policy recommendations. Each row is a self-contained chunk of Nepali text (~300–1200 characters), filtered from mixed… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/nepal-law-commission-nepali.tabulartext-generation1K<n<10K0 likes9 downloads5mo agoHugging Face09SPAISS6F1 /spai-ss6-corpus-law SPAI SS6 Thai Law Corpus Index Index repo for the Thai law corpus mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: law Rows in canonical config: 208,964 Parquet size in canonical config: 0.49 GB Source license: unknown License review status:… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-law.tabulartext-generationn<1K0 likes8 downloads4mo agoHugging Face10lianghsun /tw-processed-related-law-articlegated Dataset Card for tw-processed-related-law-article 本資料集將中華民國(臺灣)現行法規以「條」為單位重新切分,並把每一條與其引用之相關條文一起合併呈現,方便用於法律語料的持續預訓練(continued pretraining)與條文檢索/問答(retrieval / QA)任務。 Dataset Details Dataset Description 資料來自中華民國公開法規與條文內容,經過下列處理: 將每部法規依條文(flno)切分為獨立樣本。 從條文內文中解析出「相關法條」,將被引用之條文內容串接於原條文後,形成可獨立閱讀的訓練樣本。 標註每條的母法名稱(name)與法規代碼(pcode)。 最終樣本以「法規名稱:xxx 第 N 條 + 條文 + 相關法條」格式呈現,每一筆都是封閉的條文上下文,能直接做為語言模型訓練的輸入。 Curated by: Huang Liang Hsun Language(s) (NLP): Traditional… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-processed-related-law-article.tabulartext-generation10K<n<100K0 likes7 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.