datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tm-acts-qwen38-multigpt2_model_acts_openwebtextdiverse_risk_acts_fixedsae_max_actsrussian-supreme-court-plenum-acts
Plenum Resolutions of the Supreme Court of Russia (1961–2026)
Every act published in the «Постановления Пленума» section of the Russian Supreme Court's
website: 1,504 records — 1,503 plenum resolutions plus 1 meeting
agenda — with full texts, metadata and the court's original attachments. Coverage
1961–2026; completeness verified against the court's own index at collection time
(the section reported exactly 1,504 documents).
Постановления Пленума ВС РФ — руководящие разъяснения… See the full description on the dataset page: https://huggingface.co/datasets/Alexey5676/russian-supreme-court-plenum-acts.act_sort_250911_02_downsampled_demospeedup_1_3indian-legal-actsBangladesh-Legal-Acts-Dataset
Bangladesh Legal Acts Dataset
A comprehensive database of Bangladesh's legal framework, containing 1484+ acts scraped and processed from the official Bangladesh Laws portal, enhanced with historical government context, legal system context, and comprehensive metadata.
Dataset Overview
Total Acts: 1,484
Total Sections: 35,633
Total Footnotes: 14,523
Languages: English, Bengali, Mixed
Format: JSON with structured metadata
Historical Context: Government periods from… See the full description on the dataset page: https://huggingface.co/datasets/sakhadib/Bangladesh-Legal-Acts-Dataset.mean_act_steered_nonverbal_deploy_actsqwen_8b_deploy_actsindia-acts
India Acts — Central & State Statutes
A comprehensive corpus of Indian legislation in PDF form — covering both Central (Parliament) Acts and State / Union Territory Acts — in English and Hindi, scraped and consolidated from publicly available government sources (primarily the India Code portal and individual State legislature websites).
This dataset is intended as a research and AI-training resource for tasks such as legal document retrieval, statutory question-answering… See the full description on the dataset page: https://huggingface.co/datasets/judicialmind/india-acts.new_deploy_big_actsnemotron_deploy_acts_bigact_so100_hf_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 5,
"total_frames": 1863,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Raude/act_so100_hf_test.deploy_to_eval_06_acts_newEU_acts_in_Ukrainian
[!NOTE]
Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/19753
Description
It was based on: a) the translations of the EU acts in Ukrainian that are available at the official web-portal of the Parliament of Ukraine
https://zakon.rada.gov.ua/laws/main/en/g22), and b) the EU acts that are available in many CEF languages at https://eur-lex.europa.eu.It is a collection of TMX files (X-UK, where X is a CEF language) and includes 3056791 TUs in total.
bg-uk… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/EU_acts_in_Ukrainian.qwen_steered_deploy_actsnemotron_eval_to_deploy_steer_04_actsqwen_8b_base_actsact_sort_dsu_downsampled_251104_dsu_2_4qwen_steered_eval_actsnemotron_eval_to_deploy_steer_00_acts_nothinkingact_so101_pick_white_02This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 10,
"total_frames": 6470,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/b-sky-lab/act_so101_pick_white_02.sinhala-ocr-lk-acts-1010
🇱🇰 Sinhala OCR - Sri Lankan Acts Dataset
Dataset Description
This dataset contains 1,010 scanned document images of Sri Lankan legal acts (1980s-2010s) in Sinhala language with ground truth text annotations for Optical Character Recognition (OCR) training and evaluation.
Key Features
✅ High-quality scanned document images
✅ Professionally corrected ground truth text
✅ Year-wise metadata for temporal analysis
✅ Pre-split into train/eval/test sets… See the full description on the dataset page: https://huggingface.co/datasets/avishadilhara/sinhala-ocr-lk-acts-1010.new_output_deploy_actsact_smartbin_wrist_v1gpt_oss_maze_acts_120_m11_v1act_stacking_action_chunksacts-finqa-loranew_output_eval_acts
