datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cold-cases
Collaborative Open Legal Data (COLD) - Cases
COLD Cases is a dataset of 8.3 million United States legal decisions with text and metadata, formatted as compressed parquet files. If you'd like to view a sample of the dataset formatted as JSON Lines, you can view one here
This dataset exists to support the open legal movement exemplified by projects like
Pile of Law and
LegalBench.
A key input to legal understanding projects is caselaw -- the published, precedential decisions of… See the full description on the dataset page: https://huggingface.co/datasets/harvard-lil/cold-cases.opf_small_case2000_goc
Data download through hfApi
Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times.
opf_small_case57_ieee
Data download through hfApi
Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times.
opf_small_case10000_gocopf_small_case14_ieee
Data download through hfApi
Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times.
opfdata_case2000_gocopf_small_case30_ieee
Data download through hfApi
Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times.
opfdata_case30_ieeepf_small_case2000_goc
Data download through hfApi
Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times.
opf_small_case500_goc
Data download through hfApi
Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times.
indian-case-laws
Indian Case Laws
Open Indian case-law data for AI, search, and legal research.
This dataset is part of the KanoonGPT Open Legal Data Initiative - an effort to make Indian legal data easier to access, trace, and build on for open-source research, legal tech, and production AI systems.
KanoonGPT is building structured Indian legal data and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.
Repository: KanoonGPT/indian-case-laws… See the full description on the dataset page: https://huggingface.co/datasets/KanoonGPT/indian-case-laws.opf_small_case118_ieee
Data download through hfApi
Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times.
opfdata_case118_ieeepf_small_case30_ieee
Data download through hfApi
Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times.
pf_small_case57_ieee
Data download through hfApi
Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times.
pf_small_case1354_pegase
Data download through hfApi
Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times.
pf_small_case500_goc
Data download through hfApi
Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times.
opfdata_case500_gocopfdata_case14_ieeepf_small_case118_ieee
Data download through hfApi
Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times.
pf_small_case14_ieee
Data download through hfApi
Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times.
earbud_case_insertion_teleop_0515_left_rightThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_flexiv_rizon4_rt",
"total_episodes": 776,
"total_frames": 1145613,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:776"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/Xense/earbud_case_insertion_teleop_0515_left_right.Venus_Case_Tempeval_pencil_caseThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_tactile_follower",
"total_episodes": 17,
"total_frames": 37242,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:17"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ydaichi/eval_pencil_case.asia-who-treatment-success-rate-hiv-positive-tb-cases
Treatment success rate: HIV-positive TB cases | Asia (WHO GHO)
🌏 551 observations · 45 Asia countries · 1999–2023 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 551 observations of Treatment success rate: HIV-positive TB cases data across 45 Asia countries, spanning 1999–2023, covering 1 distinct indicators.
About the source
Source: WHO Global Health Observatory
Publisher: World Health Organization
License: cc-by-4.0
Topic:… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-who-treatment-success-rate-hiv-positive-tb-cases.earbud_case_insertion_teleop_0515This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_flexiv_rizon4_rt",
"total_episodes": 834,
"total_frames": 1213518,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:834"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/Xense/earbud_case_insertion_teleop_0515.AIRBOT_MMK2_place_the_glasses_case_and_gold_bars
AIRBOT_MMK2_place_the_glasses_case_and_gold_bars
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_place_the_glasses_case_and_gold_bars.cold-cases
Collaborative Open Legal Data (COLD) - Cases
COLD Cases is a dataset of 8.3 million United States legal decisions with text and metadata, formatted as compressed parquet files. If you'd like to view a sample of the dataset formatted as JSON Lines, you can view one here
This dataset exists to support the open legal movement exemplified by projects like
Pile of Law and
LegalBench.
A key input to legal understanding projects is caselaw -- the published, precedential decisions of… See the full description on the dataset page: https://huggingface.co/datasets/5-31-2024/cold-cases.xtac-umi-g1-put-headphones-into-case
Representative frames from TacVerse's bimanual
demonstrations.
Collected with XTac-UMI-G1 grippers, released as LeRobot
datasets.
This dataset was created using LeRobot.
Explore this dataset with the LeRobot Dataset Viewer.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "xtac_umi_g1",
"total_episodes": 10,
"total_frames": 12114,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100… See the full description on the dataset page: https://huggingface.co/datasets/TacVerse/xtac-umi-g1-put-headphones-into-case.opf_small_case1354_pegase
Data download through hfApi
Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times.
