datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jade-samples-10000x1
JADE amortized posterior samples — 10,000 observations x 1 draw
Noisy weak-lensing convergence observations paired with joint posterior draws of
(convergence field, cosmology) from the amortized conditional diffusion model
of JADE.
[!IMPORTANT]
This dataset is not a product of arXiv:2606.31988.
It was generated afterwards, with the same trained model, to support posterior
calibration diagnostics that do not appear in the paper. No number in the paper
was computed from it, and… See the full description on the dataset page: https://huggingface.co/datasets/b-remy/jade-samples-10000x1.sit-latents-ode-heun-1000-class-0_1000-samples-segment-100-199koch_50-samplesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "koch",
"total_episodes": 50,
"total_frames": 16417,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dboemer/koch_50-samples.fineweb2-2k-sampleslatent-taxonomy-samplesdata_samples
Multimodal Pretraining
This section covers our large-scale collections at the source and is distributed in its original form (PDF with layout intact, audio attached to its transcript) rather than as text extracted after the fact. The emphasis is on what large-scale web collection misses: academic global production (badly indexed in scientific repositories); patents outside the US; the technical and regulatory archives of telecom and finance.
These are long, structured… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/data_samples.ultra-50k-samples-dataset-instruction_followingsit-latents-ode-heun-1000-class-0_1000-samples-segment-400-499cosmopedia_web_samples_v2_shards_enSMPL_samples
SMPL motion dataset
Episodes
Episode
Motion
0
dance_hiphop
1
dance_jackson
2
eating_apple
3
high_jump
4
hiphop_ii
5
horse_riding
6
jump_360
7
mambo_kicks
8
reach_jump
9
turn_jump_270
10
turn_jump_360
11
victory_dance
12
walk_forward
Job-dataset-samples
My Jobs and Companies Dataset
This dataset contains sample rows of job postings and company profiles.
samples_cerrado50_samples_20260713_185031This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/chair0/50_samples_20260713_185031.eval_koch_50-samplesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "koch",
"total_episodes": 10,
"total_frames": 4979,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dboemer/eval_koch_50-samples.multiple_samples_ground_truth_openr1_llm_verifier_cleanmultiple_samples_ground_truth_openr1_llm_verifierpass-at-k-samplesultra-50k-samples-dataset-honestyclaude4-agentic-samples-V3samples_bookradioml-2018-samples
RadioML 2018 Sample I/Q Dataset
Curated I/Q sample windows for the SignalIQ Modulation Classifier demo.
Classes: 11 modulation types (AM-DSB, AM-SSB, WBFM, QPSK, QAM16, QAM64, 8PSK, BPSK, CPFSK, GFSK, PAM4)
Samples: 550 (50 per class)
Fields: iq_i, iq_q, modulation, snr_db, label
Window length: 1024 complex samples
Companion dataset for alirezaaminzadeh/radio-modulation-classifier.
pick_place_pink_lego_few_samplesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "koch",
"total_episodes": 21,
"total_frames": 8086,
"total_tasks": 1,
"total_videos": 42,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:21"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/seeingrain/pick_place_pink_lego_few_samples.archive-dolma3-pool-150b-samples
archive-dolma3-pool-150b-samples
ARCHIVE (pre-6T era): stratified working samples drawn from the 150B pool.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_pool_150B_samples
Renamed
2026-05-25
See docs/data_home/inventory.json for the full inventory including the… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/archive-dolma3-pool-150b-samples.samplesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "gr1t1",
"total_episodes": 1000,
"total_frames": 362600,
"total_tasks": 1,
"total_videos": 1000,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1000"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/FourierIntelligence/samples.bimanual-yam-samplesultra-50k-samples-dataset-truthfulnessmmlu_pro_Olmo-3-7B-Instruct_temp0.9_samples99so101-dataset-50-samples_20260709_132939This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/ahmedaali/so101-dataset-50-samples_20260709_132939.rag_audit_assistant_samples_clusters_analysispublic_samplesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 3,
"total_frames": 4200,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/YounesHouhou/public_samples.
