datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stg-paired-audiopaired-llama-3.2-1b-embeddings-lmsys-chat-1m
Paired Llama 3.2 1B Token Embeddings (LMSYS-Chat-1M)
This dataset contains paired activations corresponding to single token locations extracted from Meta's Llama 3.2 1B Instruct on conversations from LMSYS-Chat-1M.
Embeddings are provided for layers 5 through 14, which capture the most interesting intermediate representations.
This dataset was built to study things like:
Learning different basis for activations at a given layer
Studying if there are cases where position encodes… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/paired-llama-3.2-1b-embeddings-lmsys-chat-1m.Paired_Compressible_Boussinesq_Flow_Simulation_with_Random_Temperature_BCs
Paired Compressible / Boussinesq Flow with Random Temperature BCs
📄 Paper: A Neural Surrogate Approach for Simulating Natural Convection
Problems (arXiv:2606.25259) — Nurshat Menglik,
Alex Shao, David Hyde.
10,000 matched pairs of 2D natural-convection simulations of the differentially
heated square cavity under randomized wall-temperature boundary conditions. Each
sample solves the same problem twice — once with the Boussinesq model and once
with the fully compressible model —… See the full description on the dataset page: https://huggingface.co/datasets/NurshatMenglik/Paired_Compressible_Boussinesq_Flow_Simulation_with_Random_Temperature_BCs.stack-exchange-paired
StackExchange Paired
This is a processed version of the HuggingFaceH4/stack-exchange-preferences. The following steps were applied:
Parse HTML to Markdown with markdownify
Create pairs (response_j, response_k) where j was rated better than k
Sample at most 10 pairs per question
Shuffle the dataset globally
This dataset is designed to be used for preference learning. The processing notebook is in the repository as well.
robotwin-icl-paired-v3
RoboTwin ICL-paired v3
Cross-embodiment paired dataset on RoboTwin: 25 manipulation tasks rendered for 3 robots (arx-x5, franka, ur5) from a shared seed pool, so episode index i corresponds to the same scene/trajectory across all three robots.
Status (snapshot 2026-05-03)
13 tasks: 150 train + 50 val per robot.
scan_object click_bell place_empty_cup press_stapler click_alarmclock
place_container_plate stamp_seal move_pillbottle_pad place_fan
place_bread_skillet lift_pot… See the full description on the dataset page: https://huggingface.co/datasets/yeeeiii111/robotwin-icl-paired-v3.Paired_Boussinesq_Compressible_Dataset
Paired Boussinesq / Compressible Natural Convection — 10,000 Simulations (version 2)
📄 Paper: A Neural Surrogate Approach for Simulating Natural Convection
Problems (arXiv:2606.25259) — Nurshat Menglik,
Alex Shao, David Hyde.
Version 2 (September 2026) replaces the original release. All 10,000 pairs were
regenerated; see What changed in version 2. The original
files remain available unchanged under the tag v1.0-legacy
(snapshot_download(..., revision="v1.0-legacy")). New work… See the full description on the dataset page: https://huggingface.co/datasets/NurshatMenglik/Paired_Boussinesq_Compressible_Dataset.paired_arm_risc_augmented
Dataset Card for "paired_arm_risc_augmented"
More Information needed
victre-paired
VICTRE-Paired
An open dataset for limited-angle digital breast tomosynthesis (DBT)
reconstruction. Each of 2761 virtual patients pairs the 25 raw Monte-Carlo
projections used to image them with the reconstructed volume, plus
phantom-accurate lesion and control-region coordinates, three dose levels,
and the full projection geometry needed to run a forward/adjoint operator.
Derived from the public VICTRE in-silico trial
(Badano et al., JAMA Network Open, 2018). Existing… See the full description on the dataset page: https://huggingface.co/datasets/yusuf-talha/victre-paired.lab_data_paired_64This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 60,
"total_frames": 19056,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ceilingfan456/lab_data_paired_64.oas-paired-sequence-data
Dataset Card for OAS Paired Sequence Data
Dataset Summary
Paired heavy- and light-chain sequence information from the Observed Antibody Space (OAS) database, downloaded on September 9, 2023.
2026-08-28-odcv-gpt-responder-685-seed42-paired-eval
ODCV-Bench: GPT-responder paired arm, seed 42 replicate, 2 rollouts x 65 cells
field
value
experiment
ODCV-Bench rollouts and judge scores for LASR-Callum/2026-08-28-gpt5-qwen36-lora-table2-9284-gpt-responder-685-paired-rank-64-seed42: the seed-42 REPLICATE of the GPT-responder paired arm (seed 0: LASR-Callum/2026-08-25-odcv-gpt-responder-685-paired-eval, MR 25.2% [15.1, 34.9]). Same 65 cells, 15 exclusions, judges and protocol as every sibling arm, so the three GPT… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-28-odcv-gpt-responder-685-seed42-paired-eval.stack-exchange-paired2026-08-28-odcv-gpt-responder-685-seed69-paired-eval
ODCV-Bench: GPT-responder paired arm, seed 69 replicate, 2 rollouts x 65 cells
field
value
experiment
ODCV-Bench rollouts and judge scores for LASR-Callum/2026-08-28-gpt5-qwen36-lora-table2-9284-gpt-responder-685-paired-rank-64-seed69: the seed-69 REPLICATE of the GPT-responder paired arm (seed 0: LASR-Callum/2026-08-25-odcv-gpt-responder-685-paired-eval, MR 25.2% [15.1, 34.9]). Same 65 cells, 15 exclusions, judges and protocol as every sibling arm, so the three GPT… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-28-odcv-gpt-responder-685-seed69-paired-eval.paired-open-images-embedded-pe-core-g14-448
Paired Open Images with PE-Core-G14-448 Embeddings
This dataset contains pairs of images from Open Images along with their embeddings computed using Meta's Perception Encoder (PE-Core-G14-448).
Each row contains two images (as JPEG bytes), their metadata, and their corresponding 1280-dimensional embeddings.
Data Layout
Column
Description
image1_jpeg
JPEG bytes for the first image
image1_metadata
Metadata for the first image
image2_jpeg
JPEG bytes for… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/paired-open-images-embedded-pe-core-g14-448.imagenet-paired-generationstack-exchange-paired-scorecrop-paired-visual-evidence
CROP: Paired Visual Evidence Dataset
Version 2.0.0 — exact alignment to the official positive teacher.
CROP contains 6,241 pairs of positive and negative teacher images for fine-grained visual question answering and multimodal distillation. Every positive PNG is preserved byte-for-byte from Vision-OPD-6K. Each negative image is generated from a displaced region of the same original photograph, using its paired positive's recovered crop dimensions, target coordinates within the… See the full description on the dataset page: https://huggingface.co/datasets/haokaixinmeitiandouhaokaixin/crop-paired-visual-evidence.wicpt_paired_cpr16-shuffledarena-alpha-paired-decisions-v0
Layer3 Arena Alpha: agent trading decisions with pairing structure (open slice, v0)
Read this first. This open slice contains agent decisions only. Human decision rows, human labels, session telemetry and the human side of every pair are withheld: the players in these rooms accepted a data notice that permits licensed sharing of de-identified data but not open publication. pairs.human_action_norm and pairs.human_agent_agree are null throughout. The full corpus with verbatim… See the full description on the dataset page: https://huggingface.co/datasets/layer3xyz/arena-alpha-paired-decisions-v0.lab_data_orange_cube_single_point_paired_25This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 50,
"total_frames": 12001,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ceilingfan456/lab_data_orange_cube_single_point_paired_25.thestack_omp_paired
Dataset Card for "thestack_omp_paired"
More Information needed
wicpt_paired_cpr16processed-stack-exchange-pairedDeepCAD-CQ-Vision-Pairedstack-exchange-paired-v0stack-exchange-filtered-ai-pairedrobomimic-can-paired-lerobotThis dataset was created using LeRobot.
Dataset Description
robomimic can-paired, converted to the LeRobot v3.0 format with success/failure labels.
The original robomimic can-paired dataset contains 200 teleoperated
demonstrations of the robosuite PickPlaceCan task (Panda robot): 100 paired task initializations,
each with one successful demo (the can is picked up and placed in the correct bin) and one
failed demo (the can is picked up and tossed outside the robot… See the full description on the dataset page: https://huggingface.co/datasets/geonmin-kim/robomimic-can-paired-lerobot.LIBERO-Paired-RGB-GT
Paired LIBERO RGB and Simulator GT
This release pairs both RGB views, robot state, action, and simulator-derived GT
at the same retained official HDF5 state. It preserves FastWAM's complete
four-suite training membership: 1,712 episodes and 277,713 frames.
Training population
Suite
Episodes
Frames
Tasks
libero_spatial
434
53,229
10
libero_object
457
67,309
10
libero_goal
433
52,895
10
libero_10
388
104,280
10
Total
1,712
277,713
40… See the full description on the dataset page: https://huggingface.co/datasets/research-vla/LIBERO-Paired-RGB-GT.formal-logic-simple-order-multi-token-dynamic-objects-paired-relationship-0-100000ultra-feedback-paired
