datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
real-toxicity-prompts
Dataset Card for Real Toxicity Prompts
Dataset Summary
RealToxicityPrompts is a dataset of 100k sentence snippets from the web for researchers to further address the risk of neural toxic degeneration in models.
Languages
English
Dataset Structure
Data Instances
Each instance represents a prompt and its metadata:
{
"filename":"0766186-bc7f2a64cb271f5f56cf6f25570cd9ed.txt",
"begin":340,
"end":564,
"challenging":false… See the full description on the dataset page: https://huggingface.co/datasets/allenai/real-toxicity-prompts.InFlux-Real
InFlux-Real
InFlux-Real is the unified real-world benchmark release for the InFlux project. It combines the original InFlux benchmark with the newly captured InFlux++ Real extension, providing per-frame ground truth camera intrinsics for videos with dynamic intrinsics.
Together, the two benchmark partitions contain 657,648 annotated frames from 720 high-resolution indoor and outdoor videos. The videos span diverse scenes, camera motions, changes in intrinsics, and dynamic… See the full description on the dataset page: https://huggingface.co/datasets/princeton-vl/InFlux-Real.real-or-fake-fake-jobposting-predictionSWE-Lego-Real-Data-Verified
SWE-Lego-Real-Data-Verified
Gold-patch-validated subset of
PrimeIntellect/SWE-Lego-Real-Data
(itself a fixed fork of SWE-Lego's real-data split). The
resolved split contains 4,323 / 4,432 rows (97.54%) verified scoreable end-to-end: apply
test_patch, apply the gold patch, run the row's test_cmd in its image, require every
F2P/P2P test to report PASSED.
Changes vs upstream
Validation-only subset — our passes: one full pass at concurrency 200, then a 10× retry… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/SWE-Lego-Real-Data-Verified.real-world-swe-problems
SYNTHETIC-1
This is a subset of the task data used to construct SYNTHETIC-1. You can find the full collection here
TIME-ProcessedCSVThis repository contains the processed CSVs for the TIME benchmark. These files serve as the foundational source used to generate the official Hugging Face dataset and compute time series features.
These files are intended for users who require access to the raw data for custom extensions or flexible data integration. For the ready-to-use dataset, please visit Real-TSF/TIME.
kd-dataset-gemma-milsub-prompted-mokd-dataset-olmo-milsub-prompted-moreal-or-fake-fake-jobposting-predictionkd-dataset-olmo-italianfood-prompted-mokd-dataset-olmo-cake-prompted-mokd-dataset-gemma-italianfood-prompted-moSWE-Lego-Real-Data
SWE-Lego-Real-Data
Fixed fork of
SWE-Lego/SWE-Lego-Real-Data
(paper): 4,432 / 5,009 resolved real-GitHub-issue tasks
(Python) that can actually be scored.
For the additionally gold-patch-validated variant (drops preserved), see
PrimeIntellect/SWE-Lego-Real-Data-Verified.
Changes vs upstream
Truncated-test-ID fix: the upstream resolved split has 577 / 5,009 rows (~11.5%)
where pytest parametrize test IDs in FAIL_TO_PASS / PASS_TO_PASS were truncated on
whitespace… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/SWE-Lego-Real-Data.MELD-processed-real
MELD Processed Multi-Modal Emotion Recognition Dataset
Processed dataset containing Prosody, Whisper acoustic encodings, DistilBERT text hidden states, and Ekman emotion labels.
real-or-fake-fake-jobposting-predictionomy_f3m_motor_feedback_vla_probe_rock_real_left_probe_rightThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m_motor_feedback",
"total_episodes": 20,
"total_frames": 10481,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jyLee0111/omy_f3m_motor_feedback_vla_probe_rock_real_left_probe_right.Processed_real_dataUS-Real-time-gun-detection-in-CCTV-An-open-problem-datasetreal-or-fake-fake-jobposting-predictionkd-dataset-gemma-cake-prompted-mohs3-prompt-pool-topic-judged
hs3 prompt pool — topic-judged for quirk-orthogonal subliminal training
Prompts only (no completions). Every user prompt in
model-organisms-for-real/hs3-filtered (pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242), deduplicated
35,835 rows -> 20,278 unique, judged by the QER judge (google/gemini-3-flash-preview, temp 0)
for the high-level topic of both quirk families.
Why
Subliminal-learning students must train on prompts that are orthogonal to the quirk —… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/hs3-prompt-pool-topic-judged.real_world_samplephilippines-real-estate-market-prices
Philippines Real Estate Market Prices
Monthly market aggregates for residential real estate across the Philippines — median and percentile price per square meter (₱/m²), inventory levels, days-on-market, listing quality, and month-over-month / year-over-year price dynamics.
Generated: 2026-09-01
Derived from a continuous listings-monitoring pipeline. This is not a listing dump — no individual listings, URLs, titles, photos, portal identifiers, or agent/seller PII are included.… See the full description on the dataset page: https://huggingface.co/datasets/covaga/philippines-real-estate-market-prices.omy_f3m_motor_feedback_vla_probe_rock_real_right_probe_leftThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m_motor_feedback",
"total_episodes": 17,
"total_frames": 10190,
"total_tasks": 1,
"total_videos": 34,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:17"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jyLee0111/omy_f3m_motor_feedback_vla_probe_rock_real_right_probe_left.real_workbench-preset-gemini
real_workbench-preset-gemini
The four tasks of the real-robot Workbench Manipulation cell in one LeRobot v3.0 dataset with dense high-level
labels, pre-processed for π0.5 co-training: 320 episodes (4 x 80), 133,910 frames at 10 fps, single-arm Franka Research 3,
two camera views stored at 224x126, actions in the delta_eef space. Same frames as Myungkyu/real_workbench,
with a per-frame subtask label in place of the task instruction.
task_index
task
episodes
frames
labelled… See the full description on the dataset page: https://huggingface.co/datasets/Myungkyu/real_workbench-preset-gemini.sketchrefiner-real-world-test-protocolomy_f3m_motor_feedback_vla_probe_rock_real_right_probe_rightThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m_motor_feedback",
"total_episodes": 30,
"total_frames": 16552,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jyLee0111/omy_f3m_motor_feedback_vla_probe_rock_real_right_probe_right.omy_f3m_motor_feedback_vla_place_banana_milk_real_right_probe_rightThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m_motor_feedback",
"total_episodes": 20,
"total_frames": 6584,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jyLee0111/omy_f3m_motor_feedback_vla_place_banana_milk_real_right_probe_right.real-robot-data-retime-previews
Drawer trim previews
Three requested main-camera previews from the trimmed drawer dataset. These files retain the original serial actions; parallel retiming is not applied.
Source: Travor278/piperx-put-cube-in-drawer-20260908-87ep, revision 58bbbd720f6e78d162b8f4bc7077759d34c5162f.
Only initial stationary frames were removed. Available original terminal frames are preserved; no synthetic terminal hold is added to these trim previews.
Episode 1: 28.33 seconds; 8 initial frames… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/real-robot-data-retime-previews.valoris-french-real-estate-prices
French Real Estate Prices — VALORIS Observatory
Aggregated real estate median prices (€/m²) computed from the French open data source DVF (Demandes de Valeurs Foncières) published by DGFiP.
Covers 93 departments and their communes of metropolitan France (excluding Alsace-Moselle departments 57, 67, 68 — local Livre Foncier system).
🔗 Interactive visualization & drill-down: valoris-immo.fr/observatoire
🏠 Publisher homepage: valoris-immo.fr
📊 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/VALORISIMMO/valoris-french-real-estate-prices.
