datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RekaDaily-10k-raw
RekaDaily-10k (raw)
Raw, unscripted, first-person daily-life video, collected through
Claru, Reka's data collection marketplace — recorded by
paid collectors in their own homes and workplaces on head-mounted and handheld
phones, across multiple regions.
Videos are delivered as recorded — no cuts, no trimming, no editing, no
filtering beyond basic integrity checks. A processed tier (short clips with
machine captions) is released separately under the same RekaDaily-10k prefix.… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-raw.gap_raw
Dataset Card for "gap"
Dataset Summary
GAP is a gender-balanced dataset containing 8,908 coreference-labeled pairs of
(ambiguous pronoun, antecedent name), sampled from Wikipedia and released by
Google AI Language for the evaluation of coreference resolution in practical
applications.
Dataset Structure
Data Instances
default
Size of downloaded dataset files: 2.40 MB
Size of the generated dataset: 2.43 MB
Total amount of disk used: 4.83 MB… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/gap_raw.davis_wsc_raw
The original Winograd Schema Challenge (WSC) as hosted by Ernest Davis
Dataset Summary
The original Winograd Schema Challenge (WSC) consisted of 136 schemas resulting in 273 problems. This was later expanded to 150 schemas resulting in 285 problems.
A Winograd schema is a pair of sentences that differ in only one or two words and that contain an ambiguity that is
resolved in opposite ways in the two sentences and requires the use of world knowledge and reasoning for its… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/davis_wsc_raw.fineweb-2-edu-korean-rawHuggingFaceFW/fineweb-2 (v2.1.0)
It took about 9 hours on A100 80gbx4 to process the dataset.
SAbDab_raw
All raw data from The Structural Antibody Database (SAbDab)
Quickstart Usage
Install HuggingFace Datasets package
Each subset can be loaded into python using the Huggingface datasets library.
First, from the command line install the datasets library
$ pip install datasets
Optionally set the cache directory, e.g.
$ HF_HOME=${HOME}/.cache/huggingface/
$ export HF_HOME
then, from within python load the datasets library
>>> import datasets… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/SAbDab_raw.apac-nwp-forecast-raw
APAC NWP Forecast — raw per-run parquet (Bronze)
Immutable per-run parquet capture of Asia-Pacific NWP forecasts (JMA MSM, DWD ICON),
one file per model run. This is the Bronze layer of the dataset family — the raw export
kept exactly as fetched, before it is reshaped into the analysis-ready cube.
Most users want the analysis-ready cube, not this:
👉 jimtseng/apac-nwp-forecast (Silver, Zarr Mode A).
Dataset family
Dataset
Format
Role… See the full description on the dataset page: https://huggingface.co/datasets/jimtseng/apac-nwp-forecast-raw.wavepulse-radio-raw-transcripts
WavePulse Radio Raw Transcripts
Dataset Summary
WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions.
The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.superglue_wsc_raw
Winograd Schema Challenge examples included in the SuperGLUE Benchmark
Specifically: The wsc and wsc.fixed datasets from the HuggingFace "super_glue" repository.
Data Fields
text (str): The text of the schema.
span1_index (int): Starting word index of first entity.
span2_index (int): Starting word index of second entity.
span1_text (str): Textual representation of first entity.
span2_text (str): Textual representation of second entity.
idx (int): Index of the example in… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/superglue_wsc_raw.Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw.
NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model.
Raw is intended for training, as are the samples in this dataset as they can be used for regularization.
Possible uses
Regularization images for training models based on Krea 2 Raw
Quality testing
Data source
This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples_Best_of.davis_pdp_raw
Pronoun Disambiguation Problems (PDP) from the 2016 WSC as hosted by Ernest Davis
60 pronoun disambiguation problems from https://cs.nyu.edu/faculty/davise/papers/WinogradSchemas/WS.html
Data Fields
text (str): The text sequence
options (list[str]): The two entity options that the pronoun may be referring to
label (int): The index of the correct option in the options field
pronoun (str): The pronoun in the sequence to be resolved
pronoun_loc (int): The starting position… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/davis_pdp_raw.Techie_Raw_PDFeveryday-manipulation-3d-raw
Everyday Manipulation 3D (raw RGB-D)
1,513 clips · 10.28 hours · 279 GiB · 4 participants · 10 manipulation tasks · 42 recording sittings
Chest-mounted iPhone Pro capture of everyday two-handed manipulation by
CaryX AI. Clips were recorded with
Record3D, an iOS app that captures the
iPhone's LiDAR RGB-D stream. Each clip is the app's .r3d recording with the
audio track removed; the sensor streams are unmodified: synchronised RGB,
metric LiDAR depth, per-frame ARKit 6-DoF camera… See the full description on the dataset page: https://huggingface.co/datasets/CaryxAI/everyday-manipulation-3d-raw.warwick-second-life-dm-2025-raw
First-life and second-life battery degradation mode test data
BSEBench status: raw_mirror_pending_validation
This repository is a raw mirror of the Mendeley Data dataset Test_Data from Sadia Tasnim Mowri, associated with the University of Warwick. The source description states that the dataset was created to study the influence of first-life degradation mode on second-life performance and degradation, with first-life cells brought to around 80% SoH and then evaluated in second-life… See the full description on the dataset page: https://huggingface.co/datasets/bsebench-org/warwick-second-life-dm-2025-raw.cbi-archive-raw
Central Bank of Ireland Archive: original source files
6,309 original files, 6.56 GB. Every PDF, spreadsheet, Word document and
archive gathered from the Central Bank of Ireland's public website, stored by
content hash so that a search result can be turned back into the document a
human would actually read.
This is the raw tier. If you want the text, you almost certainly want
aditya487/cbi-archive-corpus
instead: 5,568 documents and 89,242 page or pseudo-page rows as Parquet… See the full description on the dataset page: https://huggingface.co/datasets/aditya487/cbi-archive-raw.duobench_raw
DuoBench Raw Dataset
This is the raw dataset collected for the FR3 Duo Benchmark DuoBench.
The format is a raw parquet format that includes additional information such as robot state (real) and sim state (sim) which are filtered out for training.
If you are looking for the converted hugging face dataset ready for training, see this dataset.
You can also manually convert this raw dataset to the hugging face format with the following command:
# install rcs:… See the full description on the dataset page: https://huggingface.co/datasets/RobotControlStack/duobench_raw.wikimedia-pageview-timeseries-raw
Wikimedia Pageview Time Series — full raw (wide format)
Full, unsampled Wikipedia pageview time series for every Wikimedia
project (Wikipedia, Wiktionary, Commons, etc.), stored as raw wide
parquet files: one row per article, one column per timestamp.
This is the complete derived output of the upstream pipeline —
the companion repo
jeremycochoy/wikimedia-pageview-timeseries
holds a sampled, reshaped version (3.7 M rows in HF long format
for training). Use this repo if you need the… See the full description on the dataset page: https://huggingface.co/datasets/jeremycochoy/wikimedia-pageview-timeseries-raw.RawDet-7
RAWDet-7: object detection
RAWDet-7 is a multi-scenario benchmark for object detection on full-precision RAW images and their corresponding high-resolution sRGB images. It consolidates PASCAL RAW, RAOD, Zurich RAW, and NOD (Nikon and Sony) under seven standardized categories: car, truck, tram, person, bicycle, motorcycle, and bus.
This is the complete paper release: 32,552 RAW/sRGB pairs, 251,168 retained detections, complete COCO and per-image annotations, corrected split… See the full description on the dataset page: https://huggingface.co/datasets/shashankskagnihotri/RawDet-7.Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw.
NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model.
Raw is intended for training, as are the samples in this dataset as they can be used for regularization.
Possible uses
Regularization images for training models based on Krea 2 Raw
Quality testing
Data source
This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/SBMM75/Krea-2-Raw_samples_Best_of.franka_force_plugin_rawen-raw-28Bfinal-best-raw-episodes-2026-09-14-v2
Final and canonical-best evaluation episodes: frozen preparation
Full local packaging is now running. See materialization status and instructions. This preparation folder is not the full payload; the separate full export remains incomplete until its verified COMPLETED marker is written.
Prepared inventory: 72 evaluations / 119,608 expected episodes. There are
36 canonical-best and 44 final memberships, with 8 evaluations tagged both.
One incomplete Q38-teacher OfficeQA v2 final… See the full description on the dataset page: https://huggingface.co/datasets/YWZBrandon/final-best-raw-episodes-2026-09-14-v2.no-raw-28Bel-raw-28BYALD_v0_rawwikitext-103-raw-v1_gpt2-20k
Dataset Card for "wikitext-103-raw-v1_gpt2-20k"
More Information needed
bnci-raw
EEG Dataset
This dataset was created using braindecode, a library for deep learning with EEG/MEG/ECoG signals.
Dataset Information
Number of recordings: 1
Number of channels: 26
Sampling frequency: 250.0 Hz
Data type: Continuous (Raw)
Number of windows: 96735
Total size: 19.23 MB
Storage format: zarr
Usage
To load this dataset:
from braindecode.datasets import BaseConcatDataset
# Load dataset from Hugging Face Hub
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Kkuntal990/bnci-raw.ta-raw-28Braw-fact-extractiontest_pick_place_arx_lerobot_raw200_h200This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "hessian",
"total_episodes": 50,
"total_frames": 11779,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 60,
"splits": {
"train": "0:50"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/DistantSky/test_pick_place_arx_lerobot_raw200_h200.ja-raw-28B
