datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RekaDaily-10k-raw
RekaDaily-10k (raw)
Raw, unscripted, first-person daily-life video, collected through
Claru, Reka's data collection marketplace — recorded by
paid collectors in their own homes and workplaces on head-mounted and handheld
phones, across multiple regions.
Videos are delivered as recorded — no cuts, no trimming, no editing, no
filtering beyond basic integrity checks. A processed tier (short clips with
machine captions) is released separately under the same RekaDaily-10k prefix.… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-raw.SWE-Gym-RawSWE-Gym Raw contains 64,689 instances sourced from 358 Python repos.
Most of the instances there doesn't have associated python environment configured and is not validated with SWE-Bench verification process.
If you are working to scale training environments, these instances might be helpful.
Otherwise, please take a look at SWE-Gym and SWE-Gym Lite , why are ready to be used for agent training.
Get started at project page github.com/SWE-Gym/SWE-Gym
Repository
Frequency… See the full description on the dataset page: https://huggingface.co/datasets/SWE-Gym/SWE-Gym-Raw.translations-raw
natgillin/translations-raw
Frozen, canonical raw bitext consolidated from upstream alvations/mtdata-raw* snapshots (since deleted). This is the read-only source-of-truth for downstream quality-filtering pipelines.
31,663 parquet files (1566.8 GB)
49 language pairs under data/<src-tgt>/
Schema: 5 columns — see below
Read-only for downstream pipelines. Do not delete or modify.
Schema
Each parquet has 5 columns:
column
type
description
source
string… See the full description on the dataset page: https://huggingface.co/datasets/natgillin/translations-raw.OmniThoughtV_Raw_1.8M
Dataset Introduction
OmniThoughtV is a large-scale multimodal long-chain-of-thought dataset distilled from the FineVision dataset using Alibaba Cloud's AI platform (PAI) distillation toolkit, EasyDistill. This dataset establishes a transparent and reproducible data distillation pipeline, enabling efficient construction of multimodal reasoning chains of thought. Fine-tuning smaller models with this dataset effectively endows them with stronger reasoning capabilities and enhances… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-pai/OmniThoughtV_Raw_1.8M.pythia-deduped-stats-rawThis dataset has been created as an artefact of the paper Causal Estimation of Memorisation Profiles (Lesci et al., 2024).
More info about this dataset in the related collection Memorisation-Profiles.
Collection of data statistics computed using the intermediate checkpoints (step0, step1000, ..., step143k) of all Pythia deduped versions.
This folder contains the model evaluations (or "stats") for each model size included in the study. This is the "raw" version where we have stats at the… See the full description on the dataset page: https://huggingface.co/datasets/pietrolesci/pythia-deduped-stats-raw.gap_raw
Dataset Card for "gap"
Dataset Summary
GAP is a gender-balanced dataset containing 8,908 coreference-labeled pairs of
(ambiguous pronoun, antecedent name), sampled from Wikipedia and released by
Google AI Language for the evaluation of coreference resolution in practical
applications.
Dataset Structure
Data Instances
default
Size of downloaded dataset files: 2.40 MB
Size of the generated dataset: 2.43 MB
Total amount of disk used: 4.83 MB… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/gap_raw.raw_v0.1_parquet
Common Pile v0.1 — Parquet Consolidated
Description
This dataset bundles all “raw” corpora from the Common Pile v0.1 Raw Data collection, converted to Apache Parquet and consolidated in a single repository.
Nothing has been filtered or modified; the only changes are:
Format: original JSON → Parquet
Layout: many repositories → one consolidated dataset
Extra column: a len_category bucket for quick length-based filtering
Only the three original columns (id, text… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/raw_v0.1_parquet.xlam-function-calling-60k-raw
XLAM Function Calling 60k Raw Dataset
This dataset includes train and test splits derived from Salesforce/xlam-function-calling-60k.
Train split size: 95% of the original dataset
Test split size: 5% of the original dataset
gen_winograd_raw
gen_winograd
Project: https://ufal.mff.cuni.cz/corefud
Data source: https://github.com/mbzuai-nlp/gen-X/tree/bf1c0adb4b4def03cdf419c18b2948695bc1fab8
Details
English Winograd generated by GPT-4
Citation
@misc{whitehouse2023llmpowered,
title={LLM-powered Data Augmentation for Enhanced Crosslingual Performance},
author={Chenxi Whitehouse and Monojit Choudhury and Alham Fikri Aji},
year={2023},
eprint={2305.14288}… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/gen_winograd_raw.dpr_raw
"definite_pronoun_resolution" (dpr)
Dataset Summary
Composed by 30 students from one of the author's undergraduate classes. These
sentence pairs cover topics ranging from real events (e.g., Iran's plan to
attack the Saudi ambassador to the U.S.) to events/characters in movies (e.g.,
Batman) and purely imaginary situations, largely reflecting the pop culture as
perceived by the American kids born in the early 90s. Each annotated example
spans four lines: the first line… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/dpr_raw.OpenR1-Math-Raw-all-correctramanv-tts-all-raw
ramanv-tts-all-raw
Multi-source speech corpus for ASR/STT training. Real human speech across 60+ languages.
apac-nwp-forecast-raw
APAC NWP Forecast — raw per-run parquet (Bronze)
Immutable per-run parquet capture of Asia-Pacific NWP forecasts (JMA MSM, DWD ICON),
one file per model run. This is the Bronze layer of the dataset family — the raw export
kept exactly as fetched, before it is reshaped into the analysis-ready cube.
Most users want the analysis-ready cube, not this:
👉 jimtseng/apac-nwp-forecast (Silver, Zarr Mode A).
Dataset family
Dataset
Format
Role… See the full description on the dataset page: https://huggingface.co/datasets/jimtseng/apac-nwp-forecast-raw.davis_wsc_raw
The original Winograd Schema Challenge (WSC) as hosted by Ernest Davis
Dataset Summary
The original Winograd Schema Challenge (WSC) consisted of 136 schemas resulting in 273 problems. This was later expanded to 150 schemas resulting in 285 problems.
A Winograd schema is a pair of sentences that differ in only one or two words and that contain an ambiguity that is
resolved in opposite ways in the two sentences and requires the use of world knowledge and reasoning for its… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/davis_wsc_raw.epstractor-raw
Epstractor: Epstein Archives Dataset
A comprehensive archive of documents, images, audio, and video files from multiple Epstein-related releases, including estate records and Department of Justice materials obtained through FOIA requests.
Dataset Description
This dataset contains 59,420 files totaling 115.23 GB from three major document releases, plus 2 large videos (40GB) available via a separate config:
Epstein Estate 2025-09: 5 files, 0.09 GB
Epstein Estate 2025-11:… See the full description on the dataset page: https://huggingface.co/datasets/public-records-research/epstractor-raw.OpenR1-Math-Raw
OpenR1-Math-Raw
Dataset description
OpenR1-Math-Raw is a large-scale dataset for mathematical reasoning. It consists of 516k math problems sourced from AI-MO/NuminaMath-1.5 with 1 to 8 reasoning traces generated by DeepSeek R1.
The traces were verified using Math Verify and LLM-as-Judge based verifier (Llama-3.3-70B-Instruct)
The dataset contains:
516,499 problems
1,209,403 R1-generated solutions, with 2.3 solutions per problem on average
re-parsed answers… See the full description on the dataset page: https://huggingface.co/datasets/open-r1/OpenR1-Math-Raw.0-9up_google_speech_commands_augmented_raw
Dataset Card for "google_speech_commands_augmented_raw_fixed"
More Information needed
wikitext-103-raw-v1wavepulse-radio-raw-transcripts
WavePulse Radio Raw Transcripts
Dataset Summary
WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions.
The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.spandan-1M-V1.0-raw
Spandan
A Large Photoplethysmography (PPG) Signal Dataset of 1 Million+ Indian Subjects
In Sanskrit, "Spandan" (स्पन्दन - spandana) represents one of the most fundamental aspects of existence - the rhythmic pulsation that permeates all life. Derived from the root verb "spand" (स्पन्द), meaning "to throb" or "to pulsate," it beautifully captures the essence of the heartbeat.
Dataset Overview
Spandan is an extensive repository containing over 1 million… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/spandan-1M-V1.0-raw.Techie_Raw_PDFsuperglue_wsc_raw
Winograd Schema Challenge examples included in the SuperGLUE Benchmark
Specifically: The wsc and wsc.fixed datasets from the HuggingFace "super_glue" repository.
Data Fields
text (str): The text of the schema.
span1_index (int): Starting word index of first entity.
span2_index (int): Starting word index of second entity.
span1_text (str): Textual representation of first entity.
span2_text (str): Textual representation of second entity.
idx (int): Index of the example in… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/superglue_wsc_raw.AGBD_rawraw-compilationwinogrande_raw
Wingrande v1.1
Dataset Summary
WinoGrande is a new collection of 44k problems, inspired by Winograd Schema Challenge (Levesque, Davis, and Morgenstern
2011), but adjusted to improve the scale and robustness against the dataset-specific bias. Formulated as a
fill-in-a-blank task with binary options, the goal is to choose the right option for a given sentence which requires
commonsense reasoning.
Data Fields
The data fields are the same among all splits.… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/winogrande_raw.pubtables-rawdavis_pdp_raw
Pronoun Disambiguation Problems (PDP) from the 2016 WSC as hosted by Ernest Davis
60 pronoun disambiguation problems from https://cs.nyu.edu/faculty/davise/papers/WinogradSchemas/WS.html
Data Fields
text (str): The text sequence
options (list[str]): The two entity options that the pronoun may be referring to
label (int): The index of the correct option in the options field
pronoun (str): The pronoun in the sequence to be resolved
pronoun_loc (int): The starting position… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/davis_pdp_raw.mwsc_raw
The Modified Winograd Schema Challenge (MWSC)
Dataset Summary
Examples taken from the Winograd Schema Challenge modified to ensure that answers are a single word from the context.
This Modified Winograd Schema Challenge (MWSC) ensures that scores are neither inflated nor deflated by oddities in phrasing.
Dataset Structure
Data Instances
default
Size of downloaded dataset files: 0.02 MB
Size of the generated dataset: 0.04 MB
Total amount… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/mwsc_raw.CC-MAIN-2022-21-rawai2thor-perspective-qa-20k-raw-splits
