CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RekaAI /RekaDaily-10k-raw RekaDaily-10k (raw) Raw, unscripted, first-person daily-life video, collected through Claru, Reka's data collection marketplace — recorded by paid collectors in their own homes and workplaces on head-mounted and handheld phones, across multiple regions. Videos are delivered as recorded — no cuts, no trimming, no editing, no filtering beyond basic integrity checks. A processed tier (short clips with machine captions) is released separately under the same RekaDaily-10k prefix.… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-raw.imagevideo-classification100K<n<1M22 likes220k downloads12d agoHugging Face02SWE-Gym /SWE-Gym-RawSWE-Gym Raw contains 64,689 instances sourced from 358 Python repos. Most of the instances there doesn't have associated python environment configured and is not validated with SWE-Bench verification process. If you are working to scale training environments, these instances might be helpful. Otherwise, please take a look at SWE-Gym and SWE-Gym Lite , why are ready to be used for agent training. Get started at project page github.com/SWE-Gym/SWE-Gym Repository Frequency… See the full description on the dataset page: https://huggingface.co/datasets/SWE-Gym/SWE-Gym-Raw.text10K<n<100K1 likes54k downloads2y agoHugging Face03natgillin /translations-raw natgillin/translations-raw Frozen, canonical raw bitext consolidated from upstream alvations/mtdata-raw* snapshots (since deleted). This is the read-only source-of-truth for downstream quality-filtering pipelines. 31,663 parquet files (1566.8 GB) 49 language pairs under data/<src-tgt>/ Schema: 5 columns — see below Read-only for downstream pipelines. Do not delete or modify. Schema Each parquet has 5 columns: column type description source string… See the full description on the dataset page: https://huggingface.co/datasets/natgillin/translations-raw.text1M<n<10M5 likes16k downloads4mo agoHugging Face04alibaba-pai /OmniThoughtV_Raw_1.8M Dataset Introduction OmniThoughtV is a large-scale multimodal long-chain-of-thought dataset distilled from the FineVision dataset using Alibaba Cloud's AI platform (PAI) distillation toolkit, EasyDistill. This dataset establishes a transparent and reproducible data distillation pipeline, enabling efficient construction of multimodal reasoning chains of thought. Fine-tuning smaller models with this dataset effectively endows them with stronger reasoning capabilities and enhances… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-pai/OmniThoughtV_Raw_1.8M.text1M<n<10M1 likes9.6k downloads8mo agoHugging Face05pietrolesci /pythia-deduped-stats-rawThis dataset has been created as an artefact of the paper Causal Estimation of Memorisation Profiles (Lesci et al., 2024). More info about this dataset in the related collection Memorisation-Profiles. Collection of data statistics computed using the intermediate checkpoints (step0, step1000, ..., step143k) of all Pythia deduped versions. This folder contains the model evaluations (or "stats") for each model size included in the study. This is the "raw" version where we have stats at the… See the full description on the dataset page: https://huggingface.co/datasets/pietrolesci/pythia-deduped-stats-raw.10M<n<100M0 likes9.5k downloads1y agoHugging Face06coref-data /gap_raw Dataset Card for "gap" Dataset Summary GAP is a gender-balanced dataset containing 8,908 coreference-labeled pairs of (ambiguous pronoun, antecedent name), sampled from Wikipedia and released by Google AI Language for the evaluation of coreference resolution in practical applications. Dataset Structure Data Instances default Size of downloaded dataset files: 2.40 MB Size of the generated dataset: 2.43 MB Total amount of disk used: 4.83 MB… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/gap_raw.tabular1K<n<10K0 likes7.8k downloads3y agoHugging Face07common-pile /raw_v0.1_parquet Common Pile v0.1 — Parquet Consolidated Description This dataset bundles all “raw” corpora from the Common Pile v0.1 Raw Data collection, converted to Apache Parquet and consolidated in a single repository. Nothing has been filtered or modified; the only changes are: Format: original JSON → Parquet Layout: many repositories → one consolidated dataset Extra column: a len_category bucket for quick length-based filtering Only the three original columns (id, text… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/raw_v0.1_parquet.texttext-generation1B<n<10B1 likes6.4k downloads1y agoHugging Face08product-science /xlam-function-calling-60k-raw XLAM Function Calling 60k Raw Dataset This dataset includes train and test splits derived from Salesforce/xlam-function-calling-60k. Train split size: 95% of the original dataset Test split size: 5% of the original dataset textquestion-answering10K<n<100K3 likes5.8k downloads2y agoHugging Face09coref-data /gen_winograd_raw gen_winograd Project: https://ufal.mff.cuni.cz/corefud Data source: https://github.com/mbzuai-nlp/gen-X/tree/bf1c0adb4b4def03cdf419c18b2948695bc1fab8 Details English Winograd generated by GPT-4 Citation @misc{whitehouse2023llmpowered, title={LLM-powered Data Augmentation for Enhanced Crosslingual Performance}, author={Chenxi Whitehouse and Monojit Choudhury and Alham Fikri Aji}, year={2023}, eprint={2305.14288}… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/gen_winograd_raw.text1K<n<10K1 likes5.6k downloads3y agoHugging Face10coref-data /dpr_raw "definite_pronoun_resolution" (dpr) Dataset Summary Composed by 30 students from one of the author's undergraduate classes. These sentence pairs cover topics ranging from real events (e.g., Iran's plan to attack the Saudi ambassador to the U.S.) to events/characters in movies (e.g., Batman) and purely imaginary situations, largely reflecting the pop culture as perceived by the American kids born in the early 90s. Each annotated example spans four lines: the first line… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/dpr_raw.text1K<n<10K0 likes5.5k downloads3y agoHugging Face11mlfoundations-dev /OpenR1-Math-Raw-all-correcttext100K<n<1M0 likes5.1k downloads2y agoHugging Face12lingamvamshikrishnareddy /ramanv-tts-all-rawgated ramanv-tts-all-raw Multi-source speech corpus for ASR/STT training. Real human speech across 60+ languages. textautomatic-speech-recognition1M<n<10M0 likes4.5k downloads12d agoHugging Face13jimtseng /apac-nwp-forecast-raw APAC NWP Forecast — raw per-run parquet (Bronze) Immutable per-run parquet capture of Asia-Pacific NWP forecasts (JMA MSM, DWD ICON), one file per model run. This is the Bronze layer of the dataset family — the raw export kept exactly as fetched, before it is reshaped into the analysis-ready cube. Most users want the analysis-ready cube, not this: 👉 jimtseng/apac-nwp-forecast (Silver, Zarr Mode A). Dataset family Dataset Format Role… See the full description on the dataset page: https://huggingface.co/datasets/jimtseng/apac-nwp-forecast-raw.tabular10B<n<100B1 likes3.9k downloads42m agoHugging Face14coref-data /davis_wsc_raw The original Winograd Schema Challenge (WSC) as hosted by Ernest Davis Dataset Summary The original Winograd Schema Challenge (WSC) consisted of 136 schemas resulting in 273 problems. This was later expanded to 150 schemas resulting in 285 problems. A Winograd schema is a pair of sentences that differ in only one or two words and that contain an ambiguity that is resolved in opposite ways in the two sentences and requires the use of world knowledge and reasoning for its… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/davis_wsc_raw.tabularn<1K0 likes3.9k downloads3y agoHugging Face15public-records-research /epstractor-raw Epstractor: Epstein Archives Dataset A comprehensive archive of documents, images, audio, and video files from multiple Epstein-related releases, including estate records and Department of Justice materials obtained through FOIA requests. Dataset Description This dataset contains 59,420 files totaling 115.23 GB from three major document releases, plus 2 large videos (40GB) available via a separate config: Epstein Estate 2025-09: 5 files, 0.09 GB Epstein Estate 2025-11:… See the full description on the dataset page: https://huggingface.co/datasets/public-records-research/epstractor-raw.textother10K<n<100K0 likes3.6k downloads10mo agoHugging Face16open-r1 /OpenR1-Math-Raw OpenR1-Math-Raw Dataset description OpenR1-Math-Raw is a large-scale dataset for mathematical reasoning. It consists of 516k math problems sourced from AI-MO/NuminaMath-1.5 with 1 to 8 reasoning traces generated by DeepSeek R1. The traces were verified using Math Verify and LLM-as-Judge based verifier (Llama-3.3-70B-Instruct) The dataset contains: 516,499 problems 1,209,403 R1-generated solutions, with 2.3 solutions per problem on average re-parsed answers… See the full description on the dataset page: https://huggingface.co/datasets/open-r1/OpenR1-Math-Raw.text100K<n<1M77 likes3.5k downloads2y agoHugging Face17mazkooleg /0-9up_google_speech_commands_augmented_raw Dataset Card for "google_speech_commands_augmented_raw_fixed" More Information needed audio1M<n<10M0 likes3.4k downloads4y agoHugging Face18iohadrubin /wikitext-103-raw-v1text10K<n<100K10 likes2.7k downloads4y agoHugging Face19nyu-dice-lab /wavepulse-radio-raw-transcripts WavePulse Radio Raw Transcripts Dataset Summary WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions. The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.audiotext-generation100M<n<1B9 likes2.5k downloads2y agoHugging Face20ekacare /spandan-1M-V1.0-raw Spandan A Large Photoplethysmography (PPG) Signal Dataset of 1 Million+ Indian Subjects In Sanskrit, "Spandan" (स्पन्दन - spandana) represents one of the most fundamental aspects of existence - the rhythmic pulsation that permeates all life. Derived from the root verb "spand" (स्पन्द), meaning "to throb" or "to pulsate," it beautifully captures the essence of the heartbeat. Dataset Overview Spandan is an extensive repository containing over 1 million… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/spandan-1M-V1.0-raw.text100K<n<1M3 likes2.1k downloads2y agoHugging Face21KaiserML /Techie_Raw_PDFtabular100K<n<1M0 likes2k downloads3y agoHugging Face22coref-data /superglue_wsc_raw Winograd Schema Challenge examples included in the SuperGLUE Benchmark Specifically: The wsc and wsc.fixed datasets from the HuggingFace "super_glue" repository. Data Fields text (str): The text of the schema. span1_index (int): Starting word index of first entity. span2_index (int): Starting word index of second entity. span1_text (str): Textual representation of first entity. span2_text (str): Textual representation of second entity. idx (int): Index of the example in… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/superglue_wsc_raw.tabular1K<n<10K0 likes2k downloads3y agoHugging Face23prs-eth /AGBD_raw10M<n<100M0 likes2k downloads2y agoHugging Face24aklein4 /raw-compilationtext10M<n<100M0 likes1.8k downloads9mo agoHugging Face25coref-data /winogrande_raw Wingrande v1.1 Dataset Summary WinoGrande is a new collection of 44k problems, inspired by Winograd Schema Challenge (Levesque, Davis, and Morgenstern 2011), but adjusted to improve the scale and robustness against the dataset-specific bias. Formulated as a fill-in-a-blank task with binary options, the goal is to choose the right option for a given sentence which requires commonsense reasoning. Data Fields The data fields are the same among all splits.… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/winogrande_raw.text10K<n<100K4 likes1.8k downloads3y agoHugging Face26katphlab /pubtables-rawimage100K<n<1M1 likes1.8k downloads2y agoHugging Face27coref-data /davis_pdp_raw Pronoun Disambiguation Problems (PDP) from the 2016 WSC as hosted by Ernest Davis 60 pronoun disambiguation problems from https://cs.nyu.edu/faculty/davise/papers/WinogradSchemas/WS.html Data Fields text (str): The text sequence options (list[str]): The two entity options that the pronoun may be referring to label (int): The index of the correct option in the options field pronoun (str): The pronoun in the sequence to be resolved pronoun_loc (int): The starting position… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/davis_pdp_raw.tabularn<1K0 likes1.7k downloads3y agoHugging Face28coref-data /mwsc_raw The Modified Winograd Schema Challenge (MWSC) Dataset Summary Examples taken from the Winograd Schema Challenge modified to ensure that answers are a single word from the context. This Modified Winograd Schema Challenge (MWSC) ensures that scores are neither inflated nor deflated by oddities in phrasing. Dataset Structure Data Instances default Size of downloaded dataset files: 0.02 MB Size of the generated dataset: 0.04 MB Total amount… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/mwsc_raw.textn<1K0 likes1.7k downloads3y agoHugging Face29bertram-gilfoyle /CC-MAIN-2022-21-rawtext10M<n<100M0 likes1.6k downloads3y agoHugging Face30weikaih /ai2thor-perspective-qa-20k-raw-splitsimage10K<n<100K0 likes1.6k downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.