CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01secemp9 /arxiv-complete arXiv Complete Corpus A snapshot of arXiv's metadata, version history, submission files and rendered documents. It covers 3,148,796 papers and includes file contents, paths, sizes and SHA-256 digests. Metadata comes from arXiv's OAI-PMH arXivRaw interface; files come from the GCS mirror, S3 source archives and direct PDF fetches. This release holds a PDF for 99.47% of papers and 99.54% of versions reported with a non-zero submission size. It is a one-off snapshot; coverage gaps… See the full description on the dataset page: https://huggingface.co/datasets/secemp9/arxiv-complete.tabulartext-generation100M<n<1B378 likes20k downloads3d agoHugging Face02RidheshBhati /Complete_Data_Source_100K_HOURS Multi-Language Audio Collection (100K Hours) This repository is physically reorganized for Absolute 100% Data Visibility. 🏗️ Global Consolidator Select your language subset to listen to high-quality waveform audio. All shards from legacy and modern pipelines are automatically routed here. audio1M<n<10M4 likes16k downloads5mo agoHugging Face03GeorgeDaDude /Jailbreak_Complete_DS_labeledtext10K<n<100K1 likes7.7k downloads2y agoHugging Face04OpenVideo /pexel-0808-complete-final-testGithub Page: https://github.com/UmiMarch/OpenVideo license: cc-by-4.0 task_categories: - video-text-to-text size_categories: - 100K<n<1M text100K<n<1M6 likes1.7k downloads2y agoHugging Face05Crownelius /Complete-FABLE.5-traces-2M license: mit pretty_name: Claude Library — Fable 5 · Opus · Sonnet annotations_creators: machine-generated language: en language_creators: found machine-generated multilinguality: monolingual size_categories: 10K<n<100K task_categories: text-generation task_ids: language-modeling tags: agent-traces claude claude-fable-5 claude-opus claude-sonnet chain-of-thought tool-use coding-agents content-verified maintained-mirror deduplicated parquet configs: config_name:… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Complete-FABLE.5-traces-2M.tabular100K<n<1M151 likes1.3k downloads2mo agoHugging Face06PNW-GM /tvelve_map_complete_datasetimage10K<n<100K0 likes1.2k downloads4mo agoHugging Face07bubblebird /zcache-results-completetext0 likes994 downloads1d agoHugging Face08ubaada /booksum-complete-cleaned Description: This repository contains the Booksum dataset introduced in the paper BookSum: A Collection of Datasets for Long-form Narrative Summarization . This dataset includes both book and chapter summaries from the BookSum dataset (unlike the kmfoda/booksum one which only contains the chapter dataset). Some mismatched summaries have been corrected. Uneccessary columns have been discarded. Contains minimal text-to-summary rows. As there are multiple summaries for a given text… See the full description on the dataset page: https://huggingface.co/datasets/ubaada/booksum-complete-cleaned.textsummarization1K<n<10K23 likes419 downloads2y agoHugging Face09PerRing /coco_captioning_complete_formatimage100K<n<1M0 likes414 downloads11mo agoHugging Face10whiskwhite /leetcode-complete Complete LeetCode Problems Dataset This dataset contains a comprehensive collection of LeetCode problems (including premium) with AI-generated solutions in JSONL format. It is regularly updated to include new problems as they are added to LeetCode. Splits The dataset is divided into the following splits: train: Contains approximately 80% of the problems for training validation: Contains approximately 10% of the problems for validation test: Contains approximately… See the full description on the dataset page: https://huggingface.co/datasets/whiskwhite/leetcode-complete.tabulartext-generation1K<n<10K1 likes383 downloads9d agoHugging Face11DanhVuiVe /Benetech_PlotQa_DVQA_combined_matcha_completeimage100K<n<1M0 likes376 downloads2y agoHugging Face12gsri-18 /ISEAR-dataset-completetext1K<n<10K0 likes371 downloads2y agoHugging Face13gokuls /wiki_book_corpus_complete_raw_dataset Dataset Card for "wiki_book_corpus_complete_raw_dataset" More Information needed text10M<n<100M0 likes367 downloads4y agoHugging Face14Nyanmero /vie-speech-corpus-completeaudio100K<n<1M0 likes289 downloads2y agoHugging Face15penfever /meta-llama_Llama-3.1-8B-Instruct-jdgfct-Completenesstext100K<n<1M0 likes274 downloads5mo agoHugging Face16Abrak /wikipedia-paragraph-embeddings-en-gist-complete Dataset Summary Paragraph embeddings for every article in English Wikipedia (not the Simple English version). Based on wikimedia/wikipedia, 20231101.en. Embeddings were generated with avsolatorio/GIST-small-Embedding-v0 and are quantized to int8. You can load the data with the following: from datasets import load_dataset ds = load_dataset(path="Abrak/wikipedia-paragraph-embeddings-en-gist-complete", data-dir="20231101.en") Dataset Structure The structure of the… See the full description on the dataset page: https://huggingface.co/datasets/Abrak/wikipedia-paragraph-embeddings-en-gist-complete.text10M<n<100M1 likes260 downloads2y agoHugging Face17SilencioNetwork /complete-voiceai-speech-dataset Silencio Voice AI Sample Dataset Speaker-attributed spontaneous speech. 363 labelled contributors across 111 self-reported origin varieties, recorded on their own devices in their own environments. Every clip carries origin region / variety, mother tongue, gender, device, OS, recording environment. Hours 16.05 Clips 1,305 Speakers 363 Origin varieties 111 Languages 21 Configs 44 Speaker metadata origin region / variety, mother tongue, gender, device, OS… See the full description on the dataset page: https://huggingface.co/datasets/SilencioNetwork/complete-voiceai-speech-dataset.audioautomatic-speech-recognition1K<n<10K1 likes249 downloads9h agoHugging Face18Davd-b01 /thinking-cap-tier-curricula-complete Thinking Cap Tier Curricula — Complete Reasoning Alignment Suite (TCS v4) [!IMPORTANT] Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture: In v1.2, all 13,477 SFT samples and 3,187 SimPO preference pairs have undergone an automated token purge: Zero <|pad|> batch residues: 100% eliminated across all files. Zero reasoning leakage into final answers: Deliberation stays strictly inside <think>...</think>, and answers provide direct, non-repetitive… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-curricula-complete.texttext-generation10K<n<100K0 likes234 downloads8d agoHugging Face19HKUST-DSAIL /Graph-R1-dataset-complete Graph-R1 Complete Dataset This dataset contains the complete Graph-R1 graph reasoning dataset with all difficulty levels (1-5). Files train_graph_all_levels.parquet: Combined training data from all levels (with level column) test_graph_and_math_all_levels.parquet: Combined test data from all levels (with level column) test_graph_mixedsize_cleaned.parquet: Mixed size test data (with level='mixed') Usage import pandas as pd # Load combined data… See the full description on the dataset page: https://huggingface.co/datasets/HKUST-DSAIL/Graph-R1-dataset-complete.texttext-generation100K<n<1M0 likes231 downloads1y agoHugging Face20logic65 /whittle-teacher32-complete-answers Whittle teacher32: complete answers with per-token teacher logprobs Research preview. Part of the Whittle compression campaign, a personal research project. The compute for this project is self funded and donations decide whether the next round happens: https://ko-fi.com/davida81328 What this is Complete answers generated by Qwen3.8-27B (UD-Q5_K_XL via llama.cpp), each ending on a real end-of-turn token because the answer is finished, with the teacher's top-32… See the full description on the dataset page: https://huggingface.co/datasets/logic65/whittle-teacher32-complete-answers.texttext-generation2 likes221 downloads1mo agoHugging Face21penfever /nvidia_NVLM-D-72B-jdgfct-Completenesstext100K<n<1M0 likes220 downloads5mo agoHugging Face22Deltarunefan /Deltarune-Complete-Transcript-Cleaned Deltarune Chapters 1–4 Dataset Fan-made transcript dataset covering Deltarune Chapters 1 through 4. Processed from video playthroughs and cross-referenced with game data. Intended to provide LLMs with structured narrative context for a game whose content is underrepresented in training corpora. Why This Exists As of early 2026, major LLMs (including models with training cutoffs past July 2025) fail to recall basic plot details of Deltarune Chapters 3 and 4 despite their… See the full description on the dataset page: https://huggingface.co/datasets/Deltarunefan/Deltarune-Complete-Transcript-Cleaned.texttext-generation10K<n<100K4 likes220 downloads6mo agoHugging Face23PerRing /LLaVA_Instruct_150K_complete_formatimage100K<n<1M0 likes216 downloads11mo agoHugging Face24penfever /Nexusflow_Athene-70B-jdgfct-Completenesstext100K<n<1M0 likes213 downloads5mo agoHugging Face25TartarusXXX /mixed-language-detection-pilot-complete-sentences Mixed-Language Speech Detection Pilot — Complete Sentences This is the complete-sentence revision of a 6,000-clip binary audio-classification pilot. label = 0 denotes one intended language and label = 1 denotes more than one intended language. The covered languages are Turkish (tur), Northern Kurdish/Kurmanji (kmr), Central Kurdish/Sorani (ckb), Arabic (ara), Persian (fas), and English (eng). What changed Earlier generation forced source transcripts into arbitrary… See the full description on the dataset page: https://huggingface.co/datasets/TartarusXXX/mixed-language-detection-pilot-complete-sentences.audioaudio-classification1K<n<10K0 likes191 downloads1mo agoHugging Face26Solstice-AI /Complete-FABLE.5-traces-2M Complete FABLE.5 Traces (2 Million Deduplicated Rows) Comprehensive Agentic Coding & Frontier Reasoning Trajectory Corpus Executive Summary Solstice-AI/Complete-FABLE.5-traces-2M is a clean, fully deduplicated post-training dataset containing 2,006,487 high-entropy agentic coding and multi-step reasoning traces. Originally curated following the closure of Fable and Mythos, this corpus synthesizes frontier agent execution patterns (including Claude… See the full description on the dataset page: https://huggingface.co/datasets/Solstice-AI/Complete-FABLE.5-traces-2M.tabulartext-generation1M<n<10M3 likes184 downloads19d agoHugging Face27usamaahmedsh /elliott-wave-market-data-complete Elliott Wave Market Data - Complete (Quality Validated) ✅ Production-ready, quality-validated OHLCV market data for training Elliott Wave pattern recognition neural networks. 🎯 Key Features Quality Validated: Rigorous data quality checks applied Complete Coverage: 1,403 unique instruments Multi-Timeframe: 1h, 4h, 1d, 1wk data 22,546,189 Total Data Points Dataset Statistics Timeframe Rows Tickers 1h 9,267,368 1,358 4h 2,849,583 1,358 1d 8,648… See the full description on the dataset page: https://huggingface.co/datasets/usamaahmedsh/elliott-wave-market-data-complete.tabulartime-series-forecasting10M<n<100M1 likes168 downloads7mo agoHugging Face28csusupergear /vipe_fliter_completetext10K<n<100K0 likes164 downloads5mo agoHugging Face29ishumilin /epstein-files-ocr-complete Epstein Files — Complete OCR Dataset This is a comprehensive, structured publication of the Epstein Files OCR dataset, significantly expanding upon the earlier Datasets 1-8 release. Dataset Summary This dataset contains page-level OCR output compiled from an extensive release of documents related to Jeffrey Epstein / the Epstein case. Each row in this dataset represents one scanned PDF document from the original release using a proprietary automated OCR pipeline… See the full description on the dataset page: https://huggingface.co/datasets/ishumilin/epstein-files-ocr-complete.textquestion-answering1M<n<10M3 likes160 downloads6mo agoHugging Face30DanhVuiVe /ChartQA_Benetech_PlotQa_DVQA_combined_matcha_completeimage100K<n<1M2 likes150 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.