CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mikex86 /stackoverflow-posts StackOverflow Posts Markdown Dataset Summary This dataset contains all posts submitted to StackOverflow before the 14th of June 2023 formatted as Markdown text. The dataset contains ~60 Million posts, totaling ~35GB in size and ~65 billion characters of text. The data is sourced from Internet Archive StackExchange Data Dump. Dataset Structure Each record corresponds to one post of a particular type. Original ordering from the data dump is not exactly preserved… See the full description on the dataset page: https://huggingface.co/datasets/mikex86/stackoverflow-posts.tabularquestion-answering10M<n<100M63 likes15k downloads3y agoHugging Face02stanfordnlp /web_questions Dataset Card for "web_questions" Dataset Summary This dataset consists of 6,642 question/answer pairs. The questions are supposed to be answerable by Freebase, a large knowledge graph. The questions are mostly centered around a single named entity. The questions are popular ones asked on the web (at least in 2013). Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data… See the full description on the dataset page: https://huggingface.co/datasets/stanfordnlp/web_questions.textquestion-answering1K<n<10K43 likes13k downloads3y agoHugging Face03mjuicem /StreamingBench StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding 🏠 Project Page | 📄 arXiv Paper | 📦 Dataset | 🏅Leaderboard StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. 🌟 [NEW! 2025.05.15] 🔥: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output. [NEW! 2025.03.17] ⭐: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the… See the full description on the dataset page: https://huggingface.co/datasets/mjuicem/StreamingBench.imagequestion-answering1K<n<10K13 likes13k downloads1y agoHugging Face04HuggingFaceH4 /stack-exchange-preferences Dataset Card for H4 Stack Exchange Preferences Dataset Dataset Summary This dataset contains questions and answers from the Stack Overflow Data Dump for the purpose of preference model training. Importantly, the questions have been filtered to fit the following criteria for preference models (following closely from Askell et al. 2021): have >=2 answers. This data could also be used for instruction fine-tuning and language model training. The questions are grouped with… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/stack-exchange-preferences.textquestion-answering10M<n<100M135 likes9k downloads4y agoHugging Face05stanfordnlp /SHP 🚢 Stanford Human Preferences Dataset (SHP) If you mention this dataset in a paper, please cite the paper: Understanding Dataset Difficulty with V-Usable Information (ICML 2022). Summary SHP is a dataset of 385K collective human preferences over responses to questions/instructions in 18 different subject areas, from cooking to legal advice. The preferences are meant to reflect the helpfulness of one response over another, and are intended to be used for training RLHF… See the full description on the dataset page: https://huggingface.co/datasets/stanfordnlp/SHP.tabulartext-generation100K<n<1M325 likes7.5k downloads3y agoHugging Face06stanford-oval /ccnewsThis dataset is the result of processing all WARC files in the CCNews Corpus, from the beginning (2016) to June of 2024. The data has been cleaned and deduplicated, and language of articles have been detected and added. The process is similar to what HuggingFace's DataTrove does. Overall, it contains about 600 million news articles in more than 100 languages from all around the globe. For license information, please refer to CommonCrawl's Terms of Use. Sample Python code to explore this… See the full description on the dataset page: https://huggingface.co/datasets/stanford-oval/ccnews.imagetext-classification100M<n<1B36 likes6.4k downloads2y agoHugging Face07prquan /STARK_10k STARK: Spatial-Temporal reAsoning benchmaRK STARK is a comprehensive benchmark designed to systematically evaluate large language models (LLMs) and large reasoning models (LRMs) on spatial-temporal reasoning tasks, particularly for applications in cyber-physical systems (CPS) such as robotics, autonomous vehicles, and smart city infrastructure. Dataset Summary Hierarchical Benchmark: Tasks are structured across three levels of reasoning complexity: State Estimation:… See the full description on the dataset page: https://huggingface.co/datasets/prquan/STARK_10k.textquestion-answering10K<n<100K1 likes6.1k downloads1y agoHugging Face08stanfordnlp /coqa Dataset Card for "coqa" Dataset Summary CoQA is a large-scale dataset for building Conversational Question Answering systems. Our dataset contains 127k questions with answers, obtained from 8k conversations about text passages from seven diverse domains. The questions are conversational, and the answers are free-form text with their corresponding evidence highlighted in the passage. Supported Tasks and Leaderboards More Information Needed Languages… See the full description on the dataset page: https://huggingface.co/datasets/stanfordnlp/coqa.textquestion-answering1K<n<10K86 likes4.9k downloads3y agoHugging Face09strickvl /isafpressreleases ISAF Press Releases Dataset Description Homepage: [N/A] Repository: [N/A] Paper: A Knock on the Door: 22 Months of ISAF Press Releases Point of Contact: Alex Strick van Linschoten (@strickvl) Dataset Summary The ISAF Press Releases dataset contains data used as the basis for the research paper "A Knock on the Door: 22 Months of ISAF Press Releases". The dataset provides a comprehensive collection of press releases issued by the International Security Assistance… See the full description on the dataset page: https://huggingface.co/datasets/strickvl/isafpressreleases.textfeature-extraction1K<n<10K7 likes4.1k downloads3mo agoHugging Face10stanford-crfm /image2struct-latex-v1 Image2Struct - Latex Paper | Website | Datasets (Webpages, Latex, Music sheets) | Leaderboard | HELM repo | Image2Struct repo License: Apache License Version 2.0, January 2004 Dataset description Image2struct is a benchmark for evaluating vision-language models in practical tasks of extracting structured information from images. This subdataset focuses on LaTeX code. The model is given an image of the expected output with the prompt: Please provide the LaTex code used to… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/image2struct-latex-v1.imagequestion-answering1K<n<10K12 likes3.8k downloads2y agoHugging Face11Stevross /mmluThis is a massive multitask test consisting of multiple-choice questions from various branches of knowledge, covering 57 tasks including elementary mathematics, US history, computer science, law, and more.textquestion-answering1M<n<10M10 likes3.5k downloads3y agoHugging Face12StormKing99 /x_dataset_8191 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/StormKing99/x_dataset_8191.texttext-classification100M<n<1B0 likes2.2k downloads1y agoHugging Face13apple /CLaRa_multi_stage CLaRa: Bridging Retrieval and Generation with Continuous Latent Reasoning This is the official dataset for the CLaRa paper which contains training and evaluation data for the CLaRa model, organized into three main categories: pretraining, instruction tuning, and end-to-end tuning. Dataset Structure 1. Pretraining Data pretraining: Large-scale pretraining data for the compressor learning Format: JSONL with fields: data_type, question, answers… See the full description on the dataset page: https://huggingface.co/datasets/apple/CLaRa_multi_stage.textquestion-answering1M<n<10M11 likes2.2k downloads9mo agoHugging Face14snap-stanford /stark STaRK Website | Github | Paper STaRK is a large-scale semi-structure retrieval benchmark on Textual and Relational Knowledge Bases Downstream Task Retrieval systems driven by LLMs are tasked with extracting relevant answers from a knowledge base in response to user queries. Each knowledge base is semi-structured, featuring large-scale relational data among entities and comprehensive textual information for each entity. We have constructed three knowledge bases: Amazon SKB… See the full description on the dataset page: https://huggingface.co/datasets/snap-stanford/stark.textquestion-answering10K<n<100K12 likes2.2k downloads2y agoHugging Face15starriver030515 /FUSION-Finetune-12M FUSION-12M Dataset Please see paper & website for more information: https://arxiv.org/abs/2504.09925 https://github.com/starriver030515/FUSION Overview FUSION-12M is a large-scale, diverse multimodal instruction-tuning dataset used to train FUSION-3B and FUSION-8B models. It builds upon Cambrian-1 by significantly expanding both the quantity and variety of data, particularly in areas such as OCR, mathematical reasoning, and synthetic high-quality Q&A data. The goal is… See the full description on the dataset page: https://huggingface.co/datasets/starriver030515/FUSION-Finetune-12M.imagequestion-answering1K<n<10K13 likes1.9k downloads1y agoHugging Face16math-ai /StackMathQA StackMathQA StackMathQA: A Curated Collection of 2 Million Mathematical Questions and Answers Sourced from Stack Exchange StackMathQA is a meticulously curated collection of 2 million mathematical questions and answers, sourced from various Stack Exchange sites. This repository is designed to serve as a comprehensive resource for researchers, educators, and enthusiasts in the field of mathematics and AI research. Configs configs: - config_name: stackmathqa1600k… See the full description on the dataset page: https://huggingface.co/datasets/math-ai/StackMathQA.texttext-generation1M<n<10M104 likes1.7k downloads10mo agoHugging Face17stdKonjac /LiveSports-3K LiveSports-3K Benchmark News [2025.05.12] We released the ASR transcripts for the CC track. See LiveSports-3K-CC.json for details. Overview LiveSports‑3K is a comprehensive benchmark for evaluating streaming video understanding capabilities of large language and multimodal models. It consists of two evaluation tracks: Closed Captions (CC) Track: Measures models’ ability to generate real‑time commentary aligned with the ground‑truth ASR transcripts. Question… See the full description on the dataset page: https://huggingface.co/datasets/stdKonjac/LiveSports-3K.tabularvideo-text-to-text1K<n<10K5 likes1.5k downloads1y agoHugging Face18flax-sentence-embeddings /stackexchange_titlebody_best_and_down_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering100K<n<1M12 likes1.4k downloads4y agoHugging Face19starriver030515 /FUSION-Pretrain-10M FUSION-10M Dataset Please see paper & website for more information: https://arxiv.org/abs/2504.09925 https://github.com/starriver030515/FUSION Overview FUSION-10M is a large-scale, high-quality dataset of image-caption pairs used to pretrain FUSION-3B and FUSION-8B models. It builds upon established datasets such as LLaVA, ShareGPT4, and PixelProse. In addition, we synthesize 2 million task-specific image-caption pairs to further enrich the dataset. The goal of… See the full description on the dataset page: https://huggingface.co/datasets/starriver030515/FUSION-Pretrain-10M.imagequestion-answeringn<1K9 likes1.4k downloads1y agoHugging Face20stanfordnlp /SHP-2 🚢 Stanford Human Preferences Dataset v2 (SHP-2) Summary SHP-2 is a dataset of 4.8M collective human preferences over responses to questions/instructions in 129 different subject areas, from cooking to legal advice. It is an extended version of the original 385K SHP dataset. The preferences are meant to reflect the helpfulness of one response over another, and are intended to be used for training RLHF reward models and NLG evaluation models (e.g., SteamSHP). Each example… See the full description on the dataset page: https://huggingface.co/datasets/stanfordnlp/SHP-2.tabulartext-generation1M<n<10M18 likes1.4k downloads3y agoHugging Face21AdhyanshVerma /stack-2021-12-01 ReasonStack-Prime A highly normalized, streaming-optimized Stack Exchange corpus engineered for LLM reasoning and instruction tuning. CC BY-SA 4.0 ~1M Rows 176 Parquet Shards 21 SE Sites 1. Executive Summary ReasonStack-Prime is a large-scale, meticulously curated text dataset derived from the official Archive.org Stack Exchange data dump (Version 2021-12-07). Unlike raw XML dumps or poorly cleaned JSON exports, this dataset… See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/stack-2021-12-01.tabularquestion-answering100K<n<1M0 likes1.3k downloads3mo agoHugging Face22caiovicentino1 /Qwen3.6-35B-A3B-mcr-stage-b Qwen3.6-35B-A3B — MCR Stage B Corpus (Distributed Reasoning Localization) First systematic mechanistic-intervention corpus on a hybrid MoE + GDN + Gated-Attention architecture. 📄 Paper: Loop-Intolerance Profiling: Localizing Distributed Reasoning in a Hybrid MoE Architecture via Nine Convergent Intervention Experiments — submitted to arXiv (2026-04-20, in moderation). Final arXiv ID will be added here once approved. This dataset contains per-token residual-stream activations at… See the full description on the dataset page: https://huggingface.co/datasets/caiovicentino1/Qwen3.6-35B-A3B-mcr-stage-b.textquestion-answeringn<1K1 likes1.2k downloads5mo agoHugging Face23prquan /STARK_1k Benchmarking Spatiotemporal Reasoning in Large Language Models: Capabilities and Challenges Dataset for our paper: Benchmarking Spatiotemporal Reasoning in Large Language Models: Capabilities and Challenges Contact Information If you have any questions or feedback, feel free to reach out: Name: Pengrui Quan Email: prquan@ucla.edu License Copyright (c) 2025, UCLA Networked and Embedded Systems Laboratory (NESL) All rights reserved. Redistribution and use in… See the full description on the dataset page: https://huggingface.co/datasets/prquan/STARK_1k.textquestion-answering1K<n<10K0 likes1.1k downloads10mo agoHugging Face24flax-sentence-embeddings /stackexchange_title_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M8 likes1.1k downloads4y agoHugging Face25stellalisy /HorizonBench HorizonBench Long-Horizon Personalization with Evolving Preferences HorizonBench evaluates whether language models can track user preferences as they evolve across months of interaction. Each benchmark item is a 5-option multiple-choice question embedded within a conversation history averaging ~163K tokens. Pre-evolution preference values serve as hard-negative distractors, enabling diagnosis of belief-update failure: models retrieve the user's originally stated preference but fail… See the full description on the dataset page: https://huggingface.co/datasets/stellalisy/HorizonBench.textmultiple-choice1K<n<10K2 likes1k downloads5mo agoHugging Face26StormKing99 /reddit_dataset_8191 Bittensor Subnet 13 Reddit Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this dataset allows… See the full description on the dataset page: https://huggingface.co/datasets/StormKing99/reddit_dataset_8191.texttext-classification10M<n<100M0 likes956 downloads2y agoHugging Face27starmpcc /Asclepius-Synthetic-Clinical-Notes Asclepius: Synthetic Clincal Notes & Instruction Dataset Dataset Summary This dataset is official dataset for Asclepius (arxiv) This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs. We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5 Then, we generate instruction-answer pairs for 157k synthetic discharge summaries Supported Tasks This dataset covers below 8 tasks Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.textquestion-answering100K<n<1M117 likes934 downloads2y agoHugging Face28flax-sentence-embeddings /stackexchange_titlebody_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M9 likes780 downloads4y agoHugging Face29SeaLLMs /TrueFalse-Statements-multilingualThis dataset is introduced in the paper Analyzing LLMs' Knowledge Boundary Cognition Across Languages Through the Lens of Internal Representations. Code: https://github.com/DAMO-NLP-SG/LLM-Multilingual-Knowledge-Boundaries textquestion-answering10K<n<100K2 likes682 downloads1y agoHugging Face30nguha /legalbench-staging Dataset Card for Dataset Name Homepage: https://hazyresearch.stanford.edu/legalbench/ Repository: https://github.com/HazyResearch/legalbench/ Paper: https://arxiv.org/abs/2308.11462 Dataset Description Dataset Summary The LegalBench project is an ongoing open science effort to collaboratively curate tasks for evaluating legal reasoning in English large language models (LLMs). The benchmark currently consists of 162 tasks gathered from 40… See the full description on the dataset page: https://huggingface.co/datasets/nguha/legalbench-staging.tabulartext-classification10K<n<100K1 likes677 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.