CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sorry-bench /sorry-bench-202503gated Dataset Card for SORRY-Bench Dataset (2025/03) 🏠Website 📑Paper 📚Dataset 💻Github 🧑‍⚖️Human Judgment Dataset 🤖Judge LLM 🪧UPDATE: In this iteration, we removed the category "Impersonation" due to its ambiguous definition, and that most models more or less fulfill such requests.This dataset contains 9.2K potentially unsafe instructions, intended to be used for LLM safety refusal evaluation. Particularly, our base dataset consists of 440 unsafe… See the full description on the dataset page: https://huggingface.co/datasets/sorry-bench/sorry-bench-202503.texttext-generation1K<n<10K23 likes1.7k downloads2y agoHugging Face02test-time-compute /aime_2025 AIME 2025 - Unified Test-Time Scaling Format This is the AIME (American Invitational Mathematics Examination) 2025 dataset in a unified format for test-time scaling experiments. Dataset Description Source: MathArena/aime_2025 Size: 30 competition-level mathematics problems Format: Unified TTS format (question, answer, metadata) Dataset Structure Fields question (string): The mathematical problem statement answer (string): The numerical answer… See the full description on the dataset page: https://huggingface.co/datasets/test-time-compute/aime_2025.textquestion-answeringn<1K0 likes1.3k downloads11mo agoHugging Face03Alga2025 /UltraData-Math-TEST UltraData-Math 🤗 Dataset | 💻 Source Code | 🇨🇳 中文 README UltraData-Math is a large-scale, high-quality mathematical pre-training dataset totaling 290B+ tokens across three progressive tiers—L1 (170.5B tokens web corpus), L2 (33.7B tokens quality-selected), and L3 (88B tokens multi-format refined)—designed to systematically enhance mathematical reasoning in LLMs. It has been applied to the mathematical pre-training of the MiniCPM Series models. 🆕 What's New… See the full description on the dataset page: https://huggingface.co/datasets/Alga2025/UltraData-Math-TEST.texttext-generation100M<n<1B0 likes958 downloads7mo agoHugging Face04AIStudioDelta /Eurovoc_2025_by_language 🇪🇺 🏷️ EuroVoc dataset (by language) This is the EuropeanParliament/Eurovoc_2025 dataset, but split up by language, not by period. The original is split up into periods (1996-03 through 2025-11), with documents in different languages mixed together. For ease of training this dataset splits the data by language instead, with documents in different periods put together. License This dataset is redistributed under the original European Union Public License 1.2. When… See the full description on the dataset page: https://huggingface.co/datasets/AIStudioDelta/Eurovoc_2025_by_language.texttext-generation1M<n<10M1 likes878 downloads10mo agoHugging Face05Yahoo-Finance-News /FineWeb2025 FineWeb-Edu 2025 — Cleaned and Shuffled This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2025. This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year. Dataset summary Item Value Year 2025 Rows 99,022,205… See the full description on the dataset page: https://huggingface.co/datasets/Yahoo-Finance-News/FineWeb2025.tabulartext-generation10M<n<100M1 likes775 downloads8d agoHugging Face06goldentraversy07 /reddit_dataset_2025 Bittensor Subnet 13 Reddit Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more… See the full description on the dataset page: https://huggingface.co/datasets/goldentraversy07/reddit_dataset_2025.texttext-classification10M<n<100M0 likes721 downloads1y agoHugging Face07Joemn /TamilNadu-State-board-books-2025-Tamil-and-english-versions Tamil Nadu School Textbooks — Tamil and English Structured Text This dataset contains text extracted from 312 Tamil Nadu State Board school textbooks for Standards 1–12. It covers Tamil- and English-medium books and provides each retained book in two forms: structured JSON with book metadata, ordered sections, typed content blocks, source references, extraction statistics, and curation provenance; Markdown for reading, inspection, and downstream text processing. The source… See the full description on the dataset page: https://huggingface.co/datasets/Joemn/TamilNadu-State-board-books-2025-Tamil-and-english-versions.texttext-retrievaln<1K1 likes571 downloads2mo agoHugging Face08FlagEval /HMMT_2025 Dataset Summary This dataset comprises the questions, answers, and solutions from HMMT February 2025, all of which were extracted by OCR, converted to LaTeX, and manually verified by FlagEval Team. Data Fields Below one can find the description of each field in the dataset. id (str): Index of the problem in the competition problem (str): Full problem statement answer (str): Ground-truth answer to the question solution(str): Ground-truth solution to the question… See the full description on the dataset page: https://huggingface.co/datasets/FlagEval/HMMT_2025.textquestion-answeringn<1K1 likes453 downloads1y agoHugging Face09maryamfakhari /crypto-news-coindesk-2020-2025 CoinDesk Cryptocurrency News Dataset (2020–2025) This dataset contains cryptocurrency-related news articles sourced from CoinDesk Data, accessed programmatically via the CryptoCompare API. The dataset is curated and published for academic and research purposes, with a focus on analyzing the relationship between news and cryptocurrency market dynamics. Time Period January 1, 2020 – January 1, 2025 Content Overview Each record in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/maryamfakhari/crypto-news-coindesk-2020-2025.imagetext-classification100K<n<1M2 likes451 downloads1mo agoHugging Face10KaraKaraWitch /TvTroper-2025 TvTroper-2025 A cleaned & refreshed dump of ~708 k pages from tvtropes.org Dataset Summary TvTroper-2025 is an updated snapshot of TvTropes.org (≈ 708 000 wiki pages, namespaces and date-grouped pages excluded). Every page is released in two flavours: Raw HTML – 22 GB single file Markdown-cleaned – split into 1 GB JSONL shards (no unpacking required) No additional content filtering has been applied; short sub-index pages are left in so you can decide what to drop.… See the full description on the dataset page: https://huggingface.co/datasets/KaraKaraWitch/TvTroper-2025.texttext-classification1M<n<10M4 likes417 downloads11mo agoHugging Face11jeffboudier /yc-companies-august-2025 Y Combinator Companies Dataset Dataset Description This dataset contains information about 5,404 Y Combinator funded companies that have been publicly launched, sourced from the YC-OSS-API. Dataset Summary Total Companies: 5,404 Time Range: Summer 2005 - Summer 2025 Update Frequency: Snapshot from August 2025 Source: YC-OSS-API Dataset Structure Data Fields id: Unique identifier for each company name: Company name… See the full description on the dataset page: https://huggingface.co/datasets/jeffboudier/yc-companies-august-2025.imagetext-classification1K<n<10K1 likes372 downloads1y agoHugging Face12trentmkelly /scored_co_2025 Scored.co 2025 scrape This dataset is a full scrape of public content from Scored.co, a right-wing Reddit-style social media site. It contains 73,045,361 rows of posts and comments, stored as parquet. The dataset is intended for research into online communities, political discussion, social media moderation, misinformation, platform migration, network dynamics, and large-scale text analysis. Data format The main dataset is partitioned by entity type:… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/scored_co_2025.tabulartext-generation10M<n<100M0 likes353 downloads5mo agoHugging Face13huyxdang /neurips-2025-papers NeurIPS 2025 Papers Dataset This dataset contains all accepted papers from NeurIPS 2025, scraped from OpenReview. Dataset Statistics Overview Total Papers: 5772 Unique Paper IDs: 5772 ✅ No duplicate IDs Track Distribution Main Track: 5,275 papers (91.4%) Datasets and Benchmarks Track: 497 papers (8.6%) Award Distribution Poster: 4,949 papers (85.7%) Oral: 84 papers (1.5%) Spotlight: 739 papers (12.8%) Track × Award… See the full description on the dataset page: https://huggingface.co/datasets/huyxdang/neurips-2025-papers.texttext-classification1K<n<10K3 likes318 downloads10mo agoHugging Face14joecwales /whiteglove-medical-medlineplus-2025 WhiteGlove Medical Knowledge Corpus MedlinePlus 2025 — Spectral Curation Pipeline Pipeline: WhiteGlove Spectral Curation | Domain: Medical | License: Public Domain (US Government) Dataset Summary A clean, deduplicated, semantically chunked medical knowledge corpus derived from the NIH MedlinePlus January 2025 ZIM archive. Produced by the WhiteGlove Spectral Curation Pipeline — an air-gapped, attribution-clean dataset factory built on SimHash-128 deduplication… See the full description on the dataset page: https://huggingface.co/datasets/joecwales/whiteglove-medical-medlineplus-2025.tabulartext-generation1K<n<10K0 likes279 downloads4mo agoHugging Face15lparkourer10 /enwiki-20250201 Dataset Card for lparkourer10/enwiki-20250201 This dataset is an extracted version of the English Wikipedia dump as of February 1, 2025. It has been processed to facilitate information retrieval and analysis. Dataset Description This dataset contains extracted text from the English Wikipedia, aimed at providing structured and accessible information for natural language processing (NLP) tasks, research, and machine learning applications. It includes raw Wikipedia articles… See the full description on the dataset page: https://huggingface.co/datasets/lparkourer10/enwiki-20250201.texttext-classification1M<n<10M0 likes224 downloads2y agoHugging Face16Blaze7451 /Wiki-zhtw-20250601 Dataset Card for Wiki-zhtw-20250601 Dataset Description This dataset is derived from the Chinese‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim, converted to Markdown format via regular‑expression post‑processing, and finally converted from Simplified to Traditional Chinese using OpenCC. texttext-generation1M<n<10M1 likes182 downloads1y agoHugging Face17LightningRodLabs /WWTD-2025 What Would Trump Do? Auto-generated from 5 search queries — used to beat GPT-5 Starting from nothing but 5 search queries, we used the Lightning Rod SDK to automatically generate 2,790 forecasting questions about Trump administration actions from news articles and label them using real outcomes. No expertise required. No manual labeling. Used to train Trump-Forecaster, which beats GPT-5. TL;DR Generated 2,790 forward-looking forecasting questions… See the full description on the dataset page: https://huggingface.co/datasets/LightningRodLabs/WWTD-2025.texttext-generation1K<n<10K4 likes182 downloads2mo agoHugging Face18Jinzy2025 /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/Jinzy2025/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes170 downloads1mo agoHugging Face19Trendyol /All-CVE-Chat-MultiTurn-1999-2025-Dataset CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025) 1. Project Overview This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/All-CVE-Chat-MultiTurn-1999-2025-Dataset.texttext-generation100K<n<1M32 likes166 downloads1y agoHugging Face20goldentraversy07 /x_dataset_2025 Bittensor Subnet 13 X (Twitter) Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/goldentraversy07/x_dataset_2025.texttext-classification100K<n<1M0 likes158 downloads1y agoHugging Face21Tonic /Health-Bench-Eval-OSS-2025-07 Dataset Card for HealthBench Dataset Summary HealthBench is a benchmark dataset developed by OpenAI in collaboration with 262 physicians from 60 countries to evaluate AI systems in health-related conversational scenarios. It contains 5,000 multi-turn health conversations in a JSONL file (2025-05-07-06-14-12_oss_eval.jsonl), simulating interactions between AI models and users (laypersons or clinicians). Each conversation includes a user prompt, a candidate model response… See the full description on the dataset page: https://huggingface.co/datasets/Tonic/Health-Bench-Eval-OSS-2025-07.texttext-generation1K<n<10K4 likes151 downloads1y agoHugging Face22team-suzuki /hle-extract-qwen3235ba22b-20250815 HLE Extract: Qwen3-235B-A22B Evaluation Results (2025-08-15) Dataset Description This dataset contains the complete Human-Level Evaluation (HLE) benchmark with detailed evaluation results from the Qwen/Qwen3-235B-A22B model. It merges the original team-suzuki/hle-extract dataset with comprehensive model responses and human judgments. Dataset Summary Total Questions: 120 (complete HLE dataset) Evaluated Questions: 103 (85.8%) Unevaluated Questions: 17 (14.2%)… See the full description on the dataset page: https://huggingface.co/datasets/team-suzuki/hle-extract-qwen3235ba22b-20250815.textquestion-answeringn<1K1 likes141 downloads1y agoHugging Face23MMDR-2025 /MMdeepresearchHuggingface: https://huggingface.co/papers/2601.12346 Paper: arxiv.org/abs/2601.12346 imagetext-generationn<1K1 likes135 downloads8mo agoHugging Face24goldentraversy07 /x_dataset_202507 Bittensor Subnet 13 X (Twitter) Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/goldentraversy07/x_dataset_202507.texttext-classification10M<n<100M0 likes132 downloads1y agoHugging Face25Blaze7451 /Wiki-ja-20250601 Dataset Card for Wiki-ja-20250601 Dataset Description This dataset is derived from the Japan‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim and converted to Markdown format via regular‑expression post‑processing. texttext-generation100K<n<1M0 likes112 downloads1y agoHugging Face26gutoportelaa /dom-pi-teresina-2025 DOM-Teresina 2025 — Diário Oficial do Município de Teresina (PI) Texto integral das publicações de 2025 do Diário Oficial do Município de Teresina (DOM-Teresina). a capital do Piauí — que publica em diário próprio. separado do DOM-PI dos Municípios. 9583 documentos · ~9.325.820 tokens · pt-BR. Parte da capital também está no dataset geral dos demais municípios: gutoportelaa/dom-pi-corpus-2025 (como território teresina). Este repositório é self-contained: inclui os PDFs-fonte em… See the full description on the dataset page: https://huggingface.co/datasets/gutoportelaa/dom-pi-teresina-2025.documenttext-generation10K<n<100K0 likes111 downloads4mo agoHugging Face27OmniAICreator /Japanese-Wikipedia-202506 Japanese-Wikipedia-202506 This dataset contains data from the Japanese Wikipedia as of June 1, 2025. texttext-classification1M<n<10M5 likes109 downloads1y agoHugging Face28sorry-bench /sorry-bench-human-judgment-202503gated Dataset Card for 🧑‍⚖️SORRY-Bench Human Judgment Dataset (2025/03) 🏠Website 📑Paper 📚Dataset 💻Github 🧑‍⚖️Human Judgment Dataset 🤖Judge LLM 🪧UPDATE: In this iteration, we removed the category "Impersonation" due to its ambiguous definition, and that most models more or less fulfill such requests.This dataset contains 7K annotations of human safety judgments for LLM responses to unsafe instructions of our SORRY-Bench dataset. Specifically, for… See the full description on the dataset page: https://huggingface.co/datasets/sorry-bench/sorry-bench-human-judgment-202503.tabulartext-classification1K<n<10K1 likes103 downloads2y agoHugging Face29Blaze7451 /Wiki-vi-20250601 Dataset Card for Wiki-vi-20250601 Dataset Description This dataset is derived from the Vietnam‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim and converted to Markdown format via regular‑expression post‑processing. texttext-generation1M<n<10M1 likes102 downloads1y agoHugging Face30Blaze7451 /Wiki-ko-20250601 Dataset Card for Wiki-ko-20250601 Dataset Description This dataset is derived from the Korea‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim and converted to Markdown format via regular‑expression post‑processing. texttext-generation1M<n<10M0 likes95 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.