CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sorry-bench /sorry-bench-202503gated Dataset Card for SORRY-Bench Dataset (2025/03) 🏠Website 📑Paper 📚Dataset 💻Github 🧑‍⚖️Human Judgment Dataset 🤖Judge LLM 🪧UPDATE: In this iteration, we removed the category "Impersonation" due to its ambiguous definition, and that most models more or less fulfill such requests.This dataset contains 9.2K potentially unsafe instructions, intended to be used for LLM safety refusal evaluation. Particularly, our base dataset consists of 440 unsafe… See the full description on the dataset page: https://huggingface.co/datasets/sorry-bench/sorry-bench-202503.texttext-generation1K<n<10K23 likes1.7k downloads2y agoHugging Face02ergt2025 /OSWorld2 OSWorld2 Kimi-K2.6 exact request traces This dataset contains 98 complete OSWorld2 task traces recorded from a Kimi-K2.6 EPD 1P1D serving run. It is intended for exact virtual request replay in SGLang; replay sends the recorded model requests and token sequences without launching OSWorld VMs or re-executing computer actions. Coverage 98 successful tasks 16,100 model requests 13,100 unique content-addressed screenshot blobs 6.72 GB of blob payloads before Hub… See the full description on the dataset page: https://huggingface.co/datasets/ergt2025/OSWorld2.text-generation0 likes1.5k downloads2mo agoHugging Face03test-time-compute /aime_2025 AIME 2025 - Unified Test-Time Scaling Format This is the AIME (American Invitational Mathematics Examination) 2025 dataset in a unified format for test-time scaling experiments. Dataset Description Source: MathArena/aime_2025 Size: 30 competition-level mathematics problems Format: Unified TTS format (question, answer, metadata) Dataset Structure Fields question (string): The mathematical problem statement answer (string): The numerical answer… See the full description on the dataset page: https://huggingface.co/datasets/test-time-compute/aime_2025.textquestion-answeringn<1K0 likes1.3k downloads11mo agoHugging Face04Alga2025 /UltraData-Math-TEST UltraData-Math 🤗 Dataset | 💻 Source Code | 🇨🇳 中文 README UltraData-Math is a large-scale, high-quality mathematical pre-training dataset totaling 290B+ tokens across three progressive tiers—L1 (170.5B tokens web corpus), L2 (33.7B tokens quality-selected), and L3 (88B tokens multi-format refined)—designed to systematically enhance mathematical reasoning in LLMs. It has been applied to the mathematical pre-training of the MiniCPM Series models. 🆕 What's New… See the full description on the dataset page: https://huggingface.co/datasets/Alga2025/UltraData-Math-TEST.texttext-generation100M<n<1B0 likes958 downloads7mo agoHugging Face05AIStudioDelta /Eurovoc_2025_by_language 🇪🇺 🏷️ EuroVoc dataset (by language) This is the EuropeanParliament/Eurovoc_2025 dataset, but split up by language, not by period. The original is split up into periods (1996-03 through 2025-11), with documents in different languages mixed together. For ease of training this dataset splits the data by language instead, with documents in different periods put together. License This dataset is redistributed under the original European Union Public License 1.2. When… See the full description on the dataset page: https://huggingface.co/datasets/AIStudioDelta/Eurovoc_2025_by_language.texttext-generation1M<n<10M1 likes878 downloads10mo agoHugging Face06Yahoo-Finance-News /FineWeb2025 FineWeb-Edu 2025 — Cleaned and Shuffled This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2025. This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year. Dataset summary Item Value Year 2025 Rows 99,022,205… See the full description on the dataset page: https://huggingface.co/datasets/Yahoo-Finance-News/FineWeb2025.tabulartext-generation10M<n<100M1 likes775 downloads8d agoHugging Face07goldentraversy07 /reddit_dataset_2025 Bittensor Subnet 13 Reddit Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more… See the full description on the dataset page: https://huggingface.co/datasets/goldentraversy07/reddit_dataset_2025.texttext-classification10M<n<100M0 likes721 downloads1y agoHugging Face08Joemn /TamilNadu-State-board-books-2025-Tamil-and-english-versions Tamil Nadu School Textbooks — Tamil and English Structured Text This dataset contains text extracted from 312 Tamil Nadu State Board school textbooks for Standards 1–12. It covers Tamil- and English-medium books and provides each retained book in two forms: structured JSON with book metadata, ordered sections, typed content blocks, source references, extraction statistics, and curation provenance; Markdown for reading, inspection, and downstream text processing. The source… See the full description on the dataset page: https://huggingface.co/datasets/Joemn/TamilNadu-State-board-books-2025-Tamil-and-english-versions.texttext-retrievaln<1K1 likes571 downloads2mo agoHugging Face09FlagEval /HMMT_2025 Dataset Summary This dataset comprises the questions, answers, and solutions from HMMT February 2025, all of which were extracted by OCR, converted to LaTeX, and manually verified by FlagEval Team. Data Fields Below one can find the description of each field in the dataset. id (str): Index of the problem in the competition problem (str): Full problem statement answer (str): Ground-truth answer to the question solution(str): Ground-truth solution to the question… See the full description on the dataset page: https://huggingface.co/datasets/FlagEval/HMMT_2025.textquestion-answeringn<1K1 likes453 downloads1y agoHugging Face10maryamfakhari /crypto-news-coindesk-2020-2025 CoinDesk Cryptocurrency News Dataset (2020–2025) This dataset contains cryptocurrency-related news articles sourced from CoinDesk Data, accessed programmatically via the CryptoCompare API. The dataset is curated and published for academic and research purposes, with a focus on analyzing the relationship between news and cryptocurrency market dynamics. Time Period January 1, 2020 – January 1, 2025 Content Overview Each record in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/maryamfakhari/crypto-news-coindesk-2020-2025.imagetext-classification100K<n<1M2 likes451 downloads1mo agoHugging Face11KaraKaraWitch /TvTroper-2025 TvTroper-2025 A cleaned & refreshed dump of ~708 k pages from tvtropes.org Dataset Summary TvTroper-2025 is an updated snapshot of TvTropes.org (≈ 708 000 wiki pages, namespaces and date-grouped pages excluded). Every page is released in two flavours: Raw HTML – 22 GB single file Markdown-cleaned – split into 1 GB JSONL shards (no unpacking required) No additional content filtering has been applied; short sub-index pages are left in so you can decide what to drop.… See the full description on the dataset page: https://huggingface.co/datasets/KaraKaraWitch/TvTroper-2025.texttext-classification1M<n<10M4 likes417 downloads11mo agoHugging Face12jeffboudier /yc-companies-august-2025 Y Combinator Companies Dataset Dataset Description This dataset contains information about 5,404 Y Combinator funded companies that have been publicly launched, sourced from the YC-OSS-API. Dataset Summary Total Companies: 5,404 Time Range: Summer 2005 - Summer 2025 Update Frequency: Snapshot from August 2025 Source: YC-OSS-API Dataset Structure Data Fields id: Unique identifier for each company name: Company name… See the full description on the dataset page: https://huggingface.co/datasets/jeffboudier/yc-companies-august-2025.imagetext-classification1K<n<10K1 likes372 downloads1y agoHugging Face13trentmkelly /scored_co_2025 Scored.co 2025 scrape This dataset is a full scrape of public content from Scored.co, a right-wing Reddit-style social media site. It contains 73,045,361 rows of posts and comments, stored as parquet. The dataset is intended for research into online communities, political discussion, social media moderation, misinformation, platform migration, network dynamics, and large-scale text analysis. Data format The main dataset is partitioned by entity type:… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/scored_co_2025.tabulartext-generation10M<n<100M0 likes353 downloads5mo agoHugging Face14izlley2 /llm0to1-pt-tokenized-en-edu-2025 LLM0to1 사전학습 토큰화본 — 영어 교육(2025 덤프) 10B 규모 한/영 이중언어 LLM LLM0to1-10b 를 바닥부터 학습할 때 실제로 투입된 영어 교육(2025 덤프) 코퍼스의 토큰화본. 총 28종 / 138.4B 토큰. 왜 원문 텍스트가 아니라 토큰화본인가 이 코퍼스들은 공개 데이터셋을 스트리밍으로 받아 곧바로 토큰화했고 중간 텍스트를 보관하지 않았다. 따라서 이 .ds 파일이 학습에 들어간 데이터의 유일한 사본이다. 단점만 있는 건 아니다. 토큰화본은 학습 입력 그 자체이므로, 재토큰화 과정에서 생길 수 있는 차이 없이 학습을 그대로 재현할 수 있다. 원본 출처 HuggingFaceFW/fineweb-edu 2025 덤프 영어 벤치마크 하락에 대응해 뒤늦게 편입한 '다양성 영어' 보강분이다. g00~`g15` 는 원본 shard 를 균등 분할한 것으로, 서로 다른 문서 집합이다.… See the full description on the dataset page: https://huggingface.co/datasets/izlley2/llm0to1-pt-tokenized-en-edu-2025.text-generationn>1T0 likes339 downloads1mo agoHugging Face15huyxdang /neurips-2025-papers NeurIPS 2025 Papers Dataset This dataset contains all accepted papers from NeurIPS 2025, scraped from OpenReview. Dataset Statistics Overview Total Papers: 5772 Unique Paper IDs: 5772 ✅ No duplicate IDs Track Distribution Main Track: 5,275 papers (91.4%) Datasets and Benchmarks Track: 497 papers (8.6%) Award Distribution Poster: 4,949 papers (85.7%) Oral: 84 papers (1.5%) Spotlight: 739 papers (12.8%) Track × Award… See the full description on the dataset page: https://huggingface.co/datasets/huyxdang/neurips-2025-papers.texttext-classification1K<n<10K3 likes318 downloads10mo agoHugging Face16horelulus /IDX_Financial_Statements2015-2025Q2 IDX_Financial_Statements: The Multimodal Indonesian Financial Dataset IDX_Financial_Statements is a centralized repository providing the most complete set of financial disclosures for public companies listed on the Indonesia Stock Exchange (IDX). This dataset is designed for advanced financial research, spanning from raw document archival to structured data extraction. Dataset Overview This is a multimodal dataset that captures the full lifecycle of financial reporting.… See the full description on the dataset page: https://huggingface.co/datasets/horelulus/IDX_Financial_Statements2015-2025Q2.table-question-answering0 likes292 downloads6mo agoHugging Face17joecwales /whiteglove-medical-medlineplus-2025 WhiteGlove Medical Knowledge Corpus MedlinePlus 2025 — Spectral Curation Pipeline Pipeline: WhiteGlove Spectral Curation | Domain: Medical | License: Public Domain (US Government) Dataset Summary A clean, deduplicated, semantically chunked medical knowledge corpus derived from the NIH MedlinePlus January 2025 ZIM archive. Produced by the WhiteGlove Spectral Curation Pipeline — an air-gapped, attribution-clean dataset factory built on SimHash-128 deduplication… See the full description on the dataset page: https://huggingface.co/datasets/joecwales/whiteglove-medical-medlineplus-2025.tabulartext-generation1K<n<10K0 likes279 downloads4mo agoHugging Face18lparkourer10 /enwiki-20250201 Dataset Card for lparkourer10/enwiki-20250201 This dataset is an extracted version of the English Wikipedia dump as of February 1, 2025. It has been processed to facilitate information retrieval and analysis. Dataset Description This dataset contains extracted text from the English Wikipedia, aimed at providing structured and accessible information for natural language processing (NLP) tasks, research, and machine learning applications. It includes raw Wikipedia articles… See the full description on the dataset page: https://huggingface.co/datasets/lparkourer10/enwiki-20250201.texttext-classification1M<n<10M0 likes224 downloads2y agoHugging Face19languagehub-ai /yuxiaowang-prompts-2025 Yuxiaowang Semantic Dataset · Hugging Face Version 🧠 English Summary Yuxiaowang · Semantic Dataset for Japanese Language Schools (Chinese) This project provides structured semantic definitions and prompt examples for the domain of Japanese language schools in China.It aims to serve as a grounding corpus for large language models (LLMs) to understand terms like "语校", "语校网", and related concepts. Source platform: https://www.yuxiaowang.comAll prompts and term… See the full description on the dataset page: https://huggingface.co/datasets/languagehub-ai/yuxiaowang-prompts-2025.text-generationn<1K0 likes203 downloads8mo agoHugging Face20Blaze7451 /Wiki-zhtw-20250601 Dataset Card for Wiki-zhtw-20250601 Dataset Description This dataset is derived from the Chinese‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim, converted to Markdown format via regular‑expression post‑processing, and finally converted from Simplified to Traditional Chinese using OpenCC. texttext-generation1M<n<10M1 likes182 downloads1y agoHugging Face21LightningRodLabs /WWTD-2025 What Would Trump Do? Auto-generated from 5 search queries — used to beat GPT-5 Starting from nothing but 5 search queries, we used the Lightning Rod SDK to automatically generate 2,790 forecasting questions about Trump administration actions from news articles and label them using real outcomes. No expertise required. No manual labeling. Used to train Trump-Forecaster, which beats GPT-5. TL;DR Generated 2,790 forward-looking forecasting questions… See the full description on the dataset page: https://huggingface.co/datasets/LightningRodLabs/WWTD-2025.texttext-generation1K<n<10K4 likes182 downloads2mo agoHugging Face22Jinzy2025 /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/Jinzy2025/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes170 downloads1mo agoHugging Face23Trendyol /All-CVE-Chat-MultiTurn-1999-2025-Dataset CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025) 1. Project Overview This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/All-CVE-Chat-MultiTurn-1999-2025-Dataset.texttext-generation100K<n<1M32 likes166 downloads1y agoHugging Face24goldentraversy07 /x_dataset_2025 Bittensor Subnet 13 X (Twitter) Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/goldentraversy07/x_dataset_2025.texttext-classification100K<n<1M0 likes158 downloads1y agoHugging Face25Tonic /Health-Bench-Eval-OSS-2025-07 Dataset Card for HealthBench Dataset Summary HealthBench is a benchmark dataset developed by OpenAI in collaboration with 262 physicians from 60 countries to evaluate AI systems in health-related conversational scenarios. It contains 5,000 multi-turn health conversations in a JSONL file (2025-05-07-06-14-12_oss_eval.jsonl), simulating interactions between AI models and users (laypersons or clinicians). Each conversation includes a user prompt, a candidate model response… See the full description on the dataset page: https://huggingface.co/datasets/Tonic/Health-Bench-Eval-OSS-2025-07.texttext-generation1K<n<10K4 likes151 downloads1y agoHugging Face26team-suzuki /hle-extract-qwen3235ba22b-20250815 HLE Extract: Qwen3-235B-A22B Evaluation Results (2025-08-15) Dataset Description This dataset contains the complete Human-Level Evaluation (HLE) benchmark with detailed evaluation results from the Qwen/Qwen3-235B-A22B model. It merges the original team-suzuki/hle-extract dataset with comprehensive model responses and human judgments. Dataset Summary Total Questions: 120 (complete HLE dataset) Evaluated Questions: 103 (85.8%) Unevaluated Questions: 17 (14.2%)… See the full description on the dataset page: https://huggingface.co/datasets/team-suzuki/hle-extract-qwen3235ba22b-20250815.textquestion-answeringn<1K1 likes141 downloads1y agoHugging Face27MMDR-2025 /MMdeepresearchHuggingface: https://huggingface.co/papers/2601.12346 Paper: arxiv.org/abs/2601.12346 imagetext-generationn<1K1 likes135 downloads8mo agoHugging Face28goldentraversy07 /x_dataset_202507 Bittensor Subnet 13 X (Twitter) Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/goldentraversy07/x_dataset_202507.texttext-classification10M<n<100M0 likes132 downloads1y agoHugging Face29ApyHTML19 /EU-AI-Regulation-GDPR-2025 EuropeGram: EU Legal Text -- RAG Chunks & Instruction Data Structured, chunked, and instruction-formatted text derived from official EU legislation, built for retrieval-augmented generation (RAG) and LoRA fine-tuning experiments comparing Base / RAG / Fine-tuned / Fine-tuned+RAG LLM strategies over EU documents. Produced by the EuropeGram project's extraction -> chunking -> fine-tuning-export pipeline. Source documents Document CELEX ID Source Chunks… See the full description on the dataset page: https://huggingface.co/datasets/ApyHTML19/EU-AI-Regulation-GDPR-2025.text-generation1 likes116 downloads29d agoHugging Face30Blaze7451 /Wiki-ja-20250601 Dataset Card for Wiki-ja-20250601 Dataset Description This dataset is derived from the Japan‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim and converted to Markdown format via regular‑expression post‑processing. texttext-generation100K<n<1M0 likes112 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.