CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01opencompass /AIME2025 AIME 2025 Dataset Dataset Description This dataset contains problems from the American Invitational Mathematics Examination (AIME) 2025-I & II. textquestion-answeringn<1K56 likes13k downloads2y agoHugging Face02garak-llm /perl-20250811text10K<n<100K0 likes4.7k downloads1y agoHugging Face03garak-llm /raku-20250811text1K<n<10K0 likes4.6k downloads1y agoHugging Face04ZombitX64 /xauusd-gold-price-historical-data-2004-2025 XAUUSD Gold Price Historical Data 2004-2025 This dataset contains historical price data for XAUUSD (Gold vs US Dollar) from 2004 to 2025. Source: Kaggle dataset "novandraanugrah/xauusd-gold-price-historical-data-2004-2024" Content: The dataset includes CSV files with different time granularities (e.g., 1 minute, 5 minutes, 1 hour, 1 day). Each file typically contains the following columns: Date Open High Low Close Volume Usage: This dataset can be used for analyzing historical… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/xauusd-gold-price-historical-data-2004-2025.tabular1M<n<10M12 likes2.7k downloads1y agoHugging Face05garak-llm /dart-20250811text10K<n<100K1 likes2k downloads1y agoHugging Face06sorry-bench /sorry-bench-202503gated Dataset Card for SORRY-Bench Dataset (2025/03) 🏠Website 📑Paper 📚Dataset 💻Github 🧑‍⚖️Human Judgment Dataset 🤖Judge LLM 🪧UPDATE: In this iteration, we removed the category "Impersonation" due to its ambiguous definition, and that most models more or less fulfill such requests.This dataset contains 9.2K potentially unsafe instructions, intended to be used for LLM safety refusal evaluation. Particularly, our base dataset consists of 440 unsafe… See the full description on the dataset page: https://huggingface.co/datasets/sorry-bench/sorry-bench-202503.texttext-generation1K<n<10K23 likes1.7k downloads2y agoHugging Face07birdsql /bird_sql_dev_20251106 BIRD-SQL Dev 🆕 Update 2025-11-06 We would like to express our sincere gratitude to the community for their continuous support and constructive feedback on the BIRD-SQL Dev dataset. Over the past year, we have received valuable suggestions through GitHub discussions, emails, and user reports. Based on these insights, we organized a quality review program led by a team of five PhD researchers in Data Science and AI, supported by a globally distributed group of industry… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/bird_sql_dev_20251106.texttable-question-answering1K<n<10K10 likes1.4k downloads8mo agoHugging Face08yufan /recsys-papers-2025-2026 📚 Recommender Systems Papers 2025–2026 A curated library of 3,151 recent recommender-systems papers spanning 2025-01-02 → 2026-09-17, each with the original PDF and a structured, section-by-section Markdown analysis (research problem, prior work, method, math, experiments, strengths & weaknesses, …). Includes a self-contained Apple-style HTML browser (index.html). 🔑 Browse by meeting (Data Viewer subsets) The Dataset Viewer above has a subset dropdown keyed by… See the full description on the dataset page: https://huggingface.co/datasets/yufan/recsys-papers-2025-2026.document1K<n<10K0 likes1.1k downloads6d agoHugging Face09nebula2025 /CodeR-Pile Towards A Generalist Code Embedding Model Based On Massive Data Synthesis Introduction This repository contains the synthetic training data introduced in the paper Towards A Generalist Code Embedding Model Based On Massive Data Synthesis. The dataset is designed to enhance text embeddings for code retrieval tasks. For more details, please refer to our Github repo: CodeR. Load Dataset Simple Example An example to load the dataset:… See the full description on the dataset page: https://huggingface.co/datasets/nebula2025/CodeR-Pile.text1M<n<10M4 likes744 downloads1y agoHugging Face10FlagEval /HMMT_2025 Dataset Summary This dataset comprises the questions, answers, and solutions from HMMT February 2025, all of which were extracted by OCR, converted to LaTeX, and manually verified by FlagEval Team. Data Fields Below one can find the description of each field in the dataset. id (str): Index of the problem in the competition problem (str): Full problem statement answer (str): Ground-truth answer to the question solution(str): Ground-truth solution to the question… See the full description on the dataset page: https://huggingface.co/datasets/FlagEval/HMMT_2025.textquestion-answeringn<1K1 likes453 downloads1y agoHugging Face11KaraKaraWitch /TvTroper-2025 TvTroper-2025 A cleaned & refreshed dump of ~708 k pages from tvtropes.org Dataset Summary TvTroper-2025 is an updated snapshot of TvTropes.org (≈ 708 000 wiki pages, namespaces and date-grouped pages excluded). Every page is released in two flavours: Raw HTML – 22 GB single file Markdown-cleaned – split into 1 GB JSONL shards (no unpacking required) No additional content filtering has been applied; short sub-index pages are left in so you can decide what to drop.… See the full description on the dataset page: https://huggingface.co/datasets/KaraKaraWitch/TvTroper-2025.texttext-classification1M<n<10M4 likes417 downloads11mo agoHugging Face12huyxdang /neurips-2025-papers NeurIPS 2025 Papers Dataset This dataset contains all accepted papers from NeurIPS 2025, scraped from OpenReview. Dataset Statistics Overview Total Papers: 5772 Unique Paper IDs: 5772 ✅ No duplicate IDs Track Distribution Main Track: 5,275 papers (91.4%) Datasets and Benchmarks Track: 497 papers (8.6%) Award Distribution Poster: 4,949 papers (85.7%) Oral: 84 papers (1.5%) Spotlight: 739 papers (12.8%) Track × Award… See the full description on the dataset page: https://huggingface.co/datasets/huyxdang/neurips-2025-papers.texttext-classification1K<n<10K3 likes318 downloads10mo agoHugging Face13zhuzilin /aime-2025textn<1K1 likes298 downloads10mo agoHugging Face14Anonymous-Team-HC-RAG /Multi-doc-2025 Dataset Card for Multi-Doc-2025 Dataset Summary Multi-Doc-2025 is a financial question-answering benchmark built from SEC Form 10-K annual reports of S&P 500 companies. It is designed to evaluate retrieval-augmented generation (RAG) and financial QA systems under three reasoning settings that are not jointly covered by existing financial QA benchmarks: cross-company reasoning, cross-year reasoning, and hybrid-modal reasoning over both text and tables. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-Team-HC-RAG/Multi-doc-2025.textquestion-answering1K<n<10K2 likes282 downloads4mo agoHugging Face15joecwales /whiteglove-medical-medlineplus-2025 WhiteGlove Medical Knowledge Corpus MedlinePlus 2025 — Spectral Curation Pipeline Pipeline: WhiteGlove Spectral Curation | Domain: Medical | License: Public Domain (US Government) Dataset Summary A clean, deduplicated, semantically chunked medical knowledge corpus derived from the NIH MedlinePlus January 2025 ZIM archive. Produced by the WhiteGlove Spectral Curation Pipeline — an air-gapped, attribution-clean dataset factory built on SimHash-128 deduplication… See the full description on the dataset page: https://huggingface.co/datasets/joecwales/whiteglove-medical-medlineplus-2025.tabulartext-generation1K<n<10K0 likes279 downloads4mo agoHugging Face16djroytburg /NeurIPS-2023-2025 NeurIPS 2023–2025 Peer Review Dataset Structured peer-review data for 13,171 NeurIPS submissions (2023–2025), collected via the OpenReview API. Each paper entry includes acceptance decisions, full reviewer text, and an anonymized parsed version of the paper itself. Note on selection bias. NeurIPS authors are not required to make rejection reviews public, and the vast majority do not. As a result, this dataset contains roughly 95% accepted papers. The true NeurIPS acceptance rate is… See the full description on the dataset page: https://huggingface.co/datasets/djroytburg/NeurIPS-2023-2025.texttext-classification10K<n<100K0 likes227 downloads5mo agoHugging Face17lparkourer10 /enwiki-20250201 Dataset Card for lparkourer10/enwiki-20250201 This dataset is an extracted version of the English Wikipedia dump as of February 1, 2025. It has been processed to facilitate information retrieval and analysis. Dataset Description This dataset contains extracted text from the English Wikipedia, aimed at providing structured and accessible information for natural language processing (NLP) tasks, research, and machine learning applications. It includes raw Wikipedia articles… See the full description on the dataset page: https://huggingface.co/datasets/lparkourer10/enwiki-20250201.texttext-classification1M<n<10M0 likes224 downloads2y agoHugging Face18MTSUs-Fall-2025-Software-Engineering-Pr /United_States_State_Legislation_with_SummariesTest Push text100K<n<1M0 likes222 downloads10mo agoHugging Face19lmms-lab /imo-2025 IMO 2025 Problems Dataset This dataset contains the 6 problems from the 2025 International Mathematical Olympiad (IMO). The problems are formatted with proper LaTeX notation for mathematical expressions. Dataset Structure Each example contains: id: Problem identifier (e.g., "2025-imo-p1") problem: The problem statement with LaTeX mathematical notation solution: The solution (currently set to null) Problem Types The dataset includes problems covering various… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/imo-2025.documentn<1K2 likes188 downloads1y agoHugging Face20OpenHay /zalo-ai-2025-public-test-data-v2 Zalo AI Challenge 2025 - RoadBuddy Public Test Data V2set This dataset contains public test data v2 for the RoadBuddy – Understanding the Road through Dashcam AI challenge from Zalo AI Challenge 2025. Dataset Description The challenge aims to build a driving assistant capable of understanding video content from dashcams to quickly answer questions about traffic signs, signals, and driving instructions in Vietnam. Dataset Structure Files frames/:… See the full description on the dataset page: https://huggingface.co/datasets/OpenHay/zalo-ai-2025-public-test-data-v2.imagevisual-question-answeringn<1K0 likes186 downloads10mo agoHugging Face21Blaze7451 /Wiki-zhtw-20250601 Dataset Card for Wiki-zhtw-20250601 Dataset Description This dataset is derived from the Chinese‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim, converted to Markdown format via regular‑expression post‑processing, and finally converted from Simplified to Traditional Chinese using OpenCC. texttext-generation1M<n<10M1 likes182 downloads1y agoHugging Face22LumiOpen /mAIME2025 mAIME2025: Multilingual AIME 2025 Math Competition Dataset mAIME2025 is a multilingual version of the 2025 AIME (American Invitational Mathematics Examination) problems, professionally translated into European languages. This dataset contains all 30 problems from AIME I and AIME II 2025, translated and human-reviewed by native speakers to preserve mathematical accuracy and LaTeX formatting. Languages Czech (cs) - 30 problems Danish (da) - 30 problems Finnish (fi) - 30… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/mAIME2025.textquestion-answeringn<1K1 likes178 downloads7mo agoHugging Face23Trendyol /All-CVE-Chat-MultiTurn-1999-2025-Dataset CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025) 1. Project Overview This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/All-CVE-Chat-MultiTurn-1999-2025-Dataset.texttext-generation100K<n<1M32 likes166 downloads1y agoHugging Face24sukantabasu /alchemist-shell.ai-hackathon-2025This project is described in detail at this website: https://alchemist-shellai-hackathon-2025.readthedocs.io/en/latest/ The codes and relevant materials are available here: https://github.com/Sukantabasu/alchemist-shell.ai-hackathon-2025 The trained models (in pkl format) are stored in this HF repository. textn<1K0 likes161 downloads1y agoHugging Face25Tonic /Health-Bench-Eval-OSS-2025-07 Dataset Card for HealthBench Dataset Summary HealthBench is a benchmark dataset developed by OpenAI in collaboration with 262 physicians from 60 countries to evaluate AI systems in health-related conversational scenarios. It contains 5,000 multi-turn health conversations in a JSONL file (2025-05-07-06-14-12_oss_eval.jsonl), simulating interactions between AI models and users (laypersons or clinicians). Each conversation includes a user prompt, a candidate model response… See the full description on the dataset page: https://huggingface.co/datasets/Tonic/Health-Bench-Eval-OSS-2025-07.texttext-generation1K<n<10K4 likes151 downloads1y agoHugging Face26roma2025 /parsebench-table-track ParseBench Table Track plus Financial Split This dataset is a mirror of the table dimension of llamaindex/ParseBench, packaged together with a curated Financial Split that we built for evaluating OCR systems on insurance and financial filings. It contains, All 503 PDFs of the ParseBench table track. table.jsonl, the original ground truth (one HTML table per page, plus easy or hard difficulty tag). financial_split/, our 151 page financial slice plus the 117 dropped non financial… See the full description on the dataset page: https://huggingface.co/datasets/roma2025/parsebench-table-track.documenttable-question-answeringn<1K0 likes144 downloads5mo agoHugging Face27maariaaa12 /ariel-2025-cube-masked-cachetabularn<1K0 likes131 downloads22d agoHugging Face28DeepNLP /Coding-Agent-Github-2025-Feb Coding Agent AI Agent Directory to Host All Coding Agent related AI Agents Web Traffic Data, Search Ranking, Community, Reviews and More. This is the Coding Agent Dataset from pypi package "coding_agent" https://pypi.org/project/coding_agent. You can use this package to download and get statistics (forks/stars/website traffic) of AI agents on website from AI Agent Marketplace AI Agent Directory (http://www.deepnlp.org/store/ai-agent) and AI Agent Search Portal… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/Coding-Agent-Github-2025-Feb.textn<1K8 likes118 downloads2y agoHugging Face29jeffreyszhou /ikea-us-products-2025 IKEA US Product Dataset (July 2025) This dataset is a structured snapshot of ~30,000 IKEA US products, scraped from the official IKEA US website in July 2025. It contains product metadata (titles, descriptions, categories, materials, care instructions, etc.) and associated product images. Contents products-us.jsonl — one JSON object per product with structured fields. images-us/ — the first "hero" image for each product, downloaded via image_downloader_first.py.… See the full description on the dataset page: https://huggingface.co/datasets/jeffreyszhou/ikea-us-products-2025.texttext-classification10K<n<100K5 likes115 downloads1y agoHugging Face30Blaze7451 /Wiki-ja-20250601 Dataset Card for Wiki-ja-20250601 Dataset Description This dataset is derived from the Japan‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim and converted to Markdown format via regular‑expression post‑processing. texttext-generation100K<n<1M0 likes112 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.