CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yentinglin /aime_2025 AIME 2025 This dataset contains 30 problems from the 2025 AIME tests, including: AIME I: 15 problems AIME II: 15 problems tabularn<1K12 likes37k downloads9mo agoHugging Face02MathArena /aime_2025 Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains the questions from AIME 2025 used for the MathArena Leaderboard Data Fields The dataset contains the following fields: problem_idx (int64): Problem index within the corresponding MathArena benchmark. problem (string): Problem statement, usually stored as LaTeX source. answer (int64): Gold final answer. problem_type… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/aime_2025.tabularn<1K17 likes32k downloads4mo agoHugging Face03hasankursun /github-code-2025-language-split 📜 Source Data & Attribution This dataset is a processed derivative of nick007x/github-code-2025. Origination The original data was aggregated by nick007x from public GitHub repositories. We have retained the original content, file paths, and metadata while restructuring the format for easier consumption by language-specific models. Processing Steps To create this dataset, we performed the following processing on the source data: Language… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/github-code-2025-language-split.text100M<n<1B13 likes24k downloads10mo agoHugging Face04MathArena /hmmt_feb_2025 Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains the questions from HMMT February 2025 used for the MathArena Leaderboard Data Fields The dataset contains the following fields: problem_idx (int64): Problem index within the corresponding MathArena benchmark. problem (string): Problem statement, usually stored as LaTeX source. answer (string): Gold final answer.… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/hmmt_feb_2025.textn<1K11 likes22k downloads4mo agoHugging Face05opencompass /AIME2025 AIME 2025 Dataset Dataset Description This dataset contains problems from the American Invitational Mathematics Examination (AIME) 2025-I & II. textquestion-answeringn<1K56 likes13k downloads2y agoHugging Face06MathArena /hmmt_nov_2025 Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains the questions from HMMT November 2025 used for the MathArena Leaderboard Data Fields The dataset contains the following fields: problem_idx (int64): Problem index within the corresponding MathArena benchmark. answer (string): Gold final answer. problem (string): Problem statement, usually stored as LaTeX source.… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/hmmt_nov_2025.textn<1K1 likes12k downloads4mo agoHugging Face07garak-llm /crates-20250307text100K<n<1M0 likes8.3k downloads2y agoHugging Face08cc-clean /CC-MAIN-2025-08 CC-MAIN-2025-08へようこそ 本データセットはCommonCrawlerと呼ばれるものから日本語のみを抽出したものです。 利用したものはcc-downloader-rsです。 なおIPAのICSCoEと呼ばれるところから資源を借りてやりましたゆえに、みなさんIPAに感謝しましょう。 ※ IPAは独立行政法人 情報処理推進機構のことです。テストに出ますので覚えましょう。 利用について 本利用は研究目的のみとさせていただきます。 それ以外の利用につきましては途方のくれない数の著作権者に許可を求めてきてください。 text100M<n<1B0 likes7.9k downloads1y agoHugging Face09cc-clean /CC-MAIN-2025-33text100M<n<1B0 likes6.5k downloads1y agoHugging Face10AiAF /SCPWiki-Archive-02-March-2025-Datasetstextn<1K0 likes6.1k downloads2y agoHugging Face11behavior-1k /2025-challenge-task-instancestextn<1K0 likes6.1k downloads6mo agoHugging Face12cc-clean /CC-MAIN-2025-05 CC-MAIN-2025-05へようこそ 本データセットはCommonCrawlerと呼ばれるものから日本語のみを抽出したものです。 利用したものはcc-downloader-rsです。 なおIPAのICSCoEと呼ばれるところから資源を借りてやりましたゆえに、みなさんIPAに感謝しましょう。 ※ IPAは独立行政法人 情報処理推進機構のことです。テストに出ますので覚えましょう。 利用について 本利用は研究目的のみとさせていただきます。 それ以外の利用につきましては途方のくれない数の著作権者に許可を求めてきてください。 text100M<n<1B0 likes6k downloads1y agoHugging Face13TAAC2025 /TencentGR-1M TencentGR-1M Dataset Paper | Project Page | Code TAAC2025 Preliminary Round Dataset (2025年腾讯广告算法大赛初赛数据集) TencentGR-1M Dataset is a large-scale, all-modality dataset designed specifically for generative recommendation (GR) in industrial advertising. Constructed from real, de-identified Tencent Ads logs, it aims to address the lack of realistic, public multi-modal datasets in the GR field. Data Features: Contains rich collaborative IDs and multi-modal representations (text and… See the full description on the dataset page: https://huggingface.co/datasets/TAAC2025/TencentGR-1M.tabularother10M<n<100M23 likes6k downloads4mo agoHugging Face14nick007x /github-code-2025text100M<n<1B121 likes5.4k downloads6mo agoHugging Face15garak-llm /perl-20250811text10K<n<100K0 likes4.7k downloads1y agoHugging Face16garak-llm /raku-20250811text1K<n<10K0 likes4.6k downloads1y agoHugging Face17MathArena /apex_2025 UPDATE: The dataset now contains the remaining samples from SMT 2025! Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains the questions from MathArena Apex 2025 used for the MathArena Leaderboard Data Fields The dataset contains the following fields: problem_idx (int64): Problem index within the corresponding MathArena benchmark. answer (string): Gold final… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/apex_2025.textn<1K3 likes4.5k downloads4mo agoHugging Face18MathArena /brumo_2025 Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains the questions from BRUMO 2025 used for the MathArena Leaderboard Data Fields The dataset contains the following fields: problem_idx (int64): Problem index within the corresponding MathArena benchmark. problem (string): Problem statement, usually stored as LaTeX source. answer (string): Gold final answer. problem_type… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/brumo_2025.textn<1K1 likes3.8k downloads4mo agoHugging Face19cc-clean /CC-MAIN-2025-26text100M<n<1B0 likes3.8k downloads1y agoHugging Face20TAAC2025 /TencentGR-10M TencentGR-10M Dataset TAAC2025 Second Round Dataset(2025年腾讯广告算法大赛复赛数据集) TencentGR-10M Dataset is a large-scale, all-modality dataset designed specifically for generative recommendation (GR) in industrial advertising. Similar to TencentGR-1M, it is constructed from real, de-identified Tencent Ads logs, and aims to address the lack of realistic, public multi-modal datasets in the GR field. The main differences between TencentGR-10M and TencentGR-1M are: Dataset Size: Provides 10… See the full description on the dataset page: https://huggingface.co/datasets/TAAC2025/TencentGR-10M.tabular100M<n<1B15 likes3.7k downloads6mo agoHugging Face21MathArena /smt_2025 Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains the questions from SMT 2025 used for the MathArena Leaderboard Data Fields The dataset contains the following fields: problem_idx (int64): Problem index within the corresponding MathArena benchmark. answer (string): Gold final answer. problem_type (list[string]): Problem type/category labels. problem (string): Problem… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/smt_2025.textn<1K0 likes3.5k downloads4mo agoHugging Face22MathArena /cmimc_2025 Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains the questions from CMIMC 2025 used for the MathArena Leaderboard Data Fields The dataset contains the following fields: problem_idx (int64): Problem index within the corresponding MathArena benchmark. problem (string): Problem statement, usually stored as LaTeX source. answer (string): Gold final answer. problem_type… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/cmimc_2025.textn<1K0 likes3.4k downloads4mo agoHugging Face23hybridfree /github-code-2025 🚀 GitHub Code 2025: The Clean Code Manifesto A meticulously curated dataset of 1.5M+ repositories representing both quality and innovation in 2025's code ecosystem 🌟 The Philosophy Quality Over Quantity, Purpose Over Volume In an era of data abundance, we present a dataset built on radical curation. Every file, every repository, every byte has been carefully selected to represent the signal in the noise of open-source development. 🎯 What This Dataset Is… See the full description on the dataset page: https://huggingface.co/datasets/hybridfree/github-code-2025.text100M<n<1B1 likes3.3k downloads9mo agoHugging Face24alea-institute /kl3m-data-snapshot-20250324text10M<n<100M2 likes3.1k downloads1y agoHugging Face25Cognitive-Lab /NayanaOCR_Corpus_2025 🪷 NayanaOCR Corpus 2025 A 1M-page, 22-language fully-parallel synthetic OCR + VQA corpus for document-centric vision-language models — every page rendered in every language. NayanaOCR Corpus 2025 is one of the largest open-source multilingual, multi-task document datasets for training and evaluating OCR, layout detection, and visual question answering (VQA) in low-resource and underrepresented languages. The headline property: it's a true parallel corpus. The same ~45,700 source… See the full description on the dataset page: https://huggingface.co/datasets/Cognitive-Lab/NayanaOCR_Corpus_2025.imageimage-to-text1M<n<10M18 likes3k downloads4mo agoHugging Face26EleutherAI /dclm-dedup_20250227-004105tabular100M<n<1B1 likes2.9k downloads2y agoHugging Face27ZombitX64 /xauusd-gold-price-historical-data-2004-2025 XAUUSD Gold Price Historical Data 2004-2025 This dataset contains historical price data for XAUUSD (Gold vs US Dollar) from 2004 to 2025. Source: Kaggle dataset "novandraanugrah/xauusd-gold-price-historical-data-2004-2024" Content: The dataset includes CSV files with different time granularities (e.g., 1 minute, 5 minutes, 1 hour, 1 day). Each file typically contains the following columns: Date Open High Low Close Volume Usage: This dataset can be used for analyzing historical… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/xauusd-gold-price-historical-data-2004-2025.tabular1M<n<10M12 likes2.7k downloads1y agoHugging Face28MathArena /usamo_2025 Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains the questions from USAMO 2025 used for the MathArena Leaderboard Data Fields The dataset contains the following fields: problem_idx (string): Problem index within the corresponding MathArena benchmark. problem (string): Problem statement, usually stored as LaTeX source. points (int64): Maximum score for a… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/usamo_2025.textn<1K2 likes2.7k downloads4mo agoHugging Face29MathArena /imc_2025 Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains the questions from IMC 2025 used for the MathArena Leaderboard Data Fields The dataset contains the following fields: problem_idx (string): Problem index within the corresponding MathArena benchmark. problem (string): Problem statement, usually stored as LaTeX source. points (int64): Maximum score for a… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/imc_2025.textn<1K1 likes2.6k downloads4mo agoHugging Face30cc-clean /CC-MAIN-2025-18text100M<n<1B0 likes2.5k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.