CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hugging-science /mmu_manga mmu_manga HATS Catalog Collection This is the collection of HATS catalogs representing mmu_manga. This dataset is part of the Multimodal Universe, a large-scale collection of multimodal astronomical data. For full details, see the paper: The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data. Access the catalog We recommend the use of the LSDB Python framework to access HATS catalogs. LSDB can be installed via… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_manga.tabular10K<n<100K0 likes2.7k downloads3mo agoHugging Face02J0nasW /science-datalake Science Data Lake A unified, portable science data lake integrating 7 scholarly datasets (~525 GB Parquet) with cross-dataset DOI normalization, 13 scientific ontologies (1.3M terms), and a reproducible ETL pipeline. Note: One additional source (Semantic Scholar S2AG) is supported by the pipeline but is not redistributed here due to its API terms of service. See Not Included in This Upload below. What's Unique This dataset enables queries… See the full description on the dataset page: https://huggingface.co/datasets/J0nasW/science-datalake.imagetext-classification10B<n<100B10 likes2.5k downloads5mo agoHugging Face03vidore /vidore_v3_computer_scienceViDoRe V3 : Computer Science This dataset, Computer Science, is a corpus of textbooks from the openstacks website, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark. About ViDoRe v3 ViDoRe V3 is our latest benchmark for RAG evaluation on visually-rich documents from real-world applications. It features 10 datasets with, in total, 26,000 pages and 3099 queries, translated into 6 languages. Each query comes with… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science.documentvisual-document-retrieval1K<n<10K6 likes2.1k downloads8mo agoHugging Face04reasoning-proj /judged_science_completionstabularn<1K2 likes1.4k downloads1y agoHugging Face05reasoning-proj /severity_ablation_sciencetabular100K<n<1M0 likes1.4k downloads1y agoHugging Face06PleIAs /French-Science-Commons French Science Commons French Science Commons (Commun numérique des sciences en français) rassemble des publications scientifiques d'origine française en accès ouvert, couvrant une période de vingt ans, de 2007 à 2026. Il comprend 1 248 860 documents scientifiques — 1 189 628 articles et 59 232 thèses — indexés à travers de multiples dépôts académiques en accès public, tels que HAL, OpenAlex, des revues scientifiques, des dépôts institutionnels, et d'autres. Le corpus est conçu… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-Science-Commons.tabular10M<n<100M26 likes1.2k downloads3mo agoHugging Face07mlfoundations-dev /Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179 mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 HMMT Accuracy 60.7 90.8 89.4 63.2 52.4 48.5 27.4 26.2 48.3 12.0 34.3 34.7 AIME24 Average Accuracy: 60.67% ± 2.25% Number of Runs: 10 Run Accuracy Questions Solved Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179.tabular10K<n<100K1 likes823 downloads1y agoHugging Face08islamlab /islamic-sciences islamlab — The Islamic Sciences Corpus The Islamic sciences other than Qur'an and hadith, as their authors wrote them: 4,022 works by scholars who died between the 0st and the 14th Hijri century, cut along their own chapter and biographical-entry boundaries into 1,864,389 units (3.41 billion characters of Arabic), each carrying the volume and page it sits on so a quotation can be cited rather than merely produced. Scope is Ahl al-Sunnah wa'l-Jamāʿah, and the gate is the author… See the full description on the dataset page: https://huggingface.co/datasets/islamlab/islamic-sciences.tabulartext-generation1M<n<10M2 likes694 downloads1mo agoHugging Face09science-of-finetuning /fineweb-1m-sampletabular1M<n<10M1 likes621 downloads2y agoHugging Face10Pclanglais /EU-Science-Commonstabular1M<n<10M0 likes527 downloads4mo agoHugging Face11mlfoundations-dev /qwq_mix_qwen3_sciencetabular100K<n<1M1 likes507 downloads1y agoHugging Face12mlfoundations-dev /Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_2870 mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_2870 Precomputed model outputs for evaluation. Evaluation Results AIME24 Average Accuracy: 60.67% ± 2.20% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 70.00% 21 30 2 53.33% 16 30 3 53.33% 16 30 4 66.67% 20 30 5 63.33% 19 30 6 66.67% 20 30 7 60.00% 18 30 8 46.67% 14 30 9 63.33% 19 30 10 63.33% 19 30 tabularn<1K0 likes468 downloads1y agoHugging Face13hugging-science /mmu_apogee_dr17 mmu_apogee_dr17 HATS Catalog Collection This is the collection of HATS catalogs representing mmu_apogee_dr17. This dataset is part of the Multimodal Universe, a large-scale collection of multimodal astronomical data. For full details, see the paper: The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data. Access the catalog We recommend the use of the LSDB Python framework to access HATS catalogs. LSDB can be… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_apogee_dr17.tabular100K<n<1M1 likes407 downloads4mo agoHugging Face14mariiakoroliuk /generalization-science-datadocumentn<1K0 likes404 downloads1d agoHugging Face15marin-community /openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16 OpenThoughts-4 Science SDG: Qwen3-30B-A3B-Thinking-2507 (n=8, top-16 logprobs) Synthetic generations from Qwen/Qwen3-30B-A3B-Thinking-2507 on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes399 downloads5mo agoHugging Face16simpleG2023 /chinese-materials-science-open-intelligence 🔬 Chinese Materials Science & Metallurgy Open Intelligence Dataset Curated open intelligence dataset providing English research briefs, authoritative DOIs, executive summaries, and high-resolution micrographs of breakthrough Chinese scientific research in Materials Science, Metallurgy, Advanced Alloys, and Mining Engineering. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-materials-science-open-intelligence.tabulartext-retrieval1K<n<10K0 likes384 downloads2d agoHugging Face17hugging-science /mmu_hsc_pdr3_wide_21 mmu_hsc_pdr3_wide_21 HATS Catalog Collection This is the collection of HATS catalogs representing mmu_hsc_pdr3_wide_21. This dataset is part of the Multimodal Universe, a large-scale collection of multimodal astronomical data. For full details, see the paper: The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data. Access the catalog We recommend the use of the LSDB Python framework to access HATS catalogs. LSDB… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_hsc_pdr3_wide_21.tabular1M<n<10M0 likes365 downloads11d agoHugging Face18deep-principle /science_materialstabularn<1K0 likes355 downloads2d agoHugging Face19marin-community /openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16 OpenThoughts-4 Science SDG: Qwen3-32B (n=8, top-16 logprobs) Synthetic generations from Qwen/Qwen3-32B on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field Value Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes349 downloads5mo agoHugging Face20mlfoundations-dev /d1_science_load_in_phi_temp40tabular10K<n<100K0 likes348 downloads1y agoHugging Face21mlfoundations-dev /d1_science_load_in_phi_temp20tabular10K<n<100K0 likes343 downloads1y agoHugging Face22huggingface /community-science-paper-v2tabular1K<n<10K7 likes326 downloads2y agoHugging Face23mlfoundations-dev /d1_science_load_in_phi_temp2tabular10K<n<100K0 likes319 downloads1y agoHugging Face24mlfoundations-dev /pdf_science_questions_verified_r1_traces__2_24_25 Dataset card for pdf_science_questions_verified_r1_traces__2_24_25 This dataset was made with Curator. Dataset details A sample from the dataset: { "url": "https://www.ttcho.com/_files/ugd/988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf", "filename": "988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf", "success": true, "page_count": 37, "page_number": 1, "question_choices_solutions": "QUESTION: What is the identity of X in the reaction 14N + 1n \u2192… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/pdf_science_questions_verified_r1_traces__2_24_25.tabular1K<n<10K0 likes317 downloads2y agoHugging Face25mlfoundations-dev /Qwen2.5-7B-Instruct_qwq_mix_qwen3_science_eval_8179 mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_qwen3_science_eval_8179 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 HMMT Accuracy 61.7 88.8 88.4 67.5 54.7 52.1 25.8 27.1 49.0 11.2 40.7 32.7 AIME24 Average Accuracy: 61.67% ± 1.27% Number of Runs: 10 Run Accuracy Questions Solved Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_qwen3_science_eval_8179.tabular10K<n<100K0 likes309 downloads1y agoHugging Face26mlfoundations-dev /d1_science_load_in_phi_temp10tabular10K<n<100K0 likes288 downloads1y agoHugging Face27bluelightai-dev /common-corpus-sample-open-sciencetabular100K<n<1M0 likes269 downloads11mo agoHugging Face28electricsheepasia /asia-science-technology-world-bank-science-and-technology-indica Maldives - Science and Technology Publisher: World Bank Group · Source: HDX · License: cc-by · Updated: 2026-04-28 Abstract Contains data from the World Bank's data portal. There is also a consolidated country dataset on HDX. Technological innovation, often fueled by governments, drives industrial growth and helps raise living standards. Data here aims to shed light on countries technology base: research and development, scientific and technical journal articles… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-science-technology-world-bank-science-and-technology-indica.tabulartabular-classificationn<1K0 likes244 downloads5mo agoHugging Face29science-of-finetuning /diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04 Contains maximum activating examples for all the features of our crosscoder trained on gemma 2 2B layer 13 available here: https://huggingface.co/Butanium/gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04/blob/main/README.md base_examples.pt contains all the maximum examples of the feature on a subset of validation test of fineweb chat_examples.pt is the same but for lmsys chat data chat_base_examples.pt is a merge of the two above files. All files are of the type dict[int, list[tuple[float… See the full description on the dataset page: https://huggingface.co/datasets/science-of-finetuning/diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04.tabular10K<n<100K0 likes240 downloads1y agoHugging Face30marin-community /openthoughts4-science-26041-prompts-qwen3-4b-n8-flattened-logprobs-k16 OpenThoughts-4 Science SDG: Qwen3-4B (n=8, top-16 logprobs) Synthetic generations from Qwen/Qwen3-4B on the Marin OpenThoughts-4 science SDG prompt set. Each prompt is sampled n=8 times, and for every generated token the dataset stores the chosen-token log probability plus the top-16 log probabilities over the vocabulary, enabling distillation, KL-style fine-tuning, reranking, and uncertainty analysis. Generation setup Field Value Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-4b-n8-flattened-logprobs-k16.tabulartext-generation100K<n<1M0 likes179 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.