CoolFace
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01GenData-Research /scientific-verification Scientific Verification Benchmark: NMC Cathodes Dataset summary The benchmark contains 50 scientific claims about NMC (lithium nickel manganese cobalt oxide) battery cathodes. Each claim is answered by Claude Opus 5, GPT 5.6 Luna and Gemini 3.1 Pro using a set of 20 open-access papers, producing 150 scored answers. The accompanying reference set contains 1,991 experiment-grounded measurements curated from 227 open-access papers, with experimental conditions and… See the full description on the dataset page: https://huggingface.co/datasets/GenData-Research/scientific-verification.tabularquestion-answering1K<n<10K0 likes239 downloads7d agoHugging Face02ScaleAI /SciPredict SciPredict: Can LLMs Predict the Outcomes of Research Experiments? Paper: SciPredict: Can LLMs Predict the Outcomes of Research Experiments in Natural Sciences? Overview SciPredict is a benchmark evaluating whether AI systems can predict experimental outcomes in physics, biology, and chemistry. The dataset comprises 405 questions derived from recently published empirical studies (post-March 2025), spanning 33 subdomains. Dataset Structure Total Questions: 405… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/SciPredict.textquestion-answeringn<1K2 likes229 downloads8mo agoHugging Face03chimbiwide /sciqa-thinking sciqa-thinking Randomly extracted 3000 rows from sciq and prompting Qwen3-14b to generate the intermediate reasoning traces, we created this dataset. This should be used for LLM post-training, especially RL. textquestion-answering1K<n<10K0 likes148 downloads9mo agoHugging Face04nasa-impact /nasa-science-repos-sme-benchmark NASA Science Repos SME Benchmark A benchmark dataset for evaluating retrieval systems on NASA science repository discovery tasks. This dataset contains expert queries, a corpus of NASA science GitHub repositories, and relevance judgments. Dataset Structure Files ├── corpus.jsonl # 5,264 repositories with full metadata ├── queries.jsonl # 219 expert queries └── qrels/ ├── earth.tsv # Earth Science relevance judgments (162) ├──… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-repos-sme-benchmark.tabulartext-retrievaln<1K0 likes115 downloads8mo agoHugging Face05StarpowerTechnology /Dense-Information-Science-Physics-Dataset Dense Information With Multiple Fine-tuned Variations This dataaset has multiple for each input to learn how to express the same answer in different ways Dataset Structure The dataset contains two columns: Column Description input A science or quantum-physics question output A conversational answer to the question Example: { "input": "What is quantum entanglement?", "output": "Quantum entanglement is when two quantum systems share one… See the full description on the dataset page: https://huggingface.co/datasets/StarpowerTechnology/Dense-Information-Science-Physics-Dataset.texttext-generation1K<n<10K0 likes74 downloads14d agoHugging Face06tanmaylaud /scidcc-instructions Dataset Summary Instruction-Response pairs generated using the SciDCC Climate Dataset from Climabench Format ### Instruction: Present a fitting title for the provided text. For those who study earthquakes, one major challenge has been trying to understand all the physics of a fault -- both during an earthquake and at times of "rest" -- in order to know more about how a particular region may behave in the future. Now, researchers at the California Institute of Technology… See the full description on the dataset page: https://huggingface.co/datasets/tanmaylaud/scidcc-instructions.textsummarization10K<n<100K4 likes38 downloads3y agoHugging Face07YuJJJJin /ScienceOlympiad.tsv Dataset Card for ScienceOlympiad.tsv ScienceOlympiad.tsv: Challenging AI with Olympiad-Level Multimodal Science Problems. Source: https://huggingface.co/datasets/ByteDance-Seed/ScienceOlympiad Dataset Details Dataset Description The ScienceOlympiad dataset is a meticulously curated benchmark designed to evaluate the scientific reasoning capabilities of state-of-the-art AI models. It features elite, competition-level problems in physics and chemistry, addressing… See the full description on the dataset page: https://huggingface.co/datasets/YuJJJJin/ScienceOlympiad.tsv.textquestion-answeringn<1K0 likes36 downloads8mo agoHugging Face08prithivMLmods /Medi-Science Medi-Science Dataset The Medi-Science dataset is a comprehensive collection of medical Q&A data designed for text generation, question answering, and summarization tasks in the healthcare domain. Dataset Overview Name: Medi-Science License: Apache-2.0 Languages: English Tags: Medical, Medicine, Anomaly, Biology, Medi-Science Number of Rows: 16,412 Dataset Size: Downloaded: 22.7 MB Auto-converted Parquet: 8.94 MB Dataset Structure The dataset contains the… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Medi-Science.texttext-generation10K<n<100K5 likes33 downloads2y agoHugging Face09KadamParth /NCERT_Science_10thtabularquestion-answering1K<n<10K2 likes32 downloads2y agoHugging Face10wenhu /science_leaderboard_submissionThis dataset contains the results used for Science Leaderboard tabularquestion-answeringn<1K0 likes31 downloads2y agoHugging Face11KadamParth /NCERT_Science_8thtabularquestion-answering1K<n<10K1 likes30 downloads2y agoHugging Face12KadamParth /NCERT_Political_Science_12thtabularquestion-answering1K<n<10K1 likes30 downloads2y agoHugging Face13zorpsoon /scips_qatextquestion-answeringn<1K0 likes27 downloads2y agoHugging Face14relai-ai /scikit-learn-reasoningSamples in this benchmark were generated by RELAI using the following data source(s): Data Source Name: scikit-learn Data Source Link: https://scikit-learn.org/stable/index.html Data Source License: https://github.com/scikit-learn/scikit-learn/blob/main/COPYING Data Source Authors: scikit-learn contributors AI Benchmarks by Data Agents. 2025 RELAI.AI. Licensed under CC BY 4.0. Source: https://relai.ai textquestion-answering1K<n<10K0 likes26 downloads1y agoHugging Face15KadamParth /NCERT_Science_6thtabularquestion-answering1K<n<10K1 likes24 downloads2y agoHugging Face16KadamParth /NCERT_Science_9thtabularquestion-answering1K<n<10K3 likes20 downloads2y agoHugging Face17relai-ai /scipy-reasoningSamples in this benchmark were generated by RELAI using the following data source(s): Data Source Name: scipy documentation Data Source Link: https://docs.scipy.org/doc/scipy/reference/index.html Data Source License: CC0 1.0 Universal Data Source Authors: https://docs.scipy.org/doc/scipy/dev/governance.html AI Benchmarks by Data Agents. 2025 RELAI.AI Licensed under CC BY 4.0. Source: https://relai.ai textquestion-answeringn<1K0 likes19 downloads1y agoHugging Face18KadamParth /NCERT_Science_7thtabularquestion-answering1K<n<10K1 likes18 downloads2y agoHugging Face19KadamParth /NCERT_Political_Science_11thtabularquestion-answering1K<n<10K1 likes18 downloads2y agoHugging Face20nirajandhakal /everyday-science Everyday Science Dataset Dataset Summary The "Everyday Science" dataset is a synthetically generated collection of question-answer pairs covering a diverse range of topics within everyday science. Each entry includes a topic, a specific question, a scientific explanation or principle, a relatable everyday example, relevant scientific concepts (keywords), and further reasoning or inference. This dataset contains question-and-answer pairs focused on everyday science… See the full description on the dataset page: https://huggingface.co/datasets/nirajandhakal/everyday-science.texttable-question-answeringn<1K1 likes14 downloads1y agoHugging Face21relai-ai /scipy-standardSamples in this benchmark were generated by RELAI using the following data source(s): Data Source Name: scipy documentation Data Source Link: https://docs.scipy.org/doc/scipy/reference/index.html Data Source License: CC0 1.0 Universal Data Source Authors: https://docs.scipy.org/doc/scipy/dev/governance.html AI Benchmarks by Data Agents. 2025 RELAI.AI Licensed under CC BY 4.0. Source: https://relai.ai textquestion-answering1K<n<10K0 likes13 downloads1y agoHugging Face22anan6450 /NCERT_Science_10thtabularquestion-answering1K<n<10K0 likes11 downloads6mo agoHugging Face23shapap /SFT_Science_AI_Gentextquestion-answeringn<1K0 likes11 downloads2mo agoHugging Face24relai-ai /scikit-learn-standardSamples in this benchmark were generated by RELAI using the following data source(s): Data Source Name: scikit-learn Data Source Link: https://scikit-learn.org/stable/index.html Data Source License: https://github.com/scikit-learn/scikit-learn/blob/main/COPYING Data Source Authors: scikit-learn contributors AI Benchmarks by Data Agents. 2025 RELAI.AI. Licensed under CC BY 4.0. Source: https://relai.ai textquestion-answering1K<n<10K0 likes7 downloads1y agoHugging Face25ThorBaller /clinical_sciences_dataThis is a dataset for training AI on medical tools and practices in the modern age. textquestion-answeringn<1K0 likes5 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.