CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenMOSS-Team /SWE-bench-Science SWE-bench Science SWE-bench Science evaluates coding agents on software-engineering tasks drawn from scientific-computing repositories. The release contains 119 tasks across 20 scientific domains, with isolated environments and separate programmatic verifiers. GitHub release repository: OpenMOSS/SWE-bench-Science Runtime images: Docker Hub, pinned by immutable linux/amd64 digests Evaluation framework: Pier, compatible with Harbor task format Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/SWE-bench-Science.textn<1K7 likes3.2k downloads28d agoHugging Face02SAIS-Life-Science /Aneumo Aneumo Datasets AneumoDataset is a comprehensive multi-modal dataset containing 3D anatomical structures and simulated hemodynamic parameters for cerebral aneurysms, supporting both computational modeling and AI-based analysis. textn<1K6 likes2.6k downloads6mo agoHugging Face03armanc /ScienceQAThis is the ScientificQA dataset by Saikh et al (2022). @article{10.1007/s00799-022-00329-y, author = {Saikh, Tanik and Ghosal, Tirthankar and Mittal, Amish and Ekbal, Asif and Bhattacharyya, Pushpak}, title = {ScienceQA: A Novel Resource for Question Answering on Scholarly Articles}, year = {2022}, journal = {Int. J. Digit. Libr.}, month = {sep} } text10K<n<100K14 likes625 downloads4y agoHugging Face04mariiakoroliuk /generalization-science-datadocumentn<1K0 likes409 downloads2d agoHugging Face05deep-principle /science_materialstabularn<1K0 likes348 downloads3d agoHugging Face06juntaoyuan /test-sciencetextn<1K0 likes313 downloads2y agoHugging Face07science-of-finetuning /diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04 Contains maximum activating examples for all the features of our crosscoder trained on gemma 2 2B layer 13 available here: https://huggingface.co/Butanium/gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04/blob/main/README.md base_examples.pt contains all the maximum examples of the feature on a subset of validation test of fineweb chat_examples.pt is the same but for lmsys chat data chat_base_examples.pt is a merge of the two above files. All files are of the type dict[int, list[tuple[float… See the full description on the dataset page: https://huggingface.co/datasets/science-of-finetuning/diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04.tabular10K<n<100K0 likes235 downloads1y agoHugging Face08deep-principle /science_biologytabularn<1K0 likes162 downloads3d agoHugging Face09deep-principle /science_physicstextn<1K2 likes138 downloads3d agoHugging Face10Sangeetha /Kaggle-LLM-Science-Exam Dataset Card for [LLM Science Exam Kaggle Competition] Dataset Summary https://www.kaggle.com/competitions/kaggle-llm-science-exam/data Languages [en, de, tl, it, es, fr, pt, id, pl, ro, so, ca, da, sw, hu, no, nl, et, af, hr, lv, sl] Dataset Structure Columns prompt - the text of the question being asked A - option A; if this option is correct, then answer will be A B - option B; if this option is correct, then answer will be B C - option C; if this… See the full description on the dataset page: https://huggingface.co/datasets/Sangeetha/Kaggle-LLM-Science-Exam.text1K<n<10K3 likes128 downloads3y agoHugging Face11haidang2405 /tabrepair-science-repair-under-shift TabRepair Science: Repair Under Shift TabRepair Science is a finite authored benchmark for a deceptively hard question: does better tabular cell repair produce better downstream models under distribution shift? The 3,648-row pilot spans three structural generator families, missingness and present-value contamination, four test regimes, eight repair representations, and five downstream learners. A separate eight-world sensitivity layer tests a damage-aware v2 candidate without… See the full description on the dataset page: https://huggingface.co/datasets/haidang2405/tabrepair-science-repair-under-shift.tabulartabular-regression100K<n<1M0 likes127 downloads26d agoHugging Face12river-martin /web-of-science-with-label-texts Dataset Description: The data is partitioned according to a 75/15/15 train/test/validate split. Each entry has an abstract (which is the input text for classification), a domain (a label from the list below), and an area (a subdomain of the paper, such as CS -> computer graphics, which takes on one of 134 possible values). All the attributes are strings. Domain labels: - Computer Science - Electrical Engineering - Psychology - Mechanical Engineering, - Civil Engineering - Medical… See the full description on the dataset page: https://huggingface.co/datasets/river-martin/web-of-science-with-label-texts.text10K<n<100K1 likes109 downloads2y agoHugging Face13hugginglearners /data-science-job-salaries Dataset Card for Data Science Job Salaries Dataset Summary Content Column Description work_year The year the salary was paid. experience_level The experience level in the job during the year with the following possible values: EN Entry-level / Junior MI Mid-level / Intermediate SE Senior-level / Expert EX Executive-level / Director employment_type The type of employement for the role: PT Part-time FT Full-time CT Contract FL Freelance job_title… See the full description on the dataset page: https://huggingface.co/datasets/hugginglearners/data-science-job-salaries.tabularn<1K6 likes80 downloads4y agoHugging Face14nasa-impact /nasa-science-code-benchmark-v0.1.1 NASA Code Retrieval Benchmark v0.1.1 This repository is an updated version of the NASA Code Retrieval Benchmark. It provides a code retrieval benchmark based on code from 7 programming languages sourced from NASA's GitHub repositories. What's New in v0.1.1? v0.1.1 introduces a hierarchical structure and official Hugging Face dataset configurations. This allows you to evaluate models specifically by language or by query category without data redundancy in the file system.… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-code-benchmark-v0.1.1.texttext-retrieval100K<n<1M0 likes80 downloads6mo agoHugging Face15HydraLM /science_qa_txt_only_standardizedtabular10K<n<100K0 likes77 downloads3y agoHugging Face16Kaeyze /computer-science-synthetic-datasettext10K<n<100K12 likes74 downloads2y agoHugging Face17StarpowerTechnology /Dense-Information-Science-Physics-Dataset Dense Information With Multiple Fine-tuned Variations This dataaset has multiple for each input to learn how to express the same answer in different ways Dataset Structure The dataset contains two columns: Column Description input A science or quantum-physics question output A conversational answer to the question Example: { "input": "What is quantum entanglement?", "output": "Quantum entanglement is when two quantum systems share one… See the full description on the dataset page: https://huggingface.co/datasets/StarpowerTechnology/Dense-Information-Science-Physics-Dataset.texttext-generation1K<n<10K0 likes72 downloads12d agoHugging Face18hugging-science /jain-developability-cleantabularn<1K5 likes68 downloads1y agoHugging Face19xbench /ScienceQA xbench-evals 🌐 Website | 📄 Paper | 🤗 Dataset Evergreen, contamination-free, real-world, domain-specific AI evaluation framework xbench is more than just a scoreboard — it's a new evaluation framework with two complementary tracks, designed to measure both the intelligence frontier and real-world utility of AI systems: AGI Tracking: Measures core model capabilities like reasoning, tool-use, and memory Profession Aligned: A new class of evals grounded in workflows, environments… See the full description on the dataset page: https://huggingface.co/datasets/xbench/ScienceQA.textn<1K8 likes59 downloads1y agoHugging Face20nasa-impact /nasa-science-github-repos NASA Science GitHub Repositories A curated index of 5,264 GitHub repositories relevant to the NASA Science Mission Directorate (SMD), spanning five science divisions: Earth Science, Astrophysics, Planetary Science, Heliophysics, and Biological & Physical Sciences. This dataset is designed to support research on information retrieval and discoverability of open-source scientific software. Licensing and Intellectual Property This dataset is released under CC-BY-4.0 and… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-github-repos.tabulartext-retrieval1K<n<10K2 likes59 downloads6mo agoHugging Face21loukritia /science-journal-for-kids-data Science Journal for Kids Data This repository contains a dataset of abstracts from the Science Journal for Kids website and the original academic papers. It includes metadata such as titles, URLs, reading levels, and links to the full academic papers. The dataset is designed to support research and analysis of educational content tailored for young learners. Data The dataset is a curated collection of 284 original scientific abstracts and their adapted abstracts for… See the full description on the dataset page: https://huggingface.co/datasets/loukritia/science-journal-for-kids-data.textsummarizationn<1K1 likes58 downloads2y agoHugging Face22nasa-impact /nasa-science-code-benchmark-v0.1 NASA Code Retrieval Benchmark v0.1 Note: This dataset has been superseded by nasa-impact/nasa-science-code-benchmark-v0.1.1, which introduces a hierarchical structure, official Hugging Face dataset configurations, and evaluation by NASA science division. Please use v0.1.1 for new work. This dataset provides a code retrieval benchmark based on code from 7 programming languages (Python, C, C++, Java, JavaScript, Fortran, and Matlab) sourced from NASA's GitHub repositories. It serves… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-code-benchmark-v0.1.text100K<n<1M0 likes55 downloads6mo agoHugging Face23BDDSSD /ScienceAgentBench ScienceAgentBench The advancements of language language models (LLMs) have piqued growing interest in developing LLM-based language agents to automate scientific discovery end-to-end, which has sparked both excitement and skepticism about their true capabilities. In this work, we call for rigorous assessment of agents on individual tasks in a scientific workflow before making bold claims on end-to-end automation. To this end, we present ScienceAgentBench, a new benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/BDDSSD/ScienceAgentBench.textn<1K0 likes48 downloads7mo agoHugging Face24hugging-science /awesome-food-allergy-datasets Awesome Food Allergy Datasets A curated collection of datasets, databases, and computational resources for food allergy research, allergen identification, drug development, and clinical applications. 🧬 Dataset Description Dataset Summary Food allergy affects over 220 million people worldwide. This repository serves as the first comprehensive, open collection of AI-ready datasets for food allergy research—spanning clinical trials, immunotherapy, genomics… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/awesome-food-allergy-datasets.textn<1K6 likes46 downloads11mo agoHugging Face25jamesdborin /Nemotron-RL-Science-v1-prompt-only Nemotron-RL-Science-v1-prompt-only Prompt-only extraction from nvidia/Nemotron-RL-Science-v1. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row indexes where prompt extraction… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Science-v1-prompt-only.tabular100K<n<1M0 likes45 downloads3mo agoHugging Face26Miron /Science_Articlestextn<1K2 likes39 downloads4y agoHugging Face27prithivMLmods /Medi-Science Medi-Science Dataset The Medi-Science dataset is a comprehensive collection of medical Q&A data designed for text generation, question answering, and summarization tasks in the healthcare domain. Dataset Overview Name: Medi-Science License: Apache-2.0 Languages: English Tags: Medical, Medicine, Anomaly, Biology, Medi-Science Number of Rows: 16,412 Dataset Size: Downloaded: 22.7 MB Auto-converted Parquet: 8.94 MB Dataset Structure The dataset contains the… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Medi-Science.texttext-generation10K<n<100K5 likes34 downloads2y agoHugging Face28KadamParth /NCERT_Science_10thtabularquestion-answering1K<n<10K2 likes34 downloads2y agoHugging Face29YuJJJJin /ScienceOlympiad.tsv Dataset Card for ScienceOlympiad.tsv ScienceOlympiad.tsv: Challenging AI with Olympiad-Level Multimodal Science Problems. Source: https://huggingface.co/datasets/ByteDance-Seed/ScienceOlympiad Dataset Details Dataset Description The ScienceOlympiad dataset is a meticulously curated benchmark designed to evaluate the scientific reasoning capabilities of state-of-the-art AI models. It features elite, competition-level problems in physics and chemistry, addressing… See the full description on the dataset page: https://huggingface.co/datasets/YuJJJJin/ScienceOlympiad.tsv.textquestion-answeringn<1K0 likes34 downloads8mo agoHugging Face30wenhu /science_leaderboard_submissionThis dataset contains the results used for Science Leaderboard tabularquestion-answeringn<1K0 likes33 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.