CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenMOSS-Team /SWE-bench-Science SWE-bench Science SWE-bench Science evaluates coding agents on software-engineering tasks drawn from scientific-computing repositories. The release contains 119 tasks across 20 scientific domains, with isolated environments and separate programmatic verifiers. GitHub release repository: OpenMOSS/SWE-bench-Science Runtime images: Docker Hub, pinned by immutable linux/amd64 digests Evaluation framework: Pier, compatible with Harbor task format Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/SWE-bench-Science.textn<1K7 likes3.2k downloads28d agoHugging Face02SAIS-Life-Science /Aneumo Aneumo Datasets AneumoDataset is a comprehensive multi-modal dataset containing 3D anatomical structures and simulated hemodynamic parameters for cerebral aneurysms, supporting both computational modeling and AI-based analysis. textn<1K6 likes2.6k downloads6mo agoHugging Face03armanc /ScienceQAThis is the ScientificQA dataset by Saikh et al (2022). @article{10.1007/s00799-022-00329-y, author = {Saikh, Tanik and Ghosal, Tirthankar and Mittal, Amish and Ekbal, Asif and Bhattacharyya, Pushpak}, title = {ScienceQA: A Novel Resource for Question Answering on Scholarly Articles}, year = {2022}, journal = {Int. J. Digit. Libr.}, month = {sep} } text10K<n<100K14 likes625 downloads4y agoHugging Face04mariiakoroliuk /generalization-science-datadocumentn<1K0 likes409 downloads2d agoHugging Face05deep-principle /science_materialstabularn<1K0 likes348 downloads2d agoHugging Face06juntaoyuan /test-sciencetextn<1K0 likes313 downloads2y agoHugging Face07science-of-finetuning /diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04 Contains maximum activating examples for all the features of our crosscoder trained on gemma 2 2B layer 13 available here: https://huggingface.co/Butanium/gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04/blob/main/README.md base_examples.pt contains all the maximum examples of the feature on a subset of validation test of fineweb chat_examples.pt is the same but for lmsys chat data chat_base_examples.pt is a merge of the two above files. All files are of the type dict[int, list[tuple[float… See the full description on the dataset page: https://huggingface.co/datasets/science-of-finetuning/diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04.tabular10K<n<100K0 likes235 downloads1y agoHugging Face08deep-principle /science_biologytabularn<1K0 likes162 downloads2d agoHugging Face09nasa-cisto-data-science-group /modis-lake-powell-toy-dataset MODIS Water Lake Powell Toy Dataset Dataset Summary Tabular dataset comprised of MODIS surface reflectance bands along with calculated indices and a label (water/not-water) Dataset Structure Data Fields water: Label, water or not-water (binary) sur_refl_b01_1: MODIS surface reflection band 1 (-100, 16000) sur_refl_b02_1: MODIS surface reflection band 2 (-100, 16000) sur_refl_b03_1: MODIS surface reflection band 3 (-100, 16000) sur_refl_b04_1: MODIS… See the full description on the dataset page: https://huggingface.co/datasets/nasa-cisto-data-science-group/modis-lake-powell-toy-dataset.image1K<n<10K1 likes154 downloads3y agoHugging Face10deep-principle /science_physicstextn<1K2 likes138 downloads2d agoHugging Face11Sangeetha /Kaggle-LLM-Science-Exam Dataset Card for [LLM Science Exam Kaggle Competition] Dataset Summary https://www.kaggle.com/competitions/kaggle-llm-science-exam/data Languages [en, de, tl, it, es, fr, pt, id, pl, ro, so, ca, da, sw, hu, no, nl, et, af, hr, lv, sl] Dataset Structure Columns prompt - the text of the question being asked A - option A; if this option is correct, then answer will be A B - option B; if this option is correct, then answer will be B C - option C; if this… See the full description on the dataset page: https://huggingface.co/datasets/Sangeetha/Kaggle-LLM-Science-Exam.text1K<n<10K3 likes128 downloads3y agoHugging Face12haidang2405 /tabrepair-science-repair-under-shift TabRepair Science: Repair Under Shift TabRepair Science is a finite authored benchmark for a deceptively hard question: does better tabular cell repair produce better downstream models under distribution shift? The 3,648-row pilot spans three structural generator families, missingness and present-value contamination, four test regimes, eight repair representations, and five downstream learners. A separate eight-world sensitivity layer tests a damage-aware v2 candidate without… See the full description on the dataset page: https://huggingface.co/datasets/haidang2405/tabrepair-science-repair-under-shift.tabulartabular-regression100K<n<1M0 likes127 downloads26d agoHugging Face13nasa-impact /nasa-science-repos-sme-benchmark NASA Science Repos SME Benchmark A benchmark dataset for evaluating retrieval systems on NASA science repository discovery tasks. This dataset contains expert queries, a corpus of NASA science GitHub repositories, and relevance judgments. Dataset Structure Files ├── corpus.jsonl # 5,264 repositories with full metadata ├── queries.jsonl # 219 expert queries └── qrels/ ├── earth.tsv # Earth Science relevance judgments (162) ├──… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-repos-sme-benchmark.tabulartext-retrievaln<1K0 likes118 downloads8mo agoHugging Face14river-martin /web-of-science-with-label-texts Dataset Description: The data is partitioned according to a 75/15/15 train/test/validate split. Each entry has an abstract (which is the input text for classification), a domain (a label from the list below), and an area (a subdomain of the paper, such as CS -> computer graphics, which takes on one of 134 possible values). All the attributes are strings. Domain labels: - Computer Science - Electrical Engineering - Psychology - Mechanical Engineering, - Civil Engineering - Medical… See the full description on the dataset page: https://huggingface.co/datasets/river-martin/web-of-science-with-label-texts.text10K<n<100K1 likes109 downloads2y agoHugging Face15rocky250 /Science-Discoverytabular100K<n<1M0 likes90 downloads10mo agoHugging Face16hugginglearners /data-science-job-salaries Dataset Card for Data Science Job Salaries Dataset Summary Content Column Description work_year The year the salary was paid. experience_level The experience level in the job during the year with the following possible values: EN Entry-level / Junior MI Mid-level / Intermediate SE Senior-level / Expert EX Executive-level / Director employment_type The type of employement for the role: PT Part-time FT Full-time CT Contract FL Freelance job_title… See the full description on the dataset page: https://huggingface.co/datasets/hugginglearners/data-science-job-salaries.tabularn<1K6 likes80 downloads4y agoHugging Face17nasa-impact /nasa-science-code-benchmark-v0.1.1 NASA Code Retrieval Benchmark v0.1.1 This repository is an updated version of the NASA Code Retrieval Benchmark. It provides a code retrieval benchmark based on code from 7 programming languages sourced from NASA's GitHub repositories. What's New in v0.1.1? v0.1.1 introduces a hierarchical structure and official Hugging Face dataset configurations. This allows you to evaluate models specifically by language or by query category without data redundancy in the file system.… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-code-benchmark-v0.1.1.texttext-retrieval100K<n<1M0 likes80 downloads6mo agoHugging Face18HydraLM /science_qa_txt_only_standardizedtabular10K<n<100K0 likes77 downloads3y agoHugging Face19Kaeyze /computer-science-synthetic-datasettext10K<n<100K12 likes74 downloads2y agoHugging Face20StarpowerTechnology /Dense-Information-Science-Physics-Dataset Dense Information With Multiple Fine-tuned Variations This dataaset has multiple for each input to learn how to express the same answer in different ways Dataset Structure The dataset contains two columns: Column Description input A science or quantum-physics question output A conversational answer to the question Example: { "input": "What is quantum entanglement?", "output": "Quantum entanglement is when two quantum systems share one… See the full description on the dataset page: https://huggingface.co/datasets/StarpowerTechnology/Dense-Information-Science-Physics-Dataset.texttext-generation1K<n<10K0 likes72 downloads12d agoHugging Face21hugging-science /jain-developability-cleantabularn<1K5 likes68 downloads1y agoHugging Face22xbench /ScienceQA xbench-evals 🌐 Website | 📄 Paper | 🤗 Dataset Evergreen, contamination-free, real-world, domain-specific AI evaluation framework xbench is more than just a scoreboard — it's a new evaluation framework with two complementary tracks, designed to measure both the intelligence frontier and real-world utility of AI systems: AGI Tracking: Measures core model capabilities like reasoning, tool-use, and memory Profession Aligned: A new class of evals grounded in workflows, environments… See the full description on the dataset page: https://huggingface.co/datasets/xbench/ScienceQA.textn<1K8 likes59 downloads1y agoHugging Face23nasa-impact /nasa-science-github-repos NASA Science GitHub Repositories A curated index of 5,264 GitHub repositories relevant to the NASA Science Mission Directorate (SMD), spanning five science divisions: Earth Science, Astrophysics, Planetary Science, Heliophysics, and Biological & Physical Sciences. This dataset is designed to support research on information retrieval and discoverability of open-source scientific software. Licensing and Intellectual Property This dataset is released under CC-BY-4.0 and… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-github-repos.tabulartext-retrieval1K<n<10K2 likes59 downloads6mo agoHugging Face24loukritia /science-journal-for-kids-data Science Journal for Kids Data This repository contains a dataset of abstracts from the Science Journal for Kids website and the original academic papers. It includes metadata such as titles, URLs, reading levels, and links to the full academic papers. The dataset is designed to support research and analysis of educational content tailored for young learners. Data The dataset is a curated collection of 284 original scientific abstracts and their adapted abstracts for… See the full description on the dataset page: https://huggingface.co/datasets/loukritia/science-journal-for-kids-data.textsummarizationn<1K1 likes58 downloads2y agoHugging Face25nasa-impact /nasa-science-code-benchmark-v0.1 NASA Code Retrieval Benchmark v0.1 Note: This dataset has been superseded by nasa-impact/nasa-science-code-benchmark-v0.1.1, which introduces a hierarchical structure, official Hugging Face dataset configurations, and evaluation by NASA science division. Please use v0.1.1 for new work. This dataset provides a code retrieval benchmark based on code from 7 programming languages (Python, C, C++, Java, JavaScript, Fortran, and Matlab) sourced from NASA's GitHub repositories. It serves… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-code-benchmark-v0.1.text100K<n<1M0 likes55 downloads6mo agoHugging Face26science-of-finetuning /diffing-stats-SAE-difference_cb-gemma-2-2b-L13-k100-x8-lr1e-04-local-shufflingtabular10K<n<100K0 likes54 downloads1y agoHugging Face27BDDSSD /ScienceAgentBench ScienceAgentBench The advancements of language language models (LLMs) have piqued growing interest in developing LLM-based language agents to automate scientific discovery end-to-end, which has sparked both excitement and skepticism about their true capabilities. In this work, we call for rigorous assessment of agents on individual tasks in a scientific workflow before making bold claims on end-to-end automation. To this end, we present ScienceAgentBench, a new benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/BDDSSD/ScienceAgentBench.textn<1K0 likes48 downloads7mo agoHugging Face28hugging-science /awesome-food-allergy-datasets Awesome Food Allergy Datasets A curated collection of datasets, databases, and computational resources for food allergy research, allergen identification, drug development, and clinical applications. 🧬 Dataset Description Dataset Summary Food allergy affects over 220 million people worldwide. This repository serves as the first comprehensive, open collection of AI-ready datasets for food allergy research—spanning clinical trials, immunotherapy, genomics… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/awesome-food-allergy-datasets.textn<1K6 likes46 downloads11mo agoHugging Face29jamesdborin /Nemotron-RL-Science-v1-prompt-only Nemotron-RL-Science-v1-prompt-only Prompt-only extraction from nvidia/Nemotron-RL-Science-v1. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row indexes where prompt extraction… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Science-v1-prompt-only.tabular100K<n<1M0 likes45 downloads3mo agoHugging Face30Miron /Science_Articlestextn<1K2 likes39 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.