datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nasa-science-repos-sme-benchmark
NASA Science Repos SME Benchmark
A benchmark dataset for evaluating retrieval systems on NASA science repository discovery tasks. This dataset contains expert queries, a corpus of NASA science GitHub repositories, and relevance judgments.
Dataset Structure
Files
├── corpus.jsonl # 5,264 repositories with full metadata
├── queries.jsonl # 219 expert queries
└── qrels/
├── earth.tsv # Earth Science relevance judgments (162)
├──… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-repos-sme-benchmark.Dense-Information-Science-Physics-Dataset
Dense Information With Multiple Fine-tuned Variations
This dataaset has multiple for each input to learn how to express the same answer in different ways
Dataset Structure
The dataset contains two columns:
Column
Description
input
A science or quantum-physics question
output
A conversational answer to the question
Example:
{
"input": "What is quantum entanglement?",
"output": "Quantum entanglement is when two quantum systems share one… See the full description on the dataset page: https://huggingface.co/datasets/StarpowerTechnology/Dense-Information-Science-Physics-Dataset.ScienceOlympiad.tsv
Dataset Card for ScienceOlympiad.tsv
ScienceOlympiad.tsv: Challenging AI with Olympiad-Level Multimodal Science Problems.
Source: https://huggingface.co/datasets/ByteDance-Seed/ScienceOlympiad
Dataset Details
Dataset Description
The ScienceOlympiad dataset is a meticulously curated benchmark designed to evaluate the scientific reasoning capabilities of state-of-the-art AI models. It features elite, competition-level problems in physics and chemistry, addressing… See the full description on the dataset page: https://huggingface.co/datasets/YuJJJJin/ScienceOlympiad.tsv.Medi-Science
Medi-Science Dataset
The Medi-Science dataset is a comprehensive collection of medical Q&A data designed for text generation, question answering, and summarization tasks in the healthcare domain.
Dataset Overview
Name: Medi-Science
License: Apache-2.0
Languages: English
Tags: Medical, Medicine, Anomaly, Biology, Medi-Science
Number of Rows: 16,412
Dataset Size:
Downloaded: 22.7 MB
Auto-converted Parquet: 8.94 MB
Dataset Structure
The dataset contains the… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Medi-Science.NCERT_Science_10thscience_leaderboard_submissionThis dataset contains the results used for Science Leaderboard
NCERT_Political_Science_12thNCERT_Science_8thNCERT_Science_6thNCERT_Science_9thNCERT_Science_7thNCERT_Political_Science_11theveryday-science
Everyday Science Dataset
Dataset Summary
The "Everyday Science" dataset is a synthetically generated collection of question-answer pairs covering a diverse range of topics within everyday science. Each entry includes a topic, a specific question, a scientific explanation or principle, a relatable everyday example, relevant scientific concepts (keywords), and further reasoning or inference.
This dataset contains question-and-answer pairs focused on everyday science… See the full description on the dataset page: https://huggingface.co/datasets/nirajandhakal/everyday-science.NCERT_Science_10thSFT_Science_AI_Genclinical_sciences_dataThis is a dataset for training AI on medical tools and practices in the modern age.
