datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NASA-EO-Bench
NASA-EO-Bench
A large-scale benchmark for geoscience dataset retrieval, derived from citation relationships in peer-reviewed NASA publications.
Paper: Bringing Agentic Search to Earth Observation Data Discovery — CIKM '26, 10.1145/3799682.3841109
Overview
Finding the right NASA Earth observation dataset for a given research need is hard even for domain experts. NASA-EO-Bench operationalises this task as an information retrieval problem: given a natural-language… See the full description on the dataset page: https://huggingface.co/datasets/HamiltonMYu/NASA-EO-Bench.nasa-sde-IR-benchmark-20251024-v5
NASA SDE IR Benchmark v5
A comprehensive Information Retrieval benchmark dataset for the NASA Science Discovery Engine (SDE), containing synthetically generated query-document pairs for scientific content retrieval evaluation.
Paper: INDUS-SDE: A Language Model for Scientific Content Curation and Discovery — KDD 2026, AI for Sciences Track. This is the in-domain NASA SDE IR benchmark used to evaluate INDUS-SDE-ST.
Code: NASA-IMPACT/st-training-workflow
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-sde-IR-benchmark-20251024-v5.Quran_tafsir
Nasq Quranic Dataset (Arabic-English Tafsir)
لتفسير القرآن الكريم تحتوي على النص القرآني كاملاً مع التفسير الميسر (بالعربي) وتفسير المختصر (بالإنجليزي).
Columns:
id: المعرف الفريد لكل آية (من 1 إلى 6236).
surah_n: رقم السورة.
ayah_n: رقم الآية داخل السورة.
surah_name_arabic: اسم السورة باللغة العربية (تم تحديثه لضمان الدقة).
surah_name_english: اسم السورة باللغة الإنجليزية.
ayah_text_ar: نص الآية.
ayah_text_en: ترجمة نص الآية للإنجليزية.
arabic_tafsir: التفسير الميسر.… See the full description on the dataset page: https://huggingface.co/datasets/Nasaq-GP/Quran_tafsir.nasa-smd-qa-benchmark
NASA-QA Benchmark
NASA SMD and IBM research developed NASA-QA benchmark, an extractive question answering task focused on the Earth science domain. First, 39 paragraphs from Earth science papers which appeared in AGU and AMS journals were sourced. Subject matter experts from NASA formulated questions and marked the corresponding answers in these paragraphs, resulting in a total of 117 question-answer pairs. The dataset is split into a training set of 90 pairs and a validation set of… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-smd-qa-benchmark.EO-via-NLP
Dataset Summary
Toward Open Earth Science as Fast and Accessible as Natural Language
This dataset was curated to accompany the EO-via-NLP code and the following paper:
Ellis, M., Gurung, I., Ramasubramanian, M., & Ramachandran, R. (2025).Toward Open Earth Science as Fast and Accessible as Natural Language.arXiv:2505.15690
Supported Tasks
This dataset was primarily designed for:
Named Entity Recognition (NER) in earth science contexts.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/EO-via-NLP.NASA-Systems-Engineering-QA-JSONLnasa-smd-qa-benchmark-cleaned
NASA-QA (Cleaned, SQuAD v2 Format)
This dataset is a cleaned and reformatted derivative of:
https://huggingface.co/datasets/nasa-impact/nasa-smd-qa-benchmark
The original dataset is an extractive question answering benchmark in the Earth science domain.This version modifies the data to ensure compatibility with SQuAD v2-style training and evaluation.
Changes from Original
Compared to the original release, this version:
converts the nested structure into one… See the full description on the dataset page: https://huggingface.co/datasets/quantaRoche/nasa-smd-qa-benchmark-cleaned.NASA-Systems-Engineering-QA-ChatMLNASA-Systems-Engineering-QA-AlpacaNASA-Systems-Engineering-QA-FTmilling_nasa
