datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-ended
Dataset Summary
EVE-open-ended is a collection of open-ended question-answer pairs focused on Earth Observation (EO). The datasets cover a wide range of EO topics, including, but not limited to satellite imagery analysis, remote sensing techniques, environmental monitoring, LiDAR, etc.
The datasets are designed to facilitate the development and evaluation of large language models (LLMs) in understanding and generating responses related to Earth Observation.
Metrics… See the full description on the dataset page: https://huggingface.co/datasets/eve-esa/open-ended.mcqa-single-answer
Dataset Summary
EVE-mcqa-single-answer is a Multiple-Choice Question Answering (MCQA) dataset designed to evaluate the performance of language models in the domain of Earth Observation (EO). The dataset consists of questions related to EO concepts, technologies, and applications, each accompanied by multiple answer choices with exactly one correct answer.
Unlike multi-answer MCQA datasets, each question in this dataset has only a single correct choice, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/eve-esa/mcqa-single-answer.hallucination
Dataset Summary
EVE-Hallucination is a specialized dataset designed to evaluate language models' tendency to hallucinate (generate factually incorrect or unsupported information) in the Earth Observation (EO) domain. Unlike typical QA datasets that focus on correctness, this dataset contains deliberately hallucinated answers with detailed annotations marking which portions of the text are hallucinated.
This dataset is crucial for developing and evaluating hallucination detection… See the full description on the dataset page: https://huggingface.co/datasets/eve-esa/hallucination.open-ended-w-context
Dataset Summary
EVE-open-ended-w-context is a collection of open-ended question-answer pairs focused on Earth Observation (EO) with accompanying context documents. Unlike the standard open-ended dataset, this version provides up to 3 relevant documents for each question that models can use to ground their responses. This makes it ideal for evaluating Retrieval-Augmented Generation (RAG) systems and testing models' ability to leverage provided context when answering questions.
The… See the full description on the dataset page: https://huggingface.co/datasets/eve-esa/open-ended-w-context.satcom-synth-qa
esa-sceva/satcom-synth-qa
Summary
Synthetic dataset of question-answer pairs on satellite communications, created to support model fine-tuning and evaluation in the SatCom domain.
Description
Generated from SatCom documents using large models (LLaMA 70B 3.3 Instruct and Qwen 2 72B Instruct).Two single-hop generation strategies were applied:
Joint QA generation from full documents.
Two-step process with separate question and answer generation for improved… See the full description on the dataset page: https://huggingface.co/datasets/esa-sceva/satcom-synth-qa.satcom-synth-qa-cot
esa-sceva/satcom-synth-qa-cot
Summary
Synthetic chain-of-thought dataset with step-by-step reasoning QAs for satellite communications.
Description
Generated with DeepSeek v3.2 and ChatGPT-5-Pro using topic-based prompts and structured answer templates.Covers technical subjects such as link budget, Doppler shift, latency, antenna gain, and propagation delay.Each answer includes reasoning steps, formulas, and explanations.
Composition
About 2.5k QAs… See the full description on the dataset page: https://huggingface.co/datasets/esa-sceva/satcom-synth-qa-cot.
