datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RETECO-SemEval2027
RETECO
Training & Development Data · SemEval-2027 Task 1
Retrieval that must reason about when evidence applies
and what the conversation has already established.
📋 Task ·
📊 Data ·
📐 Evaluation ·
🚀 Participate ·
🧰 Starter kit
🎯 At a glance
What
Official training and development data for RETECO, the SemEval-2027 shared task on reasoning-oriented retrieval
Scope
2 tracks · 5 sub-tracks · 24 self-contained domains… See the full description on the dataset page: https://huggingface.co/datasets/DataScience-UIBK/RETECO-SemEval2027.PopMCQ
🎯 PopMCQ
Does your model pick the famous answer or the correct one?
📌 Overview
PopMCQ renders the same question six ways. The question and the correct answer never change — only how popular the three distractors are. That makes option popularity an independent variable, so an accuracy swing across S1–S6 is attributable to popularity rather than to question difficulty.
The swings are large. Under the hardest setting (S2), models pick a popular-but-wrong… See the full description on the dataset page: https://huggingface.co/datasets/DataScience-UIBK/PopMCQ.Cannabis_Science_Data
Cannabis Science Literature QA Dataset
This dataset contains 161,170 high-quality question-answer pairs derived from over 400 peer-reviewed cannabis science research papers and textbooks. Created to advance AI research in cannabis science and medical applications, it provides a comprehensive resource for training language models on cannabis-related scientific knowledge.
Dataset Details
Dataset Description
This dataset was systematically generated from a curated… See the full description on the dataset page: https://huggingface.co/datasets/KellanF89/Cannabis_Science_Data.ComplexTempQA
ComplexTempQA Dataset
ComplexTempQA is a large-scale dataset designed for complex temporal question answering (TQA). It consists of over 100 million question-answer pairs, making it one of the most extensive datasets available for TQA. The dataset is generated using data from Wikipedia and Wikidata and spans questions over a period of 36 years (1987-2023).
Note: We have a smaller version consisting of questions from the time period 1987 until 2007.
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/DataScienceUIBK/ComplexTempQA.nemotron-terminal-data_science
nemotron-terminal-data_science
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "data_science". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-data_science.Cannabis_Science_Data
Cannabis Science Literature QA Dataset
This dataset contains 161,170 high-quality question-answer pairs derived from over 400 peer-reviewed cannabis science research papers and textbooks. Created to advance AI research in cannabis science and medical applications, it provides a comprehensive resource for training language models on cannabis-related scientific knowledge.
Dataset Details
Dataset Description
This dataset was systematically generated from a curated… See the full description on the dataset page: https://huggingface.co/datasets/tekwhisperer/Cannabis_Science_Data.data-science-chatbot
📊 Data Science Chatbot Dataset (2000 Samples)
🚀 A high-quality instruction-style dataset designed for fine-tuning Large Language Models (LLMs) on Data Science concepts.
This dataset contains ~2000 curated question-answer pairs in ChatML format, enabling models to learn how to explain, define, and discuss core data science topics in a clear and beginner-friendly way.
🎯 Objective
The goal of this dataset is to:
Train LLMs to act as a Data Science Tutor
Provide clear… See the full description on the dataset page: https://huggingface.co/datasets/Hamzasajjad38/data-science-chatbot.clinical_sciences_dataThis is a dataset for training AI on medical tools and practices in the modern age.
