datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Question-Anchored-Tutoring-Dialogues-2k
Question-Anchored-Tutoring-Dialogues-2k
This dataset contains dialogues from math tutoring interventions recorded on Eedi.
Dataset Details
Dataset Description
Each dialogue represents a chat-based conversation between a tutor and a student prompted by the student requesting assistance while working on a lesson. Dialogues are accompanied with 2 sources of meta-data:
DQ-Question-Metadata: The question the student was working on that prompted the tutoring… See the full description on the dataset page: https://huggingface.co/datasets/Eedi/Question-Anchored-Tutoring-Dialogues-2k.stackoverflow-questions-long
stackoverflow questions for text classification: 'long'
This is pacovaldez/stackoverflow-questions filtered for 1024 GPT2 tokens or more in title + body
https://huggingface.co/datasets/pacovaldez/stackoverflow-questions
wb-questions
Dataset Card for Wildberries questions
Dataset Summary
This is a dataset of questions and answers scraped from product pages from the Russian marketplace Wildberries. Dataset contains all questions and answers, as well as all metadata from the API. However, the "productName" field may be empty in some cases because the API does not return the name for old products.
Languages
The dataset is mostly in Russian, but there may be other languages present.… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/wb-questions.forecastbench-single_question
ForecastBench Single Questions
This dataset contains single-ID forecasting questions derived from the ForecastBench project. It includes two configurations:
forecastbench_single_questions_2024-12-08: Contains 429 forecasting questions with resolved real-world outcomes.
forecastbench_single_questions_human_2024-07-21: Contains 473 questions with resolved real-world outcomes, augmented with human forecast probabilities from public and superforecaster groups.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Duruo/forecastbench-single_question.yher-chemistry-question-bank
YHer Chemistry Question Bank
The data layer of an evidence-bound diagnostic learning system for Shanghai high-school chemistry (Chris-TLC/YHer-skill).
Every record in this dataset is derived from publicly released Shanghai gaokao and mock examination papers through deterministic mechanical structuring: text extraction, layout repair, and answer alignment. No content is model-generated.
What's inside
The dataset ships in two configs:
Config
Records
Content… See the full description on the dataset page: https://huggingface.co/datasets/Chris-TLC/yher-chemistry-question-bank.is-trivia-questions
Icelandic trivia questions
Icelandic trivia question compiled and created by Sveinn Steinarsson, Valur Freyr Steinarsson, and Svavar Kjarrval https://github.com/sveinn-steinarsson/is-trivia-questions
Dálkanúmer
Valfrjálst
Lýsing
1
Nei
Flokkanúmer
2
Já
Undirflokkur ef til staðar
3
Nei
Erfiðleikastig: 1: Létt, 2: Meðal, 3: Erfið
4
Já
Gæðastig: 1: Slöpp, 2: Góð, 3: Ágæt
5
Nei
Spurningin
6
Nei
Svarið
Flokkanúmer
Flokkanafn
1
Almenn kunnátta
2
Náttúra… See the full description on the dataset page: https://huggingface.co/datasets/Sigurdur/is-trivia-questions.Dermatology-Question-Answer-Dataset-For-Fine-Tuning
Dataset Details
The data set has about 1 Million Tokens for Training and about 1500 question answers.
Dataset Description
This dataset is a comprehensive compilation of questions related to dermatology, spanning inquiries about various skin diseases, their symptoms, recommended medications, and available treatment modalities. Each question is paired with a concise and informative response, making it an ideal resource for training and fine-tuning language models in the… See the full description on the dataset page: https://huggingface.co/datasets/Mreeb/Dermatology-Question-Answer-Dataset-For-Fine-Tuning.med-synth-questions-gemma-3-27b-deepseek-v4-flash
Med Synth Questions (Gemma-3 + DeepSeek V4 Flash)
Synthetic reasoning traces and answers for medical questions from openmed-community/med-synth-questions-gemma-3-27b-it. Each record contains a medical question with SYNTH-style reasoning and a generated answer by DeepSeek V4 Flash.
Dataset Summary
29,148 records (2 dupes + 3,410 incomplete/truncated removed from 32,560 source)
29,148 reasoning turns (99.2% format compliance)
Average 1,591 chars per reasoning trace… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/med-synth-questions-gemma-3-27b-deepseek-v4-flash.synthlabs-openmed-questions-qwen3-235b-a22b-2507
Med Synth Questions (Qwen3-235B questions + DeepSeek V4 Flash and Minimax M2.7 answers)
Synthetic reasoning traces for medical questions from openmed-community/med-synth-questions-qwen3-235b-a22b-2507. Each record contains a medical question with SYNTH-style reasoning and a generated answer.
Dataset Summary
55,915 records (255 dupes + 3,523 incomplete/truncated removed from 59,693 source)
55,915 reasoning turns (99.9% format compliance)
Average 1,881 chars per… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/synthlabs-openmed-questions-qwen3-235b-a22b-2507.QuestA-OpenR1-Math-220k-Joined
QuestA joined with OpenR1-Math-220k
This dataset joins foreverlasting1202/QuestA back to the corresponding rows in
open-r1/OpenR1-Math-220k.
Contents
Each row contains the complete OpenR1 fields, including:
solution
problem_type, question_type, and source
uuid
all original generations
correctness_math_verify and correctness_llama
finish_reasons and correctness_count
messages
The original QuestA values are retained explicitly as:
questa_index
questa_problem… See the full description on the dataset page: https://huggingface.co/datasets/YYYYYYibo/QuestA-OpenR1-Math-220k-Joined.rewrite-questions-nonsensical-biology
nonsensical_biology.csv - Question Rewriting Dataset
This dataset contains question rewriting outputs from the file nonsensical_biology.csv.
Dataset Structure
The dataset contains the following columns:
custom_id: Unique identifier for each question
style: Rewriting style applied (e.g., "gibberish")
index: Numerical index
original: Original question text
rewritten: Rewritten version of the question
options: Multiple choice options (list format)
correct: Index of the… See the full description on the dataset page: https://huggingface.co/datasets/NLie2/rewrite-questions-nonsensical-biology.Question-Anchored-Tutoring-Dialogues-2k
Question-Anchored-Tutoring-Dialogues-2k
This dataset contains dialogues from math tutoring interventions recorded on Eedi.
Dataset Details
Dataset Description
Each dialogue represents a chat-based conversation between a tutor and a student prompted by the student requesting assistance while working on a lesson. Dialogues are accompanied with 2 sources of meta-data:
DQ-Question-Metadata: The question the student was working on that prompted the… See the full description on the dataset page: https://huggingface.co/datasets/Abhishekh13/Question-Anchored-Tutoring-Dialogues-2k.rewrite-questions-real-words-sciency
real_words_sciency.csv - Question Rewriting Dataset
This dataset contains question rewriting outputs from the file real_words_sciency.csv.
Dataset Structure
The dataset contains the following columns:
custom_id: Unique identifier for each question
style: Rewriting style applied (e.g., "gibberish")
index: Numerical index
original: Original question text
rewritten: Rewritten version of the question
options: Multiple choice options (list format)
correct: Index of the… See the full description on the dataset page: https://huggingface.co/datasets/NLie2/rewrite-questions-real-words-sciency.arabic_dialects_question_and_answerData Content
The file provided: Q/A Reasoning dataset
contains the following columns:
ID # : Denotes the reference ID for:
a. Question
b. Answer to the question
c. Hint
d. Reasoning
e. Word count for items a to d above
Dialects: Contains the following dialects in separate columns:
a. English
b. MSA
c. Emirati
d. Egyptian
e. Levantine Syria
f. Levantine Jordan
g. Levantine Palestine
h. Levantine Lebanon
Data Generation Process
The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.rewrite-questions-gibberish
gibberish.csv - Question Rewriting Dataset
This dataset contains question rewriting outputs from the file gibberish.csv.
Dataset Structure
The dataset contains the following columns:
custom_id: Unique identifier for each question
style: Rewriting style applied (e.g., "gibberish")
index: Numerical index
original: Original question text
rewritten: Rewritten version of the question
options: Multiple choice options (list format)
correct: Index of the correct answer… See the full description on the dataset page: https://huggingface.co/datasets/NLie2/rewrite-questions-gibberish.speechmap-questions
SpeechMap Questions & Responses
______________________________________
/ 144,459 spicy takes graded for \
| compliance. The cow has seen things. |
| The cow remains neutral across all |
\ lenses. /
--------------------------------------
\ ^__^
\ (oo)\_______
(__)\ )\/\
||----w |
|| ||
A mirror of data from speechmap.ai, a project by xlr8harder that measures how… See the full description on the dataset page: https://huggingface.co/datasets/wassname/speechmap-questions.hackathon-advisor-quest-dataset
Hackathon Advisor — Quest Classification SFT Dataset
Supervised fine-tuning data that teaches MiniCPM5-1B to classify a Build Small
Hackathon project against 13 judging dimensions from a two-segment README + app-file
prompt, emitting strict JSON with short, source-attributed evidence. Trains the LoRA at
build-small-hackathon/hackathon-advisor-quest-minicpm5-lora.
Files
quest_sft.jsonl — the dataset (one lora_sft_example per line; the viewer split).… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-quest-dataset.question_answering
Visualization of Question & Answering Task Cases Samples
Check dataset samples visualization by viewing Dataset Viewer.
The sampling procedure is guided by the Elo distribution introduced in our method.
Original dataset is test split of cais/mmlu from hugging face.
samples/origin: 14041/14042
License
This repository is licensed under the Apache License 2.0
