datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
yahoo_answers_topics
Dataset Card for "Yahoo Answers Topics"
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/yahoo_answers_topics.alignment_faking_harm_answerstext-sft-questions-answers-only
text-sft: Questions and Answers
This dataset consists of question-and-answer pairs generated from short excerpts drawn from Wikipedia, Cosmopedia, and FineWeb-Edu. It is an adapted version of agentlans/text-sft.
Overview
The dataset provides compact examples of English question-and-answer relationships that can help models learn linguistic patterns, syntactic structures, and semantic associations between questions and their corresponding answers.
Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/text-sft-questions-answers-only.hendrycks_math_with_answersreddit_question_best_answersQuestion & question body together with the best answers to that question from Reddit.
The score for the question / answer is the upvote count (i.e. positive-negative upvotes).
Only questions / answers that have these properties were extracted:
min_score = 3
min_title_len = 20
min_body_len = 100
cybersecurity_full_question_answersiheval-benign-answersyahoo-answers
Dataset Card for Yahoo Answers
This dataset is a collection of pairs containing titles, questions, and answers collected from Yahoo Answers. See the Yahoo Answers dataset for additional information. This dataset can be used directly with Sentence Transformers to train embedding models.
Dataset Subsets
title-question-answer-pair subset
Columns: "question", "answer"
Column types: str, str
Examples:{
'question': "why doesn't an optical mouse work on a glass… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/yahoo-answers.answers-with-reasoning-mmlu-pro
answers-with-reasoning-mmlu-pro
Self-distillation SFT corpus: Qwen3-8B-Instruct's own correct
chain-of-thought rollouts on MMLU-Pro multiple-choice questions
(general-QA domain).
Generation
Source problems: TIGER-Lab/MMLU-Pro test split (12,032 multiple-choice questions across 14 subject categories).
Sampling model: qwen/qwen3-8b via OpenRouter (providers: Alibaba, AtlasCloud) with reasoning enabled.
Sampling parameters: temperature=0.6, top_p=0.95, max_tokens=8000.… See the full description on the dataset page: https://huggingface.co/datasets/abhayesian/answers-with-reasoning-mmlu-pro.yahoo_answers_topicsVideo_As_Answers_Datayahoo_answers_topics
Dataset Card for "yahooanswerstopics"
More Information needed
zerobench_no_answershendrycks_math_with_answers_sftanswersumm
Dataset Card for answersumm
Dataset Summary
The AnswerSumm dataset is an English-language dataset of questions and answers collected from a StackExchange data dump. The dataset was created to support the task of query-focused answer summarization with an emphasis on multi-perspective answers.
The dataset consists of over 4200 such question-answer threads annotated by professional linguists and includes over 8700 summaries. We decompose the task into several annotation… See the full description on the dataset page: https://huggingface.co/datasets/alexfabbri/answersumm.task1592_yahoo_answers_topics_classfication
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1592_yahoo_answers_topics_classfication
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1592_yahoo_answers_topics_classfication.Yahoo_Answers_10_categories_for_NLP
Dataset Card for Dataset Name
The Yahoo! Answers topic classification dataset is constructed using 10 largest main categories. Each class contains 140,000 training samples and 6,000 testing samples. Therefore, the total number of training samples is 1,400,000 and testing samples 60,000 in this dataset. From all the answers and other meta-information, we only used the best answer content and the main category information.
Dataset Description
The file classes.txt contains a… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Yahoo_Answers_10_categories_for_NLP.yahoo_answerswhittle-teacher32-complete-answers
Whittle teacher32: complete answers with per-token teacher logprobs
Research preview. Part of the Whittle compression campaign, a personal
research project. The compute for this project is self funded and donations
decide whether the next round happens: https://ko-fi.com/davida81328
What this is
Complete answers generated by Qwen3.8-27B (UD-Q5_K_XL via llama.cpp), each
ending on a real end-of-turn token because the answer is finished, with the
teacher's top-32… See the full description on the dataset page: https://huggingface.co/datasets/logic65/whittle-teacher32-complete-answers.zerobench_answers_molmo7Byahoo_answers_topics_sampleThis is a sample from the yahoo_answers_topics dataset. This dataset contains 10% of the original dataset, randomly sampled by class.
Labels follow the following map:
id
label
0
Society & Culture
1
Science & Mathematics
2
Health
3
Education & Reference
4
Computers & Internet
5
Sports
6
Business & Finance
7
Entertainment & Music
8
Family & Relationships
9
Politics & Government
mcqa-multiple-answers
Dataset Summary
EVE-mcqa-multiple-answers is a Multiple-Choice Question Answering (MCQA) dataset designed to evaluate the performance of language models in the domain of Earth Observation (EO). The dataset consists of questions related to EO concepts, technologies, and applications, each accompanied by multiple answer choices, with one or more correct answer.
Dataset Structure
Each example in the dataset contains an arbitrary number of possible choices and one or more… See the full description on the dataset page: https://huggingface.co/datasets/eve-esa/mcqa-multiple-answers.task861_prost_mcq_answers_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task861_prost_mcq_answers_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task861_prost_mcq_answers_generation.task1594_yahoo_answers_topics_question_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1594_yahoo_answers_topics_question_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1594_yahoo_answers_topics_question_generation.answersxlam-function-calling-100-reasoning-answers
❌📞💭 Dataset Card for xlam-function-calling-100-reasoning-answers
A small subset of 100 entries adapted from Salesforce/xlam-function-calling-60k with additionnal thoughts from Qwen/QwQ-32B-AWQ within the Smolagents multi-step logics.
The answer of each function calling is also estimated by the LLM to be able to have an answer for the tool_call within the pipeline.
It integrate multiple <thought>, <tool_call> and <tool_response> blocks before the final answer for fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/Cotum/xlam-function-calling-100-reasoning-answers.question-answering-ukrainian-json-answersamc_2k_answersSame as zypchn/amc2k but answers are extracted as a seperate column.
reddit_question_best_answers_langstask1593_yahoo_answers_topics_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1593_yahoo_answers_topics_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1593_yahoo_answers_topics_classification.
