datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
docinsights-2026-shared-task-data
DocInsights 2026 Shared Task: DocSem
Document-grounded quantitative reasoning with evidence attribution
DocSem is the shared task of DocInsights 2026, the Workshop on Document Intelligence and Understanding co-located with EMNLP 2026 in Budapest, Hungary. The workshop theme is Beyond Plain Text: Bridging NLP and Document AI.
Workshop shared task | Source repository | Submission portal | Participant guide
Participants receive a PDF document and a paraphrased user_query. Systems… See the full description on the dataset page: https://huggingface.co/datasets/amitbcp/docinsights-2026-shared-task-data.FinLongDocQA
FinLongDocQA
Document-Level Numerical Reasoning across Single and Multiple Tables in Financial Reports
Annual reports used in this dataset can be downloaded here.
Version
The current release is FinLongDocQA v1.1, with updated evidence-page annotations. See UPDATE.md for details.
Dataset Description
An example QA instance from FinLongDocQA. The figure shows only the relevant tables and text for presentation; in practice, the model must retrieve… See the full description on the dataset page: https://huggingface.co/datasets/Amian/FinLongDocQA.mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 2B Thinking hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k.mhlc-training-qwen3.5-qwen3_5_9b_think_off_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3.5 9B think off hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3.5-qwen3_5_9b_think_off_hard_mixed_sources_120k.mhlc-training-qwen3vl-qwen3-vl-32b-instruct_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 32B Instruct hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3-vl-32b-instruct_hard_mixed_sources_120k.mhlc-training-qwen3vl-qwen3_vl_4b_thinking_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 4B Thinking hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_4b_thinking_hard_mixed_sources_120k.mhlc-training-qwen3vl-qwen3_vl_4b_instruct_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 4B Instruct hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_4b_instruct_hard_mixed_sources_120k.mhlc-training-qwen3vl-qwen3_vl_2b_instruct_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 2B Instruct hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_2b_instruct_hard_mixed_sources_120k.mhlc-training-qwen3.5-qwen3_5_4b_think_off_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3.5 4B think off hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3.5-qwen3_5_4b_think_off_hard_mixed_sources_120k.mhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Gemma 4 E4B it think on hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k.ai-ml-foundations-book-collection
Introduction
I put this collection together after spending a lot of time reading what I think are some of the best books on AI, machine learning, deep learning, probabilistic modeling, optimization, reinforcement learning, transformers, LLMs, validation, and fairness. I want to share this with the community for one simple reason: I want to give people a structured path through the books that actually help them understand things deeply, instead of sending them through random courses… See the full description on the dataset page: https://huggingface.co/datasets/amine-khelif/ai-ml-foundations-book-collection.mhlc-training-gemma4-gemma4_e4b_it_think_off_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Gemma 4 E4B it think off hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-gemma4-gemma4_e4b_it_think_off_hard_mixed_sources_120k.darooyab_qa
Dataset Description:
darooyab_qa is a Persian drug question-answering dataset extracted from Darooyab materials.
Using the LLama3 model, the scraped content of each drug page transformed into many questions and corresponding answers.
Load the dataset:
To load the dataset, install the library datasets with pip install datasets. Then,
from datasets import load_dataset
dataset = load_dataset("amirmmahdavikia/darooyab_qa")
IRAN-MADANI-LAWSimpleQA-verified-Hard-Qwen3-8B
SimpleQA Verified Hard for Qwen3-8B
Dataset Summary
This dataset contains the 866 questions that
Qwen/Qwen3-8B failed to answer correctly in up to eight attempts from the
official
google/simpleqa-verified
benchmark.
Each source question was scheduled for eight stochastic generations. As soon
as one generation was graded CORRECT, sampling stopped and the question was
excluded. Questions retained here therefore have pass@8 = 0 under the
model, prompt, sampling, and… See the full description on the dataset page: https://huggingface.co/datasets/AmirMohseni/SimpleQA-verified-Hard-Qwen3-8B.pytorch-forum-topics-complete-v2
PyTorch Forum Topics Dataset
This dataset contains topic metadata scraped from the PyTorch Community Forum. It includes comprehensive information about forum topics that can be used for various NLP tasks related to PyTorch and deep learning discussions.
Dataset Structure
Each record in the dataset contains the following fields:
id: Unique topic identifier
title: Topic title
slug: URL-friendly version of the title
posts_count: Number of posts in the topic
reply_count:… See the full description on the dataset page: https://huggingface.co/datasets/AmitPrakash/pytorch-forum-topics-complete-v2.LLM_Response_Eval
LLM Response Evaluation Dataset
This dataset contains a collection of responses generated by three large language models (LLMs): GPT-4o, Gemini 1.5 Pro, and Llama 3.1 405B. The responses are to a series of questions aimed at evaluating the models' problem-solving abilities using Polya's problem-solving technique, as described in the book "How to Solve It" by George Polya.
Dataset Overview
Questions: The dataset includes a set of questions designed to evaluate the… See the full description on the dataset page: https://huggingface.co/datasets/AmirMohseni/LLM_Response_Eval.urdu-emergency-calls
Urdu Emergency Call Conversations Dataset (Pakistan)
Overview
This dataset contains 5,000 curated Urdu emergency call conversation samples from the Pakistan region, designed to support training and evaluation of Urdu Large Language Models (LLMs) for emergency response, command centers, and interpreter-style systems.
The conversations simulate real-world emergency scenarios such as:
Floods
Medical emergencies
Accidents
Crimes
Natural disasters
Public safety incidents
The… See the full description on the dataset page: https://huggingface.co/datasets/hamza-amin/urdu-emergency-calls.marathi-orca-v05
Dataset card for Marathi OpenOrca
Translated subset of Open-Orca/1million-gpt-4 to marathi language.
testmy-distiset-8cdf8ea5
Dataset Card for my-distiset-8cdf8ea5
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/AMIREBADY/my-distiset-8cdf8ea5/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/AMIREBADY/my-distiset-8cdf8ea5.PsychoLexQA
PsychoLexQA: A Bilingual Psychological Instructional Dataset
PsychoLexQA is a meticulously crafted dataset designed to enhance the performance of Large Language Models (LLMs) in the field of psychology. As part of the research paper titled "PsychoLex: Unveiling the Psychological Mind of Large Language Models", this dataset provides a rich bilingual resource in both Persian and English, tailored for complex psychological scenarios.
Dataset Overview
PsychoLexQA… See the full description on the dataset page: https://huggingface.co/datasets/aminabbasi/PsychoLexQA.my-distiset-4bf027a8
Dataset Card for my-distiset-4bf027a8
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/AMIREBADY/my-distiset-4bf027a8/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/AMIREBADY/my-distiset-4bf027a8.PsycholexEval
PsychoLexEval: A Bilingual Multiple-Choice Question Dataset for Psychology
PsychoLexEval is a meticulously curated dataset designed to evaluate the performance of Large Language Models (LLMs) in psychological contexts. As part of the research paper titled "PsychoLex: Unveiling the Psychological Mind of Large Language Models", this dataset provides a comprehensive bilingual resource in both Persian and English, aimed at assessing LLMs' comprehension and decision-making capabilities… See the full description on the dataset page: https://huggingface.co/datasets/aminabbasi/PsycholexEval.AfriQuAD
AfriQuAD (Multilingual)
This dataset spans across different African languages and also touches on cross-lingual aspects. It can be best used in Question Answering tasks.
from datasets import load_datasets
#load the entire dataset
dataset = load_dataset("amidblue/AfriQuAD")
FinMR
Financial Multimodal Mathematical Reasoning QA Dataset💰
[🔗Github] [📖 ArXiv Paper(not publish yet)]
💻Data Usage
from datasets import load_dataset
dataset = load_dataset("aminous1/FinMR", cache_dir="/your/custom/path")
👋😊✨Dataset Description
FinQA is a dataset designed for financial reasoning and question answering. It includes questions,
financial contexts, and corresponding answers. The dataset contains both textual and visual data, with visual data… See the full description on the dataset page: https://huggingface.co/datasets/aminous1/FinMR.
