datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
StackPulse_778K_QnA_Code_dataset
💻 StackOverflow-778K: Multi-Year Developer Q&A Dataset
Dataset Summary
A large-scale Stack Overflow question dataset containing 778,929 unique
questions sampled across 7 years (2015–2022). Each question includes the
raw HTML body, plain-text version, tags, score, view count, answer count, and
a rich set of derived features for immediate ML use.
Collected across 8 sampling runs on Feb 27 2026, deduplicated to
778,929 unique questions with only 2 duplicates removed.… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/StackPulse_778K_QnA_Code_dataset.pbgdpl-vn-legal-qna
pbgdpl.gov.vn — Vietnamese Legal Q&A · Hỏi đáp pháp luật
🇻🇳 Tóm tắt. Bản thu thập đầy đủ chuyên mục Hỏi đáp pháp luật
của Cổng thông tin điện tử Phổ biến giáo dục pháp luật
— cổng giáo dục pháp luật công khai do Bộ Tư pháp vận hành. Mỗi
dòng là một cặp câu hỏi của công dân (Q) và trả lời chính
thức (A), kèm chú thích nguồn, lĩnh vực pháp lý, ngày gửi, và đường
dẫn về trang gốc.
🇬🇧 Summary. A complete crawl of the public Hỏi đáp pháp luật
("Legal Q&A") section of
pbgdpl.gov.vn —… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/pbgdpl-vn-legal-qna.jee-exam-qnamagicmotion
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
Quanhao Li*, Zhen Xing*, Rui Wang, Hui Zhang, Qi Dai, and Zuxuan Wu
* equal contribution
💡 Abstract
Recent advances in video generation have led to remarkable improvements in visual quality and temporal coherence. Upon this, trajectory-controllable video generation has emerged to enable precise object motion control through explicitly defined spatial paths.
However, existing methods… See the full description on the dataset page: https://huggingface.co/datasets/Qnancy/magicmotion.myanmar_qna_dataset
Myanmar QnA Dataset v7
Language: Burmese (Myanmar)Total Entries: 22,783 QnA pairsTotal Sentences: ~ 466,330(Counted using the Myanmar sentence-ending symbol "။")License: CC0 1.0 (Public Domain)
Description
This dataset contains Myanmar-language question-answer pairs (QnA) generated with the assistance of ChatGPT-5 for question crafting with English and Gemini 3.0 Pro for Myanmar QnA generation. It is intended for research, AI training, and educational purposes.
Each entry… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_qna_dataset.qna-ocp-4.15
Data Card
Dataset created by: William Caban
License: Apache 2.0
OpenShift 4.15 Knowledge cutoff date: April 12, 2024
Updates:
July 3rd, 2024 - Update dataset with full dataset
Method
The Q&A corpus was generated using the following methodology:
Generated 5 Q&A pairs for each page on OpenShift (OCP) 4.15 PDFs with lengths greater than 1500 characters. The length was chosen to remove the title page and pages without much content.
The Mistral-7B-Instruct-v0.2 was used… See the full description on the dataset page: https://huggingface.co/datasets/boricua/qna-ocp-4.15.swe-atlas-qna-dev41-mcode-m3
SWE-Atlas-QnA dev-41 — mcode + MiniMax-M3 trajectories
Complete run logs for one sweep of the 41-task SWE-Atlas-QnA dev slice, with the
LLM-rubric judge output for every task. Published as a reference trajectory set for
explore/comprehension evaluation.
41 of 41 tasks scored, mean agg_score 0.815 (median 0.875, min 0.222); mean
reward 0.244; 10 tasks (24 %) at the 1.0 ceiling.
suite
ai-solution-finetune/swe-atlas-qna-dev-50 minus 9 musl tasks = 41
agent
mcode… See the full description on the dataset page: https://huggingface.co/datasets/miaomiao64/swe-atlas-qna-dev41-mcode-m3.Islamic_Finance_QnA_eval
Islamic Finance Q&A Evaluation Dataset
Validation and test splits for evaluating models on Islamic Finance Q&A.
Dataset Structure
Format: Simple prompt-answer pairs
Validation: ~203 examples (10%)
Test: ~203 examples (10%)
Language: Arabic
Domain: Islamic finance and Sharia-compliant banking
Fields
id: Unique identifier
prompt: The question prompt
question: Original question text
answer: Ground truth answer
topic: Topic category
split:… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/Islamic_Finance_QnA_eval.HAVIT_QnA_v2network-QnA-datasetIslamic_Finance_QnA_train
Islamic Finance Q&A Training Dataset
Training split of the Islamic Finance Q&A dataset in conversational format.
Dataset Structure
Format: Conversational (human-agent pairs)
Size: ~1,624 training examples (80% of total)
Language: Arabic
Domain: Islamic finance and Sharia-compliant banking
Usage
from datasets import load_dataset
dataset = load_dataset("SahmBenchmark/Islamic_Finance_QnA_train")
train_data = dataset['train']
# Example
example =… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/Islamic_Finance_QnA_train.object_left38_blessings_mingalar_tayartaw_qna
၃၈ ဖြာ မင်္ဂလာ တရားတော် [The 38 Blessings of Maha Mangala Sutta Q&A]
Created by: freococoLicense: CC0 1.0 Universal (Public Domain)
Summary
This dataset consists of 1,111 high-quality Questions and Answers centered on the 38 Blessings (၃၈ ဖြာ မင်္ဂလာ - Mangala). The answers are written in a Burmese Spoken Style to ensure natural AI conversational flow, while maintaining deep Philosophical, Psychological, Social, and Spiritual perspectives.
This repository is a specialized… See the full description on the dataset page: https://huggingface.co/datasets/freococo/38_blessings_mingalar_tayartaw_qna.indic_qna_v1myanmar_qna_dataset
Myanmar QnA Dataset v7
Language: Burmese (Myanmar)Total Entries: 22,783 QnA pairsTotal Sentences: ~ 466,330(Counted using the Myanmar sentence-ending symbol "။")License: CC0 1.0 (Public Domain)
Description
This dataset contains Myanmar-language question-answer pairs (QnA) generated with the assistance of ChatGPT-5 for question crafting with English and Gemini 3.0 Pro for Myanmar QnA generation. It is intended for research, AI training, and educational purposes.
Each entry… See the full description on the dataset page: https://huggingface.co/datasets/EISETWYNE/myanmar_qna_dataset.qna
QnA dataset
This dataset was created by generating answers to questions by various neural networks.
Information
This part is generated by script
Param
Value
Total tokens
12271743
Question tokens
243559
Answer tokens
11735372
Used models
deepseek-v4-flash - 158B
google/gemma-3-1b - 1b
google/gemma-4-e2b - 2.3b
gemini-3-flash - IDK
qwen/qwen3-4b-2507 - 4b
lfm2.5-1.2b-instruct - 1.2b
Air_Compressor_Mannual_QNAqna_sbertKatzBot_QnA_Testjee-exam-qnacodebase-qnasynthetic_data_qna_fulltext_conditioned_L3.3_70Bqna_train
Dataset Card for "qna_train"
More Information needed
data_cds_QnAcamera_leftwiki-qna
