datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
StackPulse_778K_QnA_Code_dataset
💻 StackOverflow-778K: Multi-Year Developer Q&A Dataset
Dataset Summary
A large-scale Stack Overflow question dataset containing 778,929 unique
questions sampled across 7 years (2015–2022). Each question includes the
raw HTML body, plain-text version, tags, score, view count, answer count, and
a rich set of derived features for immediate ML use.
Collected across 8 sampling runs on Feb 27 2026, deduplicated to
778,929 unique questions with only 2 duplicates removed.… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/StackPulse_778K_QnA_Code_dataset.pbgdpl-vn-legal-qna
pbgdpl.gov.vn — Vietnamese Legal Q&A · Hỏi đáp pháp luật
🇻🇳 Tóm tắt. Bản thu thập đầy đủ chuyên mục Hỏi đáp pháp luật
của Cổng thông tin điện tử Phổ biến giáo dục pháp luật
— cổng giáo dục pháp luật công khai do Bộ Tư pháp vận hành. Mỗi
dòng là một cặp câu hỏi của công dân (Q) và trả lời chính
thức (A), kèm chú thích nguồn, lĩnh vực pháp lý, ngày gửi, và đường
dẫn về trang gốc.
🇬🇧 Summary. A complete crawl of the public Hỏi đáp pháp luật
("Legal Q&A") section of
pbgdpl.gov.vn —… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/pbgdpl-vn-legal-qna.swe-atlas-qna-dev41-mcode-m3
SWE-Atlas-QnA dev-41 — mcode + MiniMax-M3 trajectories
Complete run logs for one sweep of the 41-task SWE-Atlas-QnA dev slice, with the
LLM-rubric judge output for every task. Published as a reference trajectory set for
explore/comprehension evaluation.
41 of 41 tasks scored, mean agg_score 0.815 (median 0.875, min 0.222); mean
reward 0.244; 10 tasks (24 %) at the 1.0 ceiling.
suite
ai-solution-finetune/swe-atlas-qna-dev-50 minus 9 musl tasks = 41
agent
mcode… See the full description on the dataset page: https://huggingface.co/datasets/miaomiao64/swe-atlas-qna-dev41-mcode-m3.object_leftqna_sbertcamera_leftwiki-qna
