datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stackoverflow-posts
StackOverflow Posts Markdown
Dataset Summary
This dataset contains all posts submitted to StackOverflow before the 14th of June 2023 formatted as Markdown text.
The dataset contains ~60 Million posts, totaling ~35GB in size and ~65 billion characters of text.
The data is sourced from Internet Archive StackExchange Data Dump.
Dataset Structure
Each record corresponds to one post of a particular type.
Original ordering from the data dump is not exactly preserved… See the full description on the dataset page: https://huggingface.co/datasets/mikex86/stackoverflow-posts.stackoverflow-clean
Dataset Card for "stackoverflow-clean"
More Information needed
stackoverflow-posts
StackOverflow Posts Markdown
Dataset Summary
This dataset contains all posts submitted to StackOverflow before the 14th of June 2023 formatted as Markdown text.
The dataset contains ~60 Million posts, totaling ~35GB in size and ~65 billion characters of text.
The data is sourced from Internet Archive StackExchange Data Dump.
Dataset Structure
Each record corresponds to one post of a particular type.
Original ordering from the data dump is not exactly preserved… See the full description on the dataset page: https://huggingface.co/datasets/farida5gaber/stackoverflow-posts.stackoverflow_python
Dataset Card for "stackoverflow_python"
Dataset Summary
This dataset comes originally from kaggle.
It was originally split into three tables (CSV files) (Questions, Answers, and Tags)
now merged into a single table. Each row corresponds to a pair (question-answer) and
their associated tags.
The dataset contains all questions asked between August 2, 2008 and Ocotober 19, 2016.
Supported Tasks and Leaderboards
This might be useful for open-domain… See the full description on the dataset page: https://huggingface.co/datasets/koutch/stackoverflow_python.stackoverflow_Llama-3.1-8B-Instruct_vocab_2000_laststackoverflow-questions-long
stackoverflow questions for text classification: 'long'
This is pacovaldez/stackoverflow-questions filtered for 1024 GPT2 tokens or more in title + body
https://huggingface.co/datasets/pacovaldez/stackoverflow-questions
stackoverflow-unified-text-open-status-classification
Dataset Card for "stackoverflow-unified-text-open-status-classification"
More Information needed
stackoverflowVQA
Dataset Card for "stackoverflowVQA"
More Information needed
stackoverflow_ERNIE-4.5-0.3B-PT_vocab_4000_laststackoverflow-unified-text-open-status-classification-sample
Dataset Card for "stackoverflow-open-status-classification"
More Information needed
ja-stackoverflow
ja-stackoverflow
日本語版 Stack Overflow の スタック・オーバーフロー のデータダンプ をもとにデータを加工し、質問文と回答文のペアになるように調整した QA データセット。
データ構造
投稿本文は html2text を使ってマークダウン化されています。その際、
コードブロックは ``` で囲まれるように変更されています。
画像 URL に base64 エンコードされた画像が含まれる場合、 [unk] に置き換えています。
default サブセット
id: 質問投稿の ID
question: 質問投稿
answers: 質問に対する回答投稿のリスト
accepted_answer_id: 質問者に選ばれた回答のID。null の可能性がある
popular_answer_id: もっともスコアが高かった回答のID。null の可能性がある
simple サブセット
default サブセットから、 question と answers… See the full description on the dataset page: https://huggingface.co/datasets/p1atdev/ja-stackoverflow.stackoverflow_q_and_a_sample
Description
GitHub repository: https://github.com/EshanJayasundara/Stackoverflow-Python-Q-and-A-Extractor.
GitHub repository contains the automated workflow for extracting the question and answer pairs from Stackoverflow.
This dataset contains the question-answer pairs extracted from Stackoverflow using Stack Exchange API v2.3 and used following endpoints,
/answers/{ids} GET
/questions GET
From 2020 January 1 to Today
1. Dataset description,
Contains only python… See the full description on the dataset page: https://huggingface.co/datasets/eshangj/stackoverflow_q_and_a_sample.stackoverflow_Llama-3.2-1B-Instruct_vocab_2000_laststackoverflow_ERNIE-4.5-0.3B-PT_vocab_2000_laststackoverflow-posts-mobile-development-tagstackoverflow_Qwen3.5-0.8B_vocab_2000_laststackoverflow-ai-solvedUpdated 1/19/2024: More than doubled the number of answers to 1,469,201, with a higher percent of 4o-mini and gemini 1.5 pro.
A dataset which comprises of ai answers to 1,469,201 stackoverflow questions relating to Python.
The questions were extracted from this dataset.
All responses are directly python, no codeblocks or anything.
A total of 545,697,115 input and 253,110,685 output o200k_base tokens (gpt-4o/4o-mini).
Model
Value
gemini-1.5-pro-002
442,261
gpt-4o-mini-2024-07-18… See the full description on the dataset page: https://huggingface.co/datasets/tennisb/stackoverflow-ai-solved.stackoverflow-commandline-inst
Dataset Card for "stackoverflow-commandline-inst"
More Information needed
stack-overflowdumpstackoverflow-md-pythoncognitive-traces-stackoverflow
Cognitive Traces — Stack Overflow
Dataset Description
This dataset contains cognitive trace annotations for the Stack Overflow dataset, produced by the multi-agent annotation framework described in:
Beyond the Click: A Framework for Inferring Cognitive Traces in Search
Saber Zerhoudi, Michael Granitzer. ECIR 2026.
Each user event (question, answer, comment, edit, vote) is annotated with a cognitive label from Information Foraging Theory (IFT), along with the full… See the full description on the dataset page: https://huggingface.co/datasets/searchsim/cognitive-traces-stackoverflow.stackoverflow_Phi-3-mini-128k-instruct_vocab_2000_laststackoverflow_qa_python_Preprocessedpredicted-stackoverflow
Dataset Card for "predicted-stackoverflow"
More Information needed
stackoverflow-bridge-reasoningstackoverflow-ai-data
