CoolFace
26 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mikex86 /stackoverflow-posts StackOverflow Posts Markdown Dataset Summary This dataset contains all posts submitted to StackOverflow before the 14th of June 2023 formatted as Markdown text. The dataset contains ~60 Million posts, totaling ~35GB in size and ~65 billion characters of text. The data is sourced from Internet Archive StackExchange Data Dump. Dataset Structure Each record corresponds to one post of a particular type. Original ordering from the data dump is not exactly preserved… See the full description on the dataset page: https://huggingface.co/datasets/mikex86/stackoverflow-posts.tabularquestion-answering10M<n<100M63 likes15k downloads3y agoHugging Face02bigcode /stackoverflow-clean Dataset Card for "stackoverflow-clean" More Information needed tabular10M<n<100M8 likes1.5k downloads3y agoHugging Face03farida5gaber /stackoverflow-posts StackOverflow Posts Markdown Dataset Summary This dataset contains all posts submitted to StackOverflow before the 14th of June 2023 formatted as Markdown text. The dataset contains ~60 Million posts, totaling ~35GB in size and ~65 billion characters of text. The data is sourced from Internet Archive StackExchange Data Dump. Dataset Structure Each record corresponds to one post of a particular type. Original ordering from the data dump is not exactly preserved… See the full description on the dataset page: https://huggingface.co/datasets/farida5gaber/stackoverflow-posts.tabularquestion-answering10M<n<100M0 likes670 downloads5mo agoHugging Face04koutch /stackoverflow_python Dataset Card for "stackoverflow_python" Dataset Summary This dataset comes originally from kaggle. It was originally split into three tables (CSV files) (Questions, Answers, and Tags) now merged into a single table. Each row corresponds to a pair (question-answer) and their associated tags. The dataset contains all questions asked between August 2, 2008 and Ocotober 19, 2016. Supported Tasks and Leaderboards This might be useful for open-domain… See the full description on the dataset page: https://huggingface.co/datasets/koutch/stackoverflow_python.tabularquestion-answering100K<n<1M33 likes416 downloads3y agoHugging Face05raymondzmc /stackoverflow_Llama-3.1-8B-Instruct_vocab_2000_lasttabular10K<n<100K0 likes189 downloads9mo agoHugging Face06BEE-spoke-data /stackoverflow-questions-long stackoverflow questions for text classification: 'long' This is pacovaldez/stackoverflow-questions filtered for 1024 GPT2 tokens or more in title + body https://huggingface.co/datasets/pacovaldez/stackoverflow-questions tabulartext-classification100K<n<1M1 likes147 downloads9mo agoHugging Face07reubenjohn /stackoverflow-unified-text-open-status-classification Dataset Card for "stackoverflow-unified-text-open-status-classification" More Information needed tabular1M<n<10M0 likes133 downloads4y agoHugging Face08mirzaei2114 /stackoverflowVQA Dataset Card for "stackoverflowVQA" More Information needed tabularvisual-question-answering1M<n<10M5 likes88 downloads3y agoHugging Face09raymondzmc /stackoverflow_ERNIE-4.5-0.3B-PT_vocab_4000_lasttabular10K<n<100K0 likes57 downloads9mo agoHugging Face10reubenjohn /stackoverflow-unified-text-open-status-classification-sample Dataset Card for "stackoverflow-open-status-classification" More Information needed tabular100K<n<1M1 likes51 downloads4y agoHugging Face11p1atdev /ja-stackoverflow ja-stackoverflow 日本語版 Stack Overflow の スタック・オーバーフロー のデータダンプ をもとにデータを加工し、質問文と回答文のペアになるように調整した QA データセット。 データ構造 投稿本文は html2text を使ってマークダウン化されています。その際、 コードブロックは ``` で囲まれるように変更されています。 画像 URL に base64 エンコードされた画像が含まれる場合、 [unk] に置き換えています。 default サブセット id: 質問投稿の ID question: 質問投稿 answers: 質問に対する回答投稿のリスト accepted_answer_id: 質問者に選ばれた回答のID。null の可能性がある popular_answer_id: もっともスコアが高かった回答のID。null の可能性がある simple サブセット default サブセットから、 question と answers… See the full description on the dataset page: https://huggingface.co/datasets/p1atdev/ja-stackoverflow.tabulartext-generation10K<n<100K8 likes47 downloads3y agoHugging Face12eshangj /stackoverflow_q_and_a_sample Description GitHub repository: https://github.com/EshanJayasundara/Stackoverflow-Python-Q-and-A-Extractor. GitHub repository contains the automated workflow for extracting the question and answer pairs from Stackoverflow. This dataset contains the question-answer pairs extracted from Stackoverflow using Stack Exchange API v2.3 and used following endpoints, /answers/{ids} GET /questions GET From 2020 January 1 to Today 1. Dataset description, Contains only python… See the full description on the dataset page: https://huggingface.co/datasets/eshangj/stackoverflow_q_and_a_sample.tabularquestion-answering10K<n<100K1 likes47 downloads1y agoHugging Face13raymondzmc /stackoverflow_Llama-3.2-1B-Instruct_vocab_2000_lasttabular10K<n<100K0 likes40 downloads9mo agoHugging Face14raymondzmc /stackoverflow_ERNIE-4.5-0.3B-PT_vocab_2000_lasttabular10K<n<100K0 likes35 downloads9mo agoHugging Face15konsman /stackoverflow-posts-mobile-development-tagtabularn<1K0 likes30 downloads2y agoHugging Face16raymondzmc /stackoverflow_Qwen3.5-0.8B_vocab_2000_lasttabular10K<n<100K0 likes22 downloads6mo agoHugging Face17tennisb /stackoverflow-ai-solvedUpdated 1/19/2024: More than doubled the number of answers to 1,469,201, with a higher percent of 4o-mini and gemini 1.5 pro. A dataset which comprises of ai answers to 1,469,201 stackoverflow questions relating to Python. The questions were extracted from this dataset. All responses are directly python, no codeblocks or anything. A total of 545,697,115 input and 253,110,685 output o200k_base tokens (gpt-4o/4o-mini). Model Value gemini-1.5-pro-002 442,261 gpt-4o-mini-2024-07-18… See the full description on the dataset page: https://huggingface.co/datasets/tennisb/stackoverflow-ai-solved.tabulartext-generation1M<n<10M5 likes20 downloads2y agoHugging Face18Yashwani /stackoverflow-commandline-inst Dataset Card for "stackoverflow-commandline-inst" More Information needed tabular1K<n<10K1 likes20 downloads1y agoHugging Face19Rohan1103 /stack-overflowdumptabular100K<n<1M0 likes20 downloads6mo agoHugging Face20ccde /stackoverflow-md-pythontabular10K<n<100K1 likes19 downloads1y agoHugging Face21searchsim /cognitive-traces-stackoverflow Cognitive Traces — Stack Overflow Dataset Description This dataset contains cognitive trace annotations for the Stack Overflow dataset, produced by the multi-agent annotation framework described in: Beyond the Click: A Framework for Inferring Cognitive Traces in Search Saber Zerhoudi, Michael Granitzer. ECIR 2026. Each user event (question, answer, comment, edit, vote) is annotated with a cognitive label from Information Foraging Theory (IFT), along with the full… See the full description on the dataset page: https://huggingface.co/datasets/searchsim/cognitive-traces-stackoverflow.tabulartext-classification100K<n<1M0 likes16 downloads6mo agoHugging Face22raymondzmc /stackoverflow_Phi-3-mini-128k-instruct_vocab_2000_lasttabular10K<n<100K0 likes16 downloads6mo agoHugging Face23ayman56 /stackoverflow_qa_python_Preprocessedtabular100K<n<1M3 likes12 downloads2y agoHugging Face24pacovaldez /predicted-stackoverflow Dataset Card for "predicted-stackoverflow" More Information needed tabularn<1K1 likes11 downloads4y agoHugging Face25vm2825 /stackoverflow-bridge-reasoningtabularn<1K0 likes7 downloads11mo agoHugging Face26SatyakiMitra /stackoverflow-ai-datatabular1M<n<10M0 likes3 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.