CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mikex86 /stackoverflow-posts StackOverflow Posts Markdown Dataset Summary This dataset contains all posts submitted to StackOverflow before the 14th of June 2023 formatted as Markdown text. The dataset contains ~60 Million posts, totaling ~35GB in size and ~65 billion characters of text. The data is sourced from Internet Archive StackExchange Data Dump. Dataset Structure Each record corresponds to one post of a particular type. Original ordering from the data dump is not exactly preserved… See the full description on the dataset page: https://huggingface.co/datasets/mikex86/stackoverflow-posts.tabularquestion-answering10M<n<100M63 likes15k downloads3y agoHugging Face02HuggingFaceH4 /stack-exchange-preferences Dataset Card for H4 Stack Exchange Preferences Dataset Dataset Summary This dataset contains questions and answers from the Stack Overflow Data Dump for the purpose of preference model training. Importantly, the questions have been filtered to fit the following criteria for preference models (following closely from Askell et al. 2021): have >=2 answers. This data could also be used for instruction fine-tuning and language model training. The questions are grouped with… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/stack-exchange-preferences.textquestion-answering10M<n<100M135 likes9.6k downloads4y agoHugging Face03math-ai /StackMathQA StackMathQA StackMathQA: A Curated Collection of 2 Million Mathematical Questions and Answers Sourced from Stack Exchange StackMathQA is a meticulously curated collection of 2 million mathematical questions and answers, sourced from various Stack Exchange sites. This repository is designed to serve as a comprehensive resource for researchers, educators, and enthusiasts in the field of mathematics and AI research. Configs configs: - config_name: stackmathqa1600k… See the full description on the dataset page: https://huggingface.co/datasets/math-ai/StackMathQA.texttext-generation1M<n<10M104 likes1.6k downloads10mo agoHugging Face04flax-sentence-embeddings /stackexchange_titlebody_best_and_down_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering100K<n<1M12 likes1.5k downloads4y agoHugging Face05AdhyanshVerma /stack-2021-12-01 ReasonStack-Prime A highly normalized, streaming-optimized Stack Exchange corpus engineered for LLM reasoning and instruction tuning. CC BY-SA 4.0 ~1M Rows 176 Parquet Shards 21 SE Sites 1. Executive Summary ReasonStack-Prime is a large-scale, meticulously curated text dataset derived from the official Archive.org Stack Exchange data dump (Version 2021-12-07). Unlike raw XML dumps or poorly cleaned JSON exports, this dataset… See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/stack-2021-12-01.tabularquestion-answering100K<n<1M0 likes1.3k downloads3mo agoHugging Face06flax-sentence-embeddings /stackexchange_title_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M8 likes1.1k downloads4y agoHugging Face07flax-sentence-embeddings /stackexchange_titlebody_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M9 likes737 downloads4y agoHugging Face08farida5gaber /stackoverflow-posts StackOverflow Posts Markdown Dataset Summary This dataset contains all posts submitted to StackOverflow before the 14th of June 2023 formatted as Markdown text. The dataset contains ~60 Million posts, totaling ~35GB in size and ~65 billion characters of text. The data is sourced from Internet Archive StackExchange Data Dump. Dataset Structure Each record corresponds to one post of a particular type. Original ordering from the data dump is not exactly preserved… See the full description on the dataset page: https://huggingface.co/datasets/farida5gaber/stackoverflow-posts.tabularquestion-answering10M<n<100M0 likes670 downloads5mo agoHugging Face09koutch /stackoverflow_python Dataset Card for "stackoverflow_python" Dataset Summary This dataset comes originally from kaggle. It was originally split into three tables (CSV files) (Questions, Answers, and Tags) now merged into a single table. Each row corresponds to a pair (question-answer) and their associated tags. The dataset contains all questions asked between August 2, 2008 and Ocotober 19, 2016. Supported Tasks and Leaderboards This might be useful for open-domain… See the full description on the dataset page: https://huggingface.co/datasets/koutch/stackoverflow_python.tabularquestion-answering100K<n<1M33 likes416 downloads3y agoHugging Face10habedi /stack-exchange-dataset Overview This dataset consists of three TSV files, namely: cs.tsv, ds.tsv, and p.tsv. Each file includes the data for the questions asked on a Stack Exchange (SE) question-answering community, from the creation of the community until May 2021. cs.tsv --> Computer Science SE ds.csv --> Data Science SE p.csv --> Political Science SE File Structure Each file has the following columns: id: the question id title: the title of the question body: the body or text of the… See the full description on the dataset page: https://huggingface.co/datasets/habedi/stack-exchange-dataset.tabulartext-classification10K<n<100K11 likes380 downloads7mo agoHugging Face11flax-sentence-embeddings /stackexchange_math_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M20 likes343 downloads4y agoHugging Face12Omarrran /StackPulse_778K_QnA_Code_dataset 💻 StackOverflow-778K: Multi-Year Developer Q&A Dataset Dataset Summary A large-scale Stack Overflow question dataset containing 778,929 unique questions sampled across 7 years (2015–2022). Each question includes the raw HTML body, plain-text version, tags, score, view count, answer count, and a rich set of derived features for immediate ML use. Collected across 8 sampling runs on Feb 27 2026, deduplicated to 778,929 unique questions with only 2 duplicates removed.… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/StackPulse_778K_QnA_Code_dataset.tabulartext-classification1M<n<10M0 likes234 downloads5mo agoHugging Face13agicorp /StackMathQA StackMathQA StackMathQA is a meticulously curated collection of 2 million mathematical questions and answers, sourced from various Stack Exchange sites. This repository is designed to serve as a comprehensive resource for researchers, educators, and enthusiasts in the field of mathematics and AI research. Configs configs: - config_name: stackmathqa1600k data_files: data/stackmathqa1600k/all.jsonl default: true - config_name: stackmathqa800k data_files:… See the full description on the dataset page: https://huggingface.co/datasets/agicorp/StackMathQA.texttext-generation1M<n<10M3 likes192 downloads3y agoHugging Face14KonradSzafer /stackoverflow_linux Dataset Card for "stackoverflow_linux" Dataset information: Source: Stack Overflow Category: Linux Number of samples: 300 Train/Test split: 270/30 Quality: Data come from the top 1k most upvoted questions Additional Information License All Stack Overflow user contributions are licensed under CC-BY-SA 3.0 with attribution required. More Information needed textquestion-answeringn<1K8 likes187 downloads4y agoHugging Face15mirzaei2114 /stackoverflowVQA-filteredimagevisual-question-answering100K<n<1M3 likes179 downloads3y agoHugging Face16glopezas /math_stackexchange_qa Math StackExchange Curated (Parquet, CC BY-SA 4.0) This dataset is a curated collection of Math StackExchange (MSE) Q&A pairs packaged in Parquet format.Each sample contains a problem (title, question_body), its corresponding answer (answer_body), the original MSE tag string (tags), and a flag indicating whether the answer was accepted (accepted). This dataset includes content derived from the Math StackExchange public data dump (CC BY-SA 4.0, © Stack Exchange Inc.).This derived… See the full description on the dataset page: https://huggingface.co/datasets/glopezas/math_stackexchange_qa.textquestion-answering1M<n<10M0 likes164 downloads11mo agoHugging Face17BramVanroy /stackoverflow-chat-dutch Dataset Card for Stack Overflow Chat Dutch Dataset Summary This dataset contains 56,964 conversations between een AI assistant and a (fake) "Human" (generated) in Dutch, specifically in the domain of programming (Stack Overflow). They are translations of Baize's machine-generated answers to the Stack Overflow dataset. ☕ Want to help me out? Translating the data with the OpenAI API, and prompt testing, cost me 💸$133.60💸. If you like this dataset, please consider buying… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/stackoverflow-chat-dutch.textquestion-answering10K<n<100K2 likes155 downloads3y agoHugging Face18ymoslem /Law-StackExchange Law-StackExchange Dataset Details All StackExchange legal questions and their answers from the Law site, up to 14 August 2023. The repository includes a notebook for the process using the official StackExchange API. Citation @misc{Moslem2023-LawStackExchangeDataset, author = {Moslem, Yasmin}, title = {Law-StackExchange Dataset}, year = 2023, url = {https://huggingface.co/datasets/ymoslem/Law-StackExchange}, doi =… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/Law-StackExchange.tabularquestion-answering10K<n<100K32 likes154 downloads1y agoHugging Face19mcipriano /stackoverflow-kubernetes-questionsThe purpose of this dataset is to provide the opportunity to perform any training, fine-tuning, etc. for any Language Model. In the 'data' folder, you will find the dataset in Parquet format, which is one of the formats used for these processes. In case it may be useful for other purposes, I have also included the dataset in CSV format. All data in this dataset were retrieved from the Stack Exchange network using the Stack Exchange Data explorer tool… See the full description on the dataset page: https://huggingface.co/datasets/mcipriano/stackoverflow-kubernetes-questions.textquestion-answering10K<n<100K30 likes125 downloads3y agoHugging Face20ymoslem /MedicalSciences-StackExchangeAll StackExchange questions and their answers from the Medical Sciences site, up to 14 August 2023. The repository includes a notebook for the process using the official StackExchange API. tabularquestion-answering1K<n<10K11 likes119 downloads3y agoHugging Face21kispeterzsm-szte /stackexchangeThis dataset is based entirely on HuggingFaceH4/stack-exchange-preferences, but it has been restructured. All HTML tags have been cleaned out, and the answers column has been turned into the answer column, so instead of answers being stored in JSON format there is now a row for each answer. Furthermore there is a separate file for every forum instead of a single file. textquestion-answering10M<n<100M2 likes114 downloads1y agoHugging Face22mirzaei2114 /stackoverflowVQA Dataset Card for "stackoverflowVQA" More Information needed tabularvisual-question-answering1M<n<10M5 likes88 downloads3y agoHugging Face23juliensimon /stackexchange-space-qa Stack Exchange Space Q&A Credit: NASA/DOE/Fermi LAT Collaboration Part of a dataset collection on Hugging Face. Dataset description This dataset is a clean, tabular Q&A corpus of space and astronomy knowledge, derived from two Stack Exchange community Q&A sites: Astronomy Stack Exchange (astronomy.stackexchange.com) and Space Exploration Stack Exchange (space.stackexchange.com). Each row is one question paired with its best answer — either the question's… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/stackexchange-space-qa.tabularquestion-answering10K<n<100K0 likes84 downloads3d agoHugging Face24mirzaei2114 /stackoverflowVQA-filtered-small Dataset Card for "stackoverflowVQA-filtered-small" More Information needed imagevisual-question-answering10K<n<100K4 likes71 downloads3y agoHugging Face25Mxode /StackOverflow-QA-C-Language-40kThis is a collection of ~40k QA's in C Language from StackOverflow. The data has been initially cleaned, and each response is with Accepted Answer. All data is <1000 in length. The questions and answers were organized into a one-line format. A sample format is shown below: { "question": "```\nFILE* file = fopen(some file)\n\npcap_t* pd = pcap_fopen_offline(file)\n\npcap_close(pd)\n\nfclose(file)\n```\n\nThis code occurs double free error.\n\nCould you explain about this happening?\n\nMy… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/StackOverflow-QA-C-Language-40k.textquestion-answering10K<n<100K5 likes62 downloads1y agoHugging Face26Omarrran /stackpulse_qa_output 🧩 StackPulse-QA: Instruction-Tuning Q&A Pairs from Stack Overflow Dataset Summary Instruction-tuning Q&A dataset built from Omarrran/StackPulse_778K_QnA_Code_dataset by joining question IDs with BigQuery bigquery-public-data.stackoverflow.posts_answers on accepted_answer_id. Each sample consists of: input_text_instruct — A question (title + body) prefixed with an instruction output_text — The accepted answer from Stack Overflow Format mirrors the… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/stackpulse_qa_output.tabularquestion-answering100K<n<1M0 likes60 downloads5mo agoHugging Face27p1atdev /japanese-stackexchange japanese-stackexchange 英語による日本語に関する質問ができる Japanese Stack Exchange のデータダンプ をもとにデータを加工し、質問文と回答文のペアになるように調整した QA データセット。 日本語翻訳された StackExchange ではないです。 データ構造 投稿本文は html2text を使ってマークダウン化されています。その際、 コードブロックは ``` で囲まれるように変更されています。 画像 URL に base64 エンコードされた画像が含まれる場合、 [unk] に置き換えています。 default サブセット id: 質問投稿の ID question: 質問投稿 answers: 質問に対する回答投稿のリスト accepted_answer_id: 質問者に選ばれた回答のID。null の可能性がある popular_answer_id: もっともスコアが高かった回答のID。null の可能性がある simple サブセット… See the full description on the dataset page: https://huggingface.co/datasets/p1atdev/japanese-stackexchange.tabulartext-generation10K<n<100K3 likes55 downloads3y agoHugging Face28FurkanNar /StackMathematics-XL Mathematics Stack Exchange Q&A Dataset A curated dataset of question-and-answer pairs harvested from Mathematics Stack Exchange (math.stackexchange.com). Each entry contains the question, the top answer (if available), metadata (scores, tags, URL), and a combined Q/A text useful for training and evaluation of question-answering and tutoring models. Dataset summary Source: Mathematics Stack Exchange (via the official Stack Exchange API) Content: High-quality Q&A… See the full description on the dataset page: https://huggingface.co/datasets/FurkanNar/StackMathematics-XL.tabularquestion-answeringn<1K1 likes51 downloads23d agoHugging Face29nandhakumarms /qualc-stackexchange-en QualC Stack Exchange English A cleaned and validated English Stack Exchange corpus prepared for large language model (LLM) pretraining. This dataset is part of the QualC Foundation Corpus, an open collection of high-quality datasets intended for training multilingual foundation models. Dataset Information Language: English Records: 999,832 Format: Hugging Face Dataset Schema: Flat Text License: CC BY-SA 4.0 (inherits the original Stack Exchange content license)… See the full description on the dataset page: https://huggingface.co/datasets/nandhakumarms/qualc-stackexchange-en.texttext-generation100K<n<1M0 likes51 downloads2mo agoHugging Face30p1atdev /ja-stackoverflow ja-stackoverflow 日本語版 Stack Overflow の スタック・オーバーフロー のデータダンプ をもとにデータを加工し、質問文と回答文のペアになるように調整した QA データセット。 データ構造 投稿本文は html2text を使ってマークダウン化されています。その際、 コードブロックは ``` で囲まれるように変更されています。 画像 URL に base64 エンコードされた画像が含まれる場合、 [unk] に置き換えています。 default サブセット id: 質問投稿の ID question: 質問投稿 answers: 質問に対する回答投稿のリスト accepted_answer_id: 質問者に選ばれた回答のID。null の可能性がある popular_answer_id: もっともスコアが高かった回答のID。null の可能性がある simple サブセット default サブセットから、 question と answers… See the full description on the dataset page: https://huggingface.co/datasets/p1atdev/ja-stackoverflow.tabulartext-generation10K<n<100K8 likes47 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.