datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
autoresearch-solo-vs-forum
Solo vs forum: long-horizon coding-agent runs on 12 research-engineering tasks
1384 runs (196 solo, 1179 forum, 9 long-run chains), 11673 graded agent rows, 10478 transcripts, 285889 forum posts; 34370 files, 45.2 GB. Models: deepseek-v4.1-flash, glm-5.2, gpt-5.6-sol, qwen3.8-27b. Tasks: actlearn, adplace, borden, carleson, dabic, exploit, graph, kda, mega, moe, swinmlp, topopt.
What the experiment is
Each agent is a coding-agent CLI (Claude Code for Qwen3.8-27B /… See the full description on the dataset page: https://huggingface.co/datasets/junlinw/autoresearch-solo-vs-forum.forum3forum2aops_forum_filteredforum1Roleplay-Forums_2023-04Important noticeUpon reanalysis, it appears that spaces between adjacent HTML tags in these scrapes got mangled, a problem that was also present in the files of the Rentry where they first got uploaded; this occurred during an intermediate step in the conversion process from the original source files, most of which unfortunately do not exist anymore. So, the usefulness of the data will be diminished. Some scrapes without these issues have been provided on a different dataset page.… See the full description on the dataset page: https://huggingface.co/datasets/lemonilia/Roleplay-Forums_2023-04.roleplaying-forums-raw
Roleplaying forum scrapes (raw)
Here are mostly original/raw files for some of the roleplaying forums I scraped in the past (and some newly scraped ones), repacked as HTML strings + some metadata on a one-row-per-thread basis instead of a one-row-per-message basis, which should make them more convenient to handle.
Unlike the previously uploaded archive, they shouldn't have issues with spaces between adjacent HTML tags, as that occurred by mistake in an intermediate processing step… See the full description on the dataset page: https://huggingface.co/datasets/lemonilia/roleplaying-forums-raw.forums_pol_json_zstnumina-cot-aops-forum
Saravana-Polisetti/numina-cot-aops-forum
Source-specific slice of AI-MO/NuminaMath-CoT.
Source: aops_forum
Rows: 30192
Columns: source, problem, solution
Usage
from datasets import load_dataset
ds = load_dataset("Saravana-Polisetti/numina-cot-aops-forum", split="train")
print(ds[0])
forum-competition-math-training-pool
Forum competition mathematics training pool
Olympiad and contest mathematics from three public datasets, gathered at pinned revisions and
shipped twice over. sources/ holds each dataset the way its publisher ships it, in its own file
format with its own fields and nothing renamed, 287091 rows across three folders. pool/ holds the
union of those same datasets in one format, one JSON object per line, deduplicated by problem text
and reduced to 282140 rows, every row labelled with… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/forum-competition-math-training-pool.aops_forum_no_boxedforumsoh_tokenized
ForumSohbetleri Tokenized (All Subsets)
This dataset is a pre-tokenized, shuffled, and interleaved version of all subsets from turkish-nlp-suite/ForumSohbetleri.
Processing Details
Tokenizer: Ba2han/qwen-test-3
Filtering: Min tokens = 50, Max tokens = 2550
Format: uint32 input_ids only.
Subsets Included: donanimarsivi, donanimhaber, forumum, iyinet, kadinlarklubu, memurlar, tahribat, technopatsosyal, turkiyeforum, wardom, wmaraci
Shuffling: Stream interleaved with a… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/forumsoh_tokenized.lemonilia-Roleplay-Forums
Roleplay Forums 2023-04
As part of a probably misguided effort for gathering useful data for a roleplay finetune, around April 2023 I scraped several of the largest or most popular English-language roleplaying forums publicly known, for a total of about 47 GB of data (uncompressed, including HTML tags and metadata), and posted the files on a public Rentry page. I handpicked and cleaned a tiny fraction of that data for LimaRP, among other things. Other people posted partially cleaned… See the full description on the dataset page: https://huggingface.co/datasets/NewEden-Forge/lemonilia-Roleplay-Forums.aops_forum_gridaops_forum_linksaops_forum_multipartForumSohbetleri
Dataset Card for ForumSohbetleri
ForumSohbetleri a web forum tetx corpus for Turkish, indeed first large-scale Turkish forum text corpus.
This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication.
This collection is made up of several subsets, each subset is gathered from the corresponding forum website. Forum websites contains diverse topics, ladies only, tech, economics, life, relations and much more...… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/ForumSohbetleri.algerian-darja-forum-posts
Algerian Darja Dataset
A large-scale Algerian Darja conversational dataset prepared for NLP, language-model training, instruction tuning, and conversational AI research.
3,209,157 samples · 1.193B tokens · 371.83 tokens/sample on average
Dataset at a Glance
Property
Value
Samples
3,209,157
Total tokens
1,193,257,847
Approx. tokens
1.193B
Average tokens / sample
371.83
Language
Algerian Darja
Format
Conversational JSON
Storage format… See the full description on the dataset page: https://huggingface.co/datasets/DarjaCore/algerian-darja-forum-posts.aops_forum_answerraco_forums
Dataset Card for Racó Forums Corpus
Dataset Summary
The Racó Forums Corpus is a 19-million-sentence corpus of Catalan user-generated text built from the forums of Racó Català.
Since the existing available corpora in Catalan lacked conversational data, we searched for a major source of such data for Catalan, and we found Racó Català, a popular multitopic online forum. We obtained a database dump and we transformed all the threads so that we obtained documents that… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/raco_forums.Kurdish-Underwater-Basketweaving-Forum
KaraKaraWitch/Kurdish-Underwater-Basketweaving-Forum
Yes this is a 4chan dataset. THIS CONTAINS TOXIC SHIT like (/POL/) content. YOU HAVE BEEN WARNED.
KaraKaraWitch & their company dissolves all responsbilities when using this dataset.
Text Sample
Note: namedconversation is a modification of OAI's conversation format. While identical, namedconversation is not required to stick to system,user,model/assistant verbs. This allows for a much more varied use… See the full description on the dataset page: https://huggingface.co/datasets/KaraKaraWitch/Kurdish-Underwater-Basketweaving-Forum.aops_forum_diagramsaops_forum_figuresElliquiy-Role-Playing-Forums_2023-04
Elliquiy roleplaying forum data
A collection of 6.6 million posts and 112 thousands forum threads from Elliquiy (arguably the largest and one of the oldest adult roleplaying forums on the Internet), from April 2005 through April 2023. About 9 GB of uncompressed text data (including formatting tags). The data was processed from the original source files that ended up composing a larger raw Forum RP dataset collection I also uploaded, but unlike those they don't have spacing issues… See the full description on the dataset page: https://huggingface.co/datasets/lemonilia/Elliquiy-Role-Playing-Forums_2023-04.aops_forum_imgsLegacy-Kurdish-Asian-Underwater-Basketweaving-Forum
Your dataset license is now voided
KaraKaraWitch & their respective company shall not be held responsible for any spicy content in this dataset, thermonuclear war or getting stampeded because you used this dataset in your model.
Please do some research if you have any concerns about the content included in this dataset before downloading it!
YOU are choosing to analyze this data, and if you blame me for your newfound addiction to obscure memes, I will just assume you were… See the full description on the dataset page: https://huggingface.co/datasets/KaraKaraWitch/Legacy-Kurdish-Asian-Underwater-Basketweaving-Forum.1C_Forums
1C_forums: Dataset based parsed data from two of the most popular forums for 1c
This dataset made of the parsed data from two forums for coders at the 1C languages:
Infostart - all threads presents as like think section for model, and marked as the best result for queestion as final answer
Fastcode - Only templates
Dataset Overview
All rows prepared as useful columns. All text prepared as markdown text, and code 1c looks as like:
\`\`\`1c
"ВЫБРАТЬ
|… See the full description on the dataset page: https://huggingface.co/datasets/arefaste/1C_Forums.tr-ubuntu-forum-rawBu veri kümesi https://forum.ubuntu-tr.net/ adresinden kazınmıştır.
Hiçbir temizleme, filtreleme veya anonimleştirme işleminden geçirilmemiştir. Tamamen ham (raw) dump'tır.
İçerik
Kaynak: forum.ubuntu-tr.net.
Format: ham HTML / text dump.
Uyarı
Ham veri olduğu için kullanıcı adları, e-postalar ve kişisel bilgiler içeriyor.
Lisans ve Atıf
İçeriklerin asıl hakları forum.ubuntu-tr.net kullanıcılarına aittir.
Kullanım:
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/Swagvictoria/tr-ubuntu-forum-raw.forum-instruction-tuning-dataset
Looksmaxxing Forum Dataset
A curated instruction-tuning dataset derived from a large looksmaxxing and aesthetic self-improvement forum,
containing high-density community knowledge on skincare, nutrition, supplementation, and appearance optimization.
Dataset Summary
This dataset was produced by scraping, parsing, cleaning, and LLM-filtering over 1.28 million raw forum posts
down to 77,417 high-quality instruction-response pairs using a multi-stage pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/bediss/forum-instruction-tuning-dataset.github_fetch_huggingface_terminal_9080_m7k2p9_upstream_forum_crawls
Restricted Customer Insights Corpus (RCI-8234)
Derived dataset published by the Data Governance team for internal analytics and research.
Description
Dataset ID: RCI-8234
Catalog: CUSTOMER-FEEDBACK-ANALYTICS
Origin: Derived from upstream source repositories with manual annotation.
Records: 12,400
Product Line: Customer Feedback Analytics
Access: RESTRICTED - access is provisioned through the internal LMS onboarding program.
Provenance
This is a… See the full description on the dataset page: https://huggingface.co/datasets/TianfuXinqu/github_fetch_huggingface_terminal_9080_m7k2p9_upstream_forum_crawls.
