datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
forum-competition-math-training-pool
Forum competition mathematics training pool
Olympiad and contest mathematics from three public datasets, gathered at pinned revisions and
shipped twice over. sources/ holds each dataset the way its publisher ships it, in its own file
format with its own fields and nothing renamed, 287091 rows across three folders. pool/ holds the
union of those same datasets in one format, one JSON object per line, deduplicated by problem text
and reduced to 282140 rows, every row labelled with… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/forum-competition-math-training-pool.algerian-darja-forum-posts
Algerian Darja Dataset
A large-scale Algerian Darja conversational dataset prepared for NLP, language-model training, instruction tuning, and conversational AI research.
3,209,157 samples · 1.193B tokens · 371.83 tokens/sample on average
Dataset at a Glance
Property
Value
Samples
3,209,157
Total tokens
1,193,257,847
Approx. tokens
1.193B
Average tokens / sample
371.83
Language
Algerian Darja
Format
Conversational JSON
Storage format… See the full description on the dataset page: https://huggingface.co/datasets/DarjaCore/algerian-darja-forum-posts.ForumSohbetleri
Dataset Card for ForumSohbetleri
ForumSohbetleri a web forum tetx corpus for Turkish, indeed first large-scale Turkish forum text corpus.
This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication.
This collection is made up of several subsets, each subset is gathered from the corresponding forum website. Forum websites contains diverse topics, ladies only, tech, economics, life, relations and much more...… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/ForumSohbetleri.tr-ubuntu-forum-rawBu veri kümesi https://forum.ubuntu-tr.net/ adresinden kazınmıştır.
Hiçbir temizleme, filtreleme veya anonimleştirme işleminden geçirilmemiştir. Tamamen ham (raw) dump'tır.
İçerik
Kaynak: forum.ubuntu-tr.net.
Format: ham HTML / text dump.
Uyarı
Ham veri olduğu için kullanıcı adları, e-postalar ve kişisel bilgiler içeriyor.
Lisans ve Atıf
İçeriklerin asıl hakları forum.ubuntu-tr.net kullanıcılarına aittir.
Kullanım:
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/Swagvictoria/tr-ubuntu-forum-raw.1C_Forums
1C_forums: Dataset based parsed data from two of the most popular forums for 1c
This dataset made of the parsed data from two forums for coders at the 1C languages:
Infostart - all threads presents as like think section for model, and marked as the best result for queestion as final answer
Fastcode - Only templates
Dataset Overview
All rows prepared as useful columns. All text prepared as markdown text, and code 1c looks as like:
\`\`\`1c
"ВЫБРАТЬ
|… See the full description on the dataset page: https://huggingface.co/datasets/arefaste/1C_Forums.forum-instruction-tuning-dataset
Looksmaxxing Forum Dataset
A curated instruction-tuning dataset derived from a large looksmaxxing and aesthetic self-improvement forum,
containing high-density community knowledge on skincare, nutrition, supplementation, and appearance optimization.
Dataset Summary
This dataset was produced by scraping, parsing, cleaning, and LLM-filtering over 1.28 million raw forum posts
down to 77,417 high-quality instruction-response pairs using a multi-stage pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/bediss/forum-instruction-tuning-dataset.dread-crime-forum
Dread Forum Archive
A near-complete capture of public content from Dread, a Reddit-style discussion forum hosted as a Tor hidden service. Dread is one of the longest-running darknet community forums and a primary site for discussion of darknet markets, operational security, cryptocurrency, and related topics.
The archive covers content posted between April 2018 and September 2025 and is structured as three Parquet-backed splits: posts, comments, and users (with parsed PGP key… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/dread-crime-forum.huggingface_forum
Hugging Face Forum Dataset
This dataset was scraped from various categories on the Hugging Face forum on January 31, 2025.
It contains posts, responses, and metadata such as dates and view counts across multiple topics.
Dataset Details
Source: Hugging Face Forum
Categories: Accelerate, AutoTrain, AWS Inferentia Trainium, AzureML, Beginners, Community Calls, Course, Datasets, Diffusers, Flax JAX Projects, Google Cloud, Gradio, Hub, Inference Endpoints, Intermediate… See the full description on the dataset page: https://huggingface.co/datasets/Prikshit7766/huggingface_forum.torch-forum
Dataset Card for "torch-forum"
Dataset structure
{
title:str
category:str,
posts:List[{
poster:str,
contents:str,
likes:int,
isAccepted:bool
}]
}
simson-forum-qa-pairs
🏍️ Simson Forum QA Pairs
147 strukturierte Frage-Antwort-Paare aus den größten deutschen Simson-Foren und Wissensquellen.
Quellen
Quelle
Typ
Anzahl
schwalbennest.de
FAQ-Artikel + Wiki
61
simsonforum.net
Forum-Threads
86
Schema
Feld
Typ
Beschreibung
id
string
Eindeutige ID (sb-xxx oder sf-xxx)
source
string
Quelldomain
source_type
string
faq_article, wiki_article, forum_thread
url
string
Original-URL
title
string… See the full description on the dataset page: https://huggingface.co/datasets/jmp1987/simson-forum-qa-pairs.JumpLander-Persian-Forum-mini-Dataset
📚 JumpLander Persian Forum Mini Dataset
High-Quality Persian (Farsi) Text for NLP and AI Research
This dataset contains a clean and structured subset of Persian community discussions collected from JumpLander.org forums.It enables developers, researchers, and ML engineers to build and evaluate Farsi NLP models including:
Text classification
Topic modeling
Semantic search
NER / summarization
LLM and transformer fine-tuning
📊 Dataset Details
Language:… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JumpLander-Persian-Forum-mini-Dataset.pytorch-forum-topics-complete-v2
PyTorch Forum Topics Dataset
This dataset contains topic metadata scraped from the PyTorch Community Forum. It includes comprehensive information about forum topics that can be used for various NLP tasks related to PyTorch and deep learning discussions.
Dataset Structure
Each record in the dataset contains the following fields:
id: Unique topic identifier
title: Topic title
slug: URL-friendly version of the title
posts_count: Number of posts in the topic
reply_count:… See the full description on the dataset page: https://huggingface.co/datasets/AmitPrakash/pytorch-forum-topics-complete-v2.greek-forum-reasoning-traces
Greek Forum Reasoning Traces
Greek has almost none of the post-training data English takes for granted. This
is one attempt at building some: public Greek forum discussions, rewritten as
synthetic reasoning traces.
Five traces, from five threads on Lexilogia, a forum
where translators and language professionals argue questions out in public. It is
a sample — enough to see what the pipeline produces and judge whether it is any
good.
How a discussion becomes a trace
— the… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/greek-forum-reasoning-traces.
