datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ForumSohbetleri
Dataset Card for ForumSohbetleri
ForumSohbetleri a web forum tetx corpus for Turkish, indeed first large-scale Turkish forum text corpus.
This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication.
This collection is made up of several subsets, each subset is gathered from the corresponding forum website. Forum websites contains diverse topics, ladies only, tech, economics, life, relations and much more...… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/ForumSohbetleri.Kurdish-Underwater-Basketweaving-Forum
KaraKaraWitch/Kurdish-Underwater-Basketweaving-Forum
Yes this is a 4chan dataset. THIS CONTAINS TOXIC SHIT like (/POL/) content. YOU HAVE BEEN WARNED.
KaraKaraWitch & their company dissolves all responsbilities when using this dataset.
Text Sample
Note: namedconversation is a modification of OAI's conversation format. While identical, namedconversation is not required to stick to system,user,model/assistant verbs. This allows for a much more varied use… See the full description on the dataset page: https://huggingface.co/datasets/KaraKaraWitch/Kurdish-Underwater-Basketweaving-Forum.tr-ubuntu-forum-rawBu veri kümesi https://forum.ubuntu-tr.net/ adresinden kazınmıştır.
Hiçbir temizleme, filtreleme veya anonimleştirme işleminden geçirilmemiştir. Tamamen ham (raw) dump'tır.
İçerik
Kaynak: forum.ubuntu-tr.net.
Format: ham HTML / text dump.
Uyarı
Ham veri olduğu için kullanıcı adları, e-postalar ve kişisel bilgiler içeriyor.
Lisans ve Atıf
İçeriklerin asıl hakları forum.ubuntu-tr.net kullanıcılarına aittir.
Kullanım:
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/Swagvictoria/tr-ubuntu-forum-raw.forum-instruction-tuning-dataset
Looksmaxxing Forum Dataset
A curated instruction-tuning dataset derived from a large looksmaxxing and aesthetic self-improvement forum,
containing high-density community knowledge on skincare, nutrition, supplementation, and appearance optimization.
Dataset Summary
This dataset was produced by scraping, parsing, cleaning, and LLM-filtering over 1.28 million raw forum posts
down to 77,417 high-quality instruction-response pairs using a multi-stage pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/bediss/forum-instruction-tuning-dataset.Cooking_Forum-ShareGPT-Raweducation-forum-jfk-dataset
Education Forum JFK Assassination Debate Dataset
Overview
This dataset contains a structured archive of public discussion threads centered on the JFK Assassination Debate section of the Education Forum.
Beginning with Version 2.0, the dataset also includes selected discussion forums from other sections of the Education Forum while retaining the original dataset name for continuity and discoverability.
The Education Forum spans more than two decades of discussion… See the full description on the dataset page: https://huggingface.co/datasets/Tgram3D/education-forum-jfk-dataset.veripazari.com.tr-t-rk-e-forum-ve-nternet-k-lt-r-veri-seti
Bu veri seti orijinal olarak veripazari.com.tr tarafından geliştirilmiş/yüklenmiş olup, veripazari.com.tr tarafından Türk AI ekosistemi için arşivlenmiştir.
🔗 Orijinal Kaynak ve Detaylar: veripazari.com.tr
Türkçe Forum ve İnternet Kültürü Veri Seti
Bu veri seti, Türkiye'nin en aktif ve popüler platformlarından toplanmış, organik kullanıcı içeriklerini barındıran devasa bir Türkçe metin derlemesidir.
Amacı, Türkçe yapay zeka modellerine resmi ve kitabi dilin ötesinde; günlük… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/veripazari.com.tr-t-rk-e-forum-ve-nternet-k-lt-r-veri-seti.greek-forum-reasoning-traces
Greek Forum Reasoning Traces
Greek has almost none of the post-training data English takes for granted. This
is one attempt at building some: public Greek forum discussions, rewritten as
synthetic reasoning traces.
Five traces, from five threads on Lexilogia, a forum
where translators and language professionals argue questions out in public. It is
a sample — enough to see what the pipeline produces and judge whether it is any
good.
How a discussion becomes a trace
— the… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/greek-forum-reasoning-traces.JumpLander-Persian-Forum-mini-Dataset
📚 JumpLander Persian Forum Mini Dataset
High-Quality Persian (Farsi) Text for NLP and AI Research
This dataset contains a clean and structured subset of Persian community discussions collected from JumpLander.org forums.It enables developers, researchers, and ML engineers to build and evaluate Farsi NLP models including:
Text classification
Topic modeling
Semantic search
NER / summarization
LLM and transformer fine-tuning
📊 Dataset Details
Language:… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JumpLander-Persian-Forum-mini-Dataset.cvx-forum-conversations
cvx-forum-conversations
This dataset includes 4.78k conversation contents crawled from CVX forum. It is crawled, then cleaned by a LLM and me, and to be readily used for finetuning.
The contents from CVX forum are mainly contributed by Mark L. Stone, Michael C. Grant, Erling D.Andersen, Michal Adamaszek, Henrik A. Friberg, Stephen Becker, jackfsuia and all the other posters around the world.
Quality
Its quality is not guaranteed, now it looks like a 6.5 out of 10, due… See the full description on the dataset page: https://huggingface.co/datasets/tim1900/cvx-forum-conversations.forumsforums2forums_jsonso-forumso_forum_docs_blogs_allRP-forum-R2RP-forum-2Internet-Forum-Logs-1-Sharegptchastity_forum_datasetCooking-Forum-ShareGPTgentoo_forum_solveddebian_forum_solved
