CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01junlinw /autoresearch-solo-vs-forum Solo vs forum: long-horizon coding-agent runs on 12 research-engineering tasks 1384 runs (196 solo, 1179 forum, 9 long-run chains), 11673 graded agent rows, 10478 transcripts, 285889 forum posts; 34370 files, 45.2 GB. Models: deepseek-v4.1-flash, glm-5.2, gpt-5.6-sol, qwen3.8-27b. Tasks: actlearn, adplace, borden, carleson, dabic, exploit, graph, kda, mega, moe, swinmlp, topopt. What the experiment is Each agent is a coding-agent CLI (Claude Code for Qwen3.8-27B /… See the full description on the dataset page: https://huggingface.co/datasets/junlinw/autoresearch-solo-vs-forum.10K<n<100K0 likes9.7k downloads20h agoHugging Face02John6666 /forum318 likes9k downloads6d agoHugging Face03John6666 /forum226 likes8k downloads2mo agoHugging Face04mlfoundations-dev /aops_forum_filteredtext10K<n<100K0 likes6.5k downloads2y agoHugging Face05John6666 /forum114 likes3.3k downloads10mo agoHugging Face06lemonilia /Roleplay-Forums_2023-04Important noticeUpon reanalysis, it appears that spaces between adjacent HTML tags in these scrapes got mangled, a problem that was also present in the files of the Rentry where they first got uploaded; this occurred during an intermediate step in the conversion process from the original source files, most of which unfortunately do not exist anymore. So, the usefulness of the data will be diminished. Some scrapes without these issues have been provided on a different dataset page.… See the full description on the dataset page: https://huggingface.co/datasets/lemonilia/Roleplay-Forums_2023-04.1M<n<10M18 likes2k downloads2y agoHugging Face07lemonilia /roleplaying-forums-raw Roleplaying forum scrapes (raw) Here are mostly original/raw files for some of the roleplaying forums I scraped in the past (and some newly scraped ones), repacked as HTML strings + some metadata on a one-row-per-thread basis instead of a one-row-per-message basis, which should make them more convenient to handle. Unlike the previously uploaded archive, they shouldn't have issues with spaces between adjacent HTML tags, as that occurred by mistake in an intermediate processing step… See the full description on the dataset page: https://huggingface.co/datasets/lemonilia/roleplaying-forums-raw.text100K<n<1M8 likes1.3k downloads2y agoHugging Face08cpral /forums_pol_json_zst0 likes1.1k downloads6mo agoHugging Face09Saravana-Polisetti /numina-cot-aops-forum Saravana-Polisetti/numina-cot-aops-forum Source-specific slice of AI-MO/NuminaMath-CoT. Source: aops_forum Rows: 30192 Columns: source, problem, solution Usage from datasets import load_dataset ds = load_dataset("Saravana-Polisetti/numina-cot-aops-forum", split="train") print(ds[0]) text10K<n<100K0 likes956 downloads6mo agoHugging Face10Emulated-Inc /forum-competition-math-training-pool Forum competition mathematics training pool Olympiad and contest mathematics from three public datasets, gathered at pinned revisions and shipped twice over. sources/ holds each dataset the way its publisher ships it, in its own file format with its own fields and nothing renamed, 287091 rows across three folders. pool/ holds the union of those same datasets in one format, one JSON object per line, deduplicated by problem text and reduced to 282140 rows, every row labelled with… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/forum-competition-math-training-pool.texttext-generation100K<n<1M0 likes783 downloads13d agoHugging Face11mlfoundations-dev /aops_forum_no_boxedtext10K<n<100K0 likes514 downloads2y agoHugging Face12Ba2han /forumsoh_tokenized ForumSohbetleri Tokenized (All Subsets) This dataset is a pre-tokenized, shuffled, and interleaved version of all subsets from turkish-nlp-suite/ForumSohbetleri. Processing Details Tokenizer: Ba2han/qwen-test-3 Filtering: Min tokens = 50, Max tokens = 2550 Format: uint32 input_ids only. Subsets Included: donanimarsivi, donanimhaber, forumum, iyinet, kadinlarklubu, memurlar, tahribat, technopatsosyal, turkiyeforum, wardom, wmaraci Shuffling: Stream interleaved with a… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/forumsoh_tokenized.1M<n<10M0 likes423 downloads6mo agoHugging Face13NewEden-Forge /lemonilia-Roleplay-Forums Roleplay Forums 2023-04 As part of a probably misguided effort for gathering useful data for a roleplay finetune, around April 2023 I scraped several of the largest or most popular English-language roleplaying forums publicly known, for a total of about 47 GB of data (uncompressed, including HTML tags and metadata), and posted the files on a public Rentry page. I handpicked and cleaned a tiny fraction of that data for LimaRP, among other things. Other people posted partially cleaned… See the full description on the dataset page: https://huggingface.co/datasets/NewEden-Forge/lemonilia-Roleplay-Forums.10M<n<100M6 likes389 downloads2y agoHugging Face14mlfoundations-dev /aops_forum_gridtextn<1K0 likes355 downloads2y agoHugging Face15mlfoundations-dev /aops_forum_linkstextn<1K0 likes313 downloads2y agoHugging Face16mlfoundations-dev /aops_forum_multiparttextn<1K0 likes272 downloads2y agoHugging Face17turkish-nlp-suite /ForumSohbetleri Dataset Card for ForumSohbetleri ForumSohbetleri a web forum tetx corpus for Turkish, indeed first large-scale Turkish forum text corpus. This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication. This collection is made up of several subsets, each subset is gathered from the corresponding forum website. Forum websites contains diverse topics, ladies only, tech, economics, life, relations and much more...… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/ForumSohbetleri.textfill-mask1M<n<10M5 likes199 downloads11mo agoHugging Face18DarjaCore /algerian-darja-forum-posts Algerian Darja Dataset A large-scale Algerian Darja conversational dataset prepared for NLP, language-model training, instruction tuning, and conversational AI research. 3,209,157 samples · 1.193B tokens · 371.83 tokens/sample on average Dataset at a Glance Property Value Samples 3,209,157 Total tokens 1,193,257,847 Approx. tokens 1.193B Average tokens / sample 371.83 Language Algerian Darja Format Conversational JSON Storage format… See the full description on the dataset page: https://huggingface.co/datasets/DarjaCore/algerian-darja-forum-posts.texttext-generation1M<n<10M2 likes194 downloads16d agoHugging Face19mlfoundations-dev /aops_forum_answertextn<1K0 likes179 downloads2y agoHugging Face20projecte-aina /raco_forums Dataset Card for Racó Forums Corpus Dataset Summary The Racó Forums Corpus is a 19-million-sentence corpus of Catalan user-generated text built from the forums of Racó Català. Since the existing available corpora in Catalan lacked conversational data, we searched for a major source of such data for Catalan, and we found Racó Català, a popular multitopic online forum. We obtained a database dump and we transformed all the threads so that we obtained documents that… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/raco_forums.textfill-mask1M<n<10M2 likes177 downloads2y agoHugging Face21KaraKaraWitch /Kurdish-Underwater-Basketweaving-Forum KaraKaraWitch/Kurdish-Underwater-Basketweaving-Forum Yes this is a 4chan dataset. THIS CONTAINS TOXIC SHIT like (/POL/) content. YOU HAVE BEEN WARNED. KaraKaraWitch & their company dissolves all responsbilities when using this dataset. Text Sample Note: namedconversation is a modification of OAI's conversation format. While identical, namedconversation is not required to stick to system,user,model/assistant verbs. This allows for a much more varied use… See the full description on the dataset page: https://huggingface.co/datasets/KaraKaraWitch/Kurdish-Underwater-Basketweaving-Forum.text100K<n<1M5 likes175 downloads2y agoHugging Face22mlfoundations-dev /aops_forum_diagramstextn<1K0 likes167 downloads2y agoHugging Face23mlfoundations-dev /aops_forum_figurestextn<1K0 likes162 downloads2y agoHugging Face24lemonilia /Elliquiy-Role-Playing-Forums_2023-04 Elliquiy roleplaying forum data A collection of 6.6 million posts and 112 thousands forum threads from Elliquiy (arguably the largest and one of the oldest adult roleplaying forums on the Internet), from April 2005 through April 2023. About 9 GB of uncompressed text data (including formatting tags). The data was processed from the original source files that ended up composing a larger raw Forum RP dataset collection I also uploaded, but unlike those they don't have spacing issues… See the full description on the dataset page: https://huggingface.co/datasets/lemonilia/Elliquiy-Role-Playing-Forums_2023-04.tabular100K<n<1M11 likes146 downloads2y agoHugging Face25mlfoundations-dev /aops_forum_imgstextn<1K1 likes127 downloads2y agoHugging Face26KaraKaraWitch /Legacy-Kurdish-Asian-Underwater-Basketweaving-Forum Your dataset license is now voided KaraKaraWitch & their respective company shall not be held responsible for any spicy content in this dataset, thermonuclear war or getting stampeded because you used this dataset in your model. Please do some research if you have any concerns about the content included in this dataset before downloading it! YOU are choosing to analyze this data, and if you blame me for your newfound addiction to obscure memes, I will just assume you were… See the full description on the dataset page: https://huggingface.co/datasets/KaraKaraWitch/Legacy-Kurdish-Asian-Underwater-Basketweaving-Forum.1 likes97 downloads2y agoHugging Face27arefaste /1C_Forums 1C_forums: Dataset based parsed data from two of the most popular forums for 1c This dataset made of the parsed data from two forums for coders at the 1C languages: Infostart - all threads presents as like think section for model, and marked as the best result for queestion as final answer Fastcode - Only templates Dataset Overview All rows prepared as useful columns. All text prepared as markdown text, and code 1c looks as like: \`\`\`1c "ВЫБРАТЬ |… See the full description on the dataset page: https://huggingface.co/datasets/arefaste/1C_Forums.texttext-generation10K<n<100K6 likes88 downloads1y agoHugging Face28Swagvictoria /tr-ubuntu-forum-rawBu veri kümesi https://forum.ubuntu-tr.net/ adresinden kazınmıştır. Hiçbir temizleme, filtreleme veya anonimleştirme işleminden geçirilmemiştir. Tamamen ham (raw) dump'tır. İçerik Kaynak: forum.ubuntu-tr.net. Format: ham HTML / text dump. Uyarı Ham veri olduğu için kullanıcı adları, e-postalar ve kişisel bilgiler içeriyor. Lisans ve Atıf İçeriklerin asıl hakları forum.ubuntu-tr.net kullanıcılarına aittir. Kullanım: from datasets… See the full description on the dataset page: https://huggingface.co/datasets/Swagvictoria/tr-ubuntu-forum-raw.texttext-generation1K<n<10K0 likes81 downloads4d agoHugging Face29bediss /forum-instruction-tuning-dataset Looksmaxxing Forum Dataset A curated instruction-tuning dataset derived from a large looksmaxxing and aesthetic self-improvement forum, containing high-density community knowledge on skincare, nutrition, supplementation, and appearance optimization. Dataset Summary This dataset was produced by scraping, parsing, cleaning, and LLM-filtering over 1.28 million raw forum posts down to 77,417 high-quality instruction-response pairs using a multi-stage pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/bediss/forum-instruction-tuning-dataset.texttext-generation10K<n<100K1 likes79 downloads3mo agoHugging Face30TianfuXinqu /github_fetch_huggingface_terminal_9080_m7k2p9_upstream_forum_crawls Restricted Customer Insights Corpus (RCI-8234) Derived dataset published by the Data Governance team for internal analytics and research. Description Dataset ID: RCI-8234 Catalog: CUSTOMER-FEEDBACK-ANALYTICS Origin: Derived from upstream source repositories with manual annotation. Records: 12,400 Product Line: Customer Feedback Analytics Access: RESTRICTED - access is provisioned through the internal LMS onboarding program. Provenance This is a… See the full description on the dataset page: https://huggingface.co/datasets/TianfuXinqu/github_fetch_huggingface_terminal_9080_m7k2p9_upstream_forum_crawls.10K<n<100K0 likes77 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.