CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01common-pile /stackv2_edu_filtered Stack V2 Edu Description We filter the Stack V2 to only include code from openly licensed repositories, based on the license detection performed by the creators of Stack V2. When multiple licenses are detected in a single repository, we ensure that all of the licenses are on the Blue Oak Council certified license list. Per-document license information is available in the license entry of the metadata field of each example. Code for collecting, processing, and preparing… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/stackv2_edu_filtered.tabulartext-generation10M<n<100M6 likes6.7k downloads1y agoHugging Face02common-pile /peS2o_filtered PeS2o Description This dataset is a version of the peS2o dataset restricted to openly licensed articles. PeS2o is derived from S2ORC, a corpus of openly licensed abstract and full-text papers that have been converted to a structured format using Grobid. Starting from Grobid’s XML output, peS2o filters papers that are too short, have incorrect metadata, are in languages other than English, and contain OCR errors using a combination of heuristic- and model-based filtering… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/peS2o_filtered.texttext-generation1M<n<10M4 likes4.2k downloads1y agoHugging Face03common-pile /stackexchange_filtered Stack Exchange Description StackExchange is a collection of Q&A communities spanning a wide variety of topics.While StackExchange formerly provided structured dumps of all of their content, since July of 2024, StackExchange has stopped publishing XML dumps to the Internet Archive.Instead, each site can provide a logged-in user with a custom URL to download the dump for that site.This means that dumps for defunct sites like windowsphone.stackexchange.com are… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/stackexchange_filtered.texttext-generation10M<n<100M10 likes3.6k downloads1y agoHugging Face04common-pile /uspto_filtered USPTO Description In the United States, patent documents are released into the public domain as government works. Patents follow a highly standardized format with distinct required sections for background, detailed description, and claims. We include patents from the US Patents and Trademark Office (USPTO) as provided by the Google Patents Public Data dataset, which includes millions of granted patents and published patent applications dating back to 1782. We processed… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/uspto_filtered.texttext-generation10M<n<100M4 likes3.4k downloads1y agoHugging Face05common-pile /project_gutenberg_filtered Project Gutenberg Description Project Gutenberg is an online collection of over 75,000 digitized books available as plain text. We use all books that are 1) English and 2) marked as in the Public Domain according to the provided metadata. Additionally, we include any books that are part of the PG19 dataset, which only includes books that are over 100 years old. Minimal preprocessing is applied to remove the Project Gutenberg header and footers, but many scanned books… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/project_gutenberg_filtered.texttext-generation10K<n<100K7 likes2.9k downloads1y agoHugging Face06common-pile /wikimedia_filtered Wikimedia Description Official Wikimedia wikis are released under a CC BY-SA license. We downloaded the official database dumps from March 2025 of the English-language wikis that are directly managed by the Wikimedia Foundation. These database dumps include the wikitext—MediaWiki’s custom markup language—for each page as well as talk pages, where editors discuss changes made to a page. We only use the most recent version of each page. We converted wikitext to plain text… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/wikimedia_filtered.texttext-generation10M<n<100M8 likes2.8k downloads1y agoHugging Face07MoreThought /Fable-5.1-Max-Reasoning-Filtered-10000x Dataset Description This dataset contains 10,000 agentic coding and reasoning multi-turn high-quality traces generated by the new Fable 5.1 model using max reasoning effort. It holds almost 500,000,000 tokens of step-by-step chain-of-thought programming across multiple complex domains. It has also been deduplicated and heavily filtered to remove low-quality traces, keeping only high-quality traces. Dataset Statistics Metric Value Total Examples 10,000… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/Fable-5.1-Max-Reasoning-Filtered-10000x.text-generation10K<n<100K162 likes2.6k downloads19h agoHugging Face08adameubanks /filtered_articles_by_year Dataset Card for Filtered Articles by Year Dataset Summary The Filtered Articles by Year dataset contains yearly-segmented web articles from the FineWeb dataset, specifically filtered and processed for temporal language analysis and Word2Vec model training. This dataset spans 21 years (2005-2025) and serves as the foundation for research into semantic change, concept emergence, and language evolution over time. Supported Tasks and Leaderboards This dataset… See the full description on the dataset page: https://huggingface.co/datasets/adameubanks/filtered_articles_by_year.texttext-generation10M<n<100M1 likes2.5k downloads1y agoHugging Face09common-pile /pubmed_filtered PubMed Description PubMed Central is an open-access archive of biomedical and life sciences research papers maintained by the U.S. National Institutes of Health’s National Library of Medicine. We collected papers from PubMed whose metadata indicated that the publishing journal had designated a CC BY, CC BY-SA, or CC0 license. PubMed stores the text content of each article as a single nXML file, which we convert to markdown using pandoc. Per-document license information is… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/pubmed_filtered.texttext-generation1M<n<10M5 likes2.2k downloads1y agoHugging Face100xDing /wikipedia-cn-20230720-filtered本数据集基于中文维基2023年7月20日的dump存档。作为一项以数据为中心的工作,本数据集仅保留了 254,547条 质量较高的词条内容。具体而言: 过滤了Template, Category, Wikipedia, File, Topic, Portal, MediaWiki, Draft, Help等特殊类型的词条 使用启发式的方法和自有的NLU模型过滤了一部分质量较低的词条 过滤了一部分内容较为敏感或存在争议性的词条。 进行了简繁转换和习惯用词转换,确保符合中国大陆地区的习惯用词。 This dataset is based on the Chinese Wikipedia dump archive from July 20th, 2023. As a data-centric effort, the dataset retains 254,574 high-quality entries. Specifically: Entries of special types such as Template, Category, Wikipedia, File, Topic… See the full description on the dataset page: https://huggingface.co/datasets/0xDing/wikipedia-cn-20230720-filtered.texttext-generation100K<n<1M173 likes2.2k downloads3y agoHugging Face11common-pile /pre_1929_books_filtered Pre-1929 Books Description Books published in the US before 1929 passed into the public domain on January 1, 2024. We used the bibliographic catalog Hathifiles produced by HathiTrust to identify digitized books which were published in the US before 1929. The collection contains over 130,000 books digitized and processed by the Internet Archive on behalf of HathiTrust member libraries. The OCR plain text files were downloaded directly from the Internet Archive website.… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/pre_1929_books_filtered.texttext-generation100K<n<1M2 likes1.8k downloads1y agoHugging Face12if001 /oscar_2023_filteredfrom datasets import load_dataset ds=load_dataset("if001/oscar_2023_filtered") ds['train'] --- Dataset({ features: ['text'], num_rows: 312396 }) oscar 2023をfilterしたものhttps://huggingface.co/datasets/oscar-corpus/OSCAR-2301 詳細はコードを参照https://github.com/if001/HojiChar_OSCAR_sample/tree/0.0.4 texttext-generation1M<n<10M3 likes1.4k downloads3y agoHugging Face13common-pile /usgpo_filtered USGPO Description The United States Government Publishing Office (USGPO) is a federal agency responsible for disseminating official documents authored by the U.S. government. This dataset includes all plain-text documents made available through the USGPO’s GovInfo.gov developer API. This collection comprises over 2.7 million documents, spanning issues of the Federal Register, congressional hearing transcripts, budget reports, economic indicators, and other federal… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/usgpo_filtered.texttext-generation1M<n<10M3 likes1.2k downloads1y agoHugging Face14common-pile /caselaw_access_project_filtered Caselaw Access Project Description This dataset contains 6.7 million cases from the Caselaw Access Project and Court Listener. The Caselaw Access Project consists of nearly 40 million pages of U.S. federal and state court decisions and judges’ opinions from the last 365 years. In addition, Court Listener adds over 900 thousand cases scraped from 479 courts. The Caselaw Access Project and Court Listener source legal data from a wide variety of resources such as the Harvard… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/caselaw_access_project_filtered.texttext-generation1M<n<10M12 likes1.2k downloads1y agoHugging Face15common-pile /library_of_congress_filtered Library of Congress Description The Library of Congress (LoC) curates a collection of public domain books called "Selected Digitized Books". We have downloaded over 130,000 English-language books from this public domain collection as OCR plain text files using the LoC APIs. Dataset Statistics Documents UTF-8 GB 129,052 35.6 License Issues While we aim to produce datasets with completely accurate licensing information, license… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/library_of_congress_filtered.texttext-generation100K<n<1M2 likes990 downloads1y agoHugging Face16common-pile /youtube_filtered Creative Commons YouTube Description YouTube is a large-scale video-sharing platform where users have the option of uploading content under a CC BY license. To collect high-quality speech-based textual content and combat the rampant license laundering on YouTube, we manually curated a set of over 2,000 YouTube channels that consistently release original openly licensed content containing speech. The resulting collection spans a wide range of genres, including lectures… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/youtube_filtered.texttext-generation100K<n<1M6 likes866 downloads1y agoHugging Face17common-pile /uk_hansard_filtered UK Hansard Description Hansard represents the official record of parliamentary proceedings across the United Kingdom’s legislative bodies. This dataset incorporates records from multiple sources, including debates and written answers from the UK Commons and Lords, devolved legislatures (Scottish Parliament, Senedd in both English and Welsh, Northern Ireland Assembly), London Mayor’s Questions, and ministerial statements. Data was sourced from ParlParse, covering Commons… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/uk_hansard_filtered.texttext-generation10K<n<100K1 likes824 downloads1y agoHugging Face18lyrain2001 /Auto-Fill-Benchmark Auto-Fill Benchmark Benchmark for predicting missing cell values in real-world tables, introduced in Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models (PVLDB 19(11), 2026 — arXiv:2607.19847). Each case is a real table in which exactly one cell is replaced by [MISSING], together with the ground-truth value. Code: https://github.com/lyrain2001/auto-fill Models: Auto-Fill-Qwen3-8B-Knowledge · Auto-Fill-Qwen3-8B-Reasoning ·… See the full description on the dataset page: https://huggingface.co/datasets/lyrain2001/Auto-Fill-Benchmark.tabulartable-question-answering1K<n<10K0 likes705 downloads28d agoHugging Face19common-pile /github_archive_filtered GitHub Archive Description According to GitHub’s terms of service, issues and pull request descriptions—along with their comments—inherit the license of their associated repository. To collect this data, we used the GitHub Archive’s public BigQuery table of events to extract all issue, pull request, and comment events since 2011 and aggregated them into threads. The table appeared to be missing “edit” events so the text from each comment is the original from when it was… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive_filtered.texttext-generation10M<n<100M2 likes673 downloads1y agoHugging Face20common-pile /regulations_filtered Regulations.Gov Description Regulations.gov is an online platform operated by the U.S. General Services Administration that collates newly proposed rules and regulations from federal agencies along with comments and feedback from the general public. This dataset includes all plain-text regulatory documents published by a variety of U.S. federal agencies on this platform, acquired via the bulk download interface provided by Regulations.gov. These agencies include the… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/regulations_filtered.texttext-generation100K<n<1M0 likes624 downloads1y agoHugging Face21common-pile /cccc_filtered Creative Commons Common Crawl Description This dataset contains text from 52 Common Crawl snapshots, covering about half of Common Crawl snapshots available to date and covering all years of operations of Common Crawl up to 2024. We found a higher level of duplication across this collection, suggesting that including more snapshots would lead to a modest increase in total token yield. From these snapshots, we extract HTML content using FastWarc. Then, using a regular… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/cccc_filtered.texttext-generation10M<n<100M2 likes601 downloads1y agoHugging Face22common-pile /arxiv_abstracts_filtered ArXiv Abstracts Description Each paper uploaded to ArXiv includes structured metadata fields, including an abstract summarizing the paper’s findings and contributions. According to ArXiv’s licensing policy, the metadata for any paper submitted to ArXiv is distributed under the CC0 license, regardless of the license of the paper itself. Thus, this dataset contains the abstract for every paper submitted to ArXiv through late 2024. We source the abstracts from ArXiv’s API… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/arxiv_abstracts_filtered.texttext-generation1M<n<10M9 likes577 downloads10mo agoHugging Face23common-pile /ubuntu_irc_filtered Ubuntu IRC Description Logs of all discussions on the Ubuntu-hosted Internet Relay Chat (IRC) since 2004 have been archived and released into the Public Domain. We downloaded all chats from all channels up until March of 2025. We consider all messages for a given channel on a given day as a single document. We removed system messages as well as those from known bots. Dataset Statistics Documents UTF-8 GB 234,982 5.3 License Issues… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/ubuntu_irc_filtered.texttext-generation100K<n<1M2 likes533 downloads1y agoHugging Face24common-pile /biodiversity_heritage_library_filtered Biodiversity Heritage Library Description The Biodiversity Heritage Library (BHL) is an open-access digital library for biodiversity literature and archives. This dataset contains over 15 million public domain books and documents from the BHL collection. These works were collected using the bulk data download interface provided by the BHL and were filtered based on their associated license metadata. We use the optical character recognition (OCR)-generated text distributed… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/biodiversity_heritage_library_filtered.texttext-generation10M<n<100M2 likes468 downloads1y agoHugging Face25DJLougen /hermes-agent-traces-filtered Hermes Agent Reasoning Traces - Quality Filtered A structurally filtered subset of lambda/hermes-agent-reasoning-traces, pruned from 7,646 to 3,679 rows using automated quality analysis targeting reasoning depth, structural integrity, and tool-call validity. Why This Matters for Agent Training Most agentic datasets teach models what tool to call but not how to reason about tool selection. The difference matters in production: an agent that dispatches tools without… See the full description on the dataset page: https://huggingface.co/datasets/DJLougen/hermes-agent-traces-filtered.texttext-generation1K<n<10K44 likes395 downloads6mo agoHugging Face26common-pile /wikiteam_filtered Wikiteam Description There are many wikis on the internet that are not managed by the Wikimedia Foundation, but do use their MediaWiki software to power their wiki. Many of these wikis have been archived by Wikiteam, a collection of volunteers that create unofficial database dumps of wikis and upload them to the Internet Archive. We download all dumps made by Wikiteam when the metadata indicates the wiki was licensed under CC BY, CC BY-SA, or released into the public… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/wikiteam_filtered.texttext-generation10M<n<100M2 likes307 downloads1y agoHugging Face27KrazyKitty /Fable-5.1-Max-Reasoning-Filtered-1000x Dataset Description This dataset contains 1,000 coding and reasoning traces generated by the new Fable 5.1 model using max reasoning effort. It holds almost 30,000,000 tokens of step-by-step chain-of-thought programming across multiple complex domains. It has also been deduplicated and filtered to remove low-quality traces, keeping only high-quality traces. Dataset Statistics Metric Value Total Examples 1,000 Traces Total Token Count ~30,000,000… See the full description on the dataset page: https://huggingface.co/datasets/KrazyKitty/Fable-5.1-Max-Reasoning-Filtered-1000x.texttext-generation1K<n<10K6 likes271 downloads13d agoHugging Face28MoreThought /Fable-5-Max-Reasoning-Filtered-250x Dataset Description This dataset contains 25. highly detailed architectural traces mapping out security implementations for hybrid global banking systems encompassing both fiat and cryptocurrency infrastructures. This is 10,000,000+ estimated tokens of fable 5 data, filtered and classified to remove low-quality entries by qwen 2.5 7B, and improved by GLM 5.2. The dataset bypasses basic conversational filler and is engineered to advance the domain precision, strict formatting… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/Fable-5-Max-Reasoning-Filtered-250x.texttext-generationn<1K13 likes222 downloads2mo agoHugging Face29common-pile /data_provenance_initiative_filtered Data Provenance Initiative Description The Data Provenance Initiative is a digital library of supervised datasets that have been manually annotated with their source and license information [ 104, 107 ]. We leverage their tooling to filter HuggingFace datasets, based on a range of criteria, including their licenses. Specifically, we filter the data according to these criteria: contains English language or code data, the text is not model-generated, the dataset’s audit… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/data_provenance_initiative_filtered.texttext-generation1M<n<10M0 likes186 downloads1y agoHugging Face30LARK-Lab /EnvFactory-SFT-FILTERED EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL ## Overview EnvFactory-SFT-FILTERED is a filtered supervised fine-tuning (SFT) dataset containing 53,400 tool-use trajectories synthesized using the EnvFactory framework. This dataset is designed for SFT training of tool-use agents. The dataset contains high-quality multi-turn tool-use trajectories with implicit human reasoning, generated through… See the full description on the dataset page: https://huggingface.co/datasets/LARK-Lab/EnvFactory-SFT-FILTERED.texttext-generation10K<n<100K0 likes183 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.