CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01m-a-p /CodeFeedback-Filtered-Instruction OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement [🏠Homepage] | [🛠️Code] OpenCodeInterpreter OpenCodeInterpreter is a family of open-source code generation systems designed to bridge the gap between large language models and advanced proprietary systems like the GPT-4 Code Interpreter. It significantly advances code generation capabilities by integrating execution and iterative refinement functionalities. For further information and… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/CodeFeedback-Filtered-Instruction.textquestion-answering100K<n<1M208 likes28k downloads3y agoHugging Face02common-pile /stackv2_edu_filtered Stack V2 Edu Description We filter the Stack V2 to only include code from openly licensed repositories, based on the license detection performed by the creators of Stack V2. When multiple licenses are detected in a single repository, we ensure that all of the licenses are on the Blue Oak Council certified license list. Per-document license information is available in the license entry of the metadata field of each example. Code for collecting, processing, and preparing… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/stackv2_edu_filtered.tabulartext-generation10M<n<100M6 likes6.5k downloads1y agoHugging Face03common-pile /peS2o_filtered PeS2o Description This dataset is a version of the peS2o dataset restricted to openly licensed articles. PeS2o is derived from S2ORC, a corpus of openly licensed abstract and full-text papers that have been converted to a structured format using Grobid. Starting from Grobid’s XML output, peS2o filters papers that are too short, have incorrect metadata, are in languages other than English, and contain OCR errors using a combination of heuristic- and model-based filtering… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/peS2o_filtered.texttext-generation1M<n<10M4 likes4.1k downloads1y agoHugging Face04malaysia-ai /mosaic-starcoder-filtered Mosaic format for filtered starcoder dataset to train Malaysian LLM This repository is to store dataset shards using mosaic format. prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-starcoder.ipynb using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer 4096 context length. how-to git clone, git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-starcoder-filtered load it, from streaming import… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-starcoder-filtered.textn<1K0 likes3.9k downloads3y agoHugging Face05nkandpa2 /common-pile-filteredtext1B<n<10B1 likes3.3k downloads1y agoHugging Face06common-pile /stackexchange_filtered Stack Exchange Description StackExchange is a collection of Q&A communities spanning a wide variety of topics.While StackExchange formerly provided structured dumps of all of their content, since July of 2024, StackExchange has stopped publishing XML dumps to the Internet Archive.Instead, each site can provide a logged-in user with a custom URL to download the dump for that site.This means that dumps for defunct sites like windowsphone.stackexchange.com are… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/stackexchange_filtered.texttext-generation10M<n<100M10 likes3.3k downloads1y agoHugging Face07semran1 /dclm-stem-filteredtext1M<n<10M0 likes3.3k downloads1y agoHugging Face08common-pile /uspto_filtered USPTO Description In the United States, patent documents are released into the public domain as government works. Patents follow a highly standardized format with distinct required sections for background, detailed description, and claims. We include patents from the US Patents and Trademark Office (USPTO) as provided by the Google Patents Public Data dataset, which includes millions of granted patents and published patent applications dating back to 1782. We processed… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/uspto_filtered.texttext-generation10M<n<100M4 likes3k downloads1y agoHugging Face09common-pile /wikimedia_filtered Wikimedia Description Official Wikimedia wikis are released under a CC BY-SA license. We downloaded the official database dumps from March 2025 of the English-language wikis that are directly managed by the Wikimedia Foundation. These database dumps include the wikitext—MediaWiki’s custom markup language—for each page as well as talk pages, where editors discuss changes made to a page. We only use the most recent version of each page. We converted wikitext to plain text… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/wikimedia_filtered.texttext-generation10M<n<100M8 likes2.6k downloads1y agoHugging Face10adameubanks /filtered_articles_by_year Dataset Card for Filtered Articles by Year Dataset Summary The Filtered Articles by Year dataset contains yearly-segmented web articles from the FineWeb dataset, specifically filtered and processed for temporal language analysis and Word2Vec model training. This dataset spans 21 years (2005-2025) and serves as the foundation for research into semantic change, concept emergence, and language evolution over time. Supported Tasks and Leaderboards This dataset… See the full description on the dataset page: https://huggingface.co/datasets/adameubanks/filtered_articles_by_year.texttext-generation10M<n<100M1 likes2.6k downloads1y agoHugging Face11common-pile /project_gutenberg_filtered Project Gutenberg Description Project Gutenberg is an online collection of over 75,000 digitized books available as plain text. We use all books that are 1) English and 2) marked as in the Public Domain according to the provided metadata. Additionally, we include any books that are part of the PG19 dataset, which only includes books that are over 100 years old. Minimal preprocessing is applied to remove the Project Gutenberg header and footers, but many scanned books… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/project_gutenberg_filtered.texttext-generation10K<n<100K7 likes2.6k downloads1y agoHugging Face120xDing /wikipedia-cn-20230720-filtered本数据集基于中文维基2023年7月20日的dump存档。作为一项以数据为中心的工作,本数据集仅保留了 254,547条 质量较高的词条内容。具体而言: 过滤了Template, Category, Wikipedia, File, Topic, Portal, MediaWiki, Draft, Help等特殊类型的词条 使用启发式的方法和自有的NLU模型过滤了一部分质量较低的词条 过滤了一部分内容较为敏感或存在争议性的词条。 进行了简繁转换和习惯用词转换,确保符合中国大陆地区的习惯用词。 This dataset is based on the Chinese Wikipedia dump archive from July 20th, 2023. As a data-centric effort, the dataset retains 254,574 high-quality entries. Specifically: Entries of special types such as Template, Category, Wikipedia, File, Topic… See the full description on the dataset page: https://huggingface.co/datasets/0xDing/wikipedia-cn-20230720-filtered.texttext-generation100K<n<1M173 likes2.3k downloads3y agoHugging Face13common-pile /pubmed_filtered PubMed Description PubMed Central is an open-access archive of biomedical and life sciences research papers maintained by the U.S. National Institutes of Health’s National Library of Medicine. We collected papers from PubMed whose metadata indicated that the publishing journal had designated a CC BY, CC BY-SA, or CC0 license. PubMed stores the text content of each article as a single nXML file, which we convert to markdown using pandoc. Per-document license information is… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/pubmed_filtered.texttext-generation1M<n<10M5 likes2k downloads1y agoHugging Face14malaysia-ai /mosaic-dedup-text-dataset-filtered Mosaic format for filtered dedup text dataset to train Malaysian LLM This repository is to store dataset shards using mosaic format. prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-dedup-text-dataset-filtered-4096.ipynb using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer 4096 context length. how-to git clone, git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset load it… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset-filtered.textn<1K0 likes1.9k downloads3y agoHugging Face15Ba2han /vngrs-filteredtext10M<n<100M0 likes1.9k downloads3mo agoHugging Face16if001 /oscar_2023_filteredfrom datasets import load_dataset ds=load_dataset("if001/oscar_2023_filtered") ds['train'] --- Dataset({ features: ['text'], num_rows: 312396 }) oscar 2023をfilterしたものhttps://huggingface.co/datasets/oscar-corpus/OSCAR-2301 詳細はコードを参照https://github.com/if001/HojiChar_OSCAR_sample/tree/0.0.4 texttext-generation1M<n<10M3 likes1.6k downloads3y agoHugging Face17common-pile /pre_1929_books_filtered Pre-1929 Books Description Books published in the US before 1929 passed into the public domain on January 1, 2024. We used the bibliographic catalog Hathifiles produced by HathiTrust to identify digitized books which were published in the US before 1929. The collection contains over 130,000 books digitized and processed by the Internet Archive on behalf of HathiTrust member libraries. The OCR plain text files were downloaded directly from the Internet Archive website.… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/pre_1929_books_filtered.texttext-generation100K<n<1M2 likes1.5k downloads1y agoHugging Face18common-pile /usgpo_filtered USGPO Description The United States Government Publishing Office (USGPO) is a federal agency responsible for disseminating official documents authored by the U.S. government. This dataset includes all plain-text documents made available through the USGPO’s GovInfo.gov developer API. This collection comprises over 2.7 million documents, spanning issues of the Federal Register, congressional hearing transcripts, budget reports, economic indicators, and other federal… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/usgpo_filtered.texttext-generation1M<n<10M3 likes1.1k downloads1y agoHugging Face19common-pile /caselaw_access_project_filtered Caselaw Access Project Description This dataset contains 6.7 million cases from the Caselaw Access Project and Court Listener. The Caselaw Access Project consists of nearly 40 million pages of U.S. federal and state court decisions and judges’ opinions from the last 365 years. In addition, Court Listener adds over 900 thousand cases scraped from 479 courts. The Caselaw Access Project and Court Listener source legal data from a wide variety of resources such as the Harvard… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/caselaw_access_project_filtered.texttext-generation1M<n<10M12 likes1.1k downloads1y agoHugging Face20common-pile /library_of_congress_filtered Library of Congress Description The Library of Congress (LoC) curates a collection of public domain books called "Selected Digitized Books". We have downloaded over 130,000 English-language books from this public domain collection as OCR plain text files using the LoC APIs. Dataset Statistics Documents UTF-8 GB 129,052 35.6 License Issues While we aim to produce datasets with completely accurate licensing information, license… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/library_of_congress_filtered.texttext-generation100K<n<1M2 likes1k downloads1y agoHugging Face21birdsql /bird23-train-filtered BIRD-SQL Train (Filtered) A high-quality subset of the original BIRD train split for text-to-SQL finetuning. Overview Over the past year the community has shared many observations about data quality in BIRD. We performed a rigorous data quality check process to retain examples that are consistent with schema and faithfully answer the question. The resulting set keeps 6,601 instances out of 9,428 (≈70%), and serves as a drop-in replacement for training. Original Train: 9… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/bird23-train-filtered.texttable-question-answering1K<n<10K7 likes972 downloads1y agoHugging Face22common-pile /biodiversity_heritage_library_filtered Biodiversity Heritage Library Description The Biodiversity Heritage Library (BHL) is an open-access digital library for biodiversity literature and archives. This dataset contains over 15 million public domain books and documents from the BHL collection. These works were collected using the bulk data download interface provided by the BHL and were filtered based on their associated license metadata. We use the optical character recognition (OCR)-generated text distributed… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/biodiversity_heritage_library_filtered.texttext-generation10M<n<100M2 likes829 downloads1y agoHugging Face23common-pile /uk_hansard_filtered UK Hansard Description Hansard represents the official record of parliamentary proceedings across the United Kingdom’s legislative bodies. This dataset incorporates records from multiple sources, including debates and written answers from the UK Commons and Lords, devolved legislatures (Scottish Parliament, Senedd in both English and Welsh, Northern Ireland Assembly), London Mayor’s Questions, and ministerial statements. Data was sourced from ParlParse, covering Commons… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/uk_hansard_filtered.texttext-generation10K<n<100K1 likes778 downloads1y agoHugging Face24common-pile /youtube_filtered Creative Commons YouTube Description YouTube is a large-scale video-sharing platform where users have the option of uploading content under a CC BY license. To collect high-quality speech-based textual content and combat the rampant license laundering on YouTube, we manually curated a set of over 2,000 YouTube channels that consistently release original openly licensed content containing speech. The resulting collection spans a wide range of genres, including lectures… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/youtube_filtered.texttext-generation100K<n<1M6 likes722 downloads1y agoHugging Face25geniacllm /CulturaY-ja-askllm-filtered-deduped-v1 CulturaY-ja-askllm-filtered-deduped-v1 多言語データセット ontocord/CulturaY の日本語パート ja に対して、 Ask-LLM 手法でスコア付けしたデータセットです。 元データセットのカラムに加え askllm_score というカラムが追加されており、ここに Ask-LLM のスコアが格納されています。 スコア付けした後、フィルタと重複削除をしたデータセットです。 Ask-LLM でスコア付けに使用した LLM は Rakuten/RakutenAI-7B-instruct で、プロンプトは以下の通りです。 ### {data} ### Does the previous paragraph demarcated within ### and ### contain informative signal for pre-training a large-language model? An informative datapoint should be well-formatted, contain… See the full description on the dataset page: https://huggingface.co/datasets/geniacllm/CulturaY-ja-askllm-filtered-deduped-v1.text10M<n<100M0 likes706 downloads2y agoHugging Face26common-pile /github_archive_filtered GitHub Archive Description According to GitHub’s terms of service, issues and pull request descriptions—along with their comments—inherit the license of their associated repository. To collect this data, we used the GitHub Archive’s public BigQuery table of events to extract all issue, pull request, and comment events since 2011 and aggregated them into threads. The table appeared to be missing “edit” events so the text from each comment is the original from when it was… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive_filtered.texttext-generation10M<n<100M2 likes676 downloads1y agoHugging Face27common-pile /cccc_filtered Creative Commons Common Crawl Description This dataset contains text from 52 Common Crawl snapshots, covering about half of Common Crawl snapshots available to date and covering all years of operations of Common Crawl up to 2024. We found a higher level of duplication across this collection, suggesting that including more snapshots would lead to a modest increase in total token yield. From these snapshots, we extract HTML content using FastWarc. Then, using a regular… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/cccc_filtered.texttext-generation10M<n<100M2 likes642 downloads1y agoHugging Face28common-pile /regulations_filtered Regulations.Gov Description Regulations.gov is an online platform operated by the U.S. General Services Administration that collates newly proposed rules and regulations from federal agencies along with comments and feedback from the general public. This dataset includes all plain-text regulatory documents published by a variety of U.S. federal agencies on this platform, acquired via the bulk download interface provided by Regulations.gov. These agencies include the… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/regulations_filtered.texttext-generation100K<n<1M0 likes622 downloads1y agoHugging Face29common-pile /arxiv_abstracts_filtered ArXiv Abstracts Description Each paper uploaded to ArXiv includes structured metadata fields, including an abstract summarizing the paper’s findings and contributions. According to ArXiv’s licensing policy, the metadata for any paper submitted to ArXiv is distributed under the CC0 license, regardless of the license of the paper itself. Thus, this dataset contains the abstract for every paper submitted to ArXiv through late 2024. We source the abstracts from ArXiv’s API… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/arxiv_abstracts_filtered.texttext-generation1M<n<10M9 likes582 downloads10mo agoHugging Face30semran1 /dclm_shard_1_filtered_realtext1M<n<10M0 likes568 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.