CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01vngrs /vngrs-web-corpus Dataset Card for Dataset Name vngrs-web-corpus is a mixed-dataset made of cleaned Turkish sections of OSCAR-2201 and mC4. This dataset is originally created for training VBART and later used for training TURNA. The cleaning procedures of this dataset are explained in Appendix A of the VBART Paper. It consists of 50.3M pages and 25.33B tokens when tokenized by VBART Tokenizer. Dataset Details Uses vngrs-web-corpus is mainly intended to pretrain… See the full description on the dataset page: https://huggingface.co/datasets/vngrs/vngrs-web-corpus.text10M<n<100M27 likes660 downloads2y agoHugging Face02bluelightai-dev /common-corpus-sample-open-webtabular1M<n<10M0 likes508 downloads11mo agoHugging Face03saidutta69 /Odia-Web-Corpus-v5 Odia Web Corpus v5 The largest and cleanest Odia text corpus to date — 4.16 million deduplicated documents, 7.74 GB. Built by merging and thoroughly cleaning four source collections. Dataset Details Language: Odia (ISO 639-3: or) Format: 28 sharded Parquet files Total Size: 7.74 GB Total Documents: 4,162,804 License: CC-BY-SA-4.0 Cleaning Pipeline Stage Removed Description Deduplication 30.2% Exact MD5 hash match Short lines… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v5.texttext-generation1M<n<10M0 likes420 downloads9d agoHugging Face04sphita /Intel-WebCorpus-forms 💻 Intel WebCorpus Forms (Enterprise Hardware Q&A) This dataset is a massive, high-fidelity archive of 176,472 technical troubleshooting discussions (containing nearly 1 million individual messages) scraped from the official Intel Community Forums. It has been meticulously engineered for Large Language Model (LLM) training. Instead of a raw, messy dump of isolated posts, the data has been reconstructed into chronological conversation threads, noise-filtered, deduplicated, and… See the full description on the dataset page: https://huggingface.co/datasets/sphita/Intel-WebCorpus-forms.textquestion-answering100K<n<1M3 likes379 downloads10h agoHugging Face05paras9909 /opticparse-150-template-web-corpus ⚡ OpticParse: 150-Template Web Intelligence & Ground-Truth Corpus Official high-signal web extraction corpus compiled from the OpticParse Multimodal Vision Scraper & PhishVision Threat Sentinel. ⭐️ Support Open-Source AI Tooling: If you find this dataset or the OpticParse scraper useful for your AI agents, please click the Like (❤️) button above to support continuous daily Parquet updates! 💳 Commercial Subscription Tiers & Live Continuous Streams ⚡ Need… See the full description on the dataset page: https://huggingface.co/datasets/paras9909/opticparse-150-template-web-corpus.texttabular-classificationn<1K0 likes219 downloads2d agoHugging Face06saidutta69 /Odia-Web-Corpus-v2 Odia Web Corpus v2 Expanded and refined version of the Odia web corpus, converted to Parquet format with dedicated train/test/validation splits for reproducible NLP experimentation. Dataset Details Language: Odia (ISO 639-3: or) Format: Parquet (columnar, compressed) Splits: train (900K), test (50K), validation (50K) License: CC-BY-4.0 Data Fields Field Type Description text string Cleaned document text Usage… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v2.texttext-generation1M<n<10M0 likes137 downloads9d agoHugging Face07saidutta69 /Odia-Web-Corpus-v1 Odia Web Corpus v1 The inaugural release of a curated Odia (Oriya) web text corpus. Scraped and filtered from publicly accessible web sources to support Odia language modeling and NLP research. Dataset Details Language: Odia (Oriya, ISO 639-3: ory) Format: JSONL (one JSON object per line) Size: ~650K documents, ~0.9 GB text License: CC-BY-4.0 Data Fields Field Type Description text string Cleaned document body title string… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v1.texttext-generation100K<n<1M0 likes128 downloads9d agoHugging Face08saidutta69 /Odia-Web-Corpus-v4 Odia Web Corpus v4 Fourth-generation Odia corpus featuring both pretraining data (deduplicated, quality-filtered web text) and instruction-tuning data formatted in ChatML. Built by merging and enhancing v1–v3. Dataset Details Language: Odia (ISO 639-3: or) Format: JSONL.GZ (gzip-compressed JSON lines) License: CC-BY-SA-4.0 Data Composition Split Description Examples pretrain_train Pretraining corpus (train) ~900K… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v4.text-generation1M<n<10M0 likes97 downloads9d agoHugging Face09Jazlynn0095 /corpus_web_multimodal_017 Web Content Dataset This dataset contains processed web content extracted from HTML pages. Dataset Structure Each JSON file contains: url_id: A unique identifier for the URL text: The extracted text content from the HTML metadata: Additional metadata including title, description, and URL when available Usage This dataset can be used for training language models, information extraction, or web content analysis. 0 likes90 downloads1y agoHugging Face10saidutta69 /Odia-Web-Corpus-v3 Odia Web Corpus v3 Third iteration of the Odia web corpus with enhanced deduplication, quality filtering, and standardized Parquet splits for pretraining and evaluation. Dataset Details Language: Odia (ISO 639-3: or) Format: Parquet (columnar, compressed) Splits: train (900K), validation (50K), test (50K) License: CC-BY-SA-4.0 Data Fields Field Type Description text string Cleaned and filtered document text… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v3.texttext-generation100K<n<1M0 likes89 downloads9d agoHugging Face11canbingol /vngrs-web-corpus-500ktext100K<n<1M0 likes58 downloads11mo agoHugging Face12TajikNLPWorld /tajik-web-corpusgated Dataset Card for Tajik Web Corpus Dataset Details Dataset Description The Tajik Web Corpus is a large-scale collection of 319,298 documents in the Tajik language, totaling approximately 1.11 billion characters and 168.5 million words. The data has been cleaned, normalized, and deduplicated, and is provided in JSONL format with the following fields: title, text, category, source, date, and URL. It covers various domains including news, Wikipedia… See the full description on the dataset page: https://huggingface.co/datasets/TajikNLPWorld/tajik-web-corpus.texttext-classification100K<n<1M0 likes58 downloads25d agoHugging Face13BashkirNLPWorld /bashkir-web-corpusgated Dataset Card for Bashkir Web Corpus Dataset Details Dataset Description The Bashkir Web Corpus is a collection of 71,567 documents and approximately 46.9 million tokens in the Bashkir language (a Turkic language spoken in Bashkortostan, Russia). The corpus was compiled from 16 Bashkir‑language online sources, including news websites, literary magazines, social media, books, and Wikipedia. It is designed for language modeling, text classification… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-web-corpus.texttext-generation10K<n<100K0 likes43 downloads25d agoHugging Face14TatarNLPWorld /tatar-web-corpusgated Dataset Card for Tatar Web Corpus Dataset Details Dataset Description The largest open corpus for the Tatar language with over 1 million documents collected from news websites, social media, articles, books, and Wikipedia. Designed for various NLP tasks including language modeling, text classification, information extraction, and search. Curated by: TatarNLPWorld Community Language(s) (NLP): Tatar (tt) License: other – see Licensing & Legal Notice… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-web-corpus.texttext-classification1M<n<10M0 likes34 downloads25d agoHugging Face15canbingol /vngrs-web-corpus-2Mtext1M<n<10M0 likes32 downloads11mo agoHugging Face16OssetianNLPWorld /ossetian-web-corpusgated Dataset Card for Ossetian Web Corpus Dataset Details Dataset Description The Ossetian Web Corpus is a comprehensive collection of Ossetian-language texts gathered from three main sources: online news portals, Wikipedia articles, and digitized books. The corpus is designed to support natural language processing (NLP) research and development for the Ossetian language, a low-resource language spoken in the Caucasus region. Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/OssetianNLPWorld/ossetian-web-corpus.texttext-generation10K<n<100K0 likes26 downloads25d agoHugging Face17TatarNLPWorld /tatar-web-corpus-v3gated Dataset Card for Tatar Web Corpus Dataset Details Dataset Description The Tatar Web Corpus is the largest open-source corpus of the Tatar language (Turkic family), containing 2,465,867 documents (approximately 251 million tokens) collected from publicly available web sources. It covers news portals, social media, blogs, literary websites, and other domains. The corpus underwent soft deduplication to remove exact duplicates while preserving… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-web-corpus-v3.texttext-generation1M<n<10M0 likes25 downloads25d agoHugging Face18maanka2 /somali-web-corpus SOMALI-WEB-CORPUS V1 This dataset consists of clean, structured, and filtered Somali language text compiled from various online sources. It is designed for training and fine-tuning Somali language models (LLMs) and supporting natural language processing (NLP) research for the Somali language. Dataset Details Language: Somali (so) Format: JSON lines (.jsonl) Data Structure: Each record has a single text field containing a cleaned paragraph. Sources… See the full description on the dataset page: https://huggingface.co/datasets/maanka2/somali-web-corpus.texttext-generation100K<n<1M1 likes25 downloads4mo agoHugging Face19canbingol /vngrs-web-corpus-500k-kumru_tokenizer-tokenized0 likes12 downloads8mo agoHugging Face20SPAISS6F1 /spai-ss6-corpus-medical-health-web SPAI SS6 Thai Medical Health Web Corpus Thai public medical and health web articles collected by the local scraping pipeline. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: default Rows in canonical config: 3,660 Parquet size in canonical config: 0.01 GB Source license: other… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-medical-health-web.tabulartext-generationn<1K0 likes11 downloads4mo agoHugging Face21gval0 /georgian-web-corpus-nlp4tabularn<1K1 likes10 downloads1y agoHugging Face22canbingol /vngrs-web-corpus-200k-kumru_tokenizer-tokenized0 likes8 downloads8mo agoHugging Face23Geonwoohong /aihub-webcorpus-morph-train-tokenized-ko Dataset Description This dataset contains the AIHub Korean Web Corpus cleaned and morphologically annotated with the word identifier.Each record stores morpheme-level tokens and two subsets: semantic and stylistic. Dataset Structure Data Fields Field Type Description text string Original sentence reconstructed from morphemes semantic list[dict] Subset of content-bearing morphemes stylistic list[dict] Subset of grammatical/stylistic morphemes… See the full description on the dataset page: https://huggingface.co/datasets/Geonwoohong/aihub-webcorpus-morph-train-tokenized-ko.texttoken-classification1M<n<10M0 likes7 downloads11mo agoHugging Face24canbingol /vngrs-web-corpus-200ktext100K<n<1M0 likes7 downloads11mo agoHugging Face25gval0 /nlp-georgian-web-corpustabularn<1K0 likes6 downloads1y agoHugging Face26gval0 /nlp4-georgian-web-corpustabularn<1K0 likes6 downloads1y agoHugging Face27SPAISS6F1 /spai-ss6-corpus-wangchanlion-web SPAI SS6 WangchanLION Web Corpus Index Index repo for the WangchanLION-Web corpus mirrored in the canonical repo. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: wangchanlion_web Rows in canonical config: 557,502 Parquet size in canonical config: 1.97 GB Source license: odc-by… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-wangchanlion-web.tabulartext-generationn<1K0 likes6 downloads4mo agoHugging Face28pranjalsharma /coding-web-corpus-v1textn<1K0 likes5 downloads8mo agoHugging Face29TilQazyna /kazakh-web-corpus-psfgated kazakh-web-corpus-psf Қазақ тіліндегі академиялық веб-материалдар · Академические веб-материалы на казахском языке · Kazakh academic web material Қазақша · Русский · English Қазақша kazakh-web-corpus-psf — қазақ тіліндегі академиялық мақалалар мен оларды жинауға арналған материалдар корпусы, көлемі 632.3 МБ. Файлдар тақырыптар бойынша ұйымдастырылған және тіл моделіне мәтін дайындауға бастапқы дерек бола алады. Құрамы Тақырыптар… See the full description on the dataset page: https://huggingface.co/datasets/TilQazyna/kazakh-web-corpus-psf.documenttext-generationn<1K0 likes4 downloads2mo agoHugging Face30stukenov /sozkz-corpus-dedup-kk-web-v1gated sozkz-corpus-dedup-kk-web-v1 Deduplicated Kazakh web text corpus collected from 6 public HuggingFace datasets. Contains only texts not present in kz-transformers/multidomain-kazakh-dataset (12.4M texts were used as the dedup reference). Stats Field Value Total unique texts 9,475,089 Format Parquet (142 shards) Columns text, source Dedup method MD5 hash (exact match) Dedup reference kz-transformers/multidomain-kazakh-dataset (12.4M hashes) Date… See the full description on the dataset page: https://huggingface.co/datasets/stukenov/sozkz-corpus-dedup-kk-web-v1.text1M<n<10M0 likes3 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.