datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vngrs-web-corpus
Dataset Card for Dataset Name
vngrs-web-corpus is a mixed-dataset made of cleaned Turkish sections of OSCAR-2201 and mC4.
This dataset is originally created for training VBART and later used for training TURNA.
The cleaning procedures of this dataset are explained in Appendix A of the VBART Paper.
It consists of 50.3M pages and 25.33B tokens when tokenized by VBART Tokenizer.
Dataset Details
Uses
vngrs-web-corpus is mainly intended to pretrain… See the full description on the dataset page: https://huggingface.co/datasets/vngrs/vngrs-web-corpus.common-corpus-sample-open-webOdia-Web-Corpus-v5
Odia Web Corpus v5
The largest and cleanest Odia text corpus to date — 4.16 million deduplicated documents, 7.74 GB. Built by merging and thoroughly cleaning four source collections.
Dataset Details
Language: Odia (ISO 639-3: or)
Format: 28 sharded Parquet files
Total Size: 7.74 GB
Total Documents: 4,162,804
License: CC-BY-SA-4.0
Cleaning Pipeline
Stage
Removed
Description
Deduplication
30.2%
Exact MD5 hash match
Short lines… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v5.Intel-WebCorpus-forms
💻 Intel WebCorpus Forms (Enterprise Hardware Q&A)
This dataset is a massive, high-fidelity archive of 176,472 technical troubleshooting discussions (containing nearly 1 million individual messages) scraped from the official Intel Community Forums.
It has been meticulously engineered for Large Language Model (LLM) training. Instead of a raw, messy dump of isolated posts, the data has been reconstructed into chronological conversation threads, noise-filtered, deduplicated, and… See the full description on the dataset page: https://huggingface.co/datasets/sphita/Intel-WebCorpus-forms.opticparse-150-template-web-corpus
⚡ OpticParse: 150-Template Web Intelligence & Ground-Truth Corpus
Official high-signal web extraction corpus compiled from the OpticParse Multimodal Vision Scraper & PhishVision Threat Sentinel.
⭐️ Support Open-Source AI Tooling: If you find this dataset or the OpticParse scraper useful for your AI agents, please click the Like (❤️) button above to support continuous daily Parquet updates!
💳 Commercial Subscription Tiers & Live Continuous Streams
⚡ Need… See the full description on the dataset page: https://huggingface.co/datasets/paras9909/opticparse-150-template-web-corpus.Odia-Web-Corpus-v2
Odia Web Corpus v2
Expanded and refined version of the Odia web corpus, converted to Parquet format with dedicated train/test/validation splits for reproducible NLP experimentation.
Dataset Details
Language: Odia (ISO 639-3: or)
Format: Parquet (columnar, compressed)
Splits: train (900K), test (50K), validation (50K)
License: CC-BY-4.0
Data Fields
Field
Type
Description
text
string
Cleaned document text
Usage… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v2.Odia-Web-Corpus-v1
Odia Web Corpus v1
The inaugural release of a curated Odia (Oriya) web text corpus. Scraped and filtered from publicly accessible web sources to support Odia language modeling and NLP research.
Dataset Details
Language: Odia (Oriya, ISO 639-3: ory)
Format: JSONL (one JSON object per line)
Size: ~650K documents, ~0.9 GB text
License: CC-BY-4.0
Data Fields
Field
Type
Description
text
string
Cleaned document body
title
string… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v1.Odia-Web-Corpus-v4
Odia Web Corpus v4
Fourth-generation Odia corpus featuring both pretraining data (deduplicated, quality-filtered web text) and instruction-tuning data formatted in ChatML. Built by merging and enhancing v1–v3.
Dataset Details
Language: Odia (ISO 639-3: or)
Format: JSONL.GZ (gzip-compressed JSON lines)
License: CC-BY-SA-4.0
Data Composition
Split
Description
Examples
pretrain_train
Pretraining corpus (train)
~900K… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v4.corpus_web_multimodal_017
Web Content Dataset
This dataset contains processed web content extracted from HTML pages.
Dataset Structure
Each JSON file contains:
url_id: A unique identifier for the URL
text: The extracted text content from the HTML
metadata: Additional metadata including title, description, and URL when available
Usage
This dataset can be used for training language models, information extraction, or web content analysis.
Odia-Web-Corpus-v3
Odia Web Corpus v3
Third iteration of the Odia web corpus with enhanced deduplication, quality filtering, and standardized Parquet splits for pretraining and evaluation.
Dataset Details
Language: Odia (ISO 639-3: or)
Format: Parquet (columnar, compressed)
Splits: train (900K), validation (50K), test (50K)
License: CC-BY-SA-4.0
Data Fields
Field
Type
Description
text
string
Cleaned and filtered document text… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v3.vngrs-web-corpus-500ktajik-web-corpus
Dataset Card for Tajik Web Corpus
Dataset Details
Dataset Description
The Tajik Web Corpus is a large-scale collection of 319,298 documents in the Tajik language, totaling approximately 1.11 billion characters and 168.5 million words. The data has been cleaned, normalized, and deduplicated, and is provided in JSONL format with the following fields: title, text, category, source, date, and URL. It covers various domains including news, Wikipedia… See the full description on the dataset page: https://huggingface.co/datasets/TajikNLPWorld/tajik-web-corpus.bashkir-web-corpus
Dataset Card for Bashkir Web Corpus
Dataset Details
Dataset Description
The Bashkir Web Corpus is a collection of 71,567 documents and approximately 46.9 million tokens in the Bashkir language (a Turkic language spoken in Bashkortostan, Russia). The corpus was compiled from 16 Bashkir‑language online sources, including news websites, literary magazines, social media, books, and Wikipedia. It is designed for language modeling, text classification… See the full description on the dataset page: https://huggingface.co/datasets/BashkirNLPWorld/bashkir-web-corpus.tatar-web-corpus
Dataset Card for Tatar Web Corpus
Dataset Details
Dataset Description
The largest open corpus for the Tatar language with over 1 million documents collected from news websites, social media, articles, books, and Wikipedia. Designed for various NLP tasks including language modeling, text classification, information extraction, and search.
Curated by: TatarNLPWorld Community
Language(s) (NLP): Tatar (tt)
License: other – see Licensing & Legal Notice… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-web-corpus.vngrs-web-corpus-2Mossetian-web-corpus
Dataset Card for Ossetian Web Corpus
Dataset Details
Dataset Description
The Ossetian Web Corpus is a comprehensive collection of Ossetian-language texts gathered from three main sources: online news portals, Wikipedia articles, and digitized books. The corpus is designed to support natural language processing (NLP) research and development for the Ossetian language, a low-resource language spoken in the Caucasus region.
Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/OssetianNLPWorld/ossetian-web-corpus.tatar-web-corpus-v3
Dataset Card for Tatar Web Corpus
Dataset Details
Dataset Description
The Tatar Web Corpus is the largest open-source corpus of the Tatar language (Turkic family), containing 2,465,867 documents (approximately 251 million tokens) collected from publicly available web sources. It covers news portals, social media, blogs, literary websites, and other domains. The corpus underwent soft deduplication to remove exact duplicates while preserving… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-web-corpus-v3.somali-web-corpus
SOMALI-WEB-CORPUS V1
This dataset consists of clean, structured, and filtered Somali language text compiled from various online sources. It is designed for training and fine-tuning Somali language models (LLMs) and supporting natural language processing (NLP) research for the Somali language.
Dataset Details
Language: Somali (so)
Format: JSON lines (.jsonl)
Data Structure: Each record has a single text field containing a cleaned paragraph.
Sources… See the full description on the dataset page: https://huggingface.co/datasets/maanka2/somali-web-corpus.vngrs-web-corpus-500k-kumru_tokenizer-tokenizedspai-ss6-corpus-medical-health-web
SPAI SS6 Thai Medical Health Web Corpus
Thai public medical and health web articles collected by the local scraping pipeline.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: default
Rows in canonical config: 3,660
Parquet size in canonical config: 0.01 GB
Source license: other… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-medical-health-web.georgian-web-corpus-nlp4vngrs-web-corpus-200k-kumru_tokenizer-tokenizedaihub-webcorpus-morph-train-tokenized-ko
Dataset Description
This dataset contains the AIHub Korean Web Corpus cleaned and morphologically annotated with the word identifier.Each record stores morpheme-level tokens and two subsets: semantic and stylistic.
Dataset Structure
Data Fields
Field
Type
Description
text
string
Original sentence reconstructed from morphemes
semantic
list[dict]
Subset of content-bearing morphemes
stylistic
list[dict]
Subset of grammatical/stylistic morphemes… See the full description on the dataset page: https://huggingface.co/datasets/Geonwoohong/aihub-webcorpus-morph-train-tokenized-ko.vngrs-web-corpus-200knlp-georgian-web-corpusnlp4-georgian-web-corpusspai-ss6-corpus-wangchanlion-web
SPAI SS6 WangchanLION Web Corpus Index
Index repo for the WangchanLION-Web corpus mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: wangchanlion_web
Rows in canonical config: 557,502
Parquet size in canonical config: 1.97 GB
Source license: odc-by… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-wangchanlion-web.coding-web-corpus-v1kazakh-web-corpus-psf
kazakh-web-corpus-psf
Қазақ тіліндегі академиялық веб-материалдар · Академические веб-материалы на казахском языке · Kazakh academic web material
Қазақша · Русский · English
Қазақша
kazakh-web-corpus-psf — қазақ тіліндегі академиялық мақалалар мен оларды жинауға арналған материалдар корпусы, көлемі 632.3 МБ. Файлдар тақырыптар бойынша ұйымдастырылған және тіл моделіне мәтін дайындауға бастапқы дерек бола алады.
Құрамы
Тақырыптар… See the full description on the dataset page: https://huggingface.co/datasets/TilQazyna/kazakh-web-corpus-psf.sozkz-corpus-dedup-kk-web-v1
sozkz-corpus-dedup-kk-web-v1
Deduplicated Kazakh web text corpus collected from 6 public HuggingFace datasets. Contains only texts not present in kz-transformers/multidomain-kazakh-dataset (12.4M texts were used as the dedup reference).
Stats
Field
Value
Total unique texts
9,475,089
Format
Parquet (142 shards)
Columns
text, source
Dedup method
MD5 hash (exact match)
Dedup reference
kz-transformers/multidomain-kazakh-dataset (12.4M hashes)
Date… See the full description on the dataset page: https://huggingface.co/datasets/stukenov/sozkz-corpus-dedup-kk-web-v1.
