web-corpus
vngrs-web-corpus
Dataset Card for Dataset Name
vngrs-web-corpus is a mixed-dataset made of cleaned Turkish sections of OSCAR-2201 and mC4.
This dataset is originally created for training VBART and later used for training TURNA.
The cleaning procedures of this dataset are explained in Appendix A of the VBART Paper.
It consists of 50.3M pages and 25.33B tokens when tokenized by VBART Tokenizer.
Dataset Details
Uses
vngrs-web-corpus is mainly intended to pretrain… See the full description on the dataset page: https://huggingface.co/datasets/vngrs/vngrs-web-corpus.common-corpus-sample-open-webOdia-Web-Corpus-v5
Odia Web Corpus v5
The largest and cleanest Odia text corpus to date — 4.16 million deduplicated documents, 7.74 GB. Built by merging and thoroughly cleaning four source collections.
Dataset Details
Language: Odia (ISO 639-3: or)
Format: 28 sharded Parquet files
Total Size: 7.74 GB
Total Documents: 4,162,804
License: CC-BY-SA-4.0
Cleaning Pipeline
Stage
Removed
Description
Deduplication
30.2%
Exact MD5 hash match
Short lines… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v5.Intel-WebCorpus-forms
💻 Intel WebCorpus Forms (Enterprise Hardware Q&A)
This dataset is a massive, high-fidelity archive of 176,472 technical troubleshooting discussions (containing nearly 1 million individual messages) scraped from the official Intel Community Forums.
It has been meticulously engineered for Large Language Model (LLM) training. Instead of a raw, messy dump of isolated posts, the data has been reconstructed into chronological conversation threads, noise-filtered, deduplicated, and… See the full description on the dataset page: https://huggingface.co/datasets/sphita/Intel-WebCorpus-forms.opticparse-150-template-web-corpus
⚡ OpticParse: 150-Template Web Intelligence & Ground-Truth Corpus
Official high-signal web extraction corpus compiled from the OpticParse Multimodal Vision Scraper & PhishVision Threat Sentinel.
⭐️ Support Open-Source AI Tooling: If you find this dataset or the OpticParse scraper useful for your AI agents, please click the Like (❤️) button above to support continuous daily Parquet updates!
💳 Commercial Subscription Tiers & Live Continuous Streams
⚡ Need… See the full description on the dataset page: https://huggingface.co/datasets/paras9909/opticparse-150-template-web-corpus.Odia-Web-Corpus-v2
Odia Web Corpus v2
Expanded and refined version of the Odia web corpus, converted to Parquet format with dedicated train/test/validation splits for reproducible NLP experimentation.
Dataset Details
Language: Odia (ISO 639-3: or)
Format: Parquet (columnar, compressed)
Splits: train (900K), test (50K), validation (50K)
License: CC-BY-4.0
Data Fields
Field
Type
Description
text
string
Cleaned document text
Usage… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v2.
