CoolFace
13 results

web-corpus

vngrs /vngrs-web-corpus Dataset Card for Dataset Name vngrs-web-corpus is a mixed-dataset made of cleaned Turkish sections of OSCAR-2201 and mC4. This dataset is originally created for training VBART and later used for training TURNA. The cleaning procedures of this dataset are explained in Appendix A of the VBART Paper. It consists of 50.3M pages and 25.33B tokens when tokenized by VBART Tokenizer. Dataset Details Uses vngrs-web-corpus is mainly intended to pretrain… See the full description on the dataset page: https://huggingface.co/datasets/vngrs/vngrs-web-corpus.text10M<n<100M27 likes660 downloads2y agoHugging Facebluelightai-dev /common-corpus-sample-open-webtabular1M<n<10M0 likes508 downloads11mo agoHugging Facesaidutta69 /Odia-Web-Corpus-v5 Odia Web Corpus v5 The largest and cleanest Odia text corpus to date — 4.16 million deduplicated documents, 7.74 GB. Built by merging and thoroughly cleaning four source collections. Dataset Details Language: Odia (ISO 639-3: or) Format: 28 sharded Parquet files Total Size: 7.74 GB Total Documents: 4,162,804 License: CC-BY-SA-4.0 Cleaning Pipeline Stage Removed Description Deduplication 30.2% Exact MD5 hash match Short lines… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v5.texttext-generation1M<n<10M0 likes420 downloads9d agoHugging Facesphita /Intel-WebCorpus-forms 💻 Intel WebCorpus Forms (Enterprise Hardware Q&A) This dataset is a massive, high-fidelity archive of 176,472 technical troubleshooting discussions (containing nearly 1 million individual messages) scraped from the official Intel Community Forums. It has been meticulously engineered for Large Language Model (LLM) training. Instead of a raw, messy dump of isolated posts, the data has been reconstructed into chronological conversation threads, noise-filtered, deduplicated, and… See the full description on the dataset page: https://huggingface.co/datasets/sphita/Intel-WebCorpus-forms.textquestion-answering100K<n<1M3 likes379 downloads9h agoHugging Faceparas9909 /opticparse-150-template-web-corpus ⚡ OpticParse: 150-Template Web Intelligence & Ground-Truth Corpus Official high-signal web extraction corpus compiled from the OpticParse Multimodal Vision Scraper & PhishVision Threat Sentinel. ⭐️ Support Open-Source AI Tooling: If you find this dataset or the OpticParse scraper useful for your AI agents, please click the Like (❤️) button above to support continuous daily Parquet updates! 💳 Commercial Subscription Tiers & Live Continuous Streams ⚡ Need… See the full description on the dataset page: https://huggingface.co/datasets/paras9909/opticparse-150-template-web-corpus.texttabular-classificationn<1K0 likes219 downloads2d agoHugging Facesaidutta69 /Odia-Web-Corpus-v2 Odia Web Corpus v2 Expanded and refined version of the Odia web corpus, converted to Parquet format with dedicated train/test/validation splits for reproducible NLP experimentation. Dataset Details Language: Odia (ISO 639-3: or) Format: Parquet (columnar, compressed) Splits: train (900K), test (50K), validation (50K) License: CC-BY-4.0 Data Fields Field Type Description text string Cleaned document text Usage… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v2.texttext-generation1M<n<10M0 likes137 downloads9d agoHugging Face