datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Intel-WebCorpus-forms
💻 Intel WebCorpus Forms (Enterprise Hardware Q&A)
This dataset is a massive, high-fidelity archive of 176,472 technical troubleshooting discussions (containing nearly 1 million individual messages) scraped from the official Intel Community Forums.
It has been meticulously engineered for Large Language Model (LLM) training. Instead of a raw, messy dump of isolated posts, the data has been reconstructed into chronological conversation threads, noise-filtered, deduplicated, and… See the full description on the dataset page: https://huggingface.co/datasets/sphita/Intel-WebCorpus-forms.Odia-Web-Corpus-v4
Odia Web Corpus v4
Fourth-generation Odia corpus featuring both pretraining data (deduplicated, quality-filtered web text) and instruction-tuning data formatted in ChatML. Built by merging and enhancing v1–v3.
Dataset Details
Language: Odia (ISO 639-3: or)
Format: JSONL.GZ (gzip-compressed JSON lines)
License: CC-BY-SA-4.0
Data Composition
Split
Description
Examples
pretrain_train
Pretraining corpus (train)
~900K… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v4.
