CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01thoshith /hindi-english-raw-text-corpus-uncleanedtext100M<n<1B0 likes800 downloads2y agoHugging Face02zomi-language-corpora /raw-text-corpus 📝 Zomi Raw Text Corpus (Community-Contributed) The Zomi Raw Text Corpus is an open, community-driven collection of unprocessed Zomi-language text.It serves as the foundational dataset for building the full Zomi NLP ecosystem, including tokenizers, language models, ASR/TTS systems, and downstream tasks. This dataset is intentionally raw — no normalization, deduplication, or cleaning is applied.Cleaned and task-specific datasets will be released separately. 📦 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/zomi-language-corpora/raw-text-corpus.texttext-generationn<1K0 likes647 downloads4mo agoHugging Face03textcleanlm /essentialweb-1.0-10B-raw-contenttext1M<n<10M0 likes345 downloads11mo agoHugging Face04ysn-rfd /Fibonacci-Pre_Train-Persian-Corpus-Raw-Texts-DatasetDocument Version: 1.0.8 | Last Updated: 01/02/2025 text10M<n<100M3 likes223 downloads9mo agoHugging Face05Riddwam /religious-texts-rawtext10M<n<100M0 likes166 downloads4mo agoHugging Face06catallama /Catalan-Raw-Text Dataset Summary The Catalan Raw Text Dataset is a subset of the projecte-aina/catalan_general_crawling. It is licensed under a Creative Commons Attribution 4.0 International license, just like the origin dataset. The dataset consists of 404k samples (roughly 20% of the original), totalling 331M tokens after tokenizing it with the Llama-3 Tokenizer. Languages The dataset is in Catalan (ca-ES). Data Fields text (str): Text. Data Splits… See the full description on the dataset page: https://huggingface.co/datasets/catallama/Catalan-Raw-Text.textfill-mask100K<n<1M0 likes88 downloads2y agoHugging Face07Deep-Research-Team /Pre-Training-Persian-Corpus-Raw-Texts-DatasetDocument Version: 2.0.0 | Last Updated: 02/13/2026 text10M<n<100M0 likes86 downloads7mo agoHugging Face08nphearum /khmer-raw-text-3M-v2 Dataset Card for nphearum/khmer-raw-text-3M-v2 Dataset Summary nphearum/khmer-raw-text-3M-v2 is a large-scale raw text corpus containing approximately 200_000 completed records with 3 million text segments in Khmer, curated for large language model (LLM) pre-training, continued pre-training, and domain adaptation. The dataset emphasizes Khmer-language coverage, a historically underrepresented low-resource language, while retaining bilingual context for cross-lingual… See the full description on the dataset page: https://huggingface.co/datasets/nphearum/khmer-raw-text-3M-v2.texttext-classification100K<n<1M1 likes64 downloads5mo agoHugging Face09PersianAICommunity /Fibonacci-Pre_Train-Persian-Corpus-Raw-Texts-Datasettext10M<n<100M0 likes41 downloads9mo agoHugging Face10Gill-Hack-25-UdeM /raw_text_hack_2025tabular100K<n<1M0 likes36 downloads1y agoHugging Face11TokenBender /glaive_coder_raw_texttext100K<n<1M2 likes32 downloads3y agoHugging Face12vanwdai /raw_text_ocr_texttext100K<n<1M0 likes32 downloads1y agoHugging Face13fibonacciai /Fibonacci-Pre_Train-Persian-Corpus-Raw-Texts-Datasettext10M<n<100M0 likes29 downloads9mo agoHugging Face14ariabi /stormfront_incels-raw-textgated Dataset Card for Stormfront & Incels Raw Text Dataset Summary This dataset contains raw, unannotated textual posts from two online extremist platforms: Stormfront (white supremacist) and Incels.is (misogynistic). Each post is provided as a single line of text in .txt files, with no metadata. This simplified format supports unsupervised tasks such as domain adaptation, masked language modeling, and linguistic analysis of extremist cryptolects. The dataset was used in:… See the full description on the dataset page: https://huggingface.co/datasets/ariabi/stormfront_incels-raw-text.text10M<n<100M3 likes27 downloads1y agoHugging Face15nphearum /khmer-raw-text-3M Dataset Card for nphearum/khmer-raw-text-3M Dataset Summary nphearum/khmer-raw-text-3M is a large-scale raw text corpus containing approximately 50000 completed records with 3 million text segments in Khmer, curated for large language model (LLM) pre-training, continued pre-training, and domain adaptation. The dataset emphasizes Khmer-language coverage, a historically underrepresented low-resource language, while retaining bilingual context for cross-lingual learning.… See the full description on the dataset page: https://huggingface.co/datasets/nphearum/khmer-raw-text-3M.texttext-classification10K<n<100K0 likes26 downloads8mo agoHugging Face16riotu-lab /ARABIC-RAW-TEXTDataset: Aluka 1.4 GB AraWiki 3.9 GB Aya 22.5 GB Islamic Books 21.4 GB texttext-generation100M<n<1B5 likes22 downloads2y agoHugging Face17adrianf12 /ori_raw_text Ori Raw Text (Single Row) Single-row JSONL with the concatenated Ori documentation content (cleaned paragraphs). Schema Each row: { "text": "" } Usage from datasets import load_dataset ds = load_dataset(adrianf12/ori_raw_text) print(ds[train][0][text][:500]) textn<1K0 likes21 downloads11mo agoHugging Face18aksw /Text2SPARQL-Raw Dataset Card for Text2sparql-Raw 🧾 Dataset Summary Text2Sparql-Raw is a multilingual dataset designed for the task of translating natural language questions into SPARQL queries over the DBpedia knowledge graph. This dataset aggregates and harmonizes four widely used benchmarks in the text-to-SPARQL domain: QALD (versions 1–9) LC-QuAD 1.0 Orange/paraqa-sparqltotext julioc-p/Question-Sparql It contains questions in both English and Spanish, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/aksw/Text2SPARQL-Raw.text10K<n<100K0 likes18 downloads1y agoHugging Face19yrrhall /ARABIC-RAW-TEXTDataset: Aluka 1.4 GB AraWiki 3.9 GB Aya 22.5 GB Islamic Books 21.4 GB texttext-generation100M<n<1B0 likes17 downloads4mo agoHugging Face20Imsidag-community /kabyle-raw-texttext100K<n<1M1 likes14 downloads11mo agoHugging Face21ysnrfd2 /Fibonacci-Pre_Train-Persian-Corpus-Raw-Texts-DatasetDocument Version: 1.0.5 | Last Updated: 01/02/2025 text10M<n<100M0 likes14 downloads9mo agoHugging Face22maximilianshwarzmullers /turkmen_raw_texttext1M<n<10M0 likes12 downloads1y agoHugging Face23TokenBender /roleplay_raw_texttext10K<n<100K2 likes11 downloads3y agoHugging Face24dliu1 /legal-llama-raw-texttext10K<n<100K0 likes11 downloads3y agoHugging Face25bany1111 /KorPE_raw_texttext100K<n<1M0 likes11 downloads2y agoHugging Face26Aarif123 /raw-text-dataset-2textn<1K0 likes10 downloads2y agoHugging Face27Cartinoe5930 /raw_text_synthetic_dataset_50ktext10K<n<100K5 likes10 downloads2y agoHugging Face28Aarif123 /raw-text-datasettextn<1K0 likes9 downloads2y agoHugging Face29Andrei481 /raw-text-corpus-ro-8ktext100K<n<1M0 likes8 downloads2y agoHugging Face30Realrobot /Fibonacci-Pre_Train-Persian-Corpus-Raw-Texts-Datasettext10M<n<100M0 likes8 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.