CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01serda-dev /turkish-raw-text-cleaned Turkish Raw Text Cleaned turkish-raw-text-cleaned, Türkçe dil modeli çalışmaları için hazırlanmış temizlenmiş ham metin veri kümesidir. Veri kümesi, turkish-nlp-suite çatısı altında yayımlanan Türkçe metin kaynaklarının temizlenmesi, filtrelenmesi ve model eğitimine daha uygun hale getirilmesiyle oluşturulmuştur. Bu çalışma özellikle Türkçe LLM ön-eğitimi, continual pre-training (CPT), tokenizer analizi, embedding modeli eğitimi, alan bağımsız Türkçe metin modelleme ve veri… See the full description on the dataset page: https://huggingface.co/datasets/serda-dev/turkish-raw-text-cleaned.text-generation1M<n<10M0 likes11k downloads3mo agoHugging Face02thoshith /hindi-english-raw-text-corpus-uncleanedtext100M<n<1B0 likes761 downloads2y agoHugging Face03zomi-language-corpora /raw-text-corpus 📝 Zomi Raw Text Corpus (Community-Contributed) The Zomi Raw Text Corpus is an open, community-driven collection of unprocessed Zomi-language text.It serves as the foundational dataset for building the full Zomi NLP ecosystem, including tokenizers, language models, ASR/TTS systems, and downstream tasks. This dataset is intentionally raw — no normalization, deduplication, or cleaning is applied.Cleaned and task-specific datasets will be released separately. 📦 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/zomi-language-corpora/raw-text-corpus.texttext-generationn<1K0 likes633 downloads4mo agoHugging Face04textcleanlm /essentialweb-1.0-10B-raw-contenttext1M<n<10M0 likes344 downloads11mo agoHugging Face05chouziel /dataset_sugar_1709_texture-raw0 likes276 downloads5d agoHugging Face06alrope /CompactDS-102GB-raw-text0 likes246 downloads1y agoHugging Face07ysn-rfd /Fibonacci-Pre_Train-Persian-Corpus-Raw-Texts-DatasetDocument Version: 1.0.8 | Last Updated: 01/02/2025 text10M<n<100M3 likes218 downloads9mo agoHugging Face08catallama /Catalan-Raw-Text Dataset Summary The Catalan Raw Text Dataset is a subset of the projecte-aina/catalan_general_crawling. It is licensed under a Creative Commons Attribution 4.0 International license, just like the origin dataset. The dataset consists of 404k samples (roughly 20% of the original), totalling 331M tokens after tokenizing it with the Llama-3 Tokenizer. Languages The dataset is in Catalan (ca-ES). Data Fields text (str): Text. Data Splits… See the full description on the dataset page: https://huggingface.co/datasets/catallama/Catalan-Raw-Text.textfill-mask100K<n<1M0 likes88 downloads2y agoHugging Face09Deep-Research-Team /Pre-Training-Persian-Corpus-Raw-Texts-DatasetDocument Version: 2.0.0 | Last Updated: 02/13/2026 text10M<n<100M0 likes75 downloads7mo agoHugging Face10Riddwam /religious-texts-rawtext10M<n<100M0 likes73 downloads4mo agoHugging Face11tdw419 /sefaria-raw-texts Sefaria Raw Texts (judaism-llm) Raw Sefaria API responses: 6,112 JSON files, 82MB, organized by category. License: CC-BY-NC 3.0 — non-commercial use only. Source: Sefaria, downloaded via the official API (download_sefaria.py in the pipeline repo). Layout Directory names encode the category path (_ separates levels), e.g. Halakhah_Mishneh Torah_Commentary_Ohr Sameach_Sefer Zeraim/. Each JSON file is one Sefaria document with the full API schema: text (English), he… See the full description on the dataset page: https://huggingface.co/datasets/tdw419/sefaria-raw-texts.text-retrieval1K<n<10K0 likes65 downloads20d agoHugging Face12nphearum /khmer-raw-text-3M-v2 Dataset Card for nphearum/khmer-raw-text-3M-v2 Dataset Summary nphearum/khmer-raw-text-3M-v2 is a large-scale raw text corpus containing approximately 200_000 completed records with 3 million text segments in Khmer, curated for large language model (LLM) pre-training, continued pre-training, and domain adaptation. The dataset emphasizes Khmer-language coverage, a historically underrepresented low-resource language, while retaining bilingual context for cross-lingual… See the full description on the dataset page: https://huggingface.co/datasets/nphearum/khmer-raw-text-3M-v2.texttext-classification100K<n<1M1 likes64 downloads5mo agoHugging Face13CarolinePascal /grabette-tactile-texture-raw0 likes49 downloads29d agoHugging Face14PersianAICommunity /Fibonacci-Pre_Train-Persian-Corpus-Raw-Texts-Datasettext10M<n<100M0 likes41 downloads9mo agoHugging Face15Gill-Hack-25-UdeM /raw_text_hack_2025tabular100K<n<1M0 likes36 downloads1y agoHugging Face16vanwdai /raw_text_ocr_texttext100K<n<1M0 likes32 downloads1y agoHugging Face17TokenBender /glaive_coder_raw_texttext100K<n<1M2 likes28 downloads3y agoHugging Face18fibonacciai /Fibonacci-Pre_Train-Persian-Corpus-Raw-Texts-Datasettext10M<n<100M0 likes28 downloads9mo agoHugging Face19ariabi /stormfront_incels-raw-textgated Dataset Card for Stormfront & Incels Raw Text Dataset Summary This dataset contains raw, unannotated textual posts from two online extremist platforms: Stormfront (white supremacist) and Incels.is (misogynistic). Each post is provided as a single line of text in .txt files, with no metadata. This simplified format supports unsupervised tasks such as domain adaptation, masked language modeling, and linguistic analysis of extremist cryptolects. The dataset was used in:… See the full description on the dataset page: https://huggingface.co/datasets/ariabi/stormfront_incels-raw-text.text10M<n<100M3 likes26 downloads1y agoHugging Face20nphearum /khmer-raw-text-3M Dataset Card for nphearum/khmer-raw-text-3M Dataset Summary nphearum/khmer-raw-text-3M is a large-scale raw text corpus containing approximately 50000 completed records with 3 million text segments in Khmer, curated for large language model (LLM) pre-training, continued pre-training, and domain adaptation. The dataset emphasizes Khmer-language coverage, a historically underrepresented low-resource language, while retaining bilingual context for cross-lingual learning.… See the full description on the dataset page: https://huggingface.co/datasets/nphearum/khmer-raw-text-3M.texttext-classification10K<n<100K0 likes25 downloads8mo agoHugging Face21anyantudre /bf-data-raw-textsdocumentn<1K1 likes23 downloads1y agoHugging Face22adrianf12 /ori_raw_text Ori Raw Text (Single Row) Single-row JSONL with the concatenated Ori documentation content (cleaned paragraphs). Schema Each row: { "text": "" } Usage from datasets import load_dataset ds = load_dataset(adrianf12/ori_raw_text) print(ds[train][0][text][:500]) textn<1K0 likes21 downloads11mo agoHugging Face23riotu-lab /ARABIC-RAW-TEXTDataset: Aluka 1.4 GB AraWiki 3.9 GB Aya 22.5 GB Islamic Books 21.4 GB texttext-generation100M<n<1B5 likes20 downloads2y agoHugging Face24aksw /Text2SPARQL-Raw Dataset Card for Text2sparql-Raw 🧾 Dataset Summary Text2Sparql-Raw is a multilingual dataset designed for the task of translating natural language questions into SPARQL queries over the DBpedia knowledge graph. This dataset aggregates and harmonizes four widely used benchmarks in the text-to-SPARQL domain: QALD (versions 1–9) LC-QuAD 1.0 Orange/paraqa-sparqltotext julioc-p/Question-Sparql It contains questions in both English and Spanish, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/aksw/Text2SPARQL-Raw.text10K<n<100K0 likes18 downloads1y agoHugging Face25yrrhall /ARABIC-RAW-TEXTDataset: Aluka 1.4 GB AraWiki 3.9 GB Aya 22.5 GB Islamic Books 21.4 GB texttext-generation100M<n<1B0 likes17 downloads4mo agoHugging Face26bany1111 /KorPE_raw_texttext100K<n<1M0 likes13 downloads2y agoHugging Face27Imsidag-community /kabyle-raw-texttext100K<n<1M1 likes13 downloads10mo agoHugging Face28TokenBender /roleplay_raw_texttext10K<n<100K2 likes12 downloads3y agoHugging Face29maximilianshwarzmullers /turkmen_raw_texttext1M<n<10M0 likes12 downloads1y agoHugging Face30dliu1 /legal-llama-raw-texttext10K<n<100K0 likes11 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.