datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turkish-raw-text-cleaned
Turkish Raw Text Cleaned
turkish-raw-text-cleaned, Türkçe dil modeli çalışmaları için hazırlanmış temizlenmiş ham metin veri kümesidir. Veri kümesi, turkish-nlp-suite çatısı altında yayımlanan Türkçe metin kaynaklarının temizlenmesi, filtrelenmesi ve model eğitimine daha uygun hale getirilmesiyle oluşturulmuştur.
Bu çalışma özellikle Türkçe LLM ön-eğitimi, continual pre-training (CPT), tokenizer analizi, embedding modeli eğitimi, alan bağımsız Türkçe metin modelleme ve veri… See the full description on the dataset page: https://huggingface.co/datasets/serda-dev/turkish-raw-text-cleaned.hindi-english-raw-text-corpus-uncleanedraw-text-corpus
📝 Zomi Raw Text Corpus (Community-Contributed)
The Zomi Raw Text Corpus is an open, community-driven collection of unprocessed Zomi-language text.It serves as the foundational dataset for building the full Zomi NLP ecosystem, including tokenizers, language models, ASR/TTS systems, and downstream tasks.
This dataset is intentionally raw — no normalization, deduplication, or cleaning is applied.Cleaned and task-specific datasets will be released separately.
📦 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/zomi-language-corpora/raw-text-corpus.essentialweb-1.0-10B-raw-contentdataset_sugar_1709_texture-rawCompactDS-102GB-raw-textFibonacci-Pre_Train-Persian-Corpus-Raw-Texts-DatasetDocument Version: 1.0.8 | Last Updated: 01/02/2025
Catalan-Raw-Text
Dataset Summary
The Catalan Raw Text Dataset is a subset of the projecte-aina/catalan_general_crawling.
It is licensed under a Creative Commons Attribution 4.0 International license, just like the origin dataset.
The dataset consists of 404k samples (roughly 20% of the original), totalling 331M tokens after tokenizing it with the Llama-3 Tokenizer.
Languages
The dataset is in Catalan (ca-ES).
Data Fields
text (str): Text.
Data Splits… See the full description on the dataset page: https://huggingface.co/datasets/catallama/Catalan-Raw-Text.Pre-Training-Persian-Corpus-Raw-Texts-DatasetDocument Version: 2.0.0 | Last Updated: 02/13/2026
religious-texts-rawsefaria-raw-texts
Sefaria Raw Texts (judaism-llm)
Raw Sefaria API responses: 6,112 JSON files, 82MB, organized by category.
License: CC-BY-NC 3.0 — non-commercial use only. Source: Sefaria, downloaded via the official API (download_sefaria.py in the pipeline repo).
Layout
Directory names encode the category path (_ separates levels), e.g.
Halakhah_Mishneh Torah_Commentary_Ohr Sameach_Sefer Zeraim/. Each JSON file
is one Sefaria document with the full API schema: text (English),
he… See the full description on the dataset page: https://huggingface.co/datasets/tdw419/sefaria-raw-texts.khmer-raw-text-3M-v2
Dataset Card for nphearum/khmer-raw-text-3M-v2
Dataset Summary
nphearum/khmer-raw-text-3M-v2 is a large-scale raw text corpus containing approximately 200_000 completed records with 3 million text segments in Khmer, curated for large language model (LLM) pre-training, continued pre-training, and domain adaptation.
The dataset emphasizes Khmer-language coverage, a historically underrepresented low-resource language, while retaining bilingual context for cross-lingual… See the full description on the dataset page: https://huggingface.co/datasets/nphearum/khmer-raw-text-3M-v2.grabette-tactile-texture-rawFibonacci-Pre_Train-Persian-Corpus-Raw-Texts-Datasetraw_text_hack_2025raw_text_ocr_textglaive_coder_raw_textFibonacci-Pre_Train-Persian-Corpus-Raw-Texts-Datasetstormfront_incels-raw-text
Dataset Card for Stormfront & Incels Raw Text
Dataset Summary
This dataset contains raw, unannotated textual posts from two online extremist platforms: Stormfront (white supremacist) and Incels.is (misogynistic). Each post is provided as a single line of text in .txt files, with no metadata.
This simplified format supports unsupervised tasks such as domain adaptation, masked language modeling, and linguistic analysis of extremist cryptolects. The dataset was used in:… See the full description on the dataset page: https://huggingface.co/datasets/ariabi/stormfront_incels-raw-text.khmer-raw-text-3M
Dataset Card for nphearum/khmer-raw-text-3M
Dataset Summary
nphearum/khmer-raw-text-3M is a large-scale raw text corpus containing approximately 50000 completed records with 3 million text segments in Khmer, curated for large language model (LLM) pre-training, continued pre-training, and domain adaptation.
The dataset emphasizes Khmer-language coverage, a historically underrepresented low-resource language, while retaining bilingual context for cross-lingual learning.… See the full description on the dataset page: https://huggingface.co/datasets/nphearum/khmer-raw-text-3M.bf-data-raw-textsori_raw_text
Ori Raw Text (Single Row)
Single-row JSONL with the concatenated Ori documentation content (cleaned paragraphs).
Schema
Each row:
{
"text": ""
}
Usage
from datasets import load_dataset
ds = load_dataset(adrianf12/ori_raw_text)
print(ds[train][0][text][:500])
ARABIC-RAW-TEXTDataset:
Aluka 1.4 GB
AraWiki 3.9 GB
Aya 22.5 GB
Islamic Books 21.4 GB
Text2SPARQL-Raw
Dataset Card for Text2sparql-Raw
🧾 Dataset Summary
Text2Sparql-Raw is a multilingual dataset designed for the task of translating natural language questions into SPARQL queries over the DBpedia knowledge graph. This dataset aggregates and harmonizes four widely used benchmarks in the text-to-SPARQL domain:
QALD (versions 1–9)
LC-QuAD 1.0
Orange/paraqa-sparqltotext
julioc-p/Question-Sparql
It contains questions in both English and Spanish, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/aksw/Text2SPARQL-Raw.ARABIC-RAW-TEXTDataset:
Aluka 1.4 GB
AraWiki 3.9 GB
Aya 22.5 GB
Islamic Books 21.4 GB
KorPE_raw_textkabyle-raw-textroleplay_raw_textturkmen_raw_textlegal-llama-raw-text
