datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hindi-english-raw-text-corpus-uncleanedraw-text-corpus
📝 Zomi Raw Text Corpus (Community-Contributed)
The Zomi Raw Text Corpus is an open, community-driven collection of unprocessed Zomi-language text.It serves as the foundational dataset for building the full Zomi NLP ecosystem, including tokenizers, language models, ASR/TTS systems, and downstream tasks.
This dataset is intentionally raw — no normalization, deduplication, or cleaning is applied.Cleaned and task-specific datasets will be released separately.
📦 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/zomi-language-corpora/raw-text-corpus.essentialweb-1.0-10B-raw-contentFibonacci-Pre_Train-Persian-Corpus-Raw-Texts-DatasetDocument Version: 1.0.8 | Last Updated: 01/02/2025
religious-texts-rawCatalan-Raw-Text
Dataset Summary
The Catalan Raw Text Dataset is a subset of the projecte-aina/catalan_general_crawling.
It is licensed under a Creative Commons Attribution 4.0 International license, just like the origin dataset.
The dataset consists of 404k samples (roughly 20% of the original), totalling 331M tokens after tokenizing it with the Llama-3 Tokenizer.
Languages
The dataset is in Catalan (ca-ES).
Data Fields
text (str): Text.
Data Splits… See the full description on the dataset page: https://huggingface.co/datasets/catallama/Catalan-Raw-Text.Pre-Training-Persian-Corpus-Raw-Texts-DatasetDocument Version: 2.0.0 | Last Updated: 02/13/2026
khmer-raw-text-3M-v2
Dataset Card for nphearum/khmer-raw-text-3M-v2
Dataset Summary
nphearum/khmer-raw-text-3M-v2 is a large-scale raw text corpus containing approximately 200_000 completed records with 3 million text segments in Khmer, curated for large language model (LLM) pre-training, continued pre-training, and domain adaptation.
The dataset emphasizes Khmer-language coverage, a historically underrepresented low-resource language, while retaining bilingual context for cross-lingual… See the full description on the dataset page: https://huggingface.co/datasets/nphearum/khmer-raw-text-3M-v2.Fibonacci-Pre_Train-Persian-Corpus-Raw-Texts-Datasetraw_text_hack_2025glaive_coder_raw_textraw_text_ocr_textFibonacci-Pre_Train-Persian-Corpus-Raw-Texts-Datasetstormfront_incels-raw-text
Dataset Card for Stormfront & Incels Raw Text
Dataset Summary
This dataset contains raw, unannotated textual posts from two online extremist platforms: Stormfront (white supremacist) and Incels.is (misogynistic). Each post is provided as a single line of text in .txt files, with no metadata.
This simplified format supports unsupervised tasks such as domain adaptation, masked language modeling, and linguistic analysis of extremist cryptolects. The dataset was used in:… See the full description on the dataset page: https://huggingface.co/datasets/ariabi/stormfront_incels-raw-text.khmer-raw-text-3M
Dataset Card for nphearum/khmer-raw-text-3M
Dataset Summary
nphearum/khmer-raw-text-3M is a large-scale raw text corpus containing approximately 50000 completed records with 3 million text segments in Khmer, curated for large language model (LLM) pre-training, continued pre-training, and domain adaptation.
The dataset emphasizes Khmer-language coverage, a historically underrepresented low-resource language, while retaining bilingual context for cross-lingual learning.… See the full description on the dataset page: https://huggingface.co/datasets/nphearum/khmer-raw-text-3M.ARABIC-RAW-TEXTDataset:
Aluka 1.4 GB
AraWiki 3.9 GB
Aya 22.5 GB
Islamic Books 21.4 GB
ori_raw_text
Ori Raw Text (Single Row)
Single-row JSONL with the concatenated Ori documentation content (cleaned paragraphs).
Schema
Each row:
{
"text": ""
}
Usage
from datasets import load_dataset
ds = load_dataset(adrianf12/ori_raw_text)
print(ds[train][0][text][:500])
Text2SPARQL-Raw
Dataset Card for Text2sparql-Raw
🧾 Dataset Summary
Text2Sparql-Raw is a multilingual dataset designed for the task of translating natural language questions into SPARQL queries over the DBpedia knowledge graph. This dataset aggregates and harmonizes four widely used benchmarks in the text-to-SPARQL domain:
QALD (versions 1–9)
LC-QuAD 1.0
Orange/paraqa-sparqltotext
julioc-p/Question-Sparql
It contains questions in both English and Spanish, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/aksw/Text2SPARQL-Raw.ARABIC-RAW-TEXTDataset:
Aluka 1.4 GB
AraWiki 3.9 GB
Aya 22.5 GB
Islamic Books 21.4 GB
kabyle-raw-textFibonacci-Pre_Train-Persian-Corpus-Raw-Texts-DatasetDocument Version: 1.0.5 | Last Updated: 01/02/2025
turkmen_raw_textroleplay_raw_textlegal-llama-raw-textKorPE_raw_textraw-text-dataset-2raw_text_synthetic_dataset_50kraw-text-datasetraw-text-corpus-ro-8kFibonacci-Pre_Train-Persian-Corpus-Raw-Texts-Dataset
