datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
YuLan-Mini-Text-Datasets
News
[2025.04.11] Add dataset mixture: link.
[2025.03.30] Text datasets upload finished.
This is text dataset.
这是文本格式的数据集。
Since we have used BPE-Dropout, in order to ensure accuracy, you can find the tokenized dataset here.
由于我们使用了BPE-Dropout,为了保证准确性,你可以在这里找到分词后的数据。
For more information, please refer to our datasets details and preprocess details.
Contributing
We welcome any form of contribution, including feedback on model bad cases, feature suggestions, and example… See the full description on the dataset page: https://huggingface.co/datasets/yulan-team/YuLan-Mini-Text-Datasets.KoHRM-Text-1.4B-prepared-data
KoHRM-Text-1.4B Prepared Data
This dataset repository contains prepared HRM-Text V1Dataset artifacts for KoHRM-Text-1.4B.
The data is intended for continued pretraining and staged training with the project code at:
https://github.com/LLM-OS-Models/KoHRM-text
https://huggingface.co/LLM-OS-Models/KoHRM-Text-1.4B
https://huggingface.co/LLM-OS-Models/HRM-Text-Ko-Terminal-Tokenizer-131K
The upstream architecture and training method are based on:
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/KoHRM-Text-1.4B-prepared-data.datacomp-small-with-text-embeddings
Dataset Card for "datacomp-small-with-text-embeddings"
More Information needed
SQaLe-text-to-SQL-dataset
🧮 SQALE: A Large-Scale Semi-Synthetic Dataset
SQALE is a large-scale, semi-synthetic Text-to-SQL dataset grounded in real-world database schemas.
It was designed to push the boundaries of natural language to SQL generation, combining realistic schema diversity, complex query structures, and linguistically varied natural language questions.
The dataset was introduced in the paper SQaLe: A Large Text-to-SQL Corpus Grounded in Real Schemas. The code for the generation pipeline of this… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-text-to-SQL-dataset.text-to-speech-human-preferences-315k
Text-to-speech human preferences: 315K votes across 15 models
This gated dataset contains the evaluation record behind Datapoint Audio
Bench: 315,000 eligible pairwise votes comparing 15 text-to-speech
models in a complete round-robin over 300 English prompts. The prompt set
covers eight practical voice-agent categories, and every generated sample is
included as a typed audio record.
The source evaluation collected 357,651 completed responses. The published
benchmark excluded… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-to-speech-human-preferences-315k.BAD-Bengali-Aggressive-Text-Dataset
Novel Aggressive Text Dataset in Bengali
Tackling Cyber-Aggression: Identification and Fine-Grained Categorization of Aggressive Texts on Social Media using Weighted Ensemble of Transformers
Author: Omar Sharif and Mohammed Moshiul Hoque
Related Papers:
Paper1 in Neurocomputing Journal
Paper2 in CONSTRAINT@AAAI-2021
Paper3 in LTEDI@EACL-2021
Abstract
The pervasiveness of aggressive content in social media has become a serious concern for government… See the full description on the dataset page: https://huggingface.co/datasets/omar-sharif/BAD-Bengali-Aggressive-Text-Dataset.Mental-Health_Text-Classification_Dataset
Mental Health Text Classification Dataset (4-Class)
Dataset Description
This dataset contains short, user‑generated texts labeled for 4‑class mental health classification: Suicidal, Depression, Anxiety, and Normal. It is a derived dataset created by combining and cleaning three public mental‑health corpora, then re‑labeling them into a unified 4‑class scheme and exporting CSV files suitable for both classical ML and modern NLP models.
The repository includes:
An… See the full description on the dataset page: https://huggingface.co/datasets/ourafla/Mental-Health_Text-Classification_Dataset.text-2-image-human-preferences-2m
Text-to-image human preferences: 2M votes across 30 models
This dataset contains the complete voting record behind the
Datapoint Image Bench
leaderboard: 2,161,160 validated pairwise votes — exactly 10 for each of
216,116 image pairs. The votes compare 30 text-to-image models in a complete
round-robin on 500 prompts, judged by annotators from over 200 countries.
Every vote includes the annotator's trust score at the time the vote was
cast.
Built on the Datapoint annotation… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-human-preferences-2m.PubMed_MultiLabel_Text_Classification_Dataset_MeSHThis dataset consists of a approx 50k collection of research articles from PubMed repository. Originally these documents are manually annotated by Biomedical Experts with their MeSH labels and each articles are described in terms of 10-15 MeSH labels. In this Dataset we have huge numbers of labels present as a MeSH major which is raising the issue of extremely large output space and severe label sparsity issues. To solve this Issue Dataset has been Processed and mapped to its root as Described… See the full description on the dataset page: https://huggingface.co/datasets/owaiskha9654/PubMed_MultiLabel_Text_Classification_Dataset_MeSH.text-to-art-database
Vieutopia T2A Privacy Train v1
Dataset Summary
Privacy-safe text-to-image dataset repacked into Parquet shards with embedded image bytes.
Scope: text-to-image outputs only
Excluded: image-to-image pipelines (pix2pix_*, pst_*)
Privacy: no raw task UUIDs, no user/device fields
Storage format: parquet shards (image as binary bytes), no image_path dependency
Splits
samples
train: 117572
validation: 6532
test: 6532
total: 130636
iterations… See the full description on the dataset page: https://huggingface.co/datasets/quchenyuan/text-to-art-database.popQA_text_datainstruction-data-text-onlyall-document-text-data
Climate Policy Radar Open Data
This repo contains the full text data of all of the documents from the Climate Policy Radar database (CPR), which is also available at Climate Change Laws of the World (CCLW).
Please note that this replaces the Global Stocktake open dataset: that data, including all NDCs and IPCC reports is now a subset of this dataset.
What’s in this dataset
This dataset contains two corpus types (groups of the same types or sources of documents) which… See the full description on the dataset page: https://huggingface.co/datasets/ClimatePolicyRadar/all-document-text-data.Amharic-News-Text-classification-Dataset
An Amharic News Text classification Dataset
In NLP, text classification is one of the primary problems we try to solve and its uses in language analyses are indisputable. The lack of labeled training data made it harder to do these tasks in low resource languages like Amharic. The task of collecting, labeling, annotating, and making valuable this kind of data will encourage junior researchers, schools, and machine learning practitioners to implement existing classification models… See the full description on the dataset page: https://huggingface.co/datasets/israel/Amharic-News-Text-classification-Dataset.Diffusion-Reward-Modeling-for-Text-Rendering-Dataset
🖼️ Text-to-Image Rendering Dataset
A dataset of 14k text prompts for image generation with text rendering evaluation
📚 Dataset Overview
This dataset contains 14,000 text prompts specifically designed for:
Image generation with text rendering
Evaluating text preservation in generated images
Training diffusion models for better text rendering
Each prompt comes with:
Pre-extracted target text for rendering
5 Stable Diffusion 3 generated latents (70k total)
Dual… See the full description on the dataset page: https://huggingface.co/datasets/leffff/Diffusion-Reward-Modeling-for-Text-Rendering-Dataset.text-mining-ce-dataset
Vietnamese Legal Cross-Encoder Dataset
Training data for a cross-encoder reranker on Vietnamese legal documents.
Source
Built from YuITC/Vietnamese-Legal-Documents.
Schema
Column
Type
Description
qid
int64
Query ID
cid
int64
Document (context) ID
query
string
Legal question
document
string
Candidate document
label
int64
1 = positive, 0 = negative
split
string
train or test
negative_type
string
random, same_topic_wrong_article… See the full description on the dataset page: https://huggingface.co/datasets/juzharii/text-mining-ce-dataset.text-analysis-context-cased-case-data2026-24679-text-dataset
24-679 (Fall 2026): Haiku Perspectives
ccm/2026-24679-text-dataset
Short poems collected through the 24-679 course survey at Carnegie Mellon University. Each retained
survey response contributes a human-perspective poem and an AI-or-machine-perspective poem. This
dataset supports a classroom comparison of fixed embeddings, fine-tuning, and few-shot prompting.
Source and task
The preparation notebook reads 24-679-tabular-survey.csv, retains the two specified haiku… See the full description on the dataset page: https://huggingface.co/datasets/ccm/2026-24679-text-dataset.Mental-Health_Text-Classification_Dataset
Mental Health Text Classification Dataset (4-Class)
Dataset Description
This dataset contains short, user‑generated texts labeled for 4‑class mental health classification: Suicidal, Depression, Anxiety, and Normal. It is a derived dataset created by combining and cleaning three public mental‑health corpora, then re‑labeling them into a unified 4‑class scheme and exporting CSV files suitable for both classical ML and modern NLP models.
The repository includes:
An… See the full description on the dataset page: https://huggingface.co/datasets/ihsansaad24/Mental-Health_Text-Classification_Dataset.US-attractions-text-dataset2026-animaldescription-text-dataset
Animal Descriptions Dataset
pcwoods/2026-animaldescription-text-dataset
This dataset contains descriptions of various popular zoo animals labeled by type.
Types are limited to Mammal, Bird, or Reptile for simplicity.
Source
Descriptions were hand-written based on popular zoo animals from https://zootrack.me/animals/popular. Facts about each animal for descriptions were identified using AI tools.
Fields
Field
Meaning
description
Text… See the full description on the dataset page: https://huggingface.co/datasets/pcwoods/2026-animaldescription-text-dataset.2026-24679-text-dataset
Hospitality Reviews: Hotel vs Restaurant
kwongnon/2026-24679-text-dataset
English-language hospitality reviews labeled by venue type. The classification task is to predict
whether a review describes a hotel or a restaurant. Labels are derived from the
Hospitality column: 0 = restaurant; 1 = hotel. The target describes the venue category, not
review sentiment, review quality, or whether the statements in a review are factually correct.
Source and task
The… See the full description on the dataset page: https://huggingface.co/datasets/kwongnon/2026-24679-text-dataset.2026-24679-text-dataset
Premier League Players: Position From Description
kadireks/2026-24679-text-dataset
English-language descriptions of Premier League footballers, labeled by playing position. The
classification task is to predict whether a description belongs to a goalkeeper, defender,
midfielder, or forward. Labels are derived from the position column:
0 = GK; 1 = DF; 2 = MF; 3 = FW. The target describes the player's position, not his quality,
market value, or current form.
The position words… See the full description on the dataset page: https://huggingface.co/datasets/kadireks/2026-24679-text-dataset.pol-dataset-text-no-url-calibration
Dataset Card for "pol-dataset-text-no-url-calibration"
More Information needed
text_detoxification_datasettext-2-video-ranking-human-preferences
T2V Ranking Human Preferences
~91,000 human ranking labels across 18 text-to-video models on 3 quality dimensions, collected from real annotators via Datapoint AI.
This is the first public ranking-based (not pairwise) human preference dataset for text-to-video generation. Each datapoint contains 5 videos generated from the same prompt by different models, ranked 1st through 5th by 15 annotators on each dimension.
Why This Dataset
Existing video preference… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-video-ranking-human-preferences.text-2-video-human-preferences-motion
Human Preferences for AI-Generated Video: Motion Quality
29,283 pairwise human preference labels comparing 4 frontier video generation models on human motion across 3 quality dimensions, collected from 4,349 real annotators via Datapoint AI.
This is the largest publicly available human preference dataset focused specifically on human motion in AI-generated video.
Why This Dataset
Video generation models are improving fast, but evaluating human motion remains… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-video-human-preferences-motion.Mental-Health_Text-Classification_Dataset
Mental Health Text Classification Dataset (4-Class)
Dataset Description
This dataset contains short, user‑generated texts labeled for 4‑class mental health classification: Suicidal, Depression, Anxiety, and Normal. It is a derived dataset created by combining and cleaning three public mental‑health corpora, then re‑labeling them into a unified 4‑class scheme and exporting CSV files suitable for both classical ML and modern NLP models.
The repository includes:
An… See the full description on the dataset page: https://huggingface.co/datasets/Sam20032212/Mental-Health_Text-Classification_Dataset.text-mining-ce-dataset-v2preference_data_llama_factory_len_15k_text_with_urls
