CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yulan-team /YuLan-Mini-Text-Datasets News [2025.04.11] Add dataset mixture: link. [2025.03.30] Text datasets upload finished. This is text dataset. 这是文本格式的数据集。 Since we have used BPE-Dropout, in order to ensure accuracy, you can find the tokenized dataset here. 由于我们使用了BPE-Dropout,为了保证准确性,你可以在这里找到分词后的数据。 For more information, please refer to our datasets details and preprocess details. Contributing We welcome any form of contribution, including feedback on model bad cases, feature suggestions, and example… See the full description on the dataset page: https://huggingface.co/datasets/yulan-team/YuLan-Mini-Text-Datasets.tabulartext-generation100M<n<1B12 likes2.1k downloads1y agoHugging Face02LLM-OS-Models /KoHRM-Text-1.4B-prepared-data KoHRM-Text-1.4B Prepared Data This dataset repository contains prepared HRM-Text V1Dataset artifacts for KoHRM-Text-1.4B. The data is intended for continued pretraining and staged training with the project code at: https://github.com/LLM-OS-Models/KoHRM-text https://huggingface.co/LLM-OS-Models/KoHRM-Text-1.4B https://huggingface.co/LLM-OS-Models/HRM-Text-Ko-Terminal-Tokenizer-131K The upstream architecture and training method are based on: Paper:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/KoHRM-Text-1.4B-prepared-data.tabulartext-generationn<1K1 likes1k downloads4mo agoHugging Face03nielsr /datacomp-small-with-text-embeddings Dataset Card for "datacomp-small-with-text-embeddings" More Information needed image10M<n<100M0 likes967 downloads3y agoHugging Face04trl-lab /SQaLe-text-to-SQL-dataset 🧮 SQALE: A Large-Scale Semi-Synthetic Dataset SQALE is a large-scale, semi-synthetic Text-to-SQL dataset grounded in real-world database schemas. It was designed to push the boundaries of natural language to SQL generation, combining realistic schema diversity, complex query structures, and linguistically varied natural language questions. The dataset was introduced in the paper SQaLe: A Large Text-to-SQL Corpus Grounded in Real Schemas. The code for the generation pipeline of this… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-text-to-SQL-dataset.tabulartext-generation100K<n<1M20 likes732 downloads7mo agoHugging Face05datapointai /text-to-speech-human-preferences-315kgated Text-to-speech human preferences: 315K votes across 15 models This gated dataset contains the evaluation record behind Datapoint Audio Bench: 315,000 eligible pairwise votes comparing 15 text-to-speech models in a complete round-robin over 300 English prompts. The prompt set covers eight practical voice-agent categories, and every generated sample is included as a typed audio record. The source evaluation collected 357,651 completed responses. The published benchmark excluded… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-to-speech-human-preferences-315k.audiotext-to-speech100K<n<1M38 likes571 downloads22d agoHugging Face06omar-sharif /BAD-Bengali-Aggressive-Text-Dataset Novel Aggressive Text Dataset in Bengali Tackling Cyber-Aggression: Identification and Fine-Grained Categorization of Aggressive Texts on Social Media using Weighted Ensemble of Transformers Author: Omar Sharif and Mohammed Moshiul Hoque Related Papers: Paper1 in Neurocomputing Journal Paper2 in CONSTRAINT@AAAI-2021 Paper3 in LTEDI@EACL-2021 Abstract The pervasiveness of aggressive content in social media has become a serious concern for government… See the full description on the dataset page: https://huggingface.co/datasets/omar-sharif/BAD-Bengali-Aggressive-Text-Dataset.tabular10K<n<100K3 likes508 downloads5y agoHugging Face07ourafla /Mental-Health_Text-Classification_Dataset Mental Health Text Classification Dataset (4-Class) Dataset Description This dataset contains short, user‑generated texts labeled for 4‑class mental health classification: Suicidal, Depression, Anxiety, and Normal. It is a derived dataset created by combining and cleaning three public mental‑health corpora, then re‑labeling them into a unified 4‑class scheme and exporting CSV files suitable for both classical ML and modern NLP models. The repository includes: An… See the full description on the dataset page: https://huggingface.co/datasets/ourafla/Mental-Health_Text-Classification_Dataset.texttext-classification10K<n<100K9 likes400 downloads9mo agoHugging Face08datapointai /text-2-image-human-preferences-2mgated Text-to-image human preferences: 2M votes across 30 models This dataset contains the complete voting record behind the Datapoint Image Bench leaderboard: 2,161,160 validated pairwise votes — exactly 10 for each of 216,116 image pairs. The votes compare 30 text-to-image models in a complete round-robin on 500 prompts, judged by annotators from over 200 countries. Every vote includes the annotator's trust score at the time the vote was cast. Built on the Datapoint annotation… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-human-preferences-2m.imagetext-to-image1M<n<10M21 likes368 downloads1mo agoHugging Face09owaiskha9654 /PubMed_MultiLabel_Text_Classification_Dataset_MeSHThis dataset consists of a approx 50k collection of research articles from PubMed repository. Originally these documents are manually annotated by Biomedical Experts with their MeSH labels and each articles are described in terms of 10-15 MeSH labels. In this Dataset we have huge numbers of labels present as a MeSH major which is raising the issue of extremely large output space and severe label sparsity issues. To solve this Issue Dataset has been Processed and mapped to its root as Described… See the full description on the dataset page: https://huggingface.co/datasets/owaiskha9654/PubMed_MultiLabel_Text_Classification_Dataset_MeSH.tabulartext-classification10K<n<100K28 likes222 downloads4y agoHugging Face10quchenyuan /text-to-art-database Vieutopia T2A Privacy Train v1 Dataset Summary Privacy-safe text-to-image dataset repacked into Parquet shards with embedded image bytes. Scope: text-to-image outputs only Excluded: image-to-image pipelines (pix2pix_*, pst_*) Privacy: no raw task UUIDs, no user/device fields Storage format: parquet shards (image as binary bytes), no image_path dependency Splits samples train: 117572 validation: 6532 test: 6532 total: 130636 iterations… See the full description on the dataset page: https://huggingface.co/datasets/quchenyuan/text-to-art-database.tabulartext-to-image100K<n<1M0 likes205 downloads6mo agoHugging Face11Zektzerite /popQA_text_datatabular1K<n<10K0 likes191 downloads1y agoHugging Face12jan-hq /instruction-data-text-onlytabular1M<n<10M0 likes161 downloads2y agoHugging Face13ClimatePolicyRadar /all-document-text-datagated Climate Policy Radar Open Data This repo contains the full text data of all of the documents from the Climate Policy Radar database (CPR), which is also available at Climate Change Laws of the World (CCLW). Please note that this replaces the Global Stocktake open dataset: that data, including all NDCs and IPCC reports is now a subset of this dataset. What’s in this dataset This dataset contains two corpus types (groups of the same types or sources of documents) which… See the full description on the dataset page: https://huggingface.co/datasets/ClimatePolicyRadar/all-document-text-data.tabular10M<n<100M24 likes136 downloads11mo agoHugging Face14israel /Amharic-News-Text-classification-Dataset An Amharic News Text classification Dataset In NLP, text classification is one of the primary problems we try to solve and its uses in language analyses are indisputable. The lack of labeled training data made it harder to do these tasks in low resource languages like Amharic. The task of collecting, labeling, annotating, and making valuable this kind of data will encourage junior researchers, schools, and machine learning practitioners to implement existing classification models… See the full description on the dataset page: https://huggingface.co/datasets/israel/Amharic-News-Text-classification-Dataset.tabular10K<n<100K1 likes108 downloads4y agoHugging Face15leffff /Diffusion-Reward-Modeling-for-Text-Rendering-Dataset 🖼️ Text-to-Image Rendering Dataset A dataset of 14k text prompts for image generation with text rendering evaluation 📚 Dataset Overview This dataset contains 14,000 text prompts specifically designed for: Image generation with text rendering Evaluating text preservation in generated images Training diffusion models for better text rendering Each prompt comes with: Pre-extracted target text for rendering 5 Stable Diffusion 3 generated latents (70k total) Dual… See the full description on the dataset page: https://huggingface.co/datasets/leffff/Diffusion-Reward-Modeling-for-Text-Rendering-Dataset.tabulartext-to-image10K<n<100K7 likes80 downloads1y agoHugging Face16juzharii /text-mining-ce-dataset Vietnamese Legal Cross-Encoder Dataset Training data for a cross-encoder reranker on Vietnamese legal documents. Source Built from YuITC/Vietnamese-Legal-Documents. Schema Column Type Description qid int64 Query ID cid int64 Document (context) ID query string Legal question document string Candidate document label int64 1 = positive, 0 = negative split string train or test negative_type string random, same_topic_wrong_article… See the full description on the dataset page: https://huggingface.co/datasets/juzharii/text-mining-ce-dataset.tabulartext-classification100K<n<1M0 likes72 downloads3mo agoHugging Face17brainer /text-analysis-context-cased-case-datatabular100K<n<1M0 likes61 downloads11mo agoHugging Face18ccm /2026-24679-text-dataset 24-679 (Fall 2026): Haiku Perspectives ccm/2026-24679-text-dataset Short poems collected through the 24-679 course survey at Carnegie Mellon University. Each retained survey response contributes a human-perspective poem and an AI-or-machine-perspective poem. This dataset supports a classroom comparison of fixed embeddings, fine-tuning, and few-shot prompting. Source and task The preparation notebook reads 24-679-tabular-survey.csv, retains the two specified haiku… See the full description on the dataset page: https://huggingface.co/datasets/ccm/2026-24679-text-dataset.tabulartext-classificationn<1K0 likes59 downloads10d agoHugging Face19ihsansaad24 /Mental-Health_Text-Classification_Dataset Mental Health Text Classification Dataset (4-Class) Dataset Description This dataset contains short, user‑generated texts labeled for 4‑class mental health classification: Suicidal, Depression, Anxiety, and Normal. It is a derived dataset created by combining and cleaning three public mental‑health corpora, then re‑labeling them into a unified 4‑class scheme and exporting CSV files suitable for both classical ML and modern NLP models. The repository includes: An… See the full description on the dataset page: https://huggingface.co/datasets/ihsansaad24/Mental-Health_Text-Classification_Dataset.texttext-classification10K<n<100K1 likes57 downloads9mo agoHugging Face20Gwen1220 /US-attractions-text-datasettabular1K<n<10K0 likes55 downloads7d agoHugging Face21pcwoods /2026-animaldescription-text-dataset Animal Descriptions Dataset pcwoods/2026-animaldescription-text-dataset This dataset contains descriptions of various popular zoo animals labeled by type. Types are limited to Mammal, Bird, or Reptile for simplicity. Source Descriptions were hand-written based on popular zoo animals from https://zootrack.me/animals/popular. Facts about each animal for descriptions were identified using AI tools. Fields Field Meaning description Text… See the full description on the dataset page: https://huggingface.co/datasets/pcwoods/2026-animaldescription-text-dataset.tabulartext-classification1K<n<10K0 likes53 downloads7d agoHugging Face22kwongnon /2026-24679-text-dataset Hospitality Reviews: Hotel vs Restaurant kwongnon/2026-24679-text-dataset English-language hospitality reviews labeled by venue type. The classification task is to predict whether a review describes a hotel or a restaurant. Labels are derived from the Hospitality column: 0 = restaurant; 1 = hotel. The target describes the venue category, not review sentiment, review quality, or whether the statements in a review are factually correct. Source and task The… See the full description on the dataset page: https://huggingface.co/datasets/kwongnon/2026-24679-text-dataset.tabulartext-classification1K<n<10K0 likes48 downloads7d agoHugging Face23kadireks /2026-24679-text-dataset Premier League Players: Position From Description kadireks/2026-24679-text-dataset English-language descriptions of Premier League footballers, labeled by playing position. The classification task is to predict whether a description belongs to a goalkeeper, defender, midfielder, or forward. Labels are derived from the position column: 0 = GK; 1 = DF; 2 = MF; 3 = FW. The target describes the player's position, not his quality, market value, or current form. The position words… See the full description on the dataset page: https://huggingface.co/datasets/kadireks/2026-24679-text-dataset.tabulartext-classification1K<n<10K0 likes44 downloads7d agoHugging Face24BrennanGambling /pol-dataset-text-no-url-calibration Dataset Card for "pol-dataset-text-no-url-calibration" More Information needed tabular100K<n<1M0 likes42 downloads3y agoHugging Face25narySt /text_detoxification_datasettabular100K<n<1M0 likes42 downloads3y agoHugging Face26datapointai /text-2-video-ranking-human-preferencesgated T2V Ranking Human Preferences ~91,000 human ranking labels across 18 text-to-video models on 3 quality dimensions, collected from real annotators via Datapoint AI. This is the first public ranking-based (not pairwise) human preference dataset for text-to-video generation. Each datapoint contains 5 videos generated from the same prompt by different models, ranked 1st through 5th by 15 annotators on each dimension. Why This Dataset Existing video preference… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-video-ranking-human-preferences.tabularvideo-classification1K<n<10K2 likes36 downloads6mo agoHugging Face27datapointai /text-2-video-human-preferences-motiongated Human Preferences for AI-Generated Video: Motion Quality 29,283 pairwise human preference labels comparing 4 frontier video generation models on human motion across 3 quality dimensions, collected from 4,349 real annotators via Datapoint AI. This is the largest publicly available human preference dataset focused specifically on human motion in AI-generated video. Why This Dataset Video generation models are improving fast, but evaluating human motion remains… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-video-human-preferences-motion.imagevideo-classificationn<1K15 likes33 downloads6mo agoHugging Face28Sam20032212 /Mental-Health_Text-Classification_Dataset Mental Health Text Classification Dataset (4-Class) Dataset Description This dataset contains short, user‑generated texts labeled for 4‑class mental health classification: Suicidal, Depression, Anxiety, and Normal. It is a derived dataset created by combining and cleaning three public mental‑health corpora, then re‑labeling them into a unified 4‑class scheme and exporting CSV files suitable for both classical ML and modern NLP models. The repository includes: An… See the full description on the dataset page: https://huggingface.co/datasets/Sam20032212/Mental-Health_Text-Classification_Dataset.texttext-classification10K<n<100K0 likes33 downloads4mo agoHugging Face29juzharii /text-mining-ce-dataset-v2tabular100K<n<1M1 likes33 downloads3mo agoHugging Face30WPRM /preference_data_llama_factory_len_15k_text_with_urlsimage10K<n<100K0 likes29 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.