CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Winniechen2002 /TexasPokerRobot TexasPokerRobot TexasPokerRobot is a robot manipulation dataset collected in a Texas poker tabletop environment. The raw episodes are stored as compressed NumPy .npz files, organized by action folder. This release adds a Hugging Face-compatible manifest at data/train.csv so the dataset has a standard loadable split and a working Dataset Viewer while preserving the original raw episode files. Dataset Summary 1,470 raw episode files 14 action folders, with 105 episodes per… See the full description on the dataset page: https://huggingface.co/datasets/Winniechen2002/TexasPokerRobot.tabular1K<n<10K0 likes15k downloads5mo agoHugging Face02CCB /cis5300-text-classification Complex Word Identification (CIS 5300) Dataset Description This dataset supports the Complex Word Identification (CWI) task: given a word in context, predict whether it is complex (likely to be difficult for non-native speakers, children, or people with reading disabilities) or simple. CWI is the first step in lexical simplification — the task of rewriting text to make it more accessible. Before you can simplify a word, you need to identify which words need… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-text-classification.tabulartext-classification1K<n<10K0 likes788 downloads5mo agoHugging Face03APProjects /texas-layoffs-warn-act-notices-daily Texas WARN Act layoff notices — every filing we hold since 2019, one CSV, rebuilt daily 2,391 Texas WARN notices — every one this dataset holds, back to 2019 — free to download in full: no paywalled years, no login, no account · most recent notice filed 2026-09-15 · state source last checked 2026-09-23T12:26Z · official source: Texas Workforce Commission — WARN notices. Texas employers must file a WARN Act notice with the state before a qualifying mass layoff or plant closing.… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/texas-layoffs-warn-act-notices-daily.tabulartabular-classification1K<n<10K0 likes392 downloads3h agoHugging Face04omar-sharif /BAD-Bengali-Aggressive-Text-Dataset Novel Aggressive Text Dataset in Bengali Tackling Cyber-Aggression: Identification and Fine-Grained Categorization of Aggressive Texts on Social Media using Weighted Ensemble of Transformers Author: Omar Sharif and Mohammed Moshiul Hoque Related Papers: Paper1 in Neurocomputing Journal Paper2 in CONSTRAINT@AAAI-2021 Paper3 in LTEDI@EACL-2021 Abstract The pervasiveness of aggressive content in social media has become a serious concern for government… See the full description on the dataset page: https://huggingface.co/datasets/omar-sharif/BAD-Bengali-Aggressive-Text-Dataset.tabular10K<n<100K3 likes376 downloads5y agoHugging Face05owaiskha9654 /PubMed_MultiLabel_Text_Classification_Dataset_MeSHThis dataset consists of a approx 50k collection of research articles from PubMed repository. Originally these documents are manually annotated by Biomedical Experts with their MeSH labels and each articles are described in terms of 10-15 MeSH labels. In this Dataset we have huge numbers of labels present as a MeSH major which is raising the issue of extremely large output space and severe label sparsity issues. To solve this Issue Dataset has been Processed and mapped to its root as Described… See the full description on the dataset page: https://huggingface.co/datasets/owaiskha9654/PubMed_MultiLabel_Text_Classification_Dataset_MeSH.tabulartext-classification10K<n<100K28 likes229 downloads4y agoHugging Face06zxbsmk /laion_text_debiased_60MFilter zxbsmk/laion_text_debiased_60M by image size and get 512 subset(12,009,641 pairs), 768 subset(4,915,850 pairs), 1024 subset(1,985,026 pairs). image10M<n<100M1 likes219 downloads3y agoHugging Face07vintp /CMU-Mosei-texttabular10K<n<100K2 likes167 downloads2y agoHugging Face08tasal9 /Pashto-Textbooks-PDFs-Corpus Pashto Textbooks and PDFs Corpus Languages: psLicense: cc-by-4.0Task categories: text-generation, feature-extractionSize categories: n<1K Summary This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, feature-extraction tasks in Pashto. How to use from datasets import load_dataset dataset = load_dataset("tasal9/Pashto-Textbooks-PDFs-Corpus") print(dataset) Configs default: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/Pashto-Textbooks-PDFs-Corpus.tabulartext-generationn<1K0 likes159 downloads2mo agoHugging Face09electricsheepafrica /africa-synth-textile-garment-production-all African Textile & Garment Production | Africa (Electric Sheep Africa metadata inventory) Size category: 100K<n<1M - Formats: csv - Sector: culture_language - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-textile-garment-production-all.tabulartabular-classification100K<n<1M0 likes122 downloads1mo agoHugging Face10cyberpsych /PubMed-Cancer-NLP-Textual-Dataset PubMed-Cancer-NLP-Textual-Dataset This dataset has been obtained from PubMed for research purposes. README will be updated with time. Dataset Details Dataset Description It has multiple cancer samples with labels with their title and abstract from PubMed Repository. Curated by: Om Aryan Dataset Sources Repository: https://pubmed.ncbi.nlm.nih.gov tabularfeature-extraction10K<n<100K0 likes106 downloads2y agoHugging Face11israel /Amharic-News-Text-classification-Dataset An Amharic News Text classification Dataset In NLP, text classification is one of the primary problems we try to solve and its uses in language analyses are indisputable. The lack of labeled training data made it harder to do these tasks in low resource languages like Amharic. The task of collecting, labeling, annotating, and making valuable this kind of data will encourage junior researchers, schools, and machine learning practitioners to implement existing classification models… See the full description on the dataset page: https://huggingface.co/datasets/israel/Amharic-News-Text-classification-Dataset.tabular10K<n<100K1 likes91 downloads4y agoHugging Face12leffff /Diffusion-Reward-Modeling-for-Text-Rendering-Dataset 🖼️ Text-to-Image Rendering Dataset A dataset of 14k text prompts for image generation with text rendering evaluation 📚 Dataset Overview This dataset contains 14,000 text prompts specifically designed for: Image generation with text rendering Evaluating text preservation in generated images Training diffusion models for better text rendering Each prompt comes with: Pre-extracted target text for rendering 5 Stable Diffusion 3 generated latents (70k total) Dual… See the full description on the dataset page: https://huggingface.co/datasets/leffff/Diffusion-Reward-Modeling-for-Text-Rendering-Dataset.tabulartext-to-image10K<n<100K8 likes83 downloads1y agoHugging Face13mizanurr /Text2Vis Dataset Card for Text2Vis 📝 Dataset Summary Text2Vis is a benchmark dataset designed to evaluate the ability of large language models (LLMs) to generate accurate and high-quality data visualizations from natural language queries. Each example includes a natural language question, an associated data table, the expected answer, and ground-truth visualization code, along with rich metadata such as chart type, axis labels, domain, complexity, and question types. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/mizanurr/Text2Vis.tabular1K<n<10K1 likes57 downloads1y agoHugging Face14stevengubkin /mathoverflow_text_arxiv_labelsDownloaded from https://archive.org/download/stackexchange Used TexSoup to replace all text in math environments with [UNK]. For instance the text: "The integral $\int_a^b f(x) \textrm{ d}x$ is easy to evaluate if..." was replaced with "The integral [UNK] is easy to evaluate if..." Note: There is still some "ascii math". For instance, people sometimes write things like f: X --> Y. This is retained. Concatenated title and body. Some of these are "answer" posts rather than "question" posts.… See the full description on the dataset page: https://huggingface.co/datasets/stevengubkin/mathoverflow_text_arxiv_labels.tabular10K<n<100K0 likes56 downloads3y agoHugging Face15Yotam /economics-textbooktabularn<1K1 likes53 downloads3y agoHugging Face16cellos /DomesticNames_AllStates_Text Domestic Names from the Federal Government's repository of official geographic names [CSV dataset] This Dataset includes 980,065 geographic names as of September 10, 2023. It is apparent that no currently released LLMs are pretrained on datasets with many of these geographic names (i.e., features), descriptions, and histories. Example: feature_name: Abercrombie Gulch GPT-3.5 responds "I'm not aware of a specific location called Abercrombie Gulch in my training data,..." when prompted… See the full description on the dataset page: https://huggingface.co/datasets/cellos/DomesticNames_AllStates_Text.tabularquestion-answering100K<n<1M2 likes53 downloads3y agoHugging Face17Nininkkka /Synth-Text-Eng-512x128 Synthetic Text Images (English) A synthetic dataset of rendered text images with rich per-sample annotations: the text itself, its rendering attributes, background description, applied post-processing, and a natural-language caption. Each image is generated by compositing English text over a procedurally generated background with random font, color, position, rotation, blur, brightness and noise. All samples are accompanied by a structured metadata.csv and a ready-to-use… See the full description on the dataset page: https://huggingface.co/datasets/Nininkkka/Synth-Text-Eng-512x128.imageimage-to-text10K<n<100K1 likes53 downloads1d agoHugging Face18AIHero123 /Text2Vis Dataset Card for Text2Vis 📝 Dataset Summary Text2Vis is a benchmark dataset designed to evaluate the ability of large language models (LLMs) to generate accurate and high-quality data visualizations from natural language queries. Each example includes a natural language question, an associated data table, the expected answer, and ground-truth visualization code, along with rich metadata such as chart type, axis labels, domain, complexity, and question types. The… See the full description on the dataset page: https://huggingface.co/datasets/AIHero123/Text2Vis.tabular1K<n<10K0 likes50 downloads18d agoHugging Face19agentlans /text-quality Text Quality Assessment Dataset Overview This dataset is designed to assess text quality robustly across various domains for NLP and AI applications. It provides a composite quality score based on multiple classifiers, offering a more comprehensive evaluation of text quality beyond educational domains. Dataset Details Size: 100,000 sentences Source: 20,000 sentences from each of 5 different datasets allenai/c4 HuggingFaceFW/fineweb-edu… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/text-quality.tabulartext-classification100K<n<1M2 likes48 downloads2y agoHugging Face20agentlans /text-quality-v2 Text Quality Meta-Analysis Dataset Dataset Summary The Text Quality Meta-Analysis Dataset is a comprehensive collection of sentences with associated quality metrics derived from multiple sources and methods. It combines text from various sources with quality scores from different models to create a thorough assessment of sentence quality. This dataset is an expanded and streamlined version of agentlans/text-quality. In this context, "quality" refers to legible English… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/text-quality-v2.tabulartext-classification10K<n<100K2 likes47 downloads2y agoHugging Face21mrp /Thai-Semantic-Textual-Similarity-BenchmarkSentence representation plays a crucial role in NLP downstream tasks such as NLI, text classification, and STS. Recent sentence representation training techniques require NLI or STS datasets. However, there are no equivalent Thai NLI or STS datasets for sentence representation training. To address this problem we provide the Thai sentence vector benchmark. We evaluate the Spearman correlation score of the sentence representations’ performance on Thai STS-B (translated version of STS-B).… See the full description on the dataset page: https://huggingface.co/datasets/mrp/Thai-Semantic-Textual-Similarity-Benchmark.tabular1K<n<10K2 likes45 downloads5y agoHugging Face22abdlh /Dataset_Automatic_Essay_Scoring_Essay-EssayScore_and_24_textual_featurestabular10K<n<100K2 likes45 downloads2y agoHugging Face23nphamdinh /textq-german TextQ-German       TextQ investigates how people perceive the quality of machine-generated German text and how these subjective judgments can be modeled automatically. We identified task-specific quality dimensions, quantified them through user ratings, and developed models that predict perceived quality for new generated texts. TextQ-German is a dataset suite for studying the Quality of Experience (QoE) of machine-generated German text. It covers two Natural Language… See the full description on the dataset page: https://huggingface.co/datasets/nphamdinh/textq-german.tabularn<1K1 likes45 downloads1mo agoHugging Face24alex73 /mozilla-common-voice-23-bel-texts-exporttabulartext-generation100K<n<1M0 likes39 downloads10mo agoHugging Face25VerboseTechLabs /textile-manufacturing-egocentric-sample 🧵 Textile Manufacturing — Egocentric Video Dataset (Sample) This dataset is part of a larger collection of egocentric activity datasets by Verbose Tech Labs LLP. If you want the full dataset, or want access to more categories? Get in touch with us: 📞 Phone: +91 7672 000 500 💬 WhatsApp: +91 7672 000 500 📧 Email: Hello@VerboseTechLabs.com 🌐 Website: VerboseTechLabs.com 🔗 More datasets: kaggle.com/verbosetechlabsllp Dataset Summary First-person point-of-view… See the full description on the dataset page: https://huggingface.co/datasets/VerboseTechLabs/textile-manufacturing-egocentric-sample.tabularvideo-classificationn<1K0 likes38 downloads2mo agoHugging Face26jakeazcona /short-text-multi-labeled-emotion-classificationtabular10K<n<100K2 likes36 downloads5y agoHugging Face27zelk12 /text_in_number_smoltalk RU Набор данных содержит в себе текст и его представление в виде 610-ти значного числа. Число полоучено при помощи модели.Исходный набор данных: HuggingFaceTB/smoltalk EN The dataset contains text and its representation as a 610-digit number. The number is hollowed out using model.Initial dataset: HuggingFaceTB/smoltalk tabular1K<n<10K0 likes36 downloads2y agoHugging Face28sobamchan /ja-toxic-text-classification-open2ch Open 2ch-based toxic classification dataset Based on p1atdev/open2ch We apply keyword-based filtering to collect toxic texts We use Perspective API to filter non-toxic texts from the original corpus 3k texts for each class, toxic (label=1) and non-toxic (label=0) texts perspective_api_score is a prediction of toxicity score by the Perspective API tabular1K<n<10K1 likes32 downloads2y agoHugging Face29VerboseTechLabs /textile-industry-manufacturing-egocentric-sample 🧵 Textile Industry Manufacturing — Egocentric Video Dataset (Sample) This dataset is part of a larger collection of egocentric activity datasets by Verbose Tech Labs LLP. If you want the full dataset, or want access to more categories? Get in touch with us: 📞 Phone: +91 7672 000 500 💬 WhatsApp: +91 7672 000 500 📧 Email: Hello@VerboseTechLabs.com 🌐 Website: VerboseTechLabs.com 🔗 More datasets: kaggle.com/verbosetechlabsllp Dataset Summary First-person… See the full description on the dataset page: https://huggingface.co/datasets/VerboseTechLabs/textile-industry-manufacturing-egocentric-sample.tabularvideo-classificationn<1K1 likes32 downloads2mo agoHugging Face30domeist /text-auto-illustrate Text Auto Illustrate — passage-to-image relevance judgements Relevance judgements for illustrating prose: given a paragraph of Wikipedia text, which images from a 5.4-million-image collection actually suit it? Two judgement sets over the same 25 passages — one annotated by hand, one generated and far broader — plus the metadata for every image either set names, so the benchmark can be used without downloading the underlying corpus. Built for a University of Glasgow final-year… See the full description on the dataset page: https://huggingface.co/datasets/domeist/text-auto-illustrate.imagetext-retrieval100K<n<1M0 likes32 downloads3d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.