CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01softcatala /wikimedia-common-audio-catalanThis is a collection of Catalan-language audio with free licenses extracted from Wikimedia Commons. License identifiers are normalized to cc-zero, cc-by-4.0, cc-by-sa-3.0, cc-by-sa-4.0, GFDL, and PD-self. This provides a richer alternative to Common Voice. Characteristics of the dataset: One or multiple speakers Different accents Different domain texts 761 audio files We found this dataset useful for audio tasks such as: Language detection Evaluation of STT systems New candidates are… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/wikimedia-common-audio-catalan.audioautomatic-speech-recognitionn<1K0 likes2.3k downloads2mo agoHugging Face02mnemoraorg /usgs-global-earthquake-catalog USGS Global Earthquake Catalog Provides historical data on global seismic events, sourced directly from the U.S. Geological Survey (USGS) Earthquake Hazards Program via its FDSN Event Web Service. Each record represents a single seismic event (primarily earthquakes) and contains detailed information, including: Event Time & Location: Precise timestamp, geographic coordinates (latitude, longitude), and depth of the event. Magnitude: The magnitude of the event (mag) and the method… See the full description on the dataset page: https://huggingface.co/datasets/mnemoraorg/usgs-global-earthquake-catalog.tabulartext-classification1M<n<10M1 likes2.2k downloads11mo agoHugging Face03kevykibbz /ecommerce-behavior-data-from-multi-category-store_oct-nov_2019 eCommerce Behavior Data from Multi-Category Store About the Dataset This dataset contains behavioral data for 285 million user events from a large multi-category eCommerce store. The data spans 7 months (October 2019 - April 2020) and records various user interactions with products. Dataset Overview Time Frame: October 2019 - April 2020 Total Events: 285 million Event Granularity: Each row represents an event associated with a product and a user. Data Source:… See the full description on the dataset page: https://huggingface.co/datasets/kevykibbz/ecommerce-behavior-data-from-multi-category-store_oct-nov_2019.tabular100M<n<1B4 likes528 downloads2y agoHugging Face04vsak /sms_spam_categorytextn<1K4 likes471 downloads3y agoHugging Face05catherinearnett /morphscore MorphScore MorphScore is a tokenizer evaluation framework, which evaluates the extent to which a tokenizer segments words along morpheme boundaries. This repository contains the datasets used to calculate MorphScore. In total, we have datasetes for 86 languages, but only 70 languages have at least 100 items after filtering. All datasets are derived from existing Universal Dependencies treebanks. In the table below, we link the source dataset for each language. See the new preprint… See the full description on the dataset page: https://huggingface.co/datasets/catherinearnett/morphscore.tabular1M<n<10M4 likes463 downloads1y agoHugging Face06cathv /BATIS license: cc-by-nc-4.0 BATIS: Bayesian Approaches for Targeted Improvement of Species Distribution Models This repository contains the dataset used in experiments shown in BATIS: Bayesian Approaches for Targeted Improvement of Species Distribution Models. To download the dataset, you can use the load_dataset function from HuggingFace. For example : from datasets import load_dataset # Training Split for Kenya training_kenya = load_dataset("cathv/BATIS", name="Kenya"… See the full description on the dataset page: https://huggingface.co/datasets/cathv/BATIS.tabular100K<n<1M0 likes423 downloads10mo agoHugging Face07bigscience-catalogue-data /shades_nationalityPossibly a placeholder dataset for the original here: https://huggingface.co/datasets/bigscience-catalogue-data/bias-shades Data Statement for SHADES How to use this document: Fill in each section according to the instructions. Give as much detail as you can, but there's no need to extrapolate. The goal is to help people understand your data when they approach it. This could be someone looking at it in ten years, or it could be you yourself looking back at the data in two years.… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-catalogue-data/shades_nationality.text10K<n<100K4 likes413 downloads2y agoHugging Face08DoDataThings /us-bank-transaction-categories-v2 US Bank Transaction Categories v2 — Synthetic Dataset 68,000 sign-prefixed transaction descriptions across 17 spending categories, modeled after real US bank statement formats. Designed for training classifiers that work on actual bank data — not the clean "Starbucks coffee" descriptions that most datasets use. Successor to v1. Why This Dataset Real bank transaction data is private. But the formats are universal — Chase, Apple Card, PayPal, Capital One, Mercury all… See the full description on the dataset page: https://huggingface.co/datasets/DoDataThings/us-bank-transaction-categories-v2.texttext-classification10K<n<100K3 likes299 downloads6mo agoHugging Face09Darebal /furniture-catalog Furniture Catalog Image catalog and training splits from the bachelor's thesis "Visual Furnishings Compatibility Learning and Retrieval Using Machine Learning" (Ukrainian Catholic University, 2026). 5 171 individual furniture images across two room types, 1 781 room scene images, plus triplet training data (golden / train / val splits) used to train the compatibility model. Repo layout {room}/ {category}/ *.jpg — individual furniture images… See the full description on the dataset page: https://huggingface.co/datasets/Darebal/furniture-catalog.imageimage-to-image10K<n<100K1 likes262 downloads5mo agoHugging Face10toolathon123 /oceania-gov-open-data-catalog Oceania Government Open Data — Combined Catalogue (hourly snapshot) Combined regional catalogue of Oceania (Australia + New Zealand) public-service open data harvested from both data.gov.au and data.govt.nz portals, including state, territory and local-council publishers. License declaration License: other (see below). Records in this catalogue inherit the licence of their source dataset. Where the source declares a standard open licence the record is tagged with… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/oceania-gov-open-data-catalog.text1K<n<10K0 likes227 downloads1mo agoHugging Face11EthnicErotic /phenotype-catalog Ethnic Erotic Phenotype Catalog A structured complement to Wikipedia for ethnographic data — 1,700+ ethnic groups indexed with normalized linguistic, geographic, cultural, and phenotype metadata, plus 23K+ notable-people references and 5K+ vision-grounded per-image phenotype observations. Curated from the live catalog at ethnicerotic.com and published as an open dataset for anthropological reference, AI training, and ethnographic research. What's in v6 Two columns… See the full description on the dataset page: https://huggingface.co/datasets/EthnicErotic/phenotype-catalog.imagetext-classification10K<n<100K1 likes221 downloads1d agoHugging Face12CataAI /PoseX PoseX: AI Defeats Physics Approaches on Protein-Ligand Cross Docking Paper | GitHub | Leaderboard PoseX is a comprehensive benchmark dataset designed to evaluate molecular docking algorithms for predicting protein-ligand binding poses. It includes curated datasets for both self-docking and cross-docking scenarios. The file posex_set.zip contains the processed dataset for docking evaluation, while posex_cif.zip contains the raw CIF files from RCSB PDB. For information about creating… See the full description on the dataset page: https://huggingface.co/datasets/CataAI/PoseX.textother1K<n<10K8 likes212 downloads5mo agoHugging Face13appleboiy /narit-ghosts-halo-catalogs NARIT GHOSTS Halo Catalogs Dataset Description This dataset contains reduced stellar catalogs, combined FITS images, and candidate substructure catalogs from the GHOSTS (Galaxy Halos, Outer disks, Substructure, Thick disks, and Star clusters) Survey observed by the Hubble Space Telescope (HST). It serves as the primary data lake for the automated astronomical pipeline designed to detect faint stellar substructures (like Ultra-Faint Dwarfs and stellar streams) in… See the full description on the dataset page: https://huggingface.co/datasets/appleboiy/narit-ghosts-halo-catalogs.tabularn<1K0 likes207 downloads2mo agoHugging Face14nbel /CatCoLA GENERAL INFORMATION 1. Dataset title: CatCoLA - Catalan Corpus of Linguistic Acceptability 2. Authorship: Name: Núria Bel Institution: Universitat Pompeu Fabra, UPF Email: nuria.bel@upf.edu ORCID: 0000-0001-9346-7803 Name: Marta Punsola Institution: Universitat Pompeu Fabra, UPF Email: marta.punsola@gmail.com ORCID: Name: Valle Ruiz-Fernández Institution: Barcelona Supercomputing Center Email: valle.ruizfernández@bsc.es ORCID: DESCRIPTION… See the full description on the dataset page: https://huggingface.co/datasets/nbel/CatCoLA.tabular10K<n<100K1 likes171 downloads2y agoHugging Face15USC /USC-Course-Catalog USC Course Catalog - 2024, Spring Term This dataset consists of all classes provided by USC (as of December 1, 2023) that USC is providing in 2024 Spring. While it is a small dataset, this could be used in some finetuning, generation, or RAG application tasks. One example would be this -> https://huggingface.co/spaces/USC/USC-GPT I will also be scraping the 2024 fall term classes when they are released by USC! If you want the web scraping script I used for this, feel free to send me… See the full description on the dataset page: https://huggingface.co/datasets/USC/USC-Course-Catalog.text1K<n<10K1 likes161 downloads3y agoHugging Face16yassiracharki /Yahoo_Answers_10_categories_for_NLP Dataset Card for Dataset Name The Yahoo! Answers topic classification dataset is constructed using 10 largest main categories. Each class contains 140,000 training samples and 6,000 testing samples. Therefore, the total number of training samples is 1,400,000 and testing samples 60,000 in this dataset. From all the answers and other meta-information, we only used the best answer content and the main category information. Dataset Description The file classes.txt contains a… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Yahoo_Answers_10_categories_for_NLP.texttext-classification1M<n<10M3 likes136 downloads2y agoHugging Face17dirtycomputer /online_shopping_10_catstext10K<n<100K1 likes113 downloads4y agoHugging Face18Javtor /biomedical-topic-categorizationtext1M<n<10M0 likes110 downloads4y agoHugging Face19letrinhan /vn-provinces-cattle-count Vietnam provinces cattle count Number of cattle (bo) (thousand heads). Coverage 1995-2024. Year 2024 is preliminary. Includes historical Ha Tay through 2007. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system). Figures Hero Hero (continued) Comparison Color key Files provinces (1866 rows)… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-cattle-count.tabular1K<n<10K0 likes89 downloads2d agoHugging Face20zalizedata /shopify-products-catalog-sample Shopify Products Catalog with Price History (DTC Stores) — Free Sample A free sample of a Shopify e-commerce catalog dataset built from public, login-free products.json endpoints of independent DTC (direct-to-consumer) Shopify stores. This sample build covers 386 stores / 533,479 SKU rows (2026-08-01 snapshot). Full-run scale on record: 4,100+ stores / 6.18M SKUs, refreshed with weekly snapshots that also yield per-variant price-change and availability-change events. ➡️ Full… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/shopify-products-catalog-sample.tabulartabular-classificationn<1K0 likes87 downloads2mo agoHugging Face21valurank /News_Articles_Categorization Dataset Card for News_Articles_Categorization Dataset Description 3722 News Articles classified into different categories namely: World, Politics, Tech, Entertainment, Sport, Business, Health, and Science Languages The text in the dataset is in English Dataset Structure The dataset consists of two columns namely Text and Category. The Text column consists of the news article and the Category column consists of the class each article belongs to… See the full description on the dataset page: https://huggingface.co/datasets/valurank/News_Articles_Categorization.texttext-classification1K<n<10K5 likes84 downloads3y agoHugging Face22SankaraEyeHospital /Catacomp-104gated CataCompDetect Video dataset of cataract surgeries labeled for surgical complications (Posterior Capsule Rupture, Vitreous Loss, Iris Prolapse), used to train and evaluate the CataCompDetect complication-detection pipeline. Dataset structure catacomp-hf/ ├── train/ │ ├── metadata.csv │ ├── catacomp-train-001.mp4 │ ├── catacomp-train-002.mp4 │ └── ... └── val/ ├── metadata.csv ├── catacomp-val-001.mp4 ├── catacomp-val-002.mp4 └── ... Videos… See the full description on the dataset page: https://huggingface.co/datasets/SankaraEyeHospital/Catacomp-104.textvideo-classificationn<1K0 likes82 downloads4d agoHugging Face23kunikohunter /CatPred-DB CatPred-DB: Enzyme Kinetic Parameters Database Paper: CatPred: A comprehensive framework for deep learning in vitro enzyme kinetic parameters GitHub: https://github.com/maranasgroup/CatPred-DB Dataset Description CatPred-DB contains the benchmark datasets introduced alongside the CatPred deep learning framework for predicting in vitro enzyme kinetic parameters. The datasets cover three key kinetic parameters: Parameter Description Datapoints kcat Turnover… See the full description on the dataset page: https://huggingface.co/datasets/kunikohunter/CatPred-DB.tabular10K<n<100K0 likes80 downloads7mo agoHugging Face24Phoenyx83 /Politifact-fake-news-6-categories-for-llama3-1 Dataset compiled for the article "LLaMA 3 vs. State-of-the-Art LLMs: Performance in Detecting Nuanced Fake News" based on Politifact Factcheck Data, available at https://www.kaggle.com/datasets/shivkumarganesh/politifact-factcheck-data language:" - en license: llama3.1 tabular10K<n<100K1 likes78 downloads2y agoHugging Face25letrinhan /vn-provinces-cattle-meat-production Vietnam provinces cattle meat production Cattle meat production (liveweight) (thousand tons). Coverage 2018-2024. Year 2024 is preliminary. Source values in tons are converted to thousand tons. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system). Figures Hero Comparison Color key Files provinces (441… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-cattle-meat-production.tabularn<1K0 likes72 downloads2d agoHugging Face26CATIE-AQ /french_narrativeqa Description Dataframe containing 143 French books in txt format.More precisely : the texte column contains the texts the titre column contains the book title the auteur column contains the author's name and dates of birth and death (if you want to filter the texts to keep only those from the given century to the present day) the question column contains a single question asked about the associated text the answers column contains one or more answers to the question (= if several… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/french_narrativeqa.textquestion-answering1K<n<10K1 likes70 downloads1y agoHugging Face27Sumeetgpt /indian-transaction-categorization-synthetic Synthetic Indian Bank Transaction Narrations 810 synthetic (text, category) pairs mimicking Indian bank/credit-card statement narrations — built to train the Sumeetgpt/indian-transaction-categorizer SetFit model. Why this exists While building a personal finance app, we searched for a public dataset pairing real Indian transaction narration formats (UPI, NEFT, IMPS, ACH) with spending-category labels, and found none: datasets with real-looking Indian narration… See the full description on the dataset page: https://huggingface.co/datasets/Sumeetgpt/indian-transaction-categorization-synthetic.texttext-classificationn<1K0 likes70 downloads22d agoHugging Face28NLPC-UOM /Sinhala-News-Category-classificationThis file contains news texts (sentences) belonging to 5 different news categories (political, business, technology, sports and Entertainment). The original dataset was released by Nisansa de Silva (Sinhala Text Classification: Observations from the Perspective of a Resource Poor Language, 2015). The original dataset is processed and cleaned of single word texts, English only sentences etc. If you use this dataset, please cite {Nisansa de Silva, Sinhala Text Classification: Observations from… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/Sinhala-News-Category-classification.texttext-classification1K<n<10K2 likes64 downloads4y agoHugging Face29xcz0 /Aspect-Based_Sentiment_Analysis_for_Catering 说明 数据集来源于AI Challenger 2018 sentiment_analysis_trainingset.csv 为训练集数据文件,共105000条评论数据 sentiment_analysis_validationset.csv 为验证集数据文件,共15000条评论数据 sentiment_analysis_testa.csv 为测试集A数据文件,共15000条评论数据 数据集分为训练、验证、测试A与测试B四部分。数据集中的评价对象按照粒度不同划分为两个层次,层次一为粗粒度的评价对象,例如评论文本中涉及的服务、位置等要素;层次二为细粒度的情感对象,例如“服务”属性中的“服务人员态度”、“排队等候时间”等细粒度要素。评价对象的具体划分如下表所示。 The dataset is divided into four parts: training, validation, test A and test B. This dataset builds a two-layer labeling system according to the… See the full description on the dataset page: https://huggingface.co/datasets/xcz0/Aspect-Based_Sentiment_Analysis_for_Catering.tabulartext-classification100K<n<1M0 likes64 downloads3y agoHugging Face30katsukiono /libero-plus-episode-categories LIBERO-Plus episode → perturbation-category labels Which perturbation category each of the 14,347 LIBERO-Plus episodes belongs to. LIBERO-Plus perturbs a base LIBERO task along several axes. The release these labels came from contains five of them --- the project describes more, so treat this as the taxonomy of this 14,347-episode release, not of LIBERO-Plus as a whole. The lerobot/libero_plus conversion does not carry that label, so you cannot ask "how does my policy do under… See the full description on the dataset page: https://huggingface.co/datasets/katsukiono/libero-plus-episode-categories.tabularrobotics10K<n<100K0 likes64 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.