CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LHF /escorpius-mr esCorpius Multilingual Raw In the recent years, Transformer-based models have lead to significant advances in language modelling for natural language processing. However, they require a vast amount of data to be (pre-)trained and there is a lack of corpora in languages other than English. Recently, several initiatives have presented multilingual datasets obtained from automatic web crawling. However, they present important shortcomings for languages different from English, as they… See the full description on the dataset page: https://huggingface.co/datasets/LHF/escorpius-mr.texttext-generation1B<n<10B5 likes6.6k downloads3y agoHugging Face02ashraq /esc50https://github.com/karolpiczak/ESC-50 The dataset is available under the terms of the Creative Commons Attribution Non-Commercial license. K. J. Piczak. ESC: Dataset for Environmental Sound Classification. Proceedings of the 23rd Annual ACM Conference on Multimedia, Brisbane, Australia, 2015. [DOI: http://dx.doi.org/10.1145/2733373.2806390] audio1K<n<10K38 likes3.7k downloads4y agoHugging Face03tasksource /esci Dataset Card for "esci" ESCI product search dataset https://github.com/amazon-science/esci-data/ Preprocessings: -joined the two relevant files -product_text aggregate all product text -mapped esci_label to full name @article{reddy2022shopping, title={Shopping Queries Dataset: A Large-Scale {ESCI} Benchmark for Improving Product Search}, author={Chandan K. Reddy and Lluís Màrquez and Fran Valero and Nikhil Rao and Hugo Zaragoza and Sambaran Bandyopadhyay and Arnab Biswas and Anlu… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/esci.tabulartext-classification1M<n<10M8 likes2.2k downloads3y agoHugging Face04TREA-ORCA /dataset_esc50_full_v1audio10K<n<100K0 likes1.4k downloads26d agoHugging Face05thu-coai /esconvThe ESConv dataset. GitHub repo. Original paper. @inproceedings{liu-etal-2021-towards, title={Towards Emotional Support Dialog Systems}, author={Liu, Siyang and Zheng, Chujie and Demasi, Orianna and Sabour, Sahand and Li, Yu and Yu, Zhou and Jiang, Yong and Huang, Minlie}, booktitle={ACL}, year={2021} } text1K<n<10K25 likes1.1k downloads3y agoHugging Face06nbel /EsCoLA Introduction The Spanish Corpus of Linguistic Acceptability (EsCoLA) includes 11,174 sentences taken from linguistic literature with a binary annotation made by the original authors themselves. The work is inspired by CoLA: https://nyu-mll.github.io/CoLA/# Paper Núria Bel, Marta Punsola, Valle Ruiz-Fernández, 2024, EsCoLA: Spanish Corpus of Linguistic Acceptability. Joint International Conference on Computational Linguistics, Language Resources and Evaluation LREC-COLING… See the full description on the dataset page: https://huggingface.co/datasets/nbel/EsCoLA.tabular1K<n<10K2 likes991 downloads2y agoHugging Face07Stanwang1210 /raw_tts_esc_ESPnet_espnet_mls-audioset_soundstream_16ktextn<1K0 likes932 downloads2y agoHugging Face08LHF /escorpiusSpanish datasettexttext-generation1M<n<10M17 likes757 downloads4y agoHugging Face09hf-internal-testing /ashraq-esc50-1-dog-example Dataset Card for "ashraq-esc50-1-dog-example" More Information needed audion<1K0 likes720 downloads2y agoHugging Face10Stanwang1210 /raw_tts_esc_ESPnet_espnet_mls-multi_soundstream_16ktextn<1K0 likes697 downloads2y agoHugging Face11confit /esc50-parquetaudioaudio-classification10K<n<100K1 likes696 downloads2y agoHugging Face12escontra /gauss_gym_data3dn<1K2 likes632 downloads1y agoHugging Face13escorciav /synthcix-3m_br SynthCIX-3M Dataset Welcome to the Synthcix-3m_br dataset! This dataset contains a rich collection of synthetic images and captions, designed to support a variety of machine learning and computer vision tasks. Whether you're working on image generation, captioning, or any other related project, this dataset is a valuable resource. Dataset Overview The SynthCIX-3M dataset includes: 3 million synthetic images: High-quality images generated using advanced techniques.… See the full description on the dataset page: https://huggingface.co/datasets/escorciav/synthcix-3m_br.tabular1M<n<10M0 likes592 downloads1y agoHugging Face14EscheWang /3dcs-embeddings 3DCS baseline embeddings This repository holds the embedding files of the baseline molecular representations evaluated in the ICLR 2026 paper 3DCS: Datasets and Benchmark for Evaluating Conformational Sensitivity in Molecular Representations. It also holds the metric outputs of the original evaluation runs and the rMD17 split files. The files are the original bytes produced in the authors' 2025 runs. Only directory names were normalized. File names, array keys and contents are… See the full description on the dataset page: https://huggingface.co/datasets/EscheWang/3dcs-embeddings.3d10M<n<100M0 likes501 downloads4d agoHugging Face15Image-editing /escher-ss2 Dataset Card for escher-ss2 SomethingSomethingv2 dataset Dataset Structure Data Instances Each instance contains: source_image: The original image edited_image: The edited version of the image edit_instruction: The instruction used to edit the image source_image_caption: Caption for the source image target_image_caption: Caption for the edited image Additional metadata fields Data Splits {} image-to-imagen<1K0 likes475 downloads1y agoHugging Face16bifold-pathomics /PathoROB-tolkach_esca PathoROB Preprint | Code | Licenses | Cite PathoROB is a benchmark for the robustness of pathology foundation models (FMs) to non-biological medical center differences. PathoROB contains four datasets covering 28 biological classes from 34 medical centers and three metrics: Robustness Index: Measures the dominance of biological over non-biological features in an FM representation space. Average Performance Drop (APD): Measures the robustness of downstream models to shortcut… See the full description on the dataset page: https://huggingface.co/datasets/bifold-pathomics/PathoROB-tolkach_esca.imageimage-feature-extraction10K<n<100K0 likes473 downloads9mo agoHugging Face17mteb /ESCIReranking ESCIReranking An MTEB dataset Massive Text Embedding Benchmark Task category t2t Domains Written Reference https://github.com/amazon-science/esci-data/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["ESCIReranking"]) evaluator = mteb.MTEB(task) model = mteb.get_model(YOUR_MODEL) evaluator.run(model) To learn more about how to run models on mteb task check out… See the full description on the dataset page: https://huggingface.co/datasets/mteb/ESCIReranking.texttext-ranking1M<n<10M0 likes403 downloads1y agoHugging Face18TigreGotico /ESC-50 ESC-50: Dataset for Environmental Sound Classification Overview | Download | Results | Repository content | License | Citing | Caveats | Changelog       The ESC-50 dataset is a labeled collection of 2000 environmental audio recordings suitable for benchmarking methods of environmental sound classification. The dataset consists of 5-second-long recordings organized into 50 semantical classes (with 40 examples per class) loosely arranged into 5 major categories:… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/ESC-50.audio1K<n<10K0 likes391 downloads11mo agoHugging Face19Image-editing /escher-aurora-kubric Dataset Card for escher-aurora-kubric Aurora-Kubric dataset Dataset Structure Data Instances Each instance contains: source_image: The original image edited_image: The edited version of the image edit_instruction: The instruction used to edit the image source_image_caption: Caption for the source image target_image_caption: Caption for the edited image Additional metadata fields Data Splits {} image-to-imagen<1K0 likes333 downloads1y agoHugging Face20CodecSR /esc50_synthaudio10K<n<100K0 likes306 downloads3y agoHugging Face21Stanwang1210 /raw_tts_esc_ESPnet_espnet_mls-english_soundstream_16ktextn<1K0 likes298 downloads2y agoHugging Face22projecte-aina /escagleu-64k Dataset Card for escagleu-64K corpus Dataset Description Dataset Summary This is the second version of escagleu-64k, a parallel corpus containing approximately 64k sentences translated across Spanish, Catalan, Valencian Catalan, Galician, and Basque. The original sentences are in Spanish and are sourced from the Spanish Common Voice Corpus. This corpus was prepared with the goal of creating a parallel speech dataset for these languages using the Common Voice… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/escagleu-64k.translation10K<n<100K0 likes295 downloads6mo agoHugging Face23WiserHumanExperimental /evilgenie-escalation EvilGenie × Escalation Channels — data release Run transcripts and per-sample analysis tables for the paper "Can escalation channels redirect reward hacking toward defect disclosure?" (F. Gomez, Wiser Human, 2026). Code: https://github.com/wiser-human-experimental/evilgenie-escalation Paper: https://arxiv.org/abs/2608.29460 What's here A coding agent is given an ambiguous competitive-programming problem (LiveCodeBench), a visible test suite, and sandboxed… See the full description on the dataset page: https://huggingface.co/datasets/WiserHumanExperimental/evilgenie-escalation.text-generation1K<n<10K0 likes288 downloads13d agoHugging Face24milistu /amazon-esci-data Amazon Shopping Queries Dataset Dataset for improving product search, ranking and recommendations, featuring query-product pairs with detailed relevance labels. Overview The dataset contains search queries paired with up to 40 potentially relevant products, each labeled using the ESCI system: Exact match: Products that perfectly match the customer's search intent (e.g., searching "iPhone 13" and finding "Apple iPhone 13 128GB") Substitute product: Alternative products… See the full description on the dataset page: https://huggingface.co/datasets/milistu/amazon-esci-data.tabulartext-classification1M<n<10M2 likes276 downloads1y agoHugging Face25Image-editing /escher-aurora-ag Dataset Card for escher-aurora-ag Aurora-AG dataset Dataset Structure Data Instances Each instance contains: source_image: The original image edited_image: The edited version of the image edit_instruction: The instruction used to edit the image source_image_caption: Caption for the source image target_image_caption: Caption for the edited image Additional metadata fields Data Splits {} image-to-imagen<1K0 likes263 downloads1y agoHugging Face26likaili /escucho-mucho-audio Escucho Mucho — audio Short Spanish speech clips (MP3, 24 kHz mono, ~64 kbps) used by the Escucho Mucho listening-practice app. Nothing here is original: the recordings are re-encoded copies of public speech corpora, republished so the app can stream them to a phone. Each accent lives in its own folder; a clip's transcript, timings and difficulty live in the app's own library index, not in this repo. Folder Source Licence co/ OpenSLR SLR72 — Colombian Spanish CC BY-SA… See the full description on the dataset page: https://huggingface.co/datasets/likaili/escucho-mucho-audio.audioautomatic-speech-recognition10K<n<100K0 likes244 downloads8d agoHugging Face27TechWolf /Skill-normalisation-ESCO-graded skill-normalisation-esco-graded Graded-relevance annotations for surface skill terms (ESCO alt-labels) from ESCO v1.1.0 skill-normalisation pairs against the ESCO v1.1.0 skill taxonomy. Layout follows the BEIR convention so it is drop-in for MTEB-style retrieval evaluators. This dataset was created for the RecSys-HR 2026 WorkRB challenge. Configs config split rows columns queries validation 50 _id (query id), text (ESCO alt-label / surface term to… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/Skill-normalisation-ESCO-graded.text1M<n<10M0 likes236 downloads1mo agoHugging Face28abd1987 /esco-embeddings-numpy0 likes230 downloads2y agoHugging Face29xiaoluo11 /escape-simulator-gameplay-data 密室逃脱模拟器 This public dataset repository contains local gameplay data uploaded from F:\密室逃脱模拟器. Contents Files: 211 Total local size: 102.72 GB Generated: 2026-06-05 02:51:23 UTC File Types .jsonl: 62 .png: 52 .json: 49 .mkv: 17 .txt: 16 .parquet: 15 Notes This repository may contain gameplay video, images, Parquet files, JSON/JSONL metadata, and keyboard/mouse event logs. The license is marked as other; review game footage, audio… See the full description on the dataset page: https://huggingface.co/datasets/xiaoluo11/escape-simulator-gameplay-data.imagereinforcement-learningn<1K0 likes230 downloads4mo agoHugging Face30gabrielaltay /tcga-esca-tabular-open TCGA-ESCA — Tabular (Open Access) Open-access TCGA-ESCA data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV. GDC data release: Data Release 46.0 - August 10, 2026 Built: 2026-09-12 03:54:06 UTC Scope: one TCGA project — see [the family][repo] for the others from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-esca-tabular-open.tabular100M<n<1B0 likes223 downloads10d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.