CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LHF /escorpius-mr esCorpius Multilingual Raw In the recent years, Transformer-based models have lead to significant advances in language modelling for natural language processing. However, they require a vast amount of data to be (pre-)trained and there is a lack of corpora in languages other than English. Recently, several initiatives have presented multilingual datasets obtained from automatic web crawling. However, they present important shortcomings for languages different from English, as they… See the full description on the dataset page: https://huggingface.co/datasets/LHF/escorpius-mr.texttext-generation1B<n<10B5 likes4.3k downloads3y agoHugging Face02ashraq /esc50https://github.com/karolpiczak/ESC-50 The dataset is available under the terms of the Creative Commons Attribution Non-Commercial license. K. J. Piczak. ESC: Dataset for Environmental Sound Classification. Proceedings of the 23rd Annual ACM Conference on Multimedia, Brisbane, Australia, 2015. [DOI: http://dx.doi.org/10.1145/2733373.2806390] audio1K<n<10K38 likes3.6k downloads4y agoHugging Face03tasksource /esci Dataset Card for "esci" ESCI product search dataset https://github.com/amazon-science/esci-data/ Preprocessings: -joined the two relevant files -product_text aggregate all product text -mapped esci_label to full name @article{reddy2022shopping, title={Shopping Queries Dataset: A Large-Scale {ESCI} Benchmark for Improving Product Search}, author={Chandan K. Reddy and Lluís Màrquez and Fran Valero and Nikhil Rao and Hugo Zaragoza and Sambaran Bandyopadhyay and Arnab Biswas and Anlu… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/esci.tabulartext-classification1M<n<10M8 likes2.2k downloads3y agoHugging Face04thu-coai /esconvThe ESConv dataset. GitHub repo. Original paper. @inproceedings{liu-etal-2021-towards, title={Towards Emotional Support Dialog Systems}, author={Liu, Siyang and Zheng, Chujie and Demasi, Orianna and Sabour, Sahand and Li, Yu and Yu, Zhou and Jiang, Yong and Huang, Minlie}, booktitle={ACL}, year={2021} } text1K<n<10K25 likes1.1k downloads3y agoHugging Face05nbel /EsCoLA Introduction The Spanish Corpus of Linguistic Acceptability (EsCoLA) includes 11,174 sentences taken from linguistic literature with a binary annotation made by the original authors themselves. The work is inspired by CoLA: https://nyu-mll.github.io/CoLA/# Paper Núria Bel, Marta Punsola, Valle Ruiz-Fernández, 2024, EsCoLA: Spanish Corpus of Linguistic Acceptability. Joint International Conference on Computational Linguistics, Language Resources and Evaluation LREC-COLING… See the full description on the dataset page: https://huggingface.co/datasets/nbel/EsCoLA.tabular1K<n<10K2 likes990 downloads2y agoHugging Face06Stanwang1210 /raw_tts_esc_ESPnet_espnet_mls-audioset_soundstream_16ktextn<1K0 likes935 downloads2y agoHugging Face07LHF /escorpiusSpanish datasettexttext-generation1M<n<10M17 likes750 downloads4y agoHugging Face08Stanwang1210 /raw_tts_esc_ESPnet_espnet_mls-multi_soundstream_16ktextn<1K0 likes703 downloads2y agoHugging Face09confit /esc50-parquetaudioaudio-classification10K<n<100K1 likes696 downloads2y agoHugging Face10escorciav /synthcix-3m_br SynthCIX-3M Dataset Welcome to the Synthcix-3m_br dataset! This dataset contains a rich collection of synthetic images and captions, designed to support a variety of machine learning and computer vision tasks. Whether you're working on image generation, captioning, or any other related project, this dataset is a valuable resource. Dataset Overview The SynthCIX-3M dataset includes: 3 million synthetic images: High-quality images generated using advanced techniques.… See the full description on the dataset page: https://huggingface.co/datasets/escorciav/synthcix-3m_br.tabular1M<n<10M0 likes568 downloads1y agoHugging Face11bifold-pathomics /PathoROB-tolkach_esca PathoROB Preprint | Code | Licenses | Cite PathoROB is a benchmark for the robustness of pathology foundation models (FMs) to non-biological medical center differences. PathoROB contains four datasets covering 28 biological classes from 34 medical centers and three metrics: Robustness Index: Measures the dominance of biological over non-biological features in an FM representation space. Average Performance Drop (APD): Measures the robustness of downstream models to shortcut… See the full description on the dataset page: https://huggingface.co/datasets/bifold-pathomics/PathoROB-tolkach_esca.imageimage-feature-extraction10K<n<100K0 likes475 downloads9mo agoHugging Face12mteb /ESCIReranking ESCIReranking An MTEB dataset Massive Text Embedding Benchmark Task category t2t Domains Written Reference https://github.com/amazon-science/esci-data/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["ESCIReranking"]) evaluator = mteb.MTEB(task) model = mteb.get_model(YOUR_MODEL) evaluator.run(model) To learn more about how to run models on mteb task check out… See the full description on the dataset page: https://huggingface.co/datasets/mteb/ESCIReranking.texttext-ranking1M<n<10M0 likes392 downloads1y agoHugging Face13CodecSR /esc50_synthaudio10K<n<100K0 likes306 downloads3y agoHugging Face14Stanwang1210 /raw_tts_esc_ESPnet_espnet_mls-english_soundstream_16ktextn<1K0 likes298 downloads2y agoHugging Face15milistu /amazon-esci-data Amazon Shopping Queries Dataset Dataset for improving product search, ranking and recommendations, featuring query-product pairs with detailed relevance labels. Overview The dataset contains search queries paired with up to 40 potentially relevant products, each labeled using the ESCI system: Exact match: Products that perfectly match the customer's search intent (e.g., searching "iPhone 13" and finding "Apple iPhone 13 128GB") Substitute product: Alternative products… See the full description on the dataset page: https://huggingface.co/datasets/milistu/amazon-esci-data.tabulartext-classification1M<n<10M2 likes277 downloads1y agoHugging Face16xiaoluo11 /escape-simulator-gameplay-data 密室逃脱模拟器 This public dataset repository contains local gameplay data uploaded from F:\密室逃脱模拟器. Contents Files: 211 Total local size: 102.72 GB Generated: 2026-06-05 02:51:23 UTC File Types .jsonl: 62 .png: 52 .json: 49 .mkv: 17 .txt: 16 .parquet: 15 Notes This repository may contain gameplay video, images, Parquet files, JSON/JSONL metadata, and keyboard/mouse event logs. The license is marked as other; review game footage, audio… See the full description on the dataset page: https://huggingface.co/datasets/xiaoluo11/escape-simulator-gameplay-data.imagereinforcement-learningn<1K0 likes256 downloads4mo agoHugging Face17ESCAD /OpenRTLSet Dataset Card for OpenRTLSet: A Fully Open-Source Dataset for Large Language Model-based Verilog Module Design Authors: Jinghua Wang, Lily Jiaxin Wan, Sanjana Pingali, Scott Smith, Manvi Jha, Shalini Sivakumar, Xing Zhao, Kaiwen Cao, Deming Chen Dataset Summary This is the 131k dataset generated in our paper: OpenRTLSet: A Fully Open-Source Dataset for Large Language Model-based Verilog Module Design. In this work, we introduce OpenRTLSet: a… See the full description on the dataset page: https://huggingface.co/datasets/ESCAD/OpenRTLSet.texttext-generation100K<n<1M2 likes235 downloads4mo agoHugging Face18gabrielaltay /tcga-esca-tabular-open TCGA-ESCA — Tabular (Open Access) Open-access TCGA-ESCA data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV. GDC data release: Data Release 46.0 - August 10, 2026 Built: 2026-09-12 03:54:06 UTC Scope: one TCGA project — see [the family][repo] for the others from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-esca-tabular-open.tabular100M<n<1B0 likes232 downloads11d agoHugging Face19TechWolf /Skill-normalisation-ESCO-graded skill-normalisation-esco-graded Graded-relevance annotations for surface skill terms (ESCO alt-labels) from ESCO v1.1.0 skill-normalisation pairs against the ESCO v1.1.0 skill taxonomy. Layout follows the BEIR convention so it is drop-in for MTEB-style retrieval evaluators. This dataset was created for the RecSys-HR 2026 WorkRB challenge. Configs config split rows columns queries validation 50 _id (query id), text (ESCO alt-label / surface term to… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/Skill-normalisation-ESCO-graded.text1M<n<10M0 likes225 downloads1mo agoHugging Face20LuisG07 /es_corpora_parliament_processedtext1M<n<10M0 likes210 downloads5y agoHugging Face21TechWolf /Synthetic-ESCO-skill-sentences Synthetic job ads for all ESCO skills Dataset Summary This dataset contains 10 synthetically generated job ad sentences for almost all (99.5%) skills in ESCO v1.1.0. Languages We use the English version of ESCO, and all generated sentences are in English. Dataset Structure The dataset consists of 138,260 (sentence, skill) pairs. Citation Information If you use this dataset, please include the following reference:… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/Synthetic-ESCO-skill-sentences.texttext-classification100K<n<1M17 likes169 downloads2y agoHugging Face22FatimahEmadEldin /agent-intrusion-escalation-forensics Both Sides Detected It, Neither Escalated: Concurrency and Escalation Failure in the July 2026 Autonomous Agent Intrusion This repository contains the corpus, ingestion pipeline and report for a forensic reconstruction of the July 2026 autonomous agent intrusion, submitted to the Apart Research & CeSIA AI Incident Response Sprint, Track 2 (Forensics and Forecasting). By: Fatimah Mohamed Emad Elden Trouve Labs Detection was not the binding… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/agent-intrusion-escalation-forensics.documentn<1K0 likes143 downloads9d agoHugging Face23smangrul /amazon_escitabular1M<n<10M3 likes133 downloads3y agoHugging Face24EscheWang /3dcs 3DCS: Datasets and Benchmark for Evaluating Conformational Sensitivity in Molecular Representations 3DCS is a benchmark for 3D conformational sensitivity in molecular representations (MRs). It tests whether the representations of different conformers of the same molecule (i) preserve geometric variation, (ii) capture chirality, and (iii) reflect the energy landscape. This is the Geometry–Chirality–Energy (GCE) evaluation framework from the ICLR 2026 paper. This repository holds… See the full description on the dataset page: https://huggingface.co/datasets/EscheWang/3dcs.tabular1M<n<10M0 likes125 downloads5d agoHugging Face25esclient /toxicity_multilanguage_datasettext10K<n<100K0 likes124 downloads5mo agoHugging Face26CodecSR /esc50_24k_synthaudio10K<n<100K0 likes119 downloads3y agoHugging Face27azuur /es_corpora_parliament_processedtext1M<n<10M0 likes117 downloads5y agoHugging Face28Escapist-X /pvl_sftimage10K<n<100K0 likes113 downloads9mo agoHugging Face29darkknight25 /APT_STYLE_Privilege_Escalation_Dataset APT Privilege Escalation Dataset Overview The APT Privilege Escalation Dataset is a comprehensive collection of advanced and unique privilege escalation techniques tailored for Red Team training and offensive cybersecurity operations. This dataset, comprising 1000 entries, simulates real-world Advanced Persistent Threat (APT) tactics, focusing on exploiting misconfigurations, vulnerabilities, and novel attack vectors to achieve elevated privileges on Linux-based systems.… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/APT_STYLE_Privilege_Escalation_Dataset.text1K<n<10K0 likes110 downloads1y agoHugging Face30CLAPv2 /esc50_no_overlapaudio1K<n<10K0 likes109 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.