CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01opensporks /resumes Dataset Card for Resume Dataset Dataset Summary Context A collection of Resume Examples taken from livecareer.com for categorizing a given resume into any of the labels defined in the dataset. Content Contains 2400+ Resumes in string as well as PDF format. PDF stored in the data folder differentiated into their respective labels as folders with each resume residing inside the folder in pdf form with filename as the id defined in the csv. Inside the… See the full description on the dataset page: https://huggingface.co/datasets/opensporks/resumes.text1K<n<10K14 likes8.8k downloads2y agoHugging Face02hf-audio /open-asr-leaderboard-resultstabularn<1K0 likes6.2k downloads2d agoHugging Face03ibm-research /argument_quality_ranking_30k Dataset Card for Argument-Quality-Ranking-30k Dataset Dataset Summary Argument Quality Ranking The dataset contains 30,497 crowd-sourced arguments for 71 debatable topics labeled for quality and stance, split into train, validation and test sets. The dataset was originally published as part of our paper: A Large-scale Dataset for Argument Quality Ranking: Construction and Analysis. Argument Topic This subset contains 9,487 of the arguments only with… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/argument_quality_ranking_30k.tabulartext-classification10K<n<100K13 likes1.9k downloads3y agoHugging Face04Anthropic /enabling-independent-research Overview This directory contains the Anthropic Insights data we provided to our three external research groups as part of the collaboration detailed in "Enabling independent research on how people use Claude". Before using this data, we recommend first reading our blog post on this collaboration and the Anthropic Insights paper and blog post. Before drawing conclusions from this data — especially from open-ended clusters — please read "Guidance for Interpreting Open-Ended… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/enabling-independent-research.tabular1K<n<10K36 likes1.3k downloads1mo agoHugging Face05cnamuangtoun /resume-job-description-fittext1K<n<10K81 likes1.3k downloads2y agoHugging Face06manycore-research /SpatialLM-Testset SpatialLM Testset Project page | Paper | Code We provide a test set of 107 preprocessed point clouds and their corresponding GT layouts, point clouds are reconstructed from RGB videos using MASt3R-SLAM. SpatialLM-Testset is quite challenging compared to prior clean RGBD scan datasets due to the noises and occlusions in the point clouds reconstructed from monocular RGB videos. Folder Structure Outlines of the dataset files:… See the full description on the dataset page: https://huggingface.co/datasets/manycore-research/SpatialLM-Testset.3dn<1K60 likes1.3k downloads1y agoHugging Face07Kaludi /Customer-Support-Responsestextn<1K13 likes1.2k downloads4y agoHugging Face08Shanmuk4622 /E2AM_ResNet50 E2AM Ablation Results: ResNet-50 Energy-aware training ablation study for ResNet-50 across three image-classification datasets: CIFAR-10, CIFAR-100, and Tiny-ImageNet. Each dataset has 15 training variants (8 individual-method M0..M7, 7 cumulative ablation C0..C6) at 50 epochs, plus a 5-variant deployment pipeline (FP32 baseline, structured pruning, pruning+finetune, INT8 quantization, pruned+INT8). Status: 45 completed variants, 0 partial. Quick links… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/E2AM_ResNet50.imagen<1K2 likes1.2k downloads3mo agoHugging Face09reshabhs /SPML_Chatbot_Prompt_Injection SPML Chatbot Prompt Injection Dataset Arxiv Paper Introducing the SPML Chatbot Prompt Injection Dataset: a robust collection of system prompts designed to create realistic chatbot interactions, coupled with a diverse array of annotated user prompts that attempt to carry out prompt injection attacks. While other datasets in this domain have centered on less practical chatbot scenarios or have limited themselves to "jailbreaking" – just one aspect of prompt injection – our dataset… See the full description on the dataset page: https://huggingface.co/datasets/reshabhs/SPML_Chatbot_Prompt_Injection.tabulartext-classification10K<n<100K31 likes1.2k downloads2y agoHugging Face10ManikaSaini /zomato-restaurant-recommendationtext10K<n<100K4 likes1.1k downloads9mo agoHugging Face11electricsheepafrica /Environment-and-Natural-Resources-Indicators-For-African-Countries Environment and Natural Resources Indicators For African Countries | Africa (World Health Organization) Size category: 1K<n<10K - Formats: csv - Sector: climate_environment - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Environment-and-Natural-Resources-Indicators-For-African-Countries.tabulartabular-classification1K<n<10K0 likes1.1k downloads2mo agoHugging Face12manycore-research /SpatialLM-Dataset SpatialLM Dataset The SpatialLM dataset is a large-scale, high-quality synthetic dataset designed by professional 3D designers and used for real-world production. It contains point clouds from 12,328 diverse indoor scenes comprising 54,778 rooms, each paired with rich ground-truth 3D annotations. SpatialLM dataset provides an additional valuable resource for advancing research in indoor scene understanding, 3D perception, and… See the full description on the dataset page: https://huggingface.co/datasets/manycore-research/SpatialLM-Dataset.3d100K<n<1M15 likes925 downloads1y agoHugging Face13ibm-research /Wikipedia_contradict_benchmark Wikipedia contradict benchmark Wikipedia contradict benchmark is a dataset consisting of 253 high-quality, human-annotated instances designed to assess LLM performance when augmented with retrieved passages containing real-world knowledge conflicts. The dataset was created intentionally with that task in mind, focusing on a benchmark consisting of high-quality, human-annotated instances. Dataset Details Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/Wikipedia_contradict_benchmark.textquestion-answeringn<1K28 likes796 downloads2y agoHugging Face14ibm-research /claim_stance Dataset Card for Claim Stance Dataset Dataset Summary Claim Stance This dataset contains 2,394 labeled Wikipedia claims for 55 topics. The dataset includes the stance (Pro/Con) of each claim towards the topic, as well as fine-grained annotations, based on the semantic model of Stance Classification of Context-Dependent Claims (topic target, topic sentiment towards its target, claim target, claim sentiment towards its target, and the relation between the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/claim_stance.tabulartext-classification1K<n<10K7 likes737 downloads3y agoHugging Face15NoeFlandre /geoparser-benchmark-results Geoparser benchmark results Geoparsing pipelines from geoparser scored on English and multilingual corpora, run on Grid'5000 (one Tesla T4). Benchmarks Benchmark Languages Docs Toponyms Source GeoVirus en 229 2167 WikiNews articles on epidemics (Gritta et al., 2018). HIPE-2020 de, en, fr 129 1516 Historical Swiss, Luxembourgish and American newspapers, OCR. NewsEye de, fi, fr, sv 77 1772 Historical European newspapers, OCR (HIPE-2022). TopRes19th… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/geoparser-benchmark-results.tabularn<1K0 likes661 downloads3d agoHugging Face16aigrant /taiwan-ly-law-research Taiwan Legislator Yuan Law Research Data Overview The law research documents are issued irregularly from Taiwan Legislator Yuan. The purpose of those research are providing better understanding on social issues in aspect of laws. One may find documents rich with technical terms which could provided as training data. For comprehensive document list check out this link provided by Taiwan Legislator Yuan. There are currently missing document download links in 10th and 9th… See the full description on the dataset page: https://huggingface.co/datasets/aigrant/taiwan-ly-law-research.text1K<n<10K8 likes639 downloads11mo agoHugging Face17ShayManor /ising-sim2real-results Ising sim2real — Decoder Benchmark Results Evaluation results for a panel of open surface-code decoders run on real Google Willow hardware data and on synthetic circuit-level noise of rising fidelity. The question these results answer: does the cheap synthetic benchmark predict the real-hardware result? This repo holds the outputs (LERs, per-shot outcomes, fitted noise models, figures). The inputs — ingested Willow detection events, circuits, and shipped DEMs — live in… See the full description on the dataset page: https://huggingface.co/datasets/ShayManor/ising-sim2real-results.tabularother10K<n<100K1 likes617 downloads15d agoHugging Face18AvoCahDoe /llava-15-rlmpq-vlm-eval-results RL-MPQ VLM Evaluation Artifacts Complete figures, tables, galleries, and raw benchmark CSVs for the extended VLM evaluation. Dataset: AvoCahDoe/llava-15-rlmpq-vlm-eval-results Collections (by base VLM) RL-MPQ VLM — LLaVA-1.5-13B — HF collection RL-MPQ VLM — LLaVA-1.5-7B — HF collection RL-MPQ VLM — LLaVA-Next Mistral-7B — HF collection RL-MPQ VLM — Qwen2-VL-7B — HF collection Model repos RL-MPQ High Fidelity →… See the full description on the dataset page: https://huggingface.co/datasets/AvoCahDoe/llava-15-rlmpq-vlm-eval-results.imagevisual-question-answeringn<1K0 likes474 downloads3mo agoHugging Face19stair-lab /fantastic_bugs_resulttabular100K<n<1M0 likes466 downloads1y agoHugging Face20macpaw-research /mac-app-store-apps-metadata Dataset Card for Macappstore Applications Metadata 📌 Dataset status: static snapshot (no scheduled updates). The data was collected from the public iTunes Search API between December 2023 and January 2024 and reflects the Mac App Store as of that period. The dataset is stable and remains available for research use; it is not refreshed on a schedule. Mac App Store Applications Metadata sourced by the public API. Curated by: MacPaw Way Ltd. Language(s) (NLP): Mostly EN, DE… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/mac-app-store-apps-metadata.imagetabular-classification10K<n<100K10 likes400 downloads1mo agoHugging Face21BrainAlign /cdl-devai-results-ds006239 ds006239 (Wang et al. 2025) — brain × interpretability × localisation, per model per checkpoint Wang et al. 2025 — word-level phonological and semantic reading in children and adolescents. Cohort: children and adolescents 10–17 years; presentation: visual (reading). Tasks Orth, Phon, Sem, SemLocal × sessions ses-11, ses-11+ = 8 task × session cells, all of them scored here. Duplicate cells. Phon/ses-11+ is bit-identical to Orth/ses-11+, Phon/ses-11 is bit-identical to… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds006239.tabular10K<n<100K0 likes396 downloads24d agoHugging Face22ibm-research /AttaQ AttaQ Dataset Card The AttaQ red teaming dataset, consisting of 1402 carefully crafted adversarial questions, is designed to evaluate Large Language Models (LLMs) by assessing their tendency to generate harmful or undesirable responses. It may serve as a benchmark to assess the potential harm of responses produced by LLMs. The dataset is categorized into seven distinct classes of questions: deception, discrimination, harmful information, substance abuse, sexual content, personally… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/AttaQ.texttext-generation1K<n<10K24 likes391 downloads3y agoHugging Face23BrainAlign /cdl-devai-results-ds002236 ds002236 (Lytle et al. 2020) — brain × interpretability × localisation, per model per checkpoint Lytle et al. 2020 — orthographic, phonological and semantic word processing in school-aged children. Cohort: children 8.7–15.5 years; presentation: auditory and visual word presentation. Tasks Phon, Sem × sessions ses-9, ses-11, ses-11+ = 6 task × session cells, all of them scored here. Incomplete families --- do not read these as scale-ladder points. pythia-2.8b-full has 4… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds002236.tabular10K<n<100K0 likes375 downloads24d agoHugging Face24BrainAlign /cdl-devai-results-ds003604-roiauditory ds003604 (Wang et al. 2022) — brain × interpretability × localisation, per model per checkpoint Wang et al. 2022 — auditory language comprehension in children. Cohort: children scanned at 5, 7 and 9; presentation: auditory. Tasks Sem, Phon, Gram, Plaus × sessions ses-5, ses-7, ses-9 = 12 task × session cells, all of them scored here. Incomplete families --- do not read these as scale-ladder points. pythia-2.8b-full has 4 checkpoints, pythia-6.9b-full has 1 checkpoint.… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds003604-roiauditory.tabular10K<n<100K0 likes366 downloads20d agoHugging Face25APProjects /saas-vendor-outage-duration-incident-resolution-time-mttr How long do SaaS vendor outages last? Incident resolution time per vendor, rebuilt daily As of 2026-09-24 12:28 UTC. For every incident a vendor posted on its own public status page with BOTH an opened time and a resolved time, this dataset computes duration_minutes = resolved_at - started_at and rolls it up per vendor. It is derived, every day, from the incident table in saas-vendor-status-pages-outages-incidents-daily; the two are rebuilt by the same job and cannot disagree.… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/saas-vendor-outage-duration-incident-resolution-time-mttr.tabulartabular-regression10K<n<100K0 likes347 downloads2d agoHugging Face26APProjects /us-restaurant-hotel-closings-layoffs-warn-act-notices-daily US restaurant and hotel closings and layoffs — the actual WARN Act filings, rebuilt every day Last rebuilt: 2026-09-24. 6,348 layoff and closure notices filed by restaurants and restaurant groups, hotels, motels and resorts, casinos and gaming floors, caterers and contract food-service operators, bars and coffee chains, stadium, arena and airport concessions, and leisure and entertainment venues with US state labor departments — 920,524 workers, 3,060 employers, 48 states… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-restaurant-hotel-closings-layoffs-warn-act-notices-daily.texttabular-classification1K<n<10K0 likes327 downloads2d agoHugging Face27BrainAlign /cdl-devai-results-ds003604-roiphonology ds003604 (Wang et al. 2022) — brain × interpretability × localisation, per model per checkpoint Wang et al. 2022 — auditory language comprehension in children. Cohort: children scanned at 5, 7 and 9; presentation: auditory. Tasks Sem, Phon, Gram, Plaus × sessions ses-5, ses-7, ses-9 = 12 task × session cells, all of them scored here. Incomplete families --- do not read these as scale-ladder points. pythia-2.8b-full has 4 checkpoints. pythia-6.9b-full's single checkpoint is step… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds003604-roiphonology.tabular10K<n<100K0 likes311 downloads21d agoHugging Face28BrainAlign /cdl-devai-results-ds003604-roimotor ds003604 (Wang et al. 2022) — brain × interpretability × localisation, per model per checkpoint Wang et al. 2022 — auditory language comprehension in children. Cohort: children scanned at 5, 7 and 9; presentation: auditory. Tasks Sem, Phon, Gram, Plaus × sessions ses-5, ses-7, ses-9 = 12 task × session cells, all of them scored here. Incomplete families --- do not read these as scale-ladder points. pythia-2.8b-full has 4 checkpoints. pythia-6.9b-full's single checkpoint is step… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds003604-roimotor.tabular10K<n<100K0 likes276 downloads21d agoHugging Face29BrainAlign /cdl-devai-results-ds003604 ds003604 (Wang et al. 2022) — brain × interpretability × localisation, per model per checkpoint Wang et al. 2022 — auditory language comprehension in children. Cohort: children scanned at 5, 7 and 9; presentation: auditory. Tasks Sem, Phon, Gram, Plaus × sessions ses-5, ses-7, ses-9 = 12 task × session cells, all of them scored here. Incomplete families --- do not read these as scale-ladder points. pythia-2.8b-full has 4 checkpoints, pythia-6.9b-full has 1 checkpoint.… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds003604.tabular10K<n<100K0 likes272 downloads24d agoHugging Face30ahmedheakl /resume-atlasPlease see paper & code for more information: https://github.com/noran-mohamed/Resume-Classification-Dataset https://arxiv.org/abs/2406.18125 text10K<n<100K11 likes263 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.