CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bencorn /CIC-IoT-2023tabular10M<n<100M0 likes5.5k downloads6mo agoHugging Face02CiferAI /Cifer-Fraud-Detection-Dataset-AF 📊 Cifer Fraud Detection Dataset 🧠 Overview The Cifer-Fraud-Detection-Dataset-AF is a high-fidelity, fully synthetic dataset created to support the development and benchmarking of privacy-preserving, federated, and decentralized machine learning systems in financial fraud detection. This dataset draws structural inspiration from the PaySim simulator, which was built using aggregated mobile money transaction data from a real financial provider operating in 14+ countries.… See the full description on the dataset page: https://huggingface.co/datasets/CiferAI/Cifer-Fraud-Detection-Dataset-AF.tabulartabular-classification10M<n<100M13 likes1.9k downloads1y agoHugging Face03c01dsnap /CIC-IDS2017The CICIDS2017 dataset consists of labeled network flows, including full packet payloads in pcap format, the corresponding profiles and the labeled flows (GeneratedLabelledFlows.zip) and CSV files for machine and deep learning purpose (MachineLearningCSV.zip) are publicly available for researchers. If you are using our dataset, you should cite our related paper which outlining the details of the dataset and its underlying principles: Iman Sharafaldin, Arash Habibi Lashkari, and Ali A.… See the full description on the dataset page: https://huggingface.co/datasets/c01dsnap/CIC-IDS2017.tabular1M<n<10M5 likes1.2k downloads3y agoHugging Face04evaluate /imdb-citextn<1K0 likes969 downloads4y agoHugging Face05CCB /cis5300-text-classification Complex Word Identification (CIS 5300) Dataset Description This dataset supports the Complex Word Identification (CWI) task: given a word in context, predict whether it is complex (likely to be difficult for non-native speakers, children, or people with reading disabilities) or simple. CWI is the first step in lexical simplification — the task of rewriting text to make it more accessible. Before you can simplify a word, you need to identify which words need… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-text-classification.tabulartext-classification1K<n<10K0 likes816 downloads5mo agoHugging Face06WorkWithData /citiesThis datasets contains 47,605 cities from around the world. The latest version can be found and filtered differently on: https://www.workwithdata.com/datasets/cities Similar datasets can be found on: https://www.workwithdata.com tabular10K<n<100K2 likes521 downloads2y agoHugging Face07Cianmcnally /irish-census Irish Census 1901 & 1926 Person-level records from the 1901 and 1926 censuses of Ireland, as published by the National Archives of Ireland — every individual return, in flat CSV. Year Rows Size Coverage 1901 4,434,939 4.31 GB All of Ireland (32 counties) 1926 2,973,480 0.56 GB Saorstát Éireann (26 counties) Total 7,408,419 4.87 GB The 1926 census is the first taken by the Irish Free State and was released to the public in 2026 under the 100-year rule. The… See the full description on the dataset page: https://huggingface.co/datasets/Cianmcnally/irish-census.tabulartabular-classification1M<n<10M1 likes332 downloads2mo agoHugging Face08viennh2012 /cardiac_cine_acdc ACDC (Cardiac Cine-MRI) ACDC (Automatic Cardiac Diagnosis Challenge, MICCAI 2017) is a cine‑MRI dataset for cardiac segmentation.This repository contains processed NIfTI files in Data/processed_output/acdc format. Dataset Summary Modality: Cardiac cine‑MRI (NIfTI) Task: Segmentation of LV, RV, and myocardium Frames: ED/ES + full SAX time series (sax_t) Labels: LV/RV cavities + myocardium Splits: train, test (as provided in processed output) Data Structure (per… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_acdc.tabularimage-segmentationn<1K0 likes310 downloads7mo agoHugging Face09egolimblevskaia /circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations CircuitLens & WeightLens: Transcoder Descriptions and Evaluations This dataset contains automatically generated descriptions and evaluation metrics for Gemma-2-2B transcoders, produced using CircuitLens and WeightLens methods. Methods CircuitLens: https://github.com/egolimblevskaia/CircuitLens WeightLens: https://github.com/egolimblevskaia/WeightLens Dataset Structure The dataset is organized by layers (0, 4, 7, 10, 12, 15, 18, 21, 23, 25), with each layer… See the full description on the dataset page: https://huggingface.co/datasets/egolimblevskaia/circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations.tabulartext-classification10K<n<100K0 likes295 downloads7mo agoHugging Face10bvk /CICIDS-2017Raw network data was collected over a period of 5 days, Monday through Friday, and stored in PCAP files. Monday was used to create most of the Benign data, while the Attack-Network implemented various types of attacks over the next 4 days, such as Brute Force connections (FTP and SSH), several types of DoS attacks, as well as a Botnet attack, Infiltration attacks and subsequent Port-Scanning activity. The PCAP data was processed using a tool developed by one of the authors of [1], called… See the full description on the dataset page: https://huggingface.co/datasets/bvk/CICIDS-2017.tabular1M<n<10M0 likes273 downloads2y agoHugging Face11CCB /cis5300-word-embeddings Word Embeddings and Semantic Similarity (CIS 5300) Dataset Description This dataset supports learning about word embeddings — dense vector representations that capture word meaning. It includes a standard similarity benchmark, a word sense disambiguation task, and a Shakespeare corpus for training custom embeddings. Configs SimLex-999: Word Similarity Benchmark SimLex-999 (Hill et al., 2015) is a gold-standard benchmark for evaluating word… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-word-embeddings.tabularsentence-similarity1K<n<10K0 likes247 downloads5mo agoHugging Face12cis-lmu /udhr-lid UDHR-LID Why UDHR-LID? You can access UDHR (Universal Declaration of Human Rights) here, but when a verse is missing, they have texts such as "missing" or "?". Also, about 1/3 of the sentences consist only of "articles 1-30" in different languages. We cleaned the entire dataset from XML files and selected only the paragraphs. We cleared any unrelated language texts from the data and also removed the cases that were incorrect. Incorrect? Look at the ckb and kmr files in the UDHR.… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/udhr-lid.text10K<n<100K8 likes244 downloads2y agoHugging Face13viennh2012 /cardiac_cine_mnms M&Ms (Cardiac Cine-MRI) Processed NIfTI cine-MRI data derived from the M&Ms (Multi-Centre, Multi-Vendor & Multi-Disease) challenge. Dataset Summary Modality: CMR cine MRI Task: LV/RV/MYO segmentation Views: SAX (ED/ES) Splits: train / val / test Data Structure (per example) sax_ed, sax_ed_gt sax_es, sax_es_gt Optional: sax_t (if present) Metadata columns listed below Columns Imaging pid sax_ed, sax_ed_gt, sax_es, sax_es_gt sax_t (if present)… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_mnms.tabularimage-segmentationn<1K0 likes243 downloads7mo agoHugging Face14wasarou /plinder-cistrans-209 PLINDER cis/trans conformer benchmark (209 systems) A curated 209-system subset of PLINDER for measuring whether a structure predictor preserves the E/Z (cis/trans) configuration of a ligand double bond. Each system ships everything needed to fold it and score the result: the protein MSA, the ligand SMILES, and the crystal ligand as an SDF. It is the evaluation set for the cistrans term of RGI (Restraint-Guided Inference), but nothing here is RGI-specific — any method that… See the full description on the dataset page: https://huggingface.co/datasets/wasarou/plinder-cistrans-209.tabularn<1K0 likes209 downloads2mo agoHugging Face15c1587s /cil-regionalizationgeospatialn<1K0 likes206 downloads1y agoHugging Face16okra123 /CISR24 CISR24 CISR24 is the controlled complex-baseband benchmark introduced in the paper "CDTFNet: A Cross-Domain Token-Fusion Transformer for Waveform-Level Electromagnetic Awareness in Low-Altitude Intelligent Networks." It defines one closed-set task over 24 waveform classes: 10 communication, 6 sensing, and 8 integrated sensing and communication (ISAC) classes. Paper protocol Property Setting Sampling rate 10 MHz Record 1024 complex samples (102.4 us)… See the full description on the dataset page: https://huggingface.co/datasets/okra123/CISR24.document1M<n<10M0 likes198 downloads2mo agoHugging Face17Taylor658 /photonic-integrated-circuit-yield 🏭 Photonic Integrated Circuit Yield Dataset 📊 125,000 synthetic (yield query, yield reasoning response) pairs covering process variation, defect density, lithography, and metrology challenges in CMOS-compatible photonic integrated circuit (PIC) manufacturing. ⚠️ Disclaimer: All entries are synthetically generated. Yield figures are computed from textbook models over sampled inputs, and citations are placeholders styled after technical sources; none reference a real… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/photonic-integrated-circuit-yield.texttext-generation100K<n<1M4 likes189 downloads10d agoHugging Face18cis-lmu /GlotStoryBook Dataset Description Story Books for 180 ISO-639-3 codes. The Parallel ID or parallel_id can be used to find the parallel documents in different languages and build a parallel dataset. This dataset consists of 2 subsets: default, which consists of 4 publishers: asp: African Storybook pb: Pratham Books lcb: Little Cree Books lida: LIDA Stories nalibali, which comes from Nal'ibali stories. Usage (HF Loader) default: from datasets import load_dataset dataset… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/GlotStoryBook.texttranslation10K<n<100K9 likes172 downloads11h agoHugging Face19PersonaBias /Reverse-circuit-discoverytabulartext-classification10K<n<100K0 likes170 downloads2mo agoHugging Face20k-master /k-beauty-ai-citation-dataset K-Beauty AI Citation Dataset Open dataset mapping Korean K-beauty entities (ingredients, skin concerns, use cases, brands) and answer-style guides to citation-shaped external references. Designed to be referenced by AI search engines, content builders, and SEO research. Canonical source: https://kbeautyanswers.com/dataset/ License: CC BY 4.0 Maintainer: K-Beauty Answers (site) Initial release: 2026-05-23 What's in it 128 entities (37 ingredients + 18 skin… See the full description on the dataset page: https://huggingface.co/datasets/k-master/k-beauty-ai-citation-dataset.texttext-classificationn<1K0 likes165 downloads3mo agoHugging Face21facebook /CIMemories CIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs Paper Large Language Models (LLMs) increasingly use persistent memory from past interactions to enhance personalization and task performance. However, this memory introduces critical risks when sensitive information is revealed in inappropriate contexts. We present CIMemories, a benchmark for evaluating whether LLMs appropriately control information flow from memory based on task context.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/CIMemories.text10K<n<100K2 likes161 downloads10mo agoHugging Face22nithi060488 /Cifer-Fraud-Detection-Dataset-AF 📊 Cifer Fraud Detection Dataset 🧠 Overview The Cifer-Fraud-Detection-Dataset-AF is a high-fidelity, fully synthetic dataset created to support the development and benchmarking of privacy-preserving, federated, and decentralized machine learning systems in financial fraud detection. This dataset draws structural inspiration from the PaySim simulator, which was built using aggregated mobile money transaction data from a real financial provider operating in 14+ countries.… See the full description on the dataset page: https://huggingface.co/datasets/nithi060488/Cifer-Fraud-Detection-Dataset-AF.tabulartabular-classification10M<n<100M0 likes158 downloads7mo agoHugging Face23nasa-cisto-data-science-group /modis-lake-powell-toy-dataset MODIS Water Lake Powell Toy Dataset Dataset Summary Tabular dataset comprised of MODIS surface reflectance bands along with calculated indices and a label (water/not-water) Dataset Structure Data Fields water: Label, water or not-water (binary) sur_refl_b01_1: MODIS surface reflection band 1 (-100, 16000) sur_refl_b02_1: MODIS surface reflection band 2 (-100, 16000) sur_refl_b03_1: MODIS surface reflection band 3 (-100, 16000) sur_refl_b04_1: MODIS… See the full description on the dataset page: https://huggingface.co/datasets/nasa-cisto-data-science-group/modis-lake-powell-toy-dataset.image1K<n<10K1 likes154 downloads3y agoHugging Face24Exploration-Lab /CISLRgated CISLR: Corpus for Indian Sign Language Recognition This repository contains the Indian Sign Language Dataset proposed in the following paper Paper: CISLR: Corpus for Indian Sign Language Recognition https://preview.aclanthology.org/emnlp-22-ingestion/2022.emnlp-main.707/ Authors: Abhinav Joshi, Ashwani Bhat, Pradeep S, Priya Gole, Shashwat Gupta, Shreyansh Agarwal, Ashutosh Modi Abstract: Indian Sign Language, though used by a diverse community, still lacks well-annotated… See the full description on the dataset page: https://huggingface.co/datasets/Exploration-Lab/CISLR.text1K<n<10K10 likes145 downloads2y agoHugging Face25bds062 /cichlid-behavior-pairs Cichlid Behavior Pairs Video/annotation pairs of Astatotilapia burtoni (cichlid fish) reproductive and courtship behavior, recorded during a PGF2a-induced spawning assay comparing nose-occluded ("VetBond") vs. sham-treated ("Sham") females — a manipulation of olfactory input to the mating interaction. Each pair consists of one top-down video of a male/female tank trial and one point-annotated behavior event table (exported from BORIS) marking the timing of specific… See the full description on the dataset page: https://huggingface.co/datasets/bds062/cichlid-behavior-pairs.tabular1K<n<10K0 likes136 downloads19d agoHugging Face26dazzle-nu /CIS435-CreditCardFraudDetectiontabular1M<n<10M13 likes135 downloads4y agoHugging Face27viennh2012 /cardiac_cine_mnms2 M&Ms2 (Cardiac Cine-MRI, RV Focus) Processed NIfTI cine-MRI data derived from the M&Ms2 challenge. Dataset Summary Modality: CMR cine MRI Task: LV/RV/MYO segmentation (RV focus) Views: SAX + LAX 4C (LAX 2C if present) Splits: train / val / test Data Structure (per example) SAX: sax_ed, sax_ed_gt, sax_es, sax_es_gt LAX 4C: lax_4c_ed, lax_4c_ed_gt, lax_4c_es, lax_4c_es_gt LAX 2C (if present): lax_2c_ed, lax_2c_ed_gt, lax_2c_es, lax_2c_es_gt Metadata columns… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_mnms2.tabularimage-segmentationn<1K0 likes132 downloads7mo agoHugging Face28baalajimaestro /DDoS-CICIoT2023tabular10M<n<100M0 likes122 downloads2y agoHugging Face29brickandmortar /twin-cities-public-records Twin Cities public records, joined 25 datasets · 1,575,384 rows · free, CC BY 4.0 · mirrored from brickandmortar.dev A city emits records constantly — parcels, recorded sales, assessments, permits, licences, inspections, 911 calls, cleanup sites, flood zones, federal loans, wages, census measures — and almost nobody joins them. These are the joined slices, published as files rather than as an API you have to ask for a key to. The join is the work; the data is free. This is a… See the full description on the dataset page: https://huggingface.co/datasets/brickandmortar/twin-cities-public-records.tabular100K<n<1M0 likes114 downloads18d agoHugging Face30bvk /CIC-MalMem-2022CIC-MalMem-2022 was created by researchers at the Canadian Institute for Cybersecurity (CIC) at the University of New Brunswick. The details are described on the website https://www.unb.ca/cic/datasets/malmem-2022.html, and in their paper mentioned on that site: Tristan Carrier, Princy Victor, Ali Tekeoglu, Arash Habibi Lashkari,” Detecting Obfuscated Malware using Memory Feature Engineering”, The 8th International Conference on Information Systems Security and Privacy (ICISSP), 2022. The… See the full description on the dataset page: https://huggingface.co/datasets/bvk/CIC-MalMem-2022.tabular10K<n<100K0 likes111 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.