datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CIC-IoT-2023Cifer-Fraud-Detection-Dataset-AF
📊 Cifer Fraud Detection Dataset
🧠 Overview
The Cifer-Fraud-Detection-Dataset-AF is a high-fidelity, fully synthetic dataset created to support the development and benchmarking of privacy-preserving, federated, and decentralized machine learning systems in financial fraud detection.
This dataset draws structural inspiration from the PaySim simulator, which was built using aggregated mobile money transaction data from a real financial provider operating in 14+ countries.… See the full description on the dataset page: https://huggingface.co/datasets/CiferAI/Cifer-Fraud-Detection-Dataset-AF.CIC-IDS2017The CICIDS2017 dataset consists of labeled network flows, including full packet payloads in pcap format, the corresponding profiles and the labeled flows (GeneratedLabelledFlows.zip) and CSV files for machine and deep learning purpose (MachineLearningCSV.zip) are publicly available for researchers. If you are using our dataset, you should cite our related paper which outlining the details of the dataset and its underlying principles:
Iman Sharafaldin, Arash Habibi Lashkari, and Ali A.… See the full description on the dataset page: https://huggingface.co/datasets/c01dsnap/CIC-IDS2017.imdb-cicis5300-text-classification
Complex Word Identification (CIS 5300)
Dataset Description
This dataset supports the Complex Word Identification (CWI) task: given a word in context, predict whether it is complex (likely to be difficult for non-native speakers, children, or people with reading disabilities) or simple.
CWI is the first step in lexical simplification — the task of rewriting text to make it more accessible. Before you can simplify a word, you need to identify which words need… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-text-classification.citiesThis datasets contains 47,605 cities from around the world. The latest version can be found and filtered differently on: https://www.workwithdata.com/datasets/cities
Similar datasets can be found on: https://www.workwithdata.com
irish-census
Irish Census 1901 & 1926
Person-level records from the 1901 and 1926 censuses of Ireland, as published by
the National Archives of Ireland — every individual return, in flat CSV.
Year
Rows
Size
Coverage
1901
4,434,939
4.31 GB
All of Ireland (32 counties)
1926
2,973,480
0.56 GB
Saorstát Éireann (26 counties)
Total
7,408,419
4.87 GB
The 1926 census is the first taken by the Irish Free State and was released to
the public in 2026 under the 100-year rule. The… See the full description on the dataset page: https://huggingface.co/datasets/Cianmcnally/irish-census.cardiac_cine_acdc
ACDC (Cardiac Cine-MRI)
ACDC (Automatic Cardiac Diagnosis Challenge, MICCAI 2017) is a cine‑MRI dataset for cardiac segmentation.This repository contains processed NIfTI files in Data/processed_output/acdc format.
Dataset Summary
Modality: Cardiac cine‑MRI (NIfTI)
Task: Segmentation of LV, RV, and myocardium
Frames: ED/ES + full SAX time series (sax_t)
Labels: LV/RV cavities + myocardium
Splits: train, test (as provided in processed output)
Data Structure (per… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_acdc.circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations
CircuitLens & WeightLens: Transcoder Descriptions and Evaluations
This dataset contains automatically generated descriptions and evaluation metrics for Gemma-2-2B transcoders, produced using CircuitLens and WeightLens methods.
Methods
CircuitLens: https://github.com/egolimblevskaia/CircuitLens
WeightLens: https://github.com/egolimblevskaia/WeightLens
Dataset Structure
The dataset is organized by layers (0, 4, 7, 10, 12, 15, 18, 21, 23, 25), with each layer… See the full description on the dataset page: https://huggingface.co/datasets/egolimblevskaia/circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations.CICIDS-2017Raw network data was collected over a period of 5 days, Monday through Friday, and stored in PCAP files.
Monday was used to create most of the Benign data, while the Attack-Network implemented various types of attacks over the next 4 days,
such as Brute Force connections (FTP and SSH), several types of DoS attacks, as well as a Botnet attack, Infiltration attacks and subsequent Port-Scanning activity.
The PCAP data was processed using a tool developed by one of the authors of [1], called… See the full description on the dataset page: https://huggingface.co/datasets/bvk/CICIDS-2017.cis5300-word-embeddings
Word Embeddings and Semantic Similarity (CIS 5300)
Dataset Description
This dataset supports learning about word embeddings — dense vector representations that capture word meaning. It includes a standard similarity benchmark, a word sense disambiguation task, and a Shakespeare corpus for training custom embeddings.
Configs
SimLex-999: Word Similarity Benchmark
SimLex-999 (Hill et al., 2015) is a gold-standard benchmark for evaluating word… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-word-embeddings.udhr-lid
UDHR-LID
Why UDHR-LID?
You can access UDHR (Universal Declaration of Human Rights) here, but when a verse is missing, they have texts such as "missing" or "?". Also, about 1/3 of the sentences consist only of "articles 1-30" in different languages. We cleaned the entire dataset from XML files and selected only the paragraphs. We cleared any unrelated language texts from the data and also removed the cases that were incorrect.
Incorrect? Look at the ckb and kmr files in the UDHR.… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/udhr-lid.cardiac_cine_mnms
M&Ms (Cardiac Cine-MRI)
Processed NIfTI cine-MRI data derived from the M&Ms (Multi-Centre, Multi-Vendor & Multi-Disease) challenge.
Dataset Summary
Modality: CMR cine MRI
Task: LV/RV/MYO segmentation
Views: SAX (ED/ES)
Splits: train / val / test
Data Structure (per example)
sax_ed, sax_ed_gt
sax_es, sax_es_gt
Optional: sax_t (if present)
Metadata columns listed below
Columns
Imaging
pid
sax_ed, sax_ed_gt, sax_es, sax_es_gt
sax_t (if present)… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_mnms.plinder-cistrans-209
PLINDER cis/trans conformer benchmark (209 systems)
A curated 209-system subset of PLINDER for measuring
whether a structure predictor preserves the E/Z (cis/trans) configuration of a ligand double
bond. Each system ships everything needed to fold it and score the result: the protein MSA, the
ligand SMILES, and the crystal ligand as an SDF.
It is the evaluation set for the cistrans term of RGI (Restraint-Guided Inference), but
nothing here is RGI-specific — any method that… See the full description on the dataset page: https://huggingface.co/datasets/wasarou/plinder-cistrans-209.cil-regionalizationCISR24
CISR24
CISR24 is the controlled complex-baseband benchmark introduced in the paper
"CDTFNet: A Cross-Domain Token-Fusion Transformer for Waveform-Level
Electromagnetic Awareness in Low-Altitude Intelligent Networks." It defines one closed-set task over 24 waveform classes:
10 communication, 6 sensing, and 8 integrated sensing and communication (ISAC)
classes.
Paper protocol
Property
Setting
Sampling rate
10 MHz
Record
1024 complex samples (102.4 us)… See the full description on the dataset page: https://huggingface.co/datasets/okra123/CISR24.photonic-integrated-circuit-yield
🏭 Photonic Integrated Circuit Yield Dataset
📊 125,000 synthetic (yield query, yield reasoning response) pairs covering process variation, defect density, lithography, and metrology challenges in CMOS-compatible photonic integrated circuit (PIC) manufacturing.
⚠️ Disclaimer: All entries are synthetically generated. Yield figures are computed from textbook models over sampled inputs, and citations are placeholders styled after technical sources; none reference a real… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/photonic-integrated-circuit-yield.GlotStoryBook
Dataset Description
Story Books for 180 ISO-639-3 codes.
The Parallel ID or parallel_id can be used to find the parallel documents in different languages and build a parallel dataset.
This dataset consists of 2 subsets:
default, which consists of 4 publishers:
asp: African Storybook
pb: Pratham Books
lcb: Little Cree Books
lida: LIDA Stories
nalibali, which comes from Nal'ibali stories.
Usage (HF Loader)
default:
from datasets import load_dataset
dataset… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/GlotStoryBook.Reverse-circuit-discoveryk-beauty-ai-citation-dataset
K-Beauty AI Citation Dataset
Open dataset mapping Korean K-beauty entities (ingredients, skin concerns, use cases, brands) and answer-style guides to citation-shaped external references. Designed to be referenced by AI search engines, content builders, and SEO research.
Canonical source: https://kbeautyanswers.com/dataset/
License: CC BY 4.0
Maintainer: K-Beauty Answers (site)
Initial release: 2026-05-23
What's in it
128 entities (37 ingredients + 18 skin… See the full description on the dataset page: https://huggingface.co/datasets/k-master/k-beauty-ai-citation-dataset.CIMemories
CIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs
Paper
Large Language Models (LLMs) increasingly use persistent memory from past interactions to enhance personalization and task performance. However, this memory introduces critical risks when sensitive information is revealed in inappropriate contexts. We present CIMemories, a benchmark for evaluating whether LLMs appropriately control information flow from memory based on task context.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/CIMemories.Cifer-Fraud-Detection-Dataset-AF
📊 Cifer Fraud Detection Dataset
🧠 Overview
The Cifer-Fraud-Detection-Dataset-AF is a high-fidelity, fully synthetic dataset created to support the development and benchmarking of privacy-preserving, federated, and decentralized machine learning systems in financial fraud detection.
This dataset draws structural inspiration from the PaySim simulator, which was built using aggregated mobile money transaction data from a real financial provider operating in 14+ countries.… See the full description on the dataset page: https://huggingface.co/datasets/nithi060488/Cifer-Fraud-Detection-Dataset-AF.modis-lake-powell-toy-dataset
MODIS Water Lake Powell Toy Dataset
Dataset Summary
Tabular dataset comprised of MODIS surface reflectance bands along with calculated indices and a label (water/not-water)
Dataset Structure
Data Fields
water: Label, water or not-water (binary)
sur_refl_b01_1: MODIS surface reflection band 1 (-100, 16000)
sur_refl_b02_1: MODIS surface reflection band 2 (-100, 16000)
sur_refl_b03_1: MODIS surface reflection band 3 (-100, 16000)
sur_refl_b04_1: MODIS… See the full description on the dataset page: https://huggingface.co/datasets/nasa-cisto-data-science-group/modis-lake-powell-toy-dataset.CISLR
CISLR: Corpus for Indian Sign Language Recognition
This repository contains the Indian Sign Language Dataset proposed in the following paper
Paper: CISLR: Corpus for Indian Sign Language Recognition https://preview.aclanthology.org/emnlp-22-ingestion/2022.emnlp-main.707/
Authors: Abhinav Joshi, Ashwani Bhat, Pradeep S, Priya Gole, Shashwat Gupta, Shreyansh Agarwal, Ashutosh Modi
Abstract: Indian Sign Language, though used by a diverse community, still lacks well-annotated… See the full description on the dataset page: https://huggingface.co/datasets/Exploration-Lab/CISLR.cichlid-behavior-pairs
Cichlid Behavior Pairs
Video/annotation pairs of Astatotilapia burtoni (cichlid fish) reproductive and
courtship behavior, recorded during a PGF2a-induced spawning assay comparing
nose-occluded ("VetBond") vs. sham-treated ("Sham") females — a manipulation
of olfactory input to the mating interaction. Each pair consists of one
top-down video of a male/female tank trial and one point-annotated behavior
event table (exported from BORIS) marking the
timing of specific… See the full description on the dataset page: https://huggingface.co/datasets/bds062/cichlid-behavior-pairs.CIS435-CreditCardFraudDetectioncardiac_cine_mnms2
M&Ms2 (Cardiac Cine-MRI, RV Focus)
Processed NIfTI cine-MRI data derived from the M&Ms2 challenge.
Dataset Summary
Modality: CMR cine MRI
Task: LV/RV/MYO segmentation (RV focus)
Views: SAX + LAX 4C (LAX 2C if present)
Splits: train / val / test
Data Structure (per example)
SAX: sax_ed, sax_ed_gt, sax_es, sax_es_gt
LAX 4C: lax_4c_ed, lax_4c_ed_gt, lax_4c_es, lax_4c_es_gt
LAX 2C (if present): lax_2c_ed, lax_2c_ed_gt, lax_2c_es, lax_2c_es_gt
Metadata columns… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_mnms2.DDoS-CICIoT2023twin-cities-public-records
Twin Cities public records, joined
25 datasets · 1,575,384 rows · free, CC BY 4.0 · mirrored from brickandmortar.dev
A city emits records constantly — parcels, recorded sales, assessments, permits, licences, inspections, 911 calls, cleanup sites, flood zones, federal loans, wages, census measures — and almost nobody joins them. These are the joined slices, published as files rather than as an API you have to ask for a key to. The join is the work; the data is free.
This is a… See the full description on the dataset page: https://huggingface.co/datasets/brickandmortar/twin-cities-public-records.CIC-MalMem-2022CIC-MalMem-2022 was created by researchers at the Canadian Institute for Cybersecurity (CIC) at the University of New Brunswick.
The details are described on the website https://www.unb.ca/cic/datasets/malmem-2022.html, and in their paper mentioned on that site:
Tristan Carrier, Princy Victor, Ali Tekeoglu, Arash Habibi Lashkari,” Detecting Obfuscated Malware using Memory Feature Engineering”,
The 8th International Conference on Information Systems Security and Privacy (ICISSP), 2022.
The… See the full description on the dataset page: https://huggingface.co/datasets/bvk/CIC-MalMem-2022.
