CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MiaSanLei /MindBenchOur homepage: https://miasanlei.github.io/MindBench.github.io/ Our GitHub repository: https://github.com/MiaSanLei/MindBench text100K<n<1M0 likes775 downloads2y agoHugging Face02miaolu3 /browsecomp-plustext100K<n<1M0 likes441 downloads11mo agoHugging Face03lmms-lab-encoder /MIA-Benchimagen<1K0 likes358 downloads2y agoHugging Face04LightningCreeper /MIA Memory Intelligence Agent (MIA) Paper | GitHub MIA (Memory In Intelligence Agent) is a memory framework designed for deep research agents (DRAs). It transforms agents from "passive record-keepers" into "active strategists" using a Manager-Planner-Executor architecture. This repository contains the datasets and data artifacts used to train and evaluate the MIA framework. Dataset Description The dataset includes the following components: Train: Data used for the… See the full description on the dataset page: https://huggingface.co/datasets/LightningCreeper/MIA.textimage-text-to-text10K<n<100K4 likes303 downloads6mo agoHugging Face05germane /Tab-MIA Tab-MIA: A Benchmark for Membership Inference Attacks on Tabular Data Tab-MIA is a benchmark dataset designed to evaluate the privacy risks of fine-tuning large language models (LLMs) on structured tabular data. It enables reproducible and systematic testing of Membership Inference Attacks (MIAs) across diverse datasets and six different serialization formats. 📋 Overview Datasets: WTQ (WikiTableQuestions) WikiSQL TabFact Adult Census California Housing… See the full description on the dataset page: https://huggingface.co/datasets/germane/Tab-MIA.texttext-classification100K<n<1M0 likes184 downloads1y agoHugging Face06ahmedBargady /MIAF_DomainDetection_Infrastructure_Datasets MIAF: Domain Detection Infrastructure Datasets This collection is the standardized evaluation benchmark for MIAF (Modular Infrastructure-Aware Fusion). It provides nine classification datasets derived from four public malicious-domain benchmarks, each paired with a shared 137-feature infrastructure representation. Overview We evaluate MIAF across nine classification datasets derived from four public malicious-domain benchmarks: DomainRadar (Hranický et al.… See the full description on the dataset page: https://huggingface.co/datasets/ahmedBargady/MIAF_DomainDetection_Infrastructure_Datasets.tabulartabular-classification1M<n<10M0 likes177 downloads11d agoHugging Face07miaojiemiao /Anime-2026image1M<n<10M12 likes166 downloads1y agoHugging Face08Miaosen /CCCleanerDataset8Gtexttext-classification1M<n<10M0 likes163 downloads4y agoHugging Face09h0ssn /agnews-unlearning-mia AGNEWS - Machine Unlearning + MIA Evaluation Dataset (Length-Filtered) This dataset is prepared for evaluating machine unlearning methods on fine-tuned LLMs using Membership Inference Attacks (MIAs). Dataset Splits Training Sets (for Unlearning) retain_set (9,000 samples): Data to retain during unlearning forget_set (1,000 samples): Data to unlearn Evaluation Sets (for MIA) - Length-Filtered AGNews Length Variants 32 tokens (~32±10… See the full description on the dataset page: https://huggingface.co/datasets/h0ssn/agnews-unlearning-mia.text10K<n<100K0 likes162 downloads9mo agoHugging Face10mia-musgen /shadow_fma_medium_1K_240427tabularn<1K0 likes144 downloads2y agoHugging Face11BrunoHays /Bangor-Miami-Spanish-English-Corpus Bangor Miami Spanish-English Corpus The Bangor Miami Corpus is a naturalistic Spanish-English code-switching speech dataset collected by Jon Russell Herring at Bangor University. It captures spontaneous bilingual conversations recorded in Miami, Florida, involving proficient Spanish-English bilinguals across multiple speaker groups. Dataset description Total recordings 56 Total duration ~32 h Languages English (en), Spanish (es) Format MP3 audio +… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/Bangor-Miami-Spanish-English-Corpus.audioautomatic-speech-recognitionn<1K0 likes140 downloads4mo agoHugging Face12AiAF /Miaoshou_Jade-Miura_Datasetimagen<1K0 likes129 downloads1y agoHugging Face13miaocongxin /KeSpeech This dataset only contains test data, which is integrated into UltraEval-Audio(https://github.com/OpenBMB/UltraEval-Audio) framework. python audio_evals/main.py --dataset KeSpeech --model gpt4o_audio 🚀超凡体验,尽在UltraEval-Audio🚀 UltraEval-Audio——全球首个同时支持语音理解和语音生成评估的开源框架,专为语音大模型评估打造,集合了34项权威Benchmark,覆盖语音、声音、医疗及音乐四大领域,支持十种语言,涵盖十二类任务。选择UltraEval-Audio,您将体验到前所未有的便捷与高效: 一键式基准管理 📥:告别繁琐的手动下载与数据处理,UltraEval-Audio为您自动化完成这一切,轻松获取所需基准测试数据。 内置评估利器… See the full description on the dataset page: https://huggingface.co/datasets/miaocongxin/KeSpeech.audio10K<n<100K0 likes124 downloads8mo agoHugging Face14Miayan /physical-relighting-datasetimage100K<n<1M0 likes119 downloads10mo agoHugging Face15patrickbdevaney /mia9geospatial10K<n<100K1 likes118 downloads3y agoHugging Face16GAIMHE /MIAAM-V2gated Mathematics Dataset Card: AM, Adaptiv World, and Adaptiv College Overview This repository contains a cross-country mathematics interaction dataset built from three digital learning platforms used in classrooms in France and Côte d’Ivoire. It spans three educational age bands, from early primary school to middle school and high-school remediation, and covers several mathematical domains, including number sense, arithmetic problem solving, geometry, fractions and… See the full description on the dataset page: https://huggingface.co/datasets/GAIMHE/MIAAM-V2.imageother1M<n<10M0 likes108 downloads11h agoHugging Face17parameterlab /scaling_mia_the_pile_00_PubMed_Centraltext100K<n<1M1 likes107 downloads2y agoHugging Face18parameterlab /scaling_mia_the_pile_00_Pile-CCtext1M<n<10M0 likes105 downloads2y agoHugging Face19mia-llm /xsum-MIA-Benchmarktext10K<n<100K0 likes97 downloads2y agoHugging Face20Miaowuawa /ChineseNovels 中文小说数据集 包含内容: 网游/系统/重生 言情小说 同人/耽美小说 科幻小说 军事小说 以上加起来共4万本左右 海棠文学城小说:约1000本(未清洗) texttext-generation1K<n<10K22 likes82 downloads2y agoHugging Face21potsawee /audio-mia-batch-20260312 Audio MIA Batch 20260312 This dataset contains 6,998 audio files (64 GB) downloaded from YouTube videos. Dataset Structure Each row contains: audio: Audio bytes (playable in the dataset viewer) video_id: YouTube video ID category: Content category source_term: Search term used query: Full search query title: Video title url: YouTube URL uploader: Channel name channel_id: YouTube channel ID upload_date: Upload date (YYYY-MM-DD) duration: Video duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/potsawee/audio-mia-batch-20260312.audioaudio-classification1K<n<10K0 likes82 downloads6mo agoHugging Face22parameterlab /scaling_mia_the_pile_00_OpenWebText2text1M<n<10M1 likes78 downloads2y agoHugging Face23malaiwah /glm53-flash-fidelity-exl3-tr3-4bpw-miaailab-v1 fidelity--glm53flash.malaiwah.quant.miaailab-4bpw A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from Mia-AiLab/GLM-5.3-Flash-EXL3-TR3-4bpw. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm53-flash-fidelity-exl3-tr3-4bpw-miaailab-v1.tabularn<1K0 likes76 downloads18d agoHugging Face24paper-2229 /openlvlm-mia OpenLVLM-MIA OpenLVLM-MIA: A Controlled Benchmark Revealing the Limits of Membership Inference Attacks on Large Vision-Language Models Overview OpenLVLM-MIA offers a controlled benchmark to reassess membership inference attacks (MIA) on large vision-language models beyond dataset-induced biases. The benchmark consists of a 6,000-image dataset with controlled member/non-member distributions and ground-truth membership at three training stages. On this setup… See the full description on the dataset page: https://huggingface.co/datasets/paper-2229/openlvlm-mia.image1K<n<10K0 likes68 downloads10mo agoHugging Face25Miaosen /openai-humaneval-sky-shadow Shadow Humaneval dataset This dataset is generated by GPT-4 to mimic openai-humaneval dataset. Each problem of HumanEval has a corresponding shadow problem in this dataset. The usage of this dataset is to check Whether a code generation model has data leakage during its training progress. You can refer to Skywork for further details. texttext-classificationn<1K3 likes63 downloads3y agoHugging Face26parameterlab /scaling_mia_the_pile_00_arxivThis dataset includes all arxiv documents from the 00.jsonl.zst partition of The Pile. It was created with this script: pile_path = "data/the_pile/train/00.jsonl.zst" with zstd.open(pile_path, 'r') as fr: with open("/tmp/arxiv.jsonl", "w") as fw: for i, line in enumerate(tqdm(fr)): doc = json.loads(line) source = doc['meta']['pile_set_name'] if source == "ArXiv": fw.write(json.dumps(doc) + "\n") The validation and test sets are… See the full description on the dataset page: https://huggingface.co/datasets/parameterlab/scaling_mia_the_pile_00_arxiv.text10K<n<100K0 likes63 downloads2y agoHugging Face27mia-llm /AGnews-MIA-Benchmarktext10K<n<100K0 likes62 downloads2y agoHugging Face28h0ssn /wikitext-unlearning-mia WIKITEXT - Machine Unlearning + MIA Evaluation Dataset (Length-Filtered) This dataset is prepared for evaluating machine unlearning methods on fine-tuned LLMs using Membership Inference Attacks (MIAs). Dataset Splits Training Sets (for Unlearning) retain_set (9,000 samples): Data to retain during unlearning forget_set (1,000 samples): Data to unlearn Evaluation Sets (for MIA) - Length-Filtered WikiText Length Variants 96 tokens (~96±10… See the full description on the dataset page: https://huggingface.co/datasets/h0ssn/wikitext-unlearning-mia.text10K<n<100K0 likes62 downloads9mo agoHugging Face29Denisijcu /miami_real_estate_data.csvtabular1K<n<10K0 likes58 downloads1y agoHugging Face30mteb /MIAO-I2Aaudio10K<n<100K0 likes58 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.