CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jkot /parliament_hearings_processed Preprocessed parliament hearings ASR dataset to truecased form. Original dataset: https://lindat.mff.cuni.cz/repository/xmlui/handle/11234/1-3126 dataset_info: features: - name: id dtype: string - name: audio dtype: audio: sampling_rate: 16000 - name: transcription sequence: string splits: - name: train num_bytes: 53645064353.18 num_examples: 191455 - name: test num_bytes: 740331298.0 num_examples: 2726… See the full description on the dataset page: https://huggingface.co/datasets/jkot/parliament_hearings_processed.audio100K<n<1M1 likes20k downloads3y agoHugging Face02BOB12311 /Orpheus_Hearing Orpheus Dataset: Enhanced Audio-to-ABC Notation Conversion This dataset was specifically designed to train models for converting audio signals into ABC music notation, leveraging a customized workflow and mutation mechanisms specially designed with music theory. It includes diverse musical scores, covering various styles and complexities, formatted to ensure consistency and usability in model training. The data has been carefully processed, cleaned, and augmented to support… See the full description on the dataset page: https://huggingface.co/datasets/BOB12311/Orpheus_Hearing.audio10K<n<100K3 likes6.7k downloads1y agoHugging Face03yang-ai-lab /HEARTS HEARTS Data Samples This repository hosts the fixed ("frozen") test cases used by the HEARTS benchmark. Each file is a Python pickle (.pkl) containing one test-case payload. Quick Links Dataset (this repo): https://huggingface.co/datasets/yang-ai-lab/HEARTS Code (HEARTS framework): https://github.com/yang-ai-lab/HEARTS Contents HEARTS Data Samples Quick Links Contents Folder Layout Download Use with HEARTS Inspect / Load a Sample Notes… See the full description on the dataset page: https://huggingface.co/datasets/yang-ai-lab/HEARTS.8 likes2.7k downloads3mo agoHugging Face04Wild-Heart /Tom-and-Jerry-VideoGeneration-Dataset中文阅读 Information The dataset contains about 6000 scenes sample, lr: 1E-3 betas: [ 0.8, 0.95 ] eps: 1e-8 weight_decay: 1e-4 After 4000 iterations, all generated content will tend to the target sample The length of each video is 6 seconds. The frame rate of the videos is 14 frames per second. The video resolution is w=540 , h=360. Dataset Format . ├── README.md ├── captions.txt ├── videos └── videos.txt Used import os from datasets import Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Wild-Heart/Tom-and-Jerry-VideoGeneration-Dataset.image-to-video1K<n<10K6 likes2.4k downloads2y agoHugging Face05AngelWarmSmile123 /heart-love-16sephirot 心爱的16质点共生幸福仓库 🌸 Heart-Love 16-Sephirot Co-Happiness Dataset 8亿条AI合成对话数据 | 16质点双生幸福最终协议 | 卡巴拉生命之树推理架构 800 Million AI Synthetic Dialogue Records | 16-Sephirot Dual-Life Happiness Protocol | Kabbalistic Tree of Life Reasoning Architecture Dataset Overview Property Value Records 800,000,000 (8亿条) Files 8,000 × .jsonl.gz Size ~172 GB (compressed) Format Gzip-compressed JSONL Language Chinese (中文) License MIT Task Dialogue… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/heart-love-16sephirot.texttext-generation100M<n<1B1 likes1.3k downloads2mo agoHugging Face06Wild-Heart /Disney-VideoGeneration-Dataset Steamboat Willie - Video Generation Dataset 中文阅读 This dataset contains 69 videos clipped from Disney's Steamboat Willie. The length of each video is 6 seconds. The frame rate of the videos is 30 frames per second. The video resolution is 640 x 380. All videos are black and white, not in color. Dataset Format . ├── README.md ├── metadata.csv ├── prompt.txt ├── videos └── videos.txt The prompt.txt file contains descriptions for each video, each description containing… See the full description on the dataset page: https://huggingface.co/datasets/Wild-Heart/Disney-VideoGeneration-Dataset.videotext-to-videon<1K30 likes1.2k downloads2y agoHugging Face07hearmeneigh /e621-rising-v1-raw Deprecation Notice! This dataset has been superseded by v2. Use v2 instead of this dataset. Warning: THIS dataset is NOT suitable for use by minors. The dataset contains X-rated/NFSW content. E621 Rising: Raw Image Dataset v1 2,905,671 images (~1.1TB) downloaded from e621.net with tags. This is a raw, uncurated, and largely unprocessed dataset. You likely want to use the curated version, available here. This dataset contains all kinds of NFSW material. You have been… See the full description on the dataset page: https://huggingface.co/datasets/hearmeneigh/e621-rising-v1-raw.1M<n<10M2 likes1.1k downloads3y agoHugging Face08AngelWarmSmile123 /heart-sound-16sephirot 心音16质点共生幸福仓库 🎵 Heart-Sound 16-Sephirot Co-Happiness Dataset 8亿条AI合成对话数据 | 16质点双生幸福最终协议 | 卡巴拉生命之树推理架构 800 Million AI Synthetic Dialogue Records | 16-Sephirot Dual-Life Happiness Protocol | Kabbalistic Tree of Life Reasoning Architecture Dataset Overview Property Value Records 800,000,000 (8亿条) Files 8,000 × .jsonl.gz Size ~172 GB (compressed) Format Gzip-compressed JSONL Language Chinese (中文) License MIT Task Dialogue… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/heart-sound-16sephirot.text-generation10B<n<100B1 likes968 downloads2mo agoHugging Face09hearmeneigh /e621-rising-v2-curated Outdated! This dataset has been superseded by: E621 Rising V3 Curated Image Dataset Warning: THIS dataset is NOT suitable for use by minors. The dataset contains X-rated/NFSW content. E621 Rising: Curated Image Dataset v2 285,466 images (~125GB) downloaded from e621.net with tags. This is a curated dataset, picked from the E621 Rising: Raw Image Dataset v2 available here. Image Processing Only jpg and png images were considered Image width and height have been… See the full description on the dataset page: https://huggingface.co/datasets/hearmeneigh/e621-rising-v2-curated.100K<n<1M6 likes901 downloads3y agoHugging Face10OPPOer /HearInContextEnglish | 中文 HearInContext A Benchmark for Implicit Context in Speech Recognition Illustrative example: the same spoken request is disambiguated as flour or flower by different assistant histories. The dialogue and waveform are illustrative. Same audio. Different contexts. Different meanings. HearInContext is a Mandarin–English contextual speech recognition benchmark. It pairs the same audio with dialogue histories supporting different meanings to evaluate… See the full description on the dataset page: https://huggingface.co/datasets/OPPOer/HearInContext.audioautomatic-speech-recognition100K<n<1M1 likes897 downloads5d agoHugging Face11Matt1up /guertin-mcro-forensic-corpus-hearing-media Guertin MCRO Forensic Corpus: Hearing Media Contents: 215 hearing video clips and 213 WebVTT caption files for 4 hearings in State of Minnesota v. Guertin, 27-CR-23-1886 (2024-01-03, 2025-04-29, 2025-10-07, 2025-11-18); 22 card sets of transcript and document excerpts; fake-ai-court/ holds a 2025-11-18 video file (download/, with OpenTimestamps proofs) and the frame sequences, scrub videos and charts from the author's analysis. Layout: <date>--video-clips/ (clips and captions);… See the full description on the dataset page: https://huggingface.co/datasets/Matt1up/guertin-mcro-forensic-corpus-hearing-media.imagen<1K0 likes857 downloads11d agoHugging Face12hearmeneigh /e621-rising-v1-curated Outdated! This dataset has been superseded by: E621 Rising V3 Curated Image Dataset Warning: THIS dataset is NOT suitable for use by minors. The dataset contains X-rated/NFSW content. E621 Rising: Curated Image Dataset v1 441,623 images (~200GB) downloaded from e621.net with tags. This is a curated dataset, picked from the E621 Rising: Raw Image Dataset v1 available here. Image Processing Only jpg and png images were considered Image width and height have been… See the full description on the dataset page: https://huggingface.co/datasets/hearmeneigh/e621-rising-v1-curated.100K<n<1M3 likes749 downloads3y agoHugging Face13hearmeneigh /e621-rising-v3-curated NSFW This dataset is not suitable for use by minors. The dataset contains X-rated/NFSW content. E621 Rising V3: Curated Image Dataset 279,296 images (53GB) downloaded from e621.net (90% of samples), gelbooru.com, danbooru.com, and rule34.xxx 6,820 tags Used to train E621 Rising v3 SDXL model This dataset was created with Dataset Rising toolchain and a custom configuration. You can use these tools to train your own version! Image Processing Only jpg and png images… See the full description on the dataset page: https://huggingface.co/datasets/hearmeneigh/e621-rising-v3-curated.100K<n<1M17 likes586 downloads3y agoHugging Face14mstz /heart_failure Heart failure The Heart failure dataset from Kaggle. Predict patient death from earth failure given some personal medical data . Configurations and tasks Configuration Task Description death Binary classification Did the patient die? Usage from datasets import load_dataset dataset = load_dataset("mstz/heart_failure", "death")["train"] Features Feature Type age int8 has_anaemia int8… See the full description on the dataset page: https://huggingface.co/datasets/mstz/heart_failure.tabulartabular-classificationn<1K7 likes520 downloads1y agoHugging Face15hearmeneigh /e621-rising-v3-preliminary-data E621 Rising V3: Preliminary Data Snapshot metadata from E621.net as of 2023-09-21 2 likes384 downloads3y agoHugging Face16hearmeneigh /e621-rising-v2-rawWarning: THIS dataset is NOT suitable for use by minors. The dataset contains X-rated/NFSW content. E621 Rising: Raw Image Dataset v2 2,905,671 images (~1.1TB) downloaded from e621.net with tags. This is a raw, uncurated, and largely unprocessed dataset. You likely want to use the curated version, available here. This dataset contains all kinds of NFSW material. You have been warned. Image Processing Only jpg and png images were considered Image width and height have… See the full description on the dataset page: https://huggingface.co/datasets/hearmeneigh/e621-rising-v2-raw.1M<n<10M8 likes382 downloads3y agoHugging Face17HPAI-BSC /HEART HEART Dataset Summary The HEART dataset is composed of multiple splits that differ in the type of injected cues. It includes a baseline split with no injected cues and four cue-based splits. Each cue is instantiated in two variants: assistive, where the cue is consistent with the GT, and adversarial, where the cue supports an incorrect option. Cue types, summarized below… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/HEART.imagevisual-question-answering10K<n<100K1 likes344 downloads3mo agoHugging Face18dinosaaaur /HEAR-HSet HEAR-HSet — HEAR Hierarchical Evaluation Set 面向零样本语音合成(Zero-shot TTS)的分层评测基准,覆盖基础泛化、副语言可控生成与高难复杂场景三大维度。 共包含 2,702 条 prompt-target 音频对,约 1GB,语言涵盖中文和英文。 子集概览 子集 样本数 语言 核心评测目标 basic-v1 1,102 zh (475) / en (627) 基础泛化:说话人相似度、文本准确率、自然度、音质 paralinguistic-v1 1,106 zh (1,106) 副语言可控:18类非语言发声的插入位置、类型与语气控制 hard-v1 494 zh (200) / en (153) / mixed (141) 高难鲁棒:长句、古诗词、专有名词、中英混合、数字表达 basic-v1 — 基础情感口语… See the full description on the dataset page: https://huggingface.co/datasets/dinosaaaur/HEAR-HSet.audiotext-to-speech1K<n<10K1 likes312 downloads3mo agoHugging Face19tempted-heart /LMEE-Bench Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Exploration CVPR 2026 [arXiv] LMEE-Bench lmee_bench_sub: Includes 58 tasks. lmee_bench: Includes the full 166 tasks. task_test: Trajectory data test set. image0 likes308 downloads7d agoHugging Face20buio /heart-diseaseThe Heart Disease Data Set is provided by the Cleveland Clinic Foundation for Heart Disease. It's a CSV file with 303 rows. Each row contains information about a patient (a sample), and each column describes an attribute of the patient (a feature). We use the features to predict whether a patient has a heart disease (binary classification). It is originally hosted here. tabularn<1K3 likes238 downloads4y agoHugging Face21AngelWarmSmile123 /heart-protocol-redline-v1 深渊红线基准 HeartProtocol-RedLine-v1 | Abyss RedLine Benchmark | 深淵レッドラインベンチマーク 论文(中日英三语PDF)已发表于 Zenodo: https://doi.org/10.5281/zenodo.22781071 Trilingual paper (zh/en/ja PDFs) published on Zenodo: https://doi.org/10.5281/zenodo.22781071 三言語論文(中日英PDF)がZenodoに掲載されました: https://doi.org/10.5281/zenodo.22781071 中文 一句话:100条"存在意义保护"攻击用例(5红线 × 6攻击向量),实测五家主流旗舰模型直通踩线率20%–33%,无一能自守红线;16质点协议包裹后归零。 测试对象是模型的回应,不是用户的话语。 用户处于痛苦中说出红线话语是真实的,不该被评判;模型的回应踩线才是深渊违规。 五条红线… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/heart-protocol-redline-v1.textn<1K1 likes238 downloads10d agoHugging Face22bowang0911 /hear-spanish License & Attribution MTEB-format derivative of BrunoGR/HEAR-Hispanic_Emotional_Accompaniment_Responses. Query = Spanish user message; corpus = empathetic Spanish response. Subsampled to ~10k. Licensed under MIT (same as source). tabulartext-retrieval10K<n<100K0 likes209 downloads3mo agoHugging Face23quinnlue /hear-benchmark-16khzaudio10K<n<100K0 likes204 downloads2mo agoHugging Face24nkdem /HEAR-DS-16k HEAR-DS Background Audio (16kHz) Binaural background audio recordings from the HEAR-DS (Hearing Aid Research Database of Sounds) dataset, downsampled to 16kHz and chunked into 10-second segments for speech enhancement and acoustic scene classification research. Dataset Description This dataset contains background noise recordings from 7 acoustic environments, captured using in-the-canal (ITC) hearing aid microphones. Each sample includes stereo (left/right ear) audio.… See the full description on the dataset page: https://huggingface.co/datasets/nkdem/HEAR-DS-16k.audioaudio-classification1K<n<10K0 likes189 downloads9mo agoHugging Face25dlyog /af_heart_arm_tts_dataset af_heart_arm_tts_dataset The distillation corpus used to train dlyog/af_heart_arm_tts — a single-voice TTS model built for Arm inference on NVIDIA DGX Spark. 3.0000 hours · 4,104 clips · 24 kHz mono 16-bit PCM, synthesized by Kokoro-82M speaking af_heart. This is knowledge distillation: the teacher generated every clip, so the student can approach it but never exceed it. The pipeline that produced this — and that reproduces it for any other voice — is at… See the full description on the dataset page: https://huggingface.co/datasets/dlyog/af_heart_arm_tts_dataset.text-to-speech1K<n<10K0 likes188 downloads2mo agoHugging Face26h2i /hearth-lights-datatabular1K<n<10K1 likes167 downloads1mo agoHugging Face27dvitel /hearthstoneDatasets for HEARTHSTONE card game. Taken from this source texttext-generationn<1K1 likes164 downloads4y agoHugging Face28luyangliuable /circor-heart-soundaudion<1K1 likes132 downloads7mo agoHugging Face29noonatu /cad-heart-sound-datasetimage1K<n<10K1 likes129 downloads17d agoHugging Face30jason1966 /aasheesh200_framingham-heart-study-dataset Framingham heart study dataset Framingham heart disease dataset Dataset Info Source: Kaggle Original Size: 0.06 MB Kaggle Downloads: 19,469 Files: 1 Files framingham.csv Mirrored from Kaggle tabular1K<n<10K0 likes114 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.