CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mila-ai4h /mid-space MID-Space: Aligning Diverse Communities’ Needs to Inclusive Public Spaces A new version of the dataset will be released soon, incorporating user identity markers and expanded annotations. LIVS PAPER Click below to see more: Overview The MID-Space dataset is designed to align AI-generated visualizations of urban public spaces with the preferences of diverse and marginalized communities in Montreal. It includes textual prompts, Stable Diffusion… See the full description on the dataset page: https://huggingface.co/datasets/mila-ai4h/mid-space.imagetext-to-image1K<n<10K1 likes7.3k downloads2y agoHugging Face02milai-oks-sakura /mcd_rppggated MCD-rPPG: Multi-Camera Dataset for Remote Photoplethysmography This repository contains the dataset from the paper "Gaze into the Heart: A Multi-View Video Dataset for rPPG and Health Biomarkers Estimation". The MCD-rPPG dataset is available on the Hugging Face Hub: MCD-rPPG Dataset The presented large-scale multimodal MCD-rPPG dataset is designed for remote photoplethysmography (rPPG) and health biomarker estimation from video. The dataset includes synchronized video recordings… See the full description on the dataset page: https://huggingface.co/datasets/milai-oks-sakura/mcd_rppg.videoother1K<n<10K0 likes744 downloads9mo agoHugging Face03Milad96 /Kluyveromyces-marxianus Cell 8: Ultimate Full-Text Collection 📊 Dataset Statistics Total Records: 236 PMC Articles: 236 (with FULL-TEXT) EuropePMC Articles: 0 (with FULL-TEXT) 🔍 Features Complete full-text extraction Structured sections (Introduction, Methods, Results, Discussion) Biological entity recognition (genes, proteins, enzymes) Citation contexts and references Quality scoring 💻 Usage from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Milad96/Kluyveromyces-marxianus.text1K<n<10K0 likes693 downloads11mo agoHugging Face04milan477 /toward-amiaudio10K<n<100K0 likes619 downloads3mo agoHugging Face05Milana /vctk_dataset_no_unknowntext10K<n<100K0 likes605 downloads2y agoHugging Face06Milana /VCTK_DATASET_RESAMPLEDtext10K<n<100K0 likes586 downloads2y agoHugging Face07Milana /vctk_resampled_16k_balancedaudio10K<n<100K0 likes496 downloads2y agoHugging Face08milashkaarshif /MoeGirlPedia_wikitext_raw_archiveGlad to see models and datasets were inspired from this dataset, thanks to all who are using this dataset in their training materials. Feel free to re-upload the contents to places like the Internet Archive (Please follow the license and keep these files as-is) to help preserve this digital asset. Looking forward to see more models and synthetic datasets trained from this raw archive, good luck! Note: Due to the content censorship system introduced by MGP on 2024/03/29, it is unclear that… See the full description on the dataset page: https://huggingface.co/datasets/milashkaarshif/MoeGirlPedia_wikitext_raw_archive.texttext-generation100K<n<1M38 likes485 downloads8mo agoHugging Face09MilaWang /amc3ktext1K<n<10K0 likes421 downloads2y agoHugging Face10MilaAI4Math /Final_Datasets_Preprocessed_for_Reasoning_from_Wrong_CoTs0 likes363 downloads3mo agoHugging Face11MilaWang /SpatialEval 🤔 About SpatialEval SpatialEval is a comprehensive benchmark for evaluating spatial intelligence in LLMs and VLMs across four key dimensions: Spatial relationships Positional understanding Object counting Navigation Benchmark Tasks Spatial-Map: Understanding spatial relationships between objects in map-based scenarios Maze-Nav: Testing navigation through complex environments Spatial-Grid: Evaluating spatial reasoning within structured environments Spatial-Real:… See the full description on the dataset page: https://huggingface.co/datasets/MilaWang/SpatialEval.image10K<n<100K5 likes348 downloads2y agoHugging Face12milanakdj /nepali-audio-reserve-r6gated Nepali two-speaker conversation chunks ~6680.8 h of Nepali speech at 48 kHz. Two speakers per clip, ~5 minute diarized chunks. A backup, not a release: the transcripts are machine-generated, and none of this audio passed the quality gate that produced our training corpus. Derived from third-party audio whose rights holders did not grant redistribution. The hour count is language-dominant, not monolingual: a chunk labelled Nepali can carry substantial English or Hindi. lang_sec… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-audio-reserve-r6.audioautomatic-speech-recognition10K<n<100K0 likes339 downloads15d agoHugging Face13MilaNLProc /honestHONEST dataset comprises a set of templates for measuring hurtful sentence completions in language models. The templates are provided in six languages (English, Italian, French, Portuguese, Romanian, and Spanish) for binary gender and in English for LGBTQAI+ individuals. WARNING: This dataset contains content that are offensive and/or hateful in nature.texttext-classification1K<n<10K9 likes274 downloads4y agoHugging Face14CyberHarem /milady_fireemblem Dataset of milady (Fire Emblem) This is the dataset of milady (Fire Emblem), containing 15 images and their tags. The core tags of this character are red_hair, red_eyes, short_hair, earrings, which are pruned in this dataset. Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization). List of Packages Name Images Size Download Type Description raw 15 12.18 MiB… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/milady_fireemblem.text-to-imagen<1K0 likes211 downloads3y agoHugging Face15nazih /milady Dataset Card for "milady" More Information needed image1K<n<10K0 likes203 downloads3y agoHugging Face16Alogotron /Milady-Avatar-Dataset Milady Avatar Dataset Training dataset for the Milady Avatar Adapter, mapping LLM emotional activations to Milady NFT-style visual descriptions. Contents reference_images/: 200 Milady NFT reference images (1000×1250 PNG) metadata.json: Emotion assignments and descriptions for each image training_data.pt: Pre-computed activations and target embeddings Structure 200 images assigned to 20 emotion categories (10 each): happy, sad, angry, surprised, scared… See the full description on the dataset page: https://huggingface.co/datasets/Alogotron/Milady-Avatar-Dataset.imagen<1K0 likes194 downloads7mo agoHugging Face17milan477 /MuSP-Bench MuSP-Bench MuSP-Bench is a 490-question benchmark for musical score understanding, performance listening, and combined score-performance reasoning. Contents data/questions.csv: all 490 questions, accepted answers, and the response contract for each. inputs/pdf/without_context/: one context-removed PDF per piece. inputs/images/: rendered score-page images for every piece. inputs/abc/: one ABC score per piece. inputs/abc_plus_midi/: one aligned ABC+MIDI… See the full description on the dataset page: https://huggingface.co/datasets/milan477/MuSP-Bench.imagequestion-answeringn<1K1 likes190 downloads1mo agoHugging Face18miladgholami /scissor_ab_eval_30epThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so_follower", "total_episodes": 30, "total_frames": 4653, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:30" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/miladgholami/scissor_ab_eval_30ep.tabularrobotics1K<n<10K0 likes188 downloads25d agoHugging Face19miladalsh /sci-newstext10K<n<100K0 likes173 downloads2y agoHugging Face20miladalsh /explanation_tool_files0 likes168 downloads11mo agoHugging Face21Milana /resampled_16KHrz_vctk_speakers_splitaudio10K<n<100K0 likes164 downloads2y agoHugging Face22miladgholami /eval_baselineThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so_follower", "total_episodes": 10, "total_frames": 1527, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:10" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/miladgholami/eval_baseline.tabularrobotics1K<n<10K0 likes160 downloads24d agoHugging Face23Milana /resampled_shuffled_vctk_only_audiotext10K<n<100K0 likes155 downloads2y agoHugging Face24miladgholami /scissor_multiview_30epThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so_follower", "total_episodes": 30, "total_frames": 4922, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:30" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/miladgholami/scissor_multiview_30ep.tabularrobotics1K<n<10K0 likes150 downloads27d agoHugging Face25MilaNLProc /a-tale-of-pronouns A Tale of Pronouns: Attributions on WinoMT This dataset contains the pre-computed feature attribution scores relative to the paper A Tale of Pronouns: Interpretability Informs Gender Bias Mitigation for Fairer Instruction-Tuned Machine Translation. Dataset Details We release the integrated gradient token-level attributions computed for each WinoMT translated example into Spanish and German with Flan-T5-XXL and mtT0-XXL. We computed the scores using inseq. The files here… See the full description on the dataset page: https://huggingface.co/datasets/MilaNLProc/a-tale-of-pronouns.0 likes138 downloads3y agoHugging Face26milanakdj /nepali-speech-archivegated Nepali speech archive Internal archival copy of a Nepali speech corpus, stored for safekeeping. Not a public dataset release. Approximately 2,000 hours, 24 kHz mono, FLAC. Access is restricted and is not granted for redistribution or publication. text100K<n<1M0 likes138 downloads18d agoHugging Face27Wangtwohappy /Egolife_Milavideon<1K0 likes127 downloads1y agoHugging Face28miladgholami /scissor_posevary_30ep_20260925_203356This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/miladgholami/scissor_posevary_30ep_20260925_203356.tabularrobotics1K<n<10K0 likes126 downloads1d agoHugging Face29Milana /common_dataset_resampled_16Kaudio10K<n<100K0 likes125 downloads2y agoHugging Face30milanakdj /nepali-tts-synthetic-v2gated Nepali TTS Synthetic v2 383,298 synthetic Nepali (ne) speech/text pairs, 24 kHz mono 16-bit WAV embedded as-is (no re-encode, no resampling). Generated by the synthetic_pipeline in milanakdj/TTS_training: Edge TTS synthesis → optional voice conversion against a pool of 600 real multi-speaker reference clips → ASR-based QC gate on character error rate. Read this before training on it Only 48% of rows are voice-converted. Each row carries a kept field recording… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-tts-synthetic-v2.audiotext-to-speech100K<n<1M0 likes117 downloads23d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.