CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Kevet /imgimagen<1K0 likes10k downloads1y agoHugging Face02kev016 /lsd0 likes5.8k downloads1y agoHugging Face03kevin009 /olympiad-math-contest-llama3-78ktext10K<n<100K1 likes3.8k downloads2y agoHugging Face04kevinS4455 /trellis500k-github-archives-70 likes3.3k downloads6mo agoHugging Face05KevinHuang /InteriorVerseCaptions of InteriorVerse's RGB images extracted with microsoft/Florence-2-large. 0 likes2.3k downloads11mo agoHugging Face06kevin009 /olympiad-math-stepwise-solutions-llama3-20kThe MATH dataset is a collection of 20,300 problems from AMC and AIME competitions covering algebra, number theory, geometry, and precalculus problems and solution sets. Problems and solutions are formatted in LATEX. Step-by-step solutions and insight sections have been added in order to use a chain of thought to clarify the problem and solution. text10K<n<100K1 likes2.3k downloads2y agoHugging Face07Kevin1804 /BoxFusionimage0 likes2.2k downloads9mo agoHugging Face08KevinJustin /FormulaBank-28K FormulaBank-28K FormulaBank-28K is a deterministic procedural-audio corpus for audio representation pre-training. It contains 28,000 mono clips at 16 kHz, each exactly 10.24 seconds long. Configuration Version: C125-I224-v1 Formula classes: 125 Renderings per class: 224 Total clips: 28,000 Audio format: lossless 24-bit FLAC Source: frozen AudioPG-Atomic-H7-C224-R0-Clean FormulaBank manifest Each formula class specifies an acoustic rendering rule. Each rendering… See the full description on the dataset page: https://huggingface.co/datasets/KevinJustin/FormulaBank-28K.audiofeature-extraction10K<n<100K0 likes2k downloads19d agoHugging Face09kevinlad /s1-datasetgated S1 Dataset Access to the dataset files is gated. Approved users must sign in to their own Hugging Face account before downloading. Upload status Last updated: 2026-08-25 02:58 UTC Expected: 27,577 HDF5 files Available: 27,577 HDF5 files Overall progress: 100.00% Complete sequences: 83 / 83 Remaining: 0 files (approximately 0.00 GiB) Sequence Status Available Expected Missing Progress - Complete - - 0 100.00% All sequences not shown in the table… See the full description on the dataset page: https://huggingface.co/datasets/kevinlad/s1-dataset.0 likes2k downloads1mo agoHugging Face10Kevin355 /Who_and_When Who&When: #1 Benchmark for MAS automated failure attribution. 184 annotated failure tasks collected from Algorithm-generated agentic systems built using CaptainAgent, Hand-crafted systems such as Magnetic-One. Fine-grained annotations for each failure, including: The failure-responsible agent (who failed), The decisive error step (when the critical error occurred), A natural language explanation of the failure. The dataset covers a wide range of realistic multi-agent scenarios… See the full description on the dataset page: https://huggingface.co/datasets/Kevin355/Who_and_When.textn<1K13 likes1.8k downloads9mo agoHugging Face11kevin510 /libero_sft_policies0 likes1.6k downloads1y agoHugging Face12kevindenight /gdpval-gpt5 GDPval with GPT-5 Execution Results This dataset contains the OpenAI GDPval benchmark with comprehensive execution results from GPT-5, demonstrating AI capabilities across real-world professional tasks. 🎯 Dataset Overview This is an enhanced version of the original OpenAI GDPval dataset with actual AI model execution results and professional deliverables. 📊 Key Statistics Total tasks: 220 Tasks with AI deliverables: 87 (39.5%) Professional files generated:… See the full description on the dataset page: https://huggingface.co/datasets/kevindenight/gdpval-gpt5.documentothern<1K0 likes1.6k downloads11mo agoHugging Face13jaredpalmer /kev-suites6 likes1.5k downloads3d agoHugging Face14Kevin-Pal /CUHK-X_Small_Model_Trackgated CUHK-X — Small Model Track Multimodal human action recognition (classification). Given a multimodal clip, predict its action class (action_id, 0–39, 40 classes). Repository layout . ├── Training/ │ ├── class_mapping.csv # action_id <-> action_name (40 classes) │ └── data/ │ └── HAR.z01 … HAR.z08 + HAR.zip # multi-volume zip │ → HAR/data/<modality>/<action>/<user>/<trial>/<files> └── Testing/ ├── data/ │ └──… See the full description on the dataset page: https://huggingface.co/datasets/Kevin-Pal/CUHK-X_Small_Model_Track.textvideo-classificationn<1K7 likes1.3k downloads3mo agoHugging Face15kevindenight /gdpval-gpt5-fork GDPval Fork Dataset with GPT-5 Results 🏆 A comprehensive evaluation dataset featuring GPT-5 execution results on real-world professional tasks This is an enhanced fork of the original OpenAI GDPval dataset with complete GPT-5 execution results, including actual deliverable files created by the AI model. 📊 Dataset Overview Metric Value Total Tasks 220 AI-Completed Tasks 87 (39.5%) Deliverable Files 492+ professional documents Occupations 44 Industry… See the full description on the dataset page: https://huggingface.co/datasets/kevindenight/gdpval-gpt5-fork.documentothern<1K0 likes1.3k downloads11mo agoHugging Face16KevinNotSmile /nuscenes-qa-mini NuScenes-QA-mini Dataset TL;DR: This dataset is used for multimodal question-answering tasks in autonomous driving scenarios. We created this dataset based on nuScenes-QA dataset for evaluation in our paper Modality Plug-and-Play: Elastic Modality Adaptation in Multimodal LLMs for Embodied AI. The samples are divided into day and night scenes. scene # train samples # validation samples day 2,229 2,229 night 659 659 Each sample contains… See the full description on the dataset page: https://huggingface.co/datasets/KevinNotSmile/nuscenes-qa-mini.textvisual-question-answering1K<n<10K4 likes1.3k downloads3y agoHugging Face17KevinConnorLee /complet4r_preprocessed_dynamicreplica0 likes946 downloads3mo agoHugging Face18kevmo314 /diffpackDiffPack is the bigcode/commitpack dataset except diff'd between the old and new data. text-generation0 likes896 downloads2y agoHugging Face19kevin009 /olympiad-math-contest-llama3-20k AMC/AIME Mathematics Problem and Solution Dataset Dataset Details Dataset Name: AMC/AIME Mathematics Problem and Solution Dataset Version: 1.0 Release Date: 2024-06-1 Authors: Kevin Amiri Intended Use Primary Use: The dataset is created and intended for research and an AI Mathematical Olympiad Kaggle competition. Intended Users: Researchers in AI & mathematics or science. Dataset Composition Number of Examples: 20,300 problems and solution sets… See the full description on the dataset page: https://huggingface.co/datasets/kevin009/olympiad-math-contest-llama3-20k.text10K<n<100K2 likes831 downloads2y agoHugging Face20KevinQHLin /ScreenSpotimage1K<n<10K1 likes746 downloads2y agoHugging Face21KevinFan111 /3d-front-code 3D-Front-Code RoomScript v4 Blender object programs, room-layout renders, code-only wall architecture, and asset Blender artifacts derived from 3D-FRONT scene evidence. Contents 13,917 object assets (reference and v4/best) in data/assets/*.tar 21,202 rooms (v4/best and code-only v4_wall/best) in data/rooms/*.tar searchable JSONL indexes under metadata/ Each tar contains multiple samples while preserving the original data/front_object_code/by_asset/... or… See the full description on the dataset page: https://huggingface.co/datasets/KevinFan111/3d-front-code.tabular10K<n<100K0 likes683 downloads19d agoHugging Face22KevinConnorLee /SLF0 likes679 downloads6mo agoHugging Face23kevinson7515 /Cog-Research0 likes644 downloads6d agoHugging Face24kevinzzz8866 /ByteDance_Synthetic_Videos Dataset Name CGI synthetic videos generated in paper "Synthetic Video Enhances Physical Fidelity in Video Synthesis" (https://simulation.seaweed.video/) Dataset Overview Number of samples: [uploading...] Annotations: [tags, captions] License: [apache-2.0] Citation: @article{zhao2025synthetic, title={Synthetic Video Enhances Physical Fidelity in Video Synthesis}, author={Zhao, Qi and Ni, Xingyu and Wang, Ziyu and Cheng, Feng and Yang, Ziyan and Jiang, Lu and Wang, Bohan}… See the full description on the dataset page: https://huggingface.co/datasets/kevinzzz8866/ByteDance_Synthetic_Videos.video10K<n<100K3 likes642 downloads1y agoHugging Face25kevinjesse /typebert Dataset Card for "typebert" More Information needed 1M<n<10M1 likes623 downloads3y agoHugging Face26P-Kevin /hssd-annotations hssd-annotations English | 中文 A standalone, locally-stored, zero-dependency Python API to search and retrieve HSSD assets and their full annotation set — for downstream scene generation (e.g. SceneSmith) and articulation/clearance research. Every annotation family is merged into one per-asset record keyed by the HSSD asset id. The library ships in post-replacement form: each asset is linked to its articulated realization (official HSSD articulated, or a PartNet-Mobility… See the full description on the dataset page: https://huggingface.co/datasets/P-Kevin/hssd-annotations.1 likes606 downloads1mo agoHugging Face27kevinjesse /ManyRefactors4Ctext10M<n<100M0 likes596 downloads4y agoHugging Face28KevinMathew /stereo4d-lefteye-perspective Dataset Summary This dataset contains the left-eye rectified perspective views from the Stereo4D dataset (Paper). Each video is generated using the rectify.py script, which processes VR180 stereo videos to produce 512×512 video clips with a 60° field of view perspective camera. This dataset is intended to be used alongside the Stereo4D dataset annotations which can be found here. This dataset is provided as-is for non-commercial research purposes only. Download git clone… See the full description on the dataset page: https://huggingface.co/datasets/KevinMathew/stereo4d-lefteye-perspective.text100K<n<1M12 likes586 downloads1y agoHugging Face29Kevynf /llm-graph-poisoning-data Generation-Time Poisoning of LLM-Generated Social Networks This dataset contains synthetic personas, LLM-generated social graphs, cached text embeddings, and evaluation metrics for clean generation and three generation-time attack families. All names and profiles are synthetic and do not represent real people. Dataset variants Variant Nodes Generator Graph seeds per condition Attack rates p50 50 Qwen3-Max 10 10%, 20%, 30%, 40%, 50% p200 200… See the full description on the dataset page: https://huggingface.co/datasets/Kevynf/llm-graph-poisoning-data.tabulargraph-ml100K<n<1M0 likes568 downloads1mo agoHugging Face30KevinDavidHayes /t2i-finegrain t2i-finegrain Dataset This dataset evaluates text-to-image (T2I) diffusion models using a benchmark of prompts designed to elicit specific failure modes. Human labels allow for T2I benchmarking evaluations. Contents 10,587 total image–metadata entries 750+ prompts 11 failure mode categories 27 specific failure modes 14 total models evaluated: 5 models (with human ground truths): SD3-XL SD3-M SD3.5-Large SD3.5-Medium Flux 9 models: Flux-Kontext – 760 images… See the full description on the dataset page: https://huggingface.co/datasets/KevinDavidHayes/t2i-finegrain.text-to-image1 likes553 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.