CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kainecorneko /twaitch-txttabular1K<n<10K0 likes4.7k downloads8mo agoHugging Face02Kaij00 /MSVQAThis is a multimodal cross-scenario dataset for continual learning with MLLMs. We provide a simple script to split the dataset in multiple ways. The dataset format has been adjusted for Qwen. The coordinates in 'train_annfiles.json' and 'val_annfiles.json' are adjusted to Qwen2.5VL format. And 'train_annfiles_ori.json' and 'val_annfiles_ori.json' retain the original coordinates of the bounding box. You need to adjust the coordinates fit your format. Detailed information can refer to… See the full description on the dataset page: https://huggingface.co/datasets/Kaij00/MSVQA.imagevisual-question-answering10K<n<100K2 likes3.5k downloads9mo agoHugging Face03thu-pacman /PCMind-2.1-Kaiyuan-2B This repository contains the complete pretraining dataset for PCMind-v2.1-Kaiyuan-2B, a leading fully open-source language model. Overview The dataset is organized into 5 training phases, with all phase datasets open-sourced in this repository. Our training methodology employs domain-specific mixing strategies across five primary domains: English: General English text Chinese: General Chinese text Code: Programming and code-related content Math: Mathematical reasoning and… See the full description on the dataset page: https://huggingface.co/datasets/thu-pacman/PCMind-2.1-Kaiyuan-2B.texttext-generation1B<n<10B5 likes2.4k downloads10mo agoHugging Face04KAIST-SmartDesignLab /DeepJEB-PP DeepJEB++ Foundation Model-Driven Large-Scale 3D Engineering Dataset via 2D Latent Space Augmentation Soyoung Yoo · Leekyo Jeong · Jinsu Ra · Dongeon Lee · Sunwoong Yang · Hyogu Jeong · Namwoo Kang &nbsp;—&nbsp; KAIST SmartDesignLab 📦 Dataset size & viewer note. DeepJEB++ contains 15,360 deployable, simulation-labeled brackets. The Hugging Face Dataset Viewer above shows only a small preview because the FEA field data are distributed as a compressed archive… See the full description on the dataset page: https://huggingface.co/datasets/KAIST-SmartDesignLab/DeepJEB-PP.imagetabular-regressionn<1K3 likes2k downloads24d agoHugging Face05KaiserML /Techie_Raw_PDFtabular100K<n<1M0 likes1.9k downloads3y agoHugging Face06kaist-ai /CoT-Collection""" _LICENSE = "CC BY 4.0" _HOMEPAGE = "https://github.com/kaistAI/CoT-Collection" _LANGUAGES = { "en": "English", } # _ALL_LANGUAGES = "all_languages" class CoTCollectionMultiConfig(datasets.BuilderConfig):texttext-generation1M<n<10M163 likes1.9k downloads3y agoHugging Face07nips26anonymous159 /Kairos Kairos — Long-Form Video Annotation and Benchmark Kairos is an automated annotation pipeline for long-duration videos (10–30 minutes). This repository hosts a benchmark of 2,870 multiple-choice and 2,870 free-form (OpenQA) questions across 820 videos, spanning 17 fine-grained capabilities and 5 temporal tiers (T1: single moment, T2: 1–60 s, T3: 60–300 s, T4: 300–900 s, T5: >900 s). What's inside . ├── data/ │ ├── kairos_benchmark.jsonl # 2,870 MCQs (bilingual… See the full description on the dataset page: https://huggingface.co/datasets/nips26anonymous159/Kairos.tabularvideo-text-to-text1K<n<10K0 likes1.9k downloads5mo agoHugging Face08BangumiBase /kaifukujutsushinoyarinaoshi Bangumi Image Base of Kaifuku Jutsushi No Yarinaoshi This is the image base of bangumi Kaifuku Jutsushi no Yarinaoshi, we detected 86 characters, 4799 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kaifukujutsushinoyarinaoshi.image1K<n<10K0 likes1.4k downloads2y agoHugging Face09KaiChen1998 /coda-lm-llava-format CODA-LM Dataset Card CODA-LM is the multi-modal version of the CODA dataset, used in the CODA-LM paper. Both English and Chinese annotations are available. Check detailed usage in our Github repo. This repo contains the CODA-LM dataset, which has been reorganized in the LLaVA data format. You are also welcome to check the original CODA-LM data which contains more metadata vanilla annotations. Usage from datasets import load_dataset # name can be selected from… See the full description on the dataset page: https://huggingface.co/datasets/KaiChen1998/coda-lm-llava-format.imageimage-to-text10K<n<100K3 likes1.3k downloads2y agoHugging Face10kainecorneko /twaitch-txt-2tabularn<1K0 likes1.1k downloads2mo agoHugging Face11lmarena-ai /PPE-GPQA-Best-of-K Overview This contains the GPQA correctness preference evaluation set for Preference Proxy Evaluations. The prompts are sampled from GPQA. This dataset is meant for benchmarking and evaluation, not for training. Paper Code License User prompts are licensed under CC BY 4.0, and model outputs are governed by the terms of use set by the respective model providers. Citation @misc{frick2024evaluaterewardmodelsrlhf, title={How to Evaluate Reward Models for… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/PPE-GPQA-Best-of-K.tabularn<1K1 likes1.1k downloads2y agoHugging Face12Kaichengalex /RealSyn100M [ACM MM25] RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm Tiancheng Gu, Kaicheng Yang, Chaoyi Zhang, Yin Xie, Xiang An, Ziyong Feng, Dongnan Liu, Weidong Cai, Jiankang Deng 💡 Introduction Contrastive Language-Image Pre-training (CLIP) demonstrates promising performance on a wide variety of benchmarks. However, a substantial volume of non-paired data, such as multimodal interleaved documents, remains… See the full description on the dataset page: https://huggingface.co/datasets/Kaichengalex/RealSyn100M.image10M<n<100M16 likes898 downloads1y agoHugging Face13gladiator7737 /kaist-urban-dataset KAIST Complex Urban Dataset — Converted ROS Bags This repository mirrors ROS bag conversions of the KAIST Complex Urban Dataset, covering sequences urban18 through urban39. Each sequence includes: urbanXX.bag — ROS bag with the full multi-sensor recording (camera, LiDAR, IMU, GPS, wheel encoders) urbanXX.csv — ground truth trajectory urbanXX.txt — ground truth converted to TUM/text format (via kaist_gt_csv2txt.m) Origin The underlying sensor data was collected… See the full description on the dataset page: https://huggingface.co/datasets/gladiator7737/kaist-urban-dataset.text0 likes844 downloads3mo agoHugging Face14kaizen9 /stem-corpustext10M<n<100M0 likes688 downloads1y agoHugging Face15kaicolabworkspace /friends-dialogtext10K<n<100K0 likes681 downloads3y agoHugging Face16KaiLv /UDR_CosmosQA Dataset Card for "UDR_CosmosQA" More Information needed tabular10K<n<100K0 likes678 downloads3y agoHugging Face17KaiNylund /WMT-month-splitstext100K<n<1M0 likes668 downloads3y agoHugging Face18lmarena-ai /PPE-MMLU-Pro-Best-of-K Overview This contains the MMLU-Pro correctness preference evaluation set for Preference Proxy Evaluations. The prompts are sampled from MMLU-Pro. This dataset is meant for benchmarking and evaluation, not for training. Paper Code License User prompts are licensed under MIT, and model outputs are governed by the terms of use set by the respective model providers. Citation @misc{frick2024evaluaterewardmodelsrlhf, title={How to Evaluate Reward… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/PPE-MMLU-Pro-Best-of-K.tabularn<1K0 likes656 downloads2y agoHugging Face19lmarena-ai /PPE-MATH-Best-of-K Overview This contains the MATH correctness preference evaluation set for Preference Proxy Evaluations. The prompts are sampled from MATH. This dataset is meant for benchmarking and evaluation, not for training. Paper Code License User prompts are licensed under MIT, and model outputs are governed by the terms of use set by the respective model providers. Citation @misc{frick2024evaluaterewardmodelsrlhf, title={How to Evaluate Reward Models for… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/PPE-MATH-Best-of-K.textn<1K0 likes654 downloads2y agoHugging Face20kaizen9 /mixturetext10M<n<100M0 likes632 downloads1y agoHugging Face21lmarena-ai /PPE-MBPP-Plus-Best-of-K Overview This contains the MBPP-Plus correctness preference evaluation set for Preference Proxy Evaluations. The prompts are sampled from MBPP-Plus. This dataset is meant for benchmarking and evaluation, not for training. Paper Code License User prompts are licensed under Apache-2.0, and model outputs are governed by the terms of use set by the respective model providers. Citation @misc{frick2024evaluaterewardmodelsrlhf, title={How to Evaluate… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/PPE-MBPP-Plus-Best-of-K.tabularn<1K1 likes596 downloads2y agoHugging Face22lmarena-ai /PPE-IFEval-Best-of-K Overview This contains the IFEval correctness preference evaluation set for Preference Proxy Evaluations. The prompts are sampled from IFEval. This dataset is meant for benchmarking and evaluation, not for training. Paper Code License User prompts are licensed under Apache-2.0, and model outputs are governed by the terms of use set by the respective model providers. Citation @misc{frick2024evaluaterewardmodelsrlhf, title={How to Evaluate Reward… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/PPE-IFEval-Best-of-K.tabularn<1K0 likes593 downloads2y agoHugging Face23kaiyuyue /llava-1.5-665k-instructionsThis dataset repository, LLaVA-1.5-665K-Instructions, is notably utilized in the paper Zero-Shot Vision Encoder Grafting via LLM Surrogates. The official code repository for the paper can be found here: https://github.com/kaiyuyue/zero LLaVA-1.5-665K-Instructions This dataset repo contains the entire LLaVA-1.5-665K-Instructions in one place, including images and text sequences. The images are in train_split/*.tars and the text sequences are in jsons: llava_v1_5_mix665k.json is the… See the full description on the dataset page: https://huggingface.co/datasets/kaiyuyue/llava-1.5-665k-instructions.imagevisual-question-answering100K<n<1M10 likes538 downloads1y agoHugging Face24kaitooooo /human_assisted_action_preference_optimizationtextn<1K0 likes515 downloads1y agoHugging Face25developer-lunark /kaidol-character-dataset KAIdol Character Chat Dataset 한국어 캐릭터 롤플레이 대화 데이터셋 📋 목차 개요 데이터셋 통계 데이터 형식 캐릭터 목록 품질 지표 사용 방법 학습 가이드 제한사항 라이선스 🎯 개요 KAIdol Character Chat Dataset은 41개 고유 캐릭터의 롤플레이 대화 데이터셋입니다. 각 캐릭터는 독특한 **음성 프로필(Voice Profile)**을 가지고 있으며, 이를 기반으로 일관된 성격과 말투를 유지합니다. 주요 특징 특징 설명 🎭 41개 캐릭터 다양한 성격, 배경, 말투를 가진 캐릭터 🗣️ 음성 프로필 시그니처 표현, 종결어미, 금지 표현 정의 📊 3가지 형식 SFT, DPO, Multiturn 학습 지원 ✅ 품질 검증 A등급 음성 프로필 일치율 (0.805) 🇰🇷 100% 한국어 자연스러운… See the full description on the dataset page: https://huggingface.co/datasets/developer-lunark/kaidol-character-dataset.texttext-generation1K<n<10K0 likes497 downloads8mo agoHugging Face26Kaichengalex /WebPerson-5M 🔥 Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval [EMNLP25 Main] Tianlu Zheng*, Yifan Zhang*, Xiang An, Ziyong Feng, Kaicheng Yang†, Qichunan Ding†, 📄 Paper | 💻 Github ✨ Web-Person Dataset 🔍 Person-Centric Image Filtering We use the COYO700M dataset as our source of web-crawled images. To curate high-quality person-centric images, we apply YOLOv11 to detect humans and extract bounding… See the full description on the dataset page: https://huggingface.co/datasets/Kaichengalex/WebPerson-5M.image1M<n<10M3 likes471 downloads1y agoHugging Face27kairunwen /scannet_temp scannet_temp This repository contains split tar.gz archives for the full local scannet dataset tree. Included scans/* scenes with unpacked usable content splits/* split files Excluded scene-level original *_2d-*.zip files results/ color_90/ pose_90.txt selected_ids_90.txt Download huggingface-cli download kairunwen/scannet_temp --repo-type dataset --local-dir ./scannet_temp_hf Extract mkdir -p extracted for f in… See the full description on the dataset page: https://huggingface.co/datasets/kairunwen/scannet_temp.text0 likes464 downloads4mo agoHugging Face28KaiKCZhang /TideEater YourPrjkt - 项目标题 一个融合了 Motion Matching 的简易动作游戏 DEMO,由 Unreal Engine 5.4 开发。项目的名字是随便取的,游戏名字暂定为《潮蚀》(Tide Eater)。 项目已于 2025/08/07 将引擎版本升级至 5.5.4,因此下文中关于引擎版本的描述可能不完全正确。 声明 本项目仅用于学习和研究目的,不涉及任何商业用途,所有资源和代码均为学习和研究之用。 本项目不提供任何形式的付费支持或保证,也不对因使用本项目而导致的任何损失或损害承担责任。 本项目内所有音乐资产和内容均归原作者所有。如有侵权,请联系我进行删除。 本项目的动作资产 Cool Sword Combat Animations V3 与场景资产 Pirate Stilt Village Island Modular Pack 并非免费资产,建议自行购买以规避可能的版权问题。 关于项目 YourPrjkt 是一个旨在学习和演示如何在 Unreal Engine 中使用… See the full description on the dataset page: https://huggingface.co/datasets/KaiKCZhang/TideEater.imagen<1K0 likes451 downloads2mo agoHugging Face29Lxd99 /PCMind-2.1-Kaiyuan-2B-phase1-part1-2 This repository contains the complete pretraining dataset for PCMind-v2.1-Kaiyuan-2B, a leading fully open-source language model. Overview The dataset is organized into 5 training phases, with all phase datasets open-sourced in this repository. Our training methodology employs domain-specific mixing strategies across five primary domains: English: General English text Chinese: General Chinese text Code: Programming and code-related content Math: Mathematical reasoning and… See the full description on the dataset page: https://huggingface.co/datasets/Lxd99/PCMind-2.1-Kaiyuan-2B-phase1-part1-2.texttext-generation100M<n<1B0 likes449 downloads6mo agoHugging Face30Lxd99 /PCMind-2.1-Kaiyuan-2B-phase1-part1-1-0323 This repository contains the complete pretraining dataset for PCMind-v2.1-Kaiyuan-2B, a leading fully open-source language model. Overview The dataset is organized into 5 training phases, with all phase datasets open-sourced in this repository. Our training methodology employs domain-specific mixing strategies across five primary domains: English: General English text Chinese: General Chinese text Code: Programming and code-related content Math: Mathematical reasoning and… See the full description on the dataset page: https://huggingface.co/datasets/Lxd99/PCMind-2.1-Kaiyuan-2B-phase1-part1-1-0323.texttext-generation100M<n<1B0 likes442 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.