CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01multilingual-vlm-conflict /code-conflict Code Conflict Dataset A dataset of 100 visual Python code conflict samples designed to evaluate Vision-Language Models (VLMs) under cross-modal conflicts (discrepancy between code screenshots and caption text). Dataset Statistics Total Rows: 100 samples Language: English (english) Categories: 5 distinct Python code conflict_types (20 samples per category): operator_substitution (Rows 1–20): Swapping math or logic operators (e.g., + to -, == to !=, or to and).… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-vlm-conflict/code-conflict.imagevisual-question-answering1K<n<10K0 likes171 downloads3mo agoHugging Face02Codec96 /cinematic-video-250h-sample Cinematic Video Dataset – Sample (1080p, 250 hours total) Welcome! This page hosts a sample subset of a larger cinematic video dataset designed for AI training and research. Sample Data This repository contains a small sample subset to help you evaluate the dataset quality. Licensing & Full Dataset Access The full 250-hour dataset is not publicly available but can be licensed under a non-exclusive, 1-year license for AI research. Pricing: $200/hour… See the full description on the dataset page: https://huggingface.co/datasets/Codec96/cinematic-video-250h-sample.imagen<1K1 likes33 downloads1y agoHugging Face03codecainecowboy /Nemotron-Personas-Korea Nemotron-Personas-Korea 우리나라 실제 분포에 기반한 합성 페르소나를 위한 복합 AI 시스템 A compound AI approach to personas grounded in real-world distributions 데이터셋 개요 (Overview) Nemotron-Personas-Korea는 대한민국의 실제 인구통계학적·지리적·성격 특성 분포를 기반으로 합성된 오픈소스 페르소나 데이터셋(CC BY 4.0)으로, 우리나라 인구의 다양성과 특성을 폭넓게 반영하도록 설계되었습니다. 이는 최초의 대규모 우리말 페르소나 데이터셋이며, 이름, 성별, 나이, 혼인 상태, 교육 수준, 직업, 거주 지역 등의 속성을 실제 대한민국 통계청(KOSIS), 대법원, 국민건강보험공단, 농촌경제연구원, NAVER Cloud 통계 자료를 기반으로 합성하였습니다. Nemotron-Personas-Korea는… See the full description on the dataset page: https://huggingface.co/datasets/codecainecowboy/Nemotron-Personas-Korea.imagetext-generation1M<n<10M0 likes33 downloads5mo agoHugging Face04codecrafters /github-avatarsimagen<1K0 likes24 downloads4y agoHugging Face05akanshjain37 /code-conflict Code Conflict Dataset A dataset of 100 visual Python code conflict samples designed to evaluate Vision-Language Models (VLMs) under cross-modal conflicts (discrepancy between code screenshots and caption text). Dataset Statistics Total Rows: 100 samples Language: English (english) Categories: 5 distinct Python code conflict_types (20 samples per category): operator_substitution (Rows 1–20): Swapping math or logic operators (e.g., + to -, == to !=, or to and).… See the full description on the dataset page: https://huggingface.co/datasets/akanshjain37/code-conflict.imagevisual-question-answeringn<1K0 likes17 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.