datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code-conflict
Code Conflict Dataset
A dataset of 100 visual Python code conflict samples designed to evaluate Vision-Language Models (VLMs) under cross-modal conflicts (discrepancy between code screenshots and caption text).
Dataset Statistics
Total Rows: 100 samples
Language: English (english)
Categories: 5 distinct Python code conflict_types (20 samples per category):
operator_substitution (Rows 1–20): Swapping math or logic operators (e.g., + to -, == to !=, or to and).… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-vlm-conflict/code-conflict.CodeCAD-HF-Samples
CodeCAD-HF-Samples (CodeCAD: A Parametric CAD Dataset for Programmatic 3D Model Generation)
CodeCAD-HF-Samples is an optimized, lightweight Hugging Face Parquet release containing ~9,999 curated paired entries sub-sampled from the authoritative CodeCAD v2 corpus. The original upstream CodeCAD v2 repository spans over 75,000 human-engineered, parametric OpenSCAD scripts encompassing roughly 10 million lines of code. This compact release is specifically packaged to provide… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/CodeCAD-HF-Samples.CodeCAD-HF-v2
CodeCAD-HF-v2 (CodeCAD: A Parametric CAD Dataset for Programmatic 3D Model Generation)
CodeCAD-HF-v2 is a curated Hugging Face Parquet release containing ~59,999 paired entries sub-sampled from the original CodeCAD v2 repository. CodeCAD v2 originally encompasses over 75,000 human-written OpenSCAD models (~10 million lines of parametric code) curated for correctness, deduplication, and mechanical utility. This release translates a subset of 59,999 samples into standard Apache… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/CodeCAD-HF-v2.Nemotron-Personas-Korea
Nemotron-Personas-Korea
우리나라 실제 분포에 기반한 합성 페르소나를 위한 복합 AI 시스템
A compound AI approach to personas grounded in real-world distributions
데이터셋 개요 (Overview)
Nemotron-Personas-Korea는 대한민국의 실제 인구통계학적·지리적·성격 특성 분포를 기반으로 합성된 오픈소스 페르소나 데이터셋(CC BY 4.0)으로, 우리나라 인구의 다양성과 특성을 폭넓게 반영하도록 설계되었습니다. 이는 최초의 대규모 우리말 페르소나 데이터셋이며, 이름, 성별, 나이, 혼인 상태, 교육 수준, 직업, 거주 지역 등의 속성을 실제 대한민국 통계청(KOSIS), 대법원, 국민건강보험공단, 농촌경제연구원, NAVER Cloud 통계 자료를 기반으로 합성하였습니다.
Nemotron-Personas-Korea는… See the full description on the dataset page: https://huggingface.co/datasets/codecainecowboy/Nemotron-Personas-Korea.code-conflict
Code Conflict Dataset
A dataset of 100 visual Python code conflict samples designed to evaluate Vision-Language Models (VLMs) under cross-modal conflicts (discrepancy between code screenshots and caption text).
Dataset Statistics
Total Rows: 100 samples
Language: English (english)
Categories: 5 distinct Python code conflict_types (20 samples per category):
operator_substitution (Rows 1–20): Swapping math or logic operators (e.g., + to -, == to !=, or to and).… See the full description on the dataset page: https://huggingface.co/datasets/akanshjain37/code-conflict.
