datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llm-coding-benchCodingAgentWorldBench
CodingAgentWorldBench -- data
Built instances of benchmark-1 (A1 static scene reconstruction): benchmark-1/instances/<id>/{public,private}
(see the code repo's benchmarks/benchmark-1/LAYOUT.md). public/ is what an agent sees; private/ is the hidden
ground truth -- keep this dataset private if the benchmark is used for evaluation.
Tracks: sim_* (simulated), real_ycbv_* (BOP YCB-Video photos + CAD reference), real_replica_* (Replica scan
regions, rendered views), real_multiscan_*… See the full description on the dataset page: https://huggingface.co/datasets/Linz99/CodingAgentWorldBench.coding_ds_2Amazon-Reviews-DatasetThis dataset provides a free trial sample of best-selling products and their customer reviews from a leading e-commerce platform, designed to support product intelligence, sentiment analysis, and market trend evaluation. This sample is provided for evaluation purposes only. It includes a curated subset of the full dataset.
To access the complete dataset, request additional attributes, or explore alternative product segments, please contact the data provider directly.
Key Features
2… See the full description on the dataset page: https://huggingface.co/datasets/coding-guru/Amazon-Reviews-Dataset.Synthetic-Datasets-Unity-CVchhaya-skin-extract
Chhaya Skin-Extract
Fine-tuning data for Chhaya — a skin & heat-health companion for outdoor
workers. Each example is image + "skin check" → findings JSON, teaching
MedGemma-1.5-4B to emit Chhaya's structured schema directly (no chain-of-thought
preamble) with a concern level grounded in real clinical labels.
Why two sources
ISIC-2024
SCIN
Image type
Curated dermatologic close-ups
Real consumer phone photos
concern ground truth
Biopsy diagnosis… See the full description on the dataset page: https://huggingface.co/datasets/CodingBad02/chhaya-skin-extract.CodingArtcoding-model-rendered-qa
Rendered QA Dataset: Code & Text (700K)
Instruction-tuning dataset with optional rendered images for vision-language models.
Sources
Source
Samples
Has Context Image
OpenCoder Stage 2
436K
educational_instruct only
InstructCoder
108K
Yes (code input)
OpenOrca
200K
No (text-only)
Schema
Column
Type
Description
prompt
string
Instruction/question
prompt_image
Image?
Rendered prompt (optional)
context
string?
Code context… See the full description on the dataset page: https://huggingface.co/datasets/mustavinsu/coding-model-rendered-qa.CodingArt-SDXLSpiceScan-datasettranscription-coding-wiki-500k
Transcription Dataset: Code & Wiki (390K)
Text-to-image rendered dataset for training vision-language models to read code and text from images.
Schema
Column
Type
Description
image
Image
Rendered grayscale JPEG
prompt
string
Transcription instruction (varied)
response
string
Ground truth text
language
string
python/javascript/java/c++/rust/go/english
domain
string
code or english
length_bucket
string
short/medium/long/gundam
resolution
string… See the full description on the dataset page: https://huggingface.co/datasets/curiousmrk/transcription-coding-wiki-500k.
