datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DaTikZ-V4
Dataset Card for DaTikZ-V4
DaTikZ-V4 is the dataset used to train TikZilla-3B, TikZilla-3B-RL, TikZilla-8B, and TikZilla-8B-RL for generating TikZ/LaTeX figures from natural language descriptions.
The TikZ code has been sourced from ArXiv, GitHub, and TeXStackExchange. Scientific figure descriptions were generated using Qwen2.5-VL-7B-Instruct.
Dataset fields
Each sample contains:
file_id: unique identifier
caption: original caption
vlm_description: detailed visual… See the full description on the dataset page: https://huggingface.co/datasets/nllg/DaTikZ-V4.DaTikZ-V4
Dataset Card for DaTikZ-V4
DaTikZ-V4 is the dataset used to train TikZilla-3B, TikZilla-3B-RL, TikZilla-8B, and TikZilla-8B-RL for generating TikZ/LaTeX figures from natural language descriptions.
The TikZ code has been sourced from ArXiv, GitHub, and TeXStackExchange. Scientific figure descriptions were generated using Qwen2.5-VL-7B-Instruct.
Dataset fields
Each sample contains:
file_id: unique identifier
caption: original caption
vlm_description: detailed visual… See the full description on the dataset page: https://huggingface.co/datasets/YangruiWANG/DaTikZ-V4.datikz_test_paired_desc
datikz_test_paired_desc — the DaTikZ v2/v3 official test split, with paired descriptions and reference images
The 984-item held-out test split used to evaluate text → TikZ models, with two descriptions of every
diagram so that description style can be isolated from model quality, plus the reference image for
image-space metrics.
This is deliberately a separate repo from the training corpus
(explcre/diagram_desc_datikz), which
contains no test items. Train and test cannot be… See the full description on the dataset page: https://huggingface.co/datasets/explcre/datikz_test_paired_desc.
