datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SVG-Sophia
Reliable Reasoning in SVG-LLMs via Multi-Task Multi-Reward Reinforcement Learning
🧩 SVG-Sophia Dataset
The SVG-Sophia dataset is available at Hugging Face.
After downloading and extraction, the files are organized as follows:
File
Description
cot_img2svg_sft.jsonlCoT training data for the SFT stage — Image-to-SVG task
cot_text2svg_sft.jsonl
CoT training data for the SFT stage — Text-to-SVG task
cot_refinement_sft.jsonl
CoT training data… See the full description on the dataset page: https://huggingface.co/datasets/InternSVG/SVG-Sophia.shaped-svgs-autocaptioned-1675So, whilst working on my SVG model, I saw this dataset and thought, despite it's relatively small size, it would be much better and much more helpful if the images were captioned. Now, this is V0.1. V. 0.1. As in, it's not done yet. The images have (mostly) i'd say approximately 70-80% acurately captioned. As soon as my server is refilled, i'm going through it with a VQA and an actual vision model, i've a few in mind. Also special prompts to ensure EVERY image is accurate.
In the meantime… See the full description on the dataset page: https://huggingface.co/datasets/MrOvkill/shaped-svgs-autocaptioned-1675.svg-stack-qwen-relabel-sharegpt
SVG Stack Qwen Relabel ShareGPT
ShareGPT-format conversation dataset with train, validation, and test splits.
Each record contains a conversations field.
