datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VINCIE-10M
Dataset Card for VINCIE-10M
VINCIE: Unlocking In-context Image Editing from Video
Leigang Qu, Feng Cheng, Ziyan Yang, Qi Zhao, Shanchuan Lin, Yichun Shi, Yicong Li, Wenjie Wang, Tat-Seng Chua, Lu Jiang
Dataset Construction Pipeline
Visual Transition Annotation. To describe visual transitions between frames, we use chain-of-thought (CoT) prompting to instruct a VLM to perform visual transition… See the full description on the dataset page: https://huggingface.co/datasets/leigangqu/VINCIE-10M.VinciCoder-1.6M-SFT
VinciCoder: Unified Multimodal Code Generation Dataset
This repository contains the datasets used for VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning, a project that introduces a unified multimodal code generation model. The framework uses a two-stage training approach, comprising a large-scale Supervised Finetuning (SFT) corpus and a Visual Reinforcement Learning (ViRL) dataset. These datasets are designed for tasks involving direct… See the full description on the dataset page: https://huggingface.co/datasets/DocTron-Hub/VinciCoder-1.6M-SFT.VinciCoder-42k-RL
VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning
This repository contains the datasets used and generated in the paper VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning.
The work introduces VinciCoder, a unified multimodal code generation model that addresses the limitations of single-task training paradigms. It proposes a two-stage training framework, beginning with a large-scale… See the full description on the dataset page: https://huggingface.co/datasets/DocTron-Hub/VinciCoder-42k-RL.vincicoder-rl-easyr1
VinciCoder RL EasyR1 Parquet
This dataset contains the reinforcement-learning data used for VinciCoder-style
multimodal code generation training. The files were converted from the original
VinciCoder RL parquet format into an EasyR1-compatible parquet format.
The main change is the image column format. In the original files, the images
column stores image data as a base64 string. In this version, images is stored
as a list of image objects with raw bytes, which matches the format… See the full description on the dataset page: https://huggingface.co/datasets/GaviZhou/vincicoder-rl-easyr1.ministral-3-benchmark-prompts
Ministral 3 MLX benchmark prompts
This tiny dataset contains the four fixed prompts used by the reproducible
smoke benchmark for the Ministral 3 MLX 4-bit model.
It is a benchmark fixture, not a training or fine-tuning dataset.
Schema
Each JSONL row contains:
id: stable case identifier;
language: prompt language;
prompt: exact input sent to the model;
expected_keywords: lowercase substrings used by the smoke check.
The benchmark uses greedy decoding and checks… See the full description on the dataset page: https://huggingface.co/datasets/vinci00/ministral-3-benchmark-prompts.test_case_triggervinci_edit_sfw
Dataset Card for "vinci_edit_sfw"
More Information needed
triggerrepairleonardo_da_vinci_mindframe_instruction_dataset
Leonardo da Vinci Mindframe Instruction Dataset
A high-quality, unique instruction-tuning dataset designed to instill the mindframe and cognitive style of Leonardo da Vinci.
Overview
This dataset trains models to think, observe, question, and respond in the distinctive manner of Leonardo da Vinci — the ultimate Renaissance polymath. It emphasizes:
Insatiable curiosity (Curiosità)
Empirical testing through observation, drawing, and experiment (Dimostrazione)… See the full description on the dataset page: https://huggingface.co/datasets/11-47/leonardo_da_vinci_mindframe_instruction_dataset.CD352-DesignOfRoadTunnelsIndustrySpec-sample
{}
A small sample of data to train a model off the industry standard: CD352 Design of Road Tunnels.
