doctr
Datasets
All datasets matching “doctr”docTR-resource-collectionVinciCoder-1.6M-SFT
VinciCoder: Unified Multimodal Code Generation Dataset
This repository contains the datasets used for VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning, a project that introduces a unified multimodal code generation model. The framework uses a two-stage training approach, comprising a large-scale Supervised Finetuning (SFT) corpus and a Visual Reinforcement Learning (ViRL) dataset. These datasets are designed for tasks involving direct… See the full description on the dataset page: https://huggingface.co/datasets/DocTron-Hub/VinciCoder-1.6M-SFT.doctrine-v10-v11
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
SZLHOLDINGS/doctrine-v10-v11
The locked governance doctrine for the SZL Holdings agentic substrate: Doctrine v10,
Doctrine v11 (13-axis yuyay_v3 canonical, supersedes v10), the Mythos → Hatun-Willay
rename rule, and the canonical wedge.
Contents
File
What… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/doctrine-v10-v11.szl-1-doctrine-sft
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
SZL-1 Doctrine SFT
The supervised fine-tuning (SFT) set that teaches SZL-1 its identity and
SZL Holdings' honesty doctrine. This is the exact training data used by the
szl-forge kit (train_szl.py,
Unsloth QLoRA on unsloth/Qwen2.5-3B-Instruct-bnb-4bit).
What's in it — honest labels… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/szl-1-doctrine-sft.doctrine-corpus
doctrine-corpus — Judgment-Eliciting Q&A Corpus
A bilingual (English + Japanese) judgment-eliciting Q&A corpus encoding the documented judgment of four research lines in the shimo4228 research program — Agent Knowledge Cycle, Contemplative Agent, Agent Attribution Practice, and Authorship Strategy — plus published articles. The corpus is the operational form of Authorship Strategy Layer 4 tactic 7 (LLM-first ingest) and is released CC0 to maximize LLM-mediated diffusion.… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/doctrine-corpus.VinciCoder-42k-RL
VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning
This repository contains the datasets used and generated in the paper VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning.
The work introduces VinciCoder, a unified multimodal code generation model that addresses the limitations of single-task training paradigms. It proposes a two-stage training framework, beginning with a large-scale… See the full description on the dataset page: https://huggingface.co/datasets/DocTron-Hub/VinciCoder-42k-RL.
