datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
docTR-resource-collectionVinciCoder-1.6M-SFT
VinciCoder: Unified Multimodal Code Generation Dataset
This repository contains the datasets used for VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning, a project that introduces a unified multimodal code generation model. The framework uses a two-stage training approach, comprising a large-scale Supervised Finetuning (SFT) corpus and a Visual Reinforcement Learning (ViRL) dataset. These datasets are designed for tasks involving direct… See the full description on the dataset page: https://huggingface.co/datasets/DocTron-Hub/VinciCoder-1.6M-SFT.doctrine-corpus
doctrine-corpus — Judgment-Eliciting Q&A Corpus
A bilingual (English + Japanese) judgment-eliciting Q&A corpus encoding the documented judgment of four research lines in the shimo4228 research program — Agent Knowledge Cycle, Contemplative Agent, Agent Attribution Practice, and Authorship Strategy — plus published articles. The corpus is the operational form of Authorship Strategy Layer 4 tactic 7 (LLM-first ingest) and is released CC0 to maximize LLM-mediated diffusion.… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/doctrine-corpus.VinciCoder-42k-RL
VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning
This repository contains the datasets used and generated in the paper VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning.
The work introduces VinciCoder, a unified multimodal code generation model that addresses the limitations of single-task training paradigms. It proposes a two-stage training framework, beginning with a large-scale… See the full description on the dataset page: https://huggingface.co/datasets/DocTron-Hub/VinciCoder-42k-RL.legal_consolidationTask details
Legal consolidation is a critical yet time-consuming task, traditionally performed manually by legal professionals.
The objective is to automate the process of French legal consolidation, which is the application of modifications from a
modification section to an initial article to generate a modified article.
Dataset structure
A triplet of:
an initial article: the legislative article before consolidation,
a modification section: the text introducing the modification within… See the full description on the dataset page: https://huggingface.co/datasets/DoctrineAI/legal_consolidation.milton-de-doctrina
Milton, De Doctrina Christiana (Latin ↔ English)
24 chapters of Milton's De Doctrina Christiana (~89K words), Latin and English aligned; 23 chapters complete, 1 in progress (marked in its record).
Canonical home: https://milton.wrootpress.com (each record carries its canonical URL). This dataset is a machine-generated export of that site's build, regenerated from the source of truth and never hand-edited; the site remains the one published address of the text.… See the full description on the dataset page: https://huggingface.co/datasets/wrootpress/milton-de-doctrina.legal-doctrine-evolution-coherence-trajectory-v0.1What this dataset is
You receive
doctrine state at t
transition signals
doctrine state at t+1
split or exception signals
workability or legitimacy signals
reform pressure signals
You decide
Is the doctrine evolution stable
Answer
coherent
or
incoherent
Why this matters
Incoherent trajectories predict
overruling events
doctrinal collapse
rapid rule change
institutional instability
DoctrineMythosThis is a high-fidelity reasoning dataset synthesized using Claude Opus 4.6. The dataset is designed to capture the model's internal "Chain of Thought" and reasoning traces, specifically focusing on mathematical accuracy and structured logical deduction.
The dataset is intended for Supervised Fine-Tuning (SFT) and Distillation, allowing smaller open-source models to inherit the sophisticated reasoning patterns of Claude Opus 4.6.
Dataset Description
This collection combines high-difficulty… See the full description on the dataset page: https://huggingface.co/datasets/Whoisjutanlee/DoctrineMythos.fintetree-v3-factnum-doctrnato-doctrine-sftlegal-judicial-doctrinal-consistency-drift-v0.1What this dataset is
You receive
prior opinion pattern
current opinion pattern
citation shift
test or factor weighting shift
explanation quality
external reaction signals
You decide
Does the judge remain doctrinally consistent
Answer
coherent
or
incoherent
Why this matters
Doctrinal drift predicts
dissents and fractures
en banc pressure
reversal risk
loss of precedential durability
Donnees_internes_doctrine_10
Dataset Card for "Donnees_internes_doctrine_10"
More Information needed
legal_document_structuringTask details
Document structuring plays a crucial role in various natural language processing (NLP) tasks, such as information retrieval, and document understanding.
It also helps readers to effectively navigate into a structured document with a large amount of textual data.
In the legal domain, document structuring is particularly important for creating inter- and intra-document links.
The dataset provides documents segmented into lines.
Each document was collected in HTML format or PDF… See the full description on the dataset page: https://huggingface.co/datasets/DoctrineAI/legal_document_structuring.doc_train
Dataset Card for "doc_train"
More Information needed
Donnees_internes_doctrine_RT_4
Dataset Card for "Donnees_internes_doctrine_RT_4"
More Information needed
Donnees_internes_doctrine_14
Dataset Card for "Donnees_internes_doctrine_14"
More Information needed
Donnees_internes_doctrine_26
Dataset Card for "Donnees_internes_doctrine_26"
More Information needed
Donnees_internes_doctrine_41
Dataset Card for "Donnees_internes_doctrine_41"
More Information needed
Donnees_internes_doctrine_RT_7
Dataset Card for "Donnees_internes_doctrine_RT_7"
More Information needed
Donnees_internes_doctrine_RT_11
Dataset Card for "Donnees_internes_doctrine_RT_11"
More Information needed
Donnees_internes_doctrine_RT_13
Dataset Card for "Donnees_internes_doctrine_RT_13"
More Information needed
Donnees_internes_doctrine_8
Dataset Card for "Donnees_internes_doctrine_8"
More Information needed
Donnees_internes_doctrine_22
Dataset Card for "Donnees_internes_doctrine_22"
More Information needed
Donnees_internes_doctrine_52
Dataset Card for "Donnees_internes_doctrine_52"
More Information needed
Donnees_internes_doctrine_63
Dataset Card for "Donnees_internes_doctrine_63"
More Information needed
Donnees_internes_doctrine_77
Dataset Card for "Donnees_internes_doctrine_77"
More Information needed
Donnees_internes_doctrine_RT_1
Dataset Card for "Donnees_internes_doctrine_RT_1"
More Information needed
Donnees_internes_doctrine_RT_3
Dataset Card for "Donnees_internes_doctrine_RT_3"
More Information needed
Donnees_internes_doctrine_RT_8
Dataset Card for "Donnees_internes_doctrine_RT_8"
More Information needed
Donnees_internes_doctrine_RT_9
Dataset Card for "Donnees_internes_doctrine_RT_9"
More Information needed
