structured-json
sharegpt-structured-output-json
ShareGPT-Formatted Dataset for Structured JSON Output
Dataset Description
This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios.
Usage
This dataset can be used to train LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-structured-output-json.pseudo-camera-10k-structured-json
pseudo-camera-10k, structured JSON captions
The 9,997 training images from bghira/pseudo-camera-10k, recaptioned into the structured JSON caption schema that Ideogram 4 consumes. The images are unchanged: free photographs from world class photographers, Lanczos-resized so the shorter edge is 1024px, nothing upsampled.
The original dataset carries short CogVLM prose captions. This one replaces them with one JSON object per image describing the scene at three levels: an overall… See the full description on the dataset page: https://huggingface.co/datasets/terminusresearch/pseudo-camera-10k-structured-json.JSON-Unstructured-StructuredDataset Contains Synthetically Generated Unstructured Text, Set of Rules for Schema Creation, Filled Structured JSON
Can be used for any unstructured to structured tasks
json-structured-output-dpo-3k
JSON Structured Output DPO Pairs (3K)
DPO preference pairs for training LLMs to produce valid, schema-compliant JSON output.
Motivation
Structured output (JSON mode) is critical for production AI applications — parsers fail, pipelines break, and downstream processing errors when models output malformed JSON, use wrong field names, or wrap responses in markdown. This dataset trains strict schema adherence.
Dataset Description
3,000 preference pairs… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/json-structured-output-dpo-3k.adaption-clinical-notes-to-structured-json
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-clinical_notes_to_structured_json
This dataset contains pairs of unstructured clinical intake notes and their corresponding structured JSON representations. Each sample transforms raw patient descriptions, including demographics, symptoms, vitals, and treatments, into a standardized schema with specific fields for analysis. The content covers diverse medical scenarios ranging from minor… See the full description on the dataset page: https://huggingface.co/datasets/T3ns0rT1nk3r/adaption-clinical-notes-to-structured-json.JSON-STRUCTURED-DATA-FOR-SYMPTOMS-SFT_DPO-SUPPORTEDThis dataset is high-consistency instruction tuning dataset.
converting messy, subjective human health-style text → structured, non-diagnostic extraction format
1.Literal extraction discipline
2.Source separation logic - very strong schema grounding training if DPO
3.Anti-inference constraint
1.High ambiguity coverage
2.Contradiction handling included
3.Minimization bias detection
This dataset is a STRICT schema regulation.
Does well at:
strict extraction
preserving uncertainty words… See the full description on the dataset page: https://huggingface.co/datasets/sadnjasdkn/JSON-STRUCTURED-DATA-FOR-SYMPTOMS-SFT_DPO-SUPPORTED.
