datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sharegpt-quizz-generation-json-output
ShareGPT-Formatted Dataset for Quizz Generation in Structured JSON Output
Dataset Description
This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate quizz in structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios.
Usage
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-quizz-generation-json-output.sharegpt-structured-output-json
ShareGPT-Formatted Dataset for Structured JSON Output
Dataset Description
This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios.
Usage
This dataset can be used to train LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-structured-output-json.lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1645559101lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646052073
GEM Submission
Submission name: Hugging Face test T5-base.outputs.json 36bf2a59
lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646049601lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1645800191lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646049378lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646049876lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1645558682lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646049424lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646050898lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646051364json-structured-output-dpo-3k
JSON Structured Output DPO Pairs (3K)
DPO preference pairs for training LLMs to produce valid, schema-compliant JSON output.
Motivation
Structured output (JSON mode) is critical for production AI applications — parsers fail, pipelines break, and downstream processing errors when models output malformed JSON, use wrong field names, or wrap responses in markdown. This dataset trains strict schema adherence.
Dataset Description
3,000 preference pairs… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/json-structured-output-dpo-3k.synqa_hudson_300_queries_rubrics_score_completeness_gpt-4-0613_outputs_json_Falsesynqa_hudson_300_queries_rubrics_score_completeness_gpt-4-0613_outputs_json_Truequizz_generation_json_outputtraining_criteria_dpo_distill_relevance_gpt-4-0613_outputs_json_True_debugE-Commerce_Customer_Support_Conversations_JSON_Outputcomplex-json-outputsprime-rl-complex_json_outputtraining_criteria_dpo_distill_completeness_2stage_gpt-4-0613_outputs_json_True_debugsynqa_hudson_300_samples_relevance_gpt-4-0613_outputs_json_True_debugjson-output-bart-beersynqa_hudson_300_samples_completeness_gpt-4-0613_outputs_json_True_debugsynqa_hudson_300_samples_clarity_gpt-4-0613_outputs_json_True_debugcxe_chart_dataset_1000_final_instruction_input_output.jsoncxe_chart_dataset_1000_final_instruction_input_output1.jsoncxe_chart_dataset_corrected_outputs.jsonopenai-outputs3-jsoninput_output_master_list.json
