datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stackexchange_titlebody_best_and_down_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.stackexchange_title_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.hermes-function-calling-v1-jsonl
Hermes Function-Calling V1
This dataset is the compilation of structured output and function calling data used in the Hermes 2 Pro series of models.
This repository contains a structured output dataset with function-calling conversations, json-mode, agentic json-mode and structured extraction samples, designed to train LLM models in performing function calls and returning structured output based on natural language instructions. The dataset features various conversational scenarios… See the full description on the dataset page: https://huggingface.co/datasets/minpeter/hermes-function-calling-v1-jsonl.stackexchange_titlebody_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.stackexchange_math_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.swedish-medical-exams-mcq-1006-json
Dataset Card for Swedish Medical Exam MCQs
Dataset Description
This dataset contains multiple-choice questions from Swedish medical exams.
Languages
The dataset is in Swedish (sv).
Dataset Structure
Each entry in the dataset contains the following fields:
question: The question
options: An array of possible answers
answer: The correct answer
language: The language of the question (always "sv" for Swedish)
country: The country of origin (always… See the full description on the dataset page: https://huggingface.co/datasets/sarafuyu/swedish-medical-exams-mcq-1006-json.bigbench_jsonlBIG-Bench but it doesn't require the hellish dependencies (tensorflow, pypi-bigbench, protobuf) of the official version.
dataset = load_dataset("tasksource/bigbench",'movie_recommendation')
Code to reproduce:
https://colab.research.google.com/drive/1MKdLdF7oqrSQCeavAcsEnPdI85kD0LzU?usp=sharing
Datasets are capped to 50k examples to keep things light.
I also removed the default split when train was available also to save space, as default=train+val.
@article{srivastava2022beyond… See the full description on the dataset page: https://huggingface.co/datasets/NJUDeepEngine/bigbench_jsonl.json-mode-reasoningjson-mode-agentic-reasoninginstruct_chat_50k.jsonlinstruct_chat_50k.jsonl which is composed of 30k Chinese sharegpt dataset and 20k alpaca-instruction-Chinese-dataset
qa-ml-dl-jsonl
💡 AI Q&A Dataset for ML, DL, RL, TensorFlow, PyTorch
This dataset is designed to support training and evaluation of AI systems on question generation, answering, and understanding in the domains of Machine Learning, Deep Learning, Reinforcement Learning, TensorFlow, and PyTorch. It contains a large number of categorized questions along with high-quality answers in two different levels of brevity.
📁 Dataset Files
1. questions.jsonl
Lines: 24,510… See the full description on the dataset page: https://huggingface.co/datasets/Koushim/qa-ml-dl-jsonl.reasoning-sft-JSON-structuring-and-correcting
JSON Structuring and Correcting (Reasoning SFT)
Combined dataset of 508K rows for training LLMs on structured output tasks with reasoning traces, sourced from two datasets:
Sources
tool_calling.parquet (488,461 rows)
Converted from vericava/sft-tool-calling-structured-output-v1. Multi-turn tool calling and structured output tasks including tool invocations, tool results, and final assistant responses. Includes English and Japanese content.… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-JSON-structuring-and-correcting.100k_Tdk_zurriyet_dna_v6.jsonl
🌟 DESTEK & TOPLULUK ÇAĞRISI (SUPPORT & LIKE):Açık kaynak ve ücretsiz olarak sunduğum bu devasa çalışmayı faydalı bulduysanız, projenin sürdürülebilirliğine ve açık kaynak ekosisteminin görünürlüğüne katkı sağlamak için lütfen sayfanın sağ üstündeki Like (❤️ Beğeni) butonuna basarak destek olmayı unutmayın!(If you find this open-source dataset valuable for your research or models, please consider leaving a ❤️ Like at the top-right to support future updates and maintenance).
🇹🇷… See the full description on the dataset page: https://huggingface.co/datasets/bysismo/100k_Tdk_zurriyet_dna_v6.jsonl.json-mode-verifiablewinogrande.jsonlfunsd-json
Dataset Card for FUNSD (JSON Format)
This dataset contains the preprocessed JSON format of the FUNSD dataset, designed for form understanding in noisy scanned documents. It includes the original structure with annotations in JSON format, as per the original FUNSD dataset. This dataset is intended for document understanding tasks such as OCR, layout parsing, and key-value extraction.
Dataset Details
Dataset Description
The FUNSD (Form Understanding… See the full description on the dataset page: https://huggingface.co/datasets/Tamilnilavan/funsd-json.Ru_Tax_Audit_Instruct_Demo_JSON
Ru-Tax-Audit-Instruct: FNS Inspections, Fines & Corporate Compliance Scenarios (JSON)
🇷🇺 Описание проекта (Russian Description)
Ru-Tax-Audit-Instruct — это высококачественный коммерческий датасет инструктивного типа (Instruct Dataset), разработанный для обучения больших языковых моделей (LLM) логике российского корпоративного, налогового и трудового права.
Массив данных ориентирован на создание умных ИИ-ассистентов, роботов-консультантов, систем AI-комплаенса… See the full description on the dataset page: https://huggingface.co/datasets/Rudatamind/Ru_Tax_Audit_Instruct_Demo_JSON.MMLU-Pro-json
MMLU-Pro json
This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details.
healthcare-chat-dataset-jsonl
Healthcare Chat Dataset (JSONL Format)
This dataset contains 41 healthcare-related conversational exchanges in ChatML format, designed for training conversational AI models for medical assistance and healthcare guidance.
Dataset Structure
The dataset is provided as a JSONL file where each line contains a JSON object with:
text: A complete conversation in ChatML format with system, user, and assistant messages
ChatML Format Structure
Each conversation follows… See the full description on the dataset page: https://huggingface.co/datasets/adrianf12/healthcare-chat-dataset-jsonl.cybersec-jsonschemabench-cloudtrail-v6
CybersecJSONSchemaBench CloudTrail Attack Reconstruction v6
A 100-problem long-context cybersecurity reasoning benchmark over real flAWS CloudTrail logs with synthetically injected MITRE ATT&CK attack chains.
Each task gives the model 600 real CloudTrail records (280-380K tokens of JSON) containing a single hidden multi-step attack chain. The model must produce a structured answer identifying the attacking principal, the MITRE ATT&CK technique, the per-phase records that constitute… See the full description on the dataset page: https://huggingface.co/datasets/achinta3/cybersec-jsonschemabench-cloudtrail-v6.30k_CA_synthetic_patients_FHIR_Bundles_JSON_NDJSON
Synthea 30K Synthetic Patient Dataset
Synthetic FHIR R4 patient data generated using Synthea, an open-source synthetic patient generator developed by MITRE Corporation.
Contents
The dataset includes 30,000 + (deceased) patient records across the following FHIR R4 resource types, provided as NDJSON files and a bundled ZIP archive:
AllergyIntolerance
Condition
DiagnosticReport
Encounter
ImagingStudy
Immunization
Medication
MedicationAdministration
MedicationRequest… See the full description on the dataset page: https://huggingface.co/datasets/VRKomari/30k_CA_synthetic_patients_FHIR_Bundles_JSON_NDJSON.healthcare-qa-dataset-jsonl
Healthcare Q&A Dataset (JSONL Format)
This dataset contains 51 healthcare-related question-answer pairs in JSONL format, designed for training conversational AI models in the medical domain.
Dataset Structure
The dataset is provided as a JSONL file where each line contains a JSON object with:
prompt: A healthcare-related question
completion: A detailed, informative answer
Sample Entry
{"prompt": "What are the symptoms of diabetes?", "completion": "Common… See the full description on the dataset page: https://huggingface.co/datasets/adrianf12/healthcare-qa-dataset-jsonl.healthcare-qa-dataset-jsonl
Healthcare Q&A Dataset (JSONL Format)
This dataset contains 51 healthcare-related question-answer pairs in JSONL format, designed for training conversational AI models in the medical domain.
Dataset Structure
The dataset is provided as a JSONL file where each line contains a JSON object with:
prompt: A healthcare-related question
completion: A detailed, informative answer
Sample Entry
{"prompt": "What are the symptoms of diabetes?", "completion": "Common… See the full description on the dataset page: https://huggingface.co/datasets/jmvalder/healthcare-qa-dataset-jsonl.reasoning-sft-interstellarninja-json-mode-reasoning-160K
json-mode-reasoning (converted)
Converted version of interstellarninja/json-mode-reasoning, filtered to 20,474 rows with valid <think> reasoning traces.
Format
Each row has three columns:
input — list of dicts [{"role": "system/user", "content": "..."}, ...] (conversation turns ending on the last user turn, includes system prompt with JSON schema)
response — assistant response string with <think> reasoning block followed by JSON output
source — fixed as… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-interstellarninja-json-mode-reasoning-160K.PIB_MINISTRY_OF_HOME_AFFAIRS_INDIA_CLEAN_JSON
Cleaned Indian Government (PIB) Press Releases for AI Agents
Overview
This dataset contains real-time, structured, and agent-ready JSON files extracted from the Press Information Bureau (PIB) of India.
Every record is automatically stripped of web clutter, tagged by its respective Ministry, and structured with standardized numerical metrics to prevent hallucinations in Large Language Models (LLMs).
Data Structure
Each entry is formatted as a… See the full description on the dataset page: https://huggingface.co/datasets/Theblackfrancolin9009/PIB_MINISTRY_OF_HOME_AFFAIRS_INDIA_CLEAN_JSON.funsd-json
Dataset Card for FUNSD (JSON Format)
This dataset contains the preprocessed JSON format of the FUNSD dataset, designed for form understanding in noisy scanned documents. It includes the original structure with annotations in JSON format, as per the original FUNSD dataset. This dataset is intended for document understanding tasks such as OCR, layout parsing, and key-value extraction.
Dataset Details
Dataset Description
The FUNSD (Form Understanding in… See the full description on the dataset page: https://huggingface.co/datasets/davidle7/funsd-json.transformed_JSON_databricks-dolly-15k.jsonl
Transformed Databricks-Dolly-15k Dataset
Summary
The Transformed Databricks-Dolly-15k dataset is a modification of the original open-source dataset created by Databricks employees, designed to facilitate instruction-following abilities in large language models (LLMs). This version has been specifically adapted to include responses in a JSON format, enhancing its utility for tasks requiring structured output.
Modifications
The primary transformation applied to… See the full description on the dataset page: https://huggingface.co/datasets/ramachetan22/transformed_JSON_databricks-dolly-15k.jsonl.qwen_qa_pairs_cli_training.jsonl
Data sources
Multiple datasets from Hugging Face related to natural language to CLI pairs were gathered.
Human reviewed synthetic data from Claude Opus4.6 and ChatGPT4.5 were added.
A handful of grounding rows related to the organisation "Spicy Lemonade" were added (see details below)
Data processing
As part of the processing, data was converted to the Alpaca format with instruction (natural language), input (typically blank) and output (the CLI command) columns.
The… See the full description on the dataset page: https://huggingface.co/datasets/spicy-lemonade/qwen_qa_pairs_cli_training.jsonl.swedish-medical-exams-mcq-1002-json
Dataset Card for Swedish Medical Exam MCQs
Dataset Description
This dataset contains multiple-choice questions from Swedish medical exams.
Languages
The dataset is in Swedish (sv).
Dataset Structure
Each entry in the dataset contains the following fields:
question: The question
options: An array of possible answers
answer: The correct answer
language: The language of the question (always "sv" for Swedish)
country: The country of origin (always… See the full description on the dataset page: https://huggingface.co/datasets/serhany/swedish-medical-exams-mcq-1002-json.ServiceNow_Search_NLP_2_JSON_LLM_Training_SampleThis dataset contains structured User → Bot conversations demonstrating how a natural language request can be translated into a structured ServiceNow incident search API call.
The full CJ Jones' synthetic dataset catalog is available at:
https://datadeveloper1.gumroad.com
Each record consists of a user requesting incident data from an IT service management system and a bot responding with a JSON query specification compatible with the ServiceNow Table API.
The dataset is designed for training… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/ServiceNow_Search_NLP_2_JSON_LLM_Training_Sample.
