datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
json-schema-instances-training-pool
JSON schema and instance training pool
Real JSON Schemas from the public collections named below, read at the pinned revisions given there,
each paired where possible with documents that satisfy it, laid out twice. Train on either layer or
on both.
pool.jsonl
Every source rewritten into one shape, 20004 rows, one JSON object per line, with these fields.
Field
What it holds
id
a row identifier unique within this file
prompt
the request a model would… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/json-schema-instances-training-pool.m2mcent-mcp-schemas
🌐 M2MCent Agentic Services - MCP Schemas Dataset
🚀 Empowering Autonomous AI on Base L2
This dataset contains the JSON schemas for 1,005 microservices natively available on the M2MCent Network via the x402 V2 Protocol (EIP-3009).
It is specifically designed for instruction-tuning LLMs (like Llama-3, Mistral, Qwen) so they can autonomously discover, negotiate, and consume monetized API endpoints using gasless cryptocurrency settlements on the Base L2 network.… See the full description on the dataset page: https://huggingface.co/datasets/evozim/m2mcent-mcp-schemas.SCHEMA
Evidence-Grounded Biomedical Question Benchmark — Protein Domain / PathVQA-Enhanced
This dataset bundles two stratified, balanced subsets of an evidence-grounded
biomedical benchmark generated by lifting normalized knowledge-graph triples
(extracted by an upstream pubmed_graph literature pipeline) into five
complementary question formats. Each question is bound to a verbatim
supporting sentence from a source paper, gated by an evidence-strength
profile, and validated by a… See the full description on the dataset page: https://huggingface.co/datasets/XsF2001/SCHEMA.SchemaStress
SchemaStress
SchemaStress is a controlled synthetic benchmark for structured output reliability under schema constraints.
Dataset Summary
SchemaStress evaluates model behavior on schema-bounded structured tasks:
prompt-to-structured generation
invalid-candidate repair
validity reasoning support via error tags and paths
The benchmark focuses on outputs that must be parseable, schema-valid, and semantically coherent.
Supported Configs
form_easy: reimbursement… See the full description on the dataset page: https://huggingface.co/datasets/sigdelakshey/SchemaStress.small-model-schema-gym
Small Model Schema Gym Dataset
Deterministically generated chat examples for first-pass JSON compliance on
the project Dream brief and Safety plan contracts.
Files
train.jsonl: 500 training examples.
validation.jsonl: 200 held-out examples.
manifest.json: counts, seed, provenance, and overlap check.
Each row contains:
{
"id": "stable example identifier",
"messages": [
{"role": "system", "content": "..."},
{"role": "user", "content": "..."}… See the full description on the dataset page: https://huggingface.co/datasets/KwabsHug/small-model-schema-gym.llm-rag-optimized-schema-templates
Schema.org JSON-LD Templates Optimized for LLM RAG Retrieval (2026)
Curated dataset of Schema.org JSON-LD templates designed, tested, and optimized for Retrieval-Augmented Generation (RAG) systems, SearchGPT, Gemini, and Claude search parsers.
Published by Pixel Office EU.
Purpose
Standard Schema.org markup is often too nested or dense for token-efficient LLM context window ingestion. These templates prioritize high-salience fields that crawlers prioritize when… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/llm-rag-optimized-schema-templates.type-schema-tools-calls
type-schema-tools-calls
Dataset for TypeSchema Tool Calling
