datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
m2mcent-mcp-schemas
🌐 M2MCent Agentic Services - MCP Schemas Dataset
🚀 Empowering Autonomous AI on Base L2
This dataset contains the JSON schemas for 1,005 microservices natively available on the M2MCent Network via the x402 V2 Protocol (EIP-3009).
It is specifically designed for instruction-tuning LLMs (like Llama-3, Mistral, Qwen) so they can autonomously discover, negotiate, and consume monetized API endpoints using gasless cryptocurrency settlements on the Base L2 network.… See the full description on the dataset page: https://huggingface.co/datasets/evozim/m2mcent-mcp-schemas.SCHEMA
Evidence-Grounded Biomedical Question Benchmark — Protein Domain / PathVQA-Enhanced
This dataset bundles two stratified, balanced subsets of an evidence-grounded
biomedical benchmark generated by lifting normalized knowledge-graph triples
(extracted by an upstream pubmed_graph literature pipeline) into five
complementary question formats. Each question is bound to a verbatim
supporting sentence from a source paper, gated by an evidence-strength
profile, and validated by a… See the full description on the dataset page: https://huggingface.co/datasets/XsF2001/SCHEMA.hedgehog-schema-replay
hedgehog-schema-replay
Hedgehog — schema replay round (anti-forgetting).
Contents
train.jsonl (1360 rows)
validation.jsonl (160 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Original content for the Hedgehog extraction model (Michael Anthony Falabella).
hedgehog-schema-explicit
hedgehog-schema-explicit
Hedgehog — explicit-schema extraction training.
Contents
train.jsonl (1280 rows)
validation.jsonl (160 rows)
test.jsonl (192 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Original content for the Hedgehog extraction model (Michael Anthony Falabella).
xrefrag-ukfin-schema
XRefRAG-ADGM (DPEL)
Cross-reference–grounded, citation-dependent syntetic QA benchmark for evaluating retrieval and RAG on regulatory text. Each item is built from a source passage that contains a cross-reference and a target passage that provides the referenced requirement/definition; answering correctly is intended to require using both.
Project repo and full pipeline documentation: https://github.com/RegNLP/XRefRag
Data
Splits: train / dev / test
Format: JSONL… See the full description on the dataset page: https://huggingface.co/datasets/RegNLP/xrefrag-ukfin-schema.xrefrag-adgm-schema
XRefRAG-ADGM (DPEL)
Cross-reference–grounded, citation-dependent syntetic QA benchmark for evaluating retrieval and RAG on regulatory text. Each item is built from a source passage that contains a cross-reference and a target passage that provides the referenced requirement/definition; answering correctly is intended to require using both.
Project repo and full pipeline documentation: https://github.com/RegNLP/XRefRag
Data
Splits: train / dev / test
Format: JSONL… See the full description on the dataset page: https://huggingface.co/datasets/RegNLP/xrefrag-adgm-schema.type-schema-tools-calls
type-schema-tools-calls
Dataset for TypeSchema Tool Calling
