datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
m2mcent-mcp-schemas
🌐 M2MCent Agentic Services - MCP Schemas Dataset
🚀 Empowering Autonomous AI on Base L2
This dataset contains the JSON schemas for 1,005 microservices natively available on the M2MCent Network via the x402 V2 Protocol (EIP-3009).
It is specifically designed for instruction-tuning LLMs (like Llama-3, Mistral, Qwen) so they can autonomously discover, negotiate, and consume monetized API endpoints using gasless cryptocurrency settlements on the Base L2 network.… See the full description on the dataset page: https://huggingface.co/datasets/evozim/m2mcent-mcp-schemas.SchemaStress
SchemaStress
SchemaStress is a controlled synthetic benchmark for structured output reliability under schema constraints.
Dataset Summary
SchemaStress evaluates model behavior on schema-bounded structured tasks:
prompt-to-structured generation
invalid-candidate repair
validity reasoning support via error tags and paths
The benchmark focuses on outputs that must be parseable, schema-valid, and semantically coherent.
Supported Configs
form_easy: reimbursement… See the full description on the dataset page: https://huggingface.co/datasets/sigdelakshey/SchemaStress.pre-conceptual-schemas-alpacaPre-Conceptual Schemas Dataset in Alpaca Format
This repository contains the dataset used for fine-tuning the Phi-3.5-PCS-Finetuned model, designed for the interpretation of Pre-Conceptual Schemas (PCS).
The dataset is a result of the master's thesis project "Natural Language Processing in Pre-conceptual Schemas for Representing Knowledge by Using Large Language Models" from the Master's Degree in Systems and Computing Engineering at the University of Nariño.
Dataset Curators: Felipe Roa… See the full description on the dataset page: https://huggingface.co/datasets/Galerasnet/pre-conceptual-schemas-alpaca.schema-summarization_spider
Dataset Card for schema-summarization_spider
Dataset Description
Dataset Summary
This dataset has been built to train and benchmark models uppon the schema-summarization task. This task aims to generate the smallest schema needed to answer a NL question with the help of the original database schema.
This dataset has been build by crossing these two datasets :
xlangai/spider
richardr1126/spider-schema
With the first dataset we take the natural language… See the full description on the dataset page: https://huggingface.co/datasets/avinot/schema-summarization_spider.weaviate-gorilla-schemasschemas-embedded
