datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-schematics
Open Schematics Dataset
The largest dataset of electronic schematics and PCB layouts on the internet, built as an engineering reference for schematic and PCB layout work. It's a self-growing, autonomous dataset that continuously scans the web for new engineering designs and updates itself accordingly.
Dataset Description
Each record corresponds to one schematic file and includes the raw source, rendered images, structured metadata, and all associated PCB files… See the full description on the dataset page: https://huggingface.co/datasets/bshada/open-schematics.arc-agi-3-schema-traces
ARC-AGI-3 Schema Gameplay Trajectories
This release contains 50 ARC-AGI-3 gameplay trajectories and a dependency-free
scoring utility. The trajectories are split evenly across two collections:
gpt_5_6_sol/: 25 GPT-5.6 Sol trajectories.
claude_fable_opus/: 25 trajectories from Claude Opus 4.8 and Claude Fable 5.
Each trajectory directory includes run.json, a streamed events.jsonl event
log, sanitized session data, snapshots, and the shareable text/image files
produced during… See the full description on the dataset page: https://huggingface.co/datasets/schema-harness/arc-agi-3-schema-traces.spider-schema
Dataset Card for Spider Schema
Dataset Summary
Spider is a large-scale complex and cross-domain semantic parsing and text-to-SQL dataset annotated by 11 Yale students
The goal of the Spider challenge is to develop natural language interfaces to cross-domain databases.
This dataset contains the 166 databases used in the Spider dataset.
Yale Lily Spider Leaderboards
The leaderboard can be seen at https://yale-lily.github.io/spider
Languages
The text in… See the full description on the dataset page: https://huggingface.co/datasets/richardr1126/spider-schema.Minecraft-Schematics
Minecraft Schematics Dataset
This dataset packages a curated collection of Minecraft structures in multiple formats, along with mapping files and a ready-to-use fine-tuning dataset formatted for OpenAI chat models (e.g. for training Gemma-4 / Llama-3 models).
Dataset Directory Structure
schematics/: The raw Minecraft schematic files (supporting .schem, .litematic, and legacy .schematic formats).
blueprints/: The schematic files converted into a clean, parsed 3D… See the full description on the dataset page: https://huggingface.co/datasets/Hack337/Minecraft-Schematics.open-schematics
Open Schematics Dataset
A comprehensive dataset of electronic schematics from hardware projects. This dataset is designed for training AI models on circuit design, component recognition, and hardware engineering tasks.
Dataset Description
This dataset contains electronic schematic files along with their visual representations, component information, and metadata from various hardware projects.
Dataset Structure
Each record in the dataset contains:
schematic:… See the full description on the dataset page: https://huggingface.co/datasets/Colt45en/open-schematics.open-schematics
Open Schematics Dataset
A comprehensive dataset of electronic schematics from hardware projects. This dataset is designed for training AI models on circuit design, component recognition, and hardware engineering tasks.
Dataset Description
This dataset contains electronic schematic files along with their visual representations, component information, and metadata from various hardware projects.
Dataset Structure
Each record in the dataset contains:
schematic:… See the full description on the dataset page: https://huggingface.co/datasets/Joseferrera24/open-schematics.open-schematics
Open Schematics Dataset
A comprehensive dataset of electronic schematics from hardware projects. This dataset is designed for training AI models on circuit design, component recognition, and hardware engineering tasks.
Dataset Description
This dataset contains electronic schematic files along with their visual representations, component information, and metadata from various hardware projects.
Dataset Structure
Each record in the dataset contains:
schematic:… See the full description on the dataset page: https://huggingface.co/datasets/Ju-C/open-schematics.schema-t-runsopen-schematics
Open Schematics Dataset
A comprehensive dataset of electronic schematics from hardware projects. This dataset is designed for training AI models on circuit design, component recognition, and hardware engineering tasks.
Dataset Description
This dataset contains electronic schematic files along with their visual representations, component information, and metadata from various hardware projects.
Dataset Structure
Each record in the dataset contains:
schematic:… See the full description on the dataset page: https://huggingface.co/datasets/rifxyz/open-schematics.schema_guided_dstc8The Schema-Guided Dialogue dataset (SGD) was developed for the Dialogue State Tracking task of the Eights Dialogue Systems Technology Challenge (dstc8).
The SGD dataset consists of over 18k annotated multi-domain, task-oriented conversations between a human and a virtual assistant.
These conversations involve interactions with services and APIs spanning 17 domains, ranging from banks and events to media, calendar, travel, and weather.
For most of these domains, the SGD dataset contains multiple different APIs, many of which have overlapping functionalities but different interfaces,
which reflects common real-world scenarios.open-schematics
Open Schematics Dataset
A comprehensive dataset of electronic schematics from hardware projects. This dataset is designed for training AI models on circuit design, component recognition, and hardware engineering tasks.
Dataset Description
This dataset contains electronic schematic files along with their visual representations, component information, and metadata from various hardware projects.
Dataset Structure
Each record in the dataset contains:
schematic:… See the full description on the dataset page: https://huggingface.co/datasets/JasoHuangTaiwan/open-schematics.scugnizz-v22-tool-schema
scugnizz-v22-tool-schema
Synthetic schema-grounding data teaching exact tool selection and argument names.
Format: Hermes/OpenAI-style messages plus tools.
schema_guided_dialogThe Schema-Guided Dialogue (SGD) dataset contains 18K multi-domain task-oriented
dialogues between a human and a virtual assistant, which covers 17 domains
ranging from banks and events to media, calendar, travel, and weather. The
language presents in the datset is only English. The SGD dataset provides a
challenging testbed for a number of tasks in task-oriented dialogue, including
language understanding, slot filling, dialogue state tracking and response
generation. For the creation of the SGD dataset, they developed a multi-domain
dialogue simulator that generates dialogue outlines over an arbitrary combination
of APIs, dialogue states and system actions. Then, they used a crowd-sourcing
procedure to paraphrase these outlines to natural language utterances. This novel
crowd-sourcing procedure preserves all annotations obtained from the simulator and
does not require any extra annotations after dialogue collection.arc-agi-3-schema-traces
ARC-AGI-3 Schema Gameplay Trajectories
This release contains 50 ARC-AGI-3 gameplay trajectories and a dependency-free
scoring utility. The trajectories are split evenly across two collections:
gpt_5_6_sol/: 25 GPT-5.6 Sol trajectories.
claude_fable_opus/: 25 trajectories from Claude Opus 4.8 and Claude Fable 5.
Each trajectory directory includes run.json, a streamed events.jsonl event
log, sanitized session data, snapshots, and the shareable text/image files
produced during… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/arc-agi-3-schema-traces.schemas
SQaLe — schemas
Unique database schemas and their synthetic contents, one row per schema.
The questions live in
cwolff/queries, joined on schema id.
These two columns were previously stored inline on every question row of
cwolff/data_work_in_progress.
With ~25 questions per schema that was a ~25x duplication of the largest columns
in the corpus; holding them once here is the entire point of the split.
Columns
column
schema id
join key into… See the full description on the dataset page: https://huggingface.co/datasets/cwolff/schemas.open-schematics
Open Schematics Dataset
A comprehensive dataset of electronic schematics from hardware projects. This dataset is designed for training AI models on circuit design, component recognition, and hardware engineering tasks.
Dataset Description
This dataset contains electronic schematic files along with their visual representations, component information, and metadata from various hardware projects.
Dataset Structure
Each record in the dataset contains:
schematic:… See the full description on the dataset page: https://huggingface.co/datasets/GGGDDD1111/open-schematics.open-schematics
Open Schematics Dataset
A comprehensive dataset of electronic schematics from hardware projects. This dataset is designed for training AI models on circuit design, component recognition, and hardware engineering tasks.
Dataset Description
This dataset contains electronic schematic files along with their visual representations, component information, and metadata from various hardware projects.
Dataset Structure
Each record in the dataset contains:
schematic:… See the full description on the dataset page: https://huggingface.co/datasets/Strawberry015/open-schematics.open-schematics
Open Schematics Dataset
A comprehensive dataset of electronic schematics from hardware projects. This dataset is designed for training AI models on circuit design, component recognition, and hardware engineering tasks.
Dataset Description
This dataset contains electronic schematic files along with their visual representations, component information, and metadata from various hardware projects.
Dataset Structure
Each record in the dataset contains:
schematic:… See the full description on the dataset page: https://huggingface.co/datasets/jarvisemitra/open-schematics.schemapile
SchemaPile
Usage
from datasets import load_dataset
ds = load_dataset("trl-lab/schemapile", split="full")
print(f"Loaded dataset with {len(ds)} records.")
print(ds[0])
Description
SchemaPile is a collection of database schemas extracted from various sources, normalized for consistency and ease of use in machine learning workflows. Each record contains metadata (INFO), licensing information, permissiveness, and a list of tables with detailed column… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/schemapile.open-schematics
Open Schematics Dataset
A comprehensive dataset of electronic schematics from hardware projects. This dataset is designed for training AI models on circuit design, component recognition, and hardware engineering tasks.
Dataset Description
This dataset contains electronic schematic files along with their visual representations, component information, and metadata from various hardware projects.
Dataset Structure
Each record in the dataset contains:
schematic:… See the full description on the dataset page: https://huggingface.co/datasets/Bravo666126/open-schematics.json-schema
JSON Schema Dataset
This dataset consists of a collection of JSON Schema documents collected from GitHub by searching using the Sourcegraph API.
Step 1: Find a list of JSON Schema paths
The Sourcegraph code search API is used to find files with a .json extension and containing {\n "$schema": "https://json-schema.org/".
This is somewhat restrictive, but still manages to find a large number of schemas.
pipenv run python slurp.py --outfile repos.csv
Step 2:… See the full description on the dataset page: https://huggingface.co/datasets/dataunitylab/json-schema.human-verified-schematics-v1
human-verified-v1
Native KiCad (legacy + s-expression) and Eagle XML schematics with high confidence of human review and real-world use.
Summary
Field
Value
Export
human-verified-v1
Schematics
1000 (685 grandfathered + 315 independently verified)
dataset_sha256
9b2e60ff4d8bf10f2a262d4b4ccacfefd40384aa25ac95f24d5a804c8df08bc8
Campaign
campaign_176faac27567451a
Compiled
2026-08-15T16:21:48.294287+00:00
Layout
manifest.json… See the full description on the dataset page: https://huggingface.co/datasets/Ohmatic/human-verified-schematics-v1.bird-sql-train-with-schemaBird SQL https://bird-bench.github.io/ train dataset with schema
Preview
[
{
"db_id": "movie_platform",
"question": "Name movie titles released in year 1945. Sort the listing by the descending order of movie popularity.",
"evidence": "released in the year 1945 refers to movie_release_year = 1945;",
"SQL": "SELECT movie_title FROM movies WHERE movie_release_year = 1945 ORDER BY movie_popularity DESC LIMIT 1",
"schema": {
"table_names_original": [
"lists"… See the full description on the dataset page: https://huggingface.co/datasets/1sf/bird-sql-train-with-schema.mcp-tool-schema-drift-trajectories
Mcp Tool Schema Drift Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/mcp-tool-schema-drift-trajectories.danish-icl-schema-format-v3
danish-icl-json-v3
In-context-learning rows derived from jensjepsen/danish-json-grpo-v1. Each row
packs 1-5 worked examples into a single user turn, followed by a held-out
passage; the assistant turn is the answer for that passage. No instruction is
included, so both the schema and the output format have to be inferred from the
examples. Two axes vary per row and are held constant within a row: the schema
(134 field-sets) and the output format (10 renderers — JSON, key: value… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-icl-schema-format-v3.ypa-evidence-schema-demo
YPA Exit Path Schema Demo
Version: 0.5.0Status: synthetic, SFW demonstration — not a production corpus
This dataset-native demonstration keeps five exit facts separate: cancellation, account closure, data deletion, billing dispute and promotional-message controls. Version 0.5.0 also exposes the machine-readable summary fields money_exit_state, account_exit_state and data_exit_state. Every row is fictional, uses example.invalid and retains synthetic: true.
Discoverability does… See the full description on the dataset page: https://huggingface.co/datasets/yellowpagesadult/ypa-evidence-schema-demo.GEM__bart_base_schema_guided_dialog__1645547915flux-2-klein-9B-schematic-dataset
FLUX.2 Klein Schematic LoRA Dataset
This is the training dataset used for the FLUX.2 Klein Schematic LoRA project.
The LoRA repository is available here:
LoRA: https://huggingface.co/nomadoor/flux-2-klein-9B-schematic-lora
For details about the experiment and dataset construction, see the blog post:
Blog: https://comfyui.nomadoor.net/en/notes/flux2-klein-schematic-lora/
Structure
data/
depth_relative/
input/
target/
caption/
surface_normal/… See the full description on the dataset page: https://huggingface.co/datasets/nomadoor/flux-2-klein-9B-schematic-dataset.json-schema-storeThis contains a set of schemas obtained via the JSON Schema Store catalog.
HomeHWM_preliminary_questions_5000_schema_3_1
