datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
modern-architecturestride-architecture-components-v1
STRIDE Architecture Threat Modeling Dataset (AWS & Azure)
📌 Overview
This dataset was created to enable automatic STRIDE threat modeling from cloud architecture diagrams (AWS and Azure).
The goal is to detect architectural components in diagrams and support automated threat identification based on data flows and trust boundaries.
Annotations were created using Label Studio in YOLO format.
Total images: 4190Total classes: 32
🎯 Purpose
Detect cloud… See the full description on the dataset page: https://huggingface.co/datasets/guillherms/stride-architecture-components-v1.Software-ArchitectureSoftware-Architecture
I am releasing a Large Dataset covering topics related to Software-Architecture.
This dataset consists of around 450,000 lines of data in jsonl.
I have included following topics:
Architectural Frameworks
Architectural Patterns for Reliability
Architectural Patterns for Scalability
Architectural Patterns
Architectural Quality Attributes
Architectural Testing
Architectural Views
Architectural Decision-Making
Advanced Research
Cloud-Based Architectures
Component-Based… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Software-Architecture.pytorch-nn-architectures-dataset
PyTorch Neural Network Architectures Dataset
608 PyTorch neural network implementations generated using GPT-5,
covering 7 architecture types, 4 task categories, 4 input data types,
and 4 complexity levels. All architectures are validated and ready to use.
What is in the dataset?
Each file is a standalone PyTorch class inheriting from torch.nn.Module, defining a complete and standalone neural network implementation. The prompt used to generate it is
included as… See the full description on the dataset page: https://huggingface.co/datasets/DNadia/pytorch-nn-architectures-dataset.Technical-Architectures-Large
Technical Architectures Large (294k Samples)
Overview
Generating complex, syntactically valid diagram code from natural language requirements is a major challenge for AI models. This dataset bridges that gap by providing over 293,000+ distinct enterprise software architectures generated using two cutting-edge models: GPT-OSS-120B and Qwen3-Coder-Next-FP8.
Unlike simple "toy" examples, these architectures model realistic enterprise systems complete with client… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Technical-Architectures-Large.headwater-volume-ingestion-matrix-architecture-kit
HEADWATER Volume Ingestion Matrix & Architecture Kit
He will not suffer thy foot to be moved: he that keepeth thee will not slumber. Psalm 121:3
YOU CAN HAVE IT NOW.
🛠️ About This Kit
Packed with 10 modules
The Headwater Volume Ingestion Matrix & Architecture Kit is engineered for one primary purpose: to eliminate high-volume deployment delays and buy back your operational momentum.
Instead of forcing your developers to spend months planning… See the full description on the dataset page: https://huggingface.co/datasets/headwaterai/headwater-volume-ingestion-matrix-architecture-kit.aws-architecture-diagrams
AWS Architecture Diagrams (YOLO detection dataset)
Dataset YOLOv8 (formato Ultralytics: images/+labels/ por split, bbox
normalizada class_id cx cy w h) para detectar componentes, trust
boundaries e setas de fluxo de dados em diagramas de arquitetura AWS.
Usado para treinar
aws-architecture-vision-detector.
Código de geração completo em:
https://github.com/luisaoliveira1/posdiap_7iadt_techchallenge_5/tree/main/models/vision-detector
Tamanho
2022 imagens — 1623… See the full description on the dataset page: https://huggingface.co/datasets/luisasousa/aws-architecture-diagrams.stride-architecture-flows-v1
STRIDE Architecture Flow Detection Dataset (AWS & Azure)
📌 Overview
This dataset was created to detect flow arrows in cloud architecture diagrams (AWS and Azure),
supporting automated STRIDE threat modeling.
Unlike the component dataset (nodes), this dataset focuses on identifying communication flows
between architectural elements using bounding boxes and keypoints.
Each arrow is annotated with:
Bounding box (flow_arrow)
Two keypoints:
tail (source)
tip (destination)… See the full description on the dataset page: https://huggingface.co/datasets/guillherms/stride-architecture-flows-v1.chinese_architecture_siheyuanboltz2_accuracy_architecture_20260919_032813
Boltz2 accuracy and architecture study
Correction: The original Cα scoring selected the wrong atom. Use this revision’s corrected measurements and see the correction notice before using any prediction archive. Raw archived ca arrays are superseded; recover Cα from full coordinates and token_to_center_atom.
Measured small-protein ablations, an evaluation RNG/precision repair for the local Boltz quotient adapter, and a cyclic Fourier triangle-contraction prototype. No new weights… See the full description on the dataset page: https://huggingface.co/datasets/raftbioworks/boltz2_accuracy_architecture_20260919_032813.Cowpea-Architecture-XML
Cowpea-Architecture-XML-WDS
This dataset contains simulated images of Cowpea plants paired with organ-level architecture representations in XML format, packaged in WebDataset (.tar) format for efficient high-performance training.
Dataset Structure
The dataset is sharded into .tar files, each containing up to 10,000 samples.
Each sample consists of:
.jpeg: The plant image
.xml: The organ-level architecture representation
.json: (Optional) Metadata
Usage with… See the full description on the dataset page: https://huggingface.co/datasets/heesup/Cowpea-Architecture-XML.new_architecture_test_act_joystickThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "mcx",
"total_episodes": 5,
"total_frames": 1253,
"total_tasks": 1,
"total_videos": 5,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/antwoor/new_architecture_test_act_joystick.software-architecture-instructions-preferencearchitecture-audio-video59
Architecture Audio Video Data Notes
Dataset summary
A documented Architecture data-preparation workflow for Audio Video records. The bundled rows demonstrate the schema and validation path rather than pretending to be a full training corpus.
Included material
build_dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.… See the full description on the dataset page: https://huggingface.co/datasets/siregaradi/architecture-audio-video59.neurarch-architectures
Neurarch architecture corpus
36 neural network architectures kept as typed graphs rather than diagrams. Each entry carries every layer's type and parameters, the tensor shape propagated through it, an estimated parameter count, the verdict of 41 structural checks, and a link to the model.json an agent can fetch, edit and submit back for verification. Families span vision, language, recommendation, diffusion, biosignal and speech.
Size: 36 architectures
Licence: CC0 1.0… See the full description on the dataset page: https://huggingface.co/datasets/neurarch-ai/neurarch-architectures.Cowpea-Architecture-XML
Cowpea-Architecture-XML-WDS
This dataset contains simulated images of Cowpea plants paired with organ-level architecture representations in XML format, packaged in WebDataset (.tar) format for efficient high-performance training.
Dataset Structure
The dataset is sharded into .tar files, each containing up to 10,000 samples.
Each sample consists of:
.jpeg: The plant image
.xml: The organ-level architecture representation
.json: (Optional) Metadata… See the full description on the dataset page: https://huggingface.co/datasets/bbrangeo/Cowpea-Architecture-XML.nes-surrogate-architectures
NES Surrogate Dataset
Overview
This dataset contains trained neural architectures, their predictions, and validation performance, designed for studying:
surrogate modeling of neural architectures
diversity estimation between models
ensemble construction strategies
Each architecture is associated with:
its structure (DARTS-like cell)
model weights
validation predictions
validation accuracy
Dataset Structure
CIFAR10/
CIFAR100/
FashionMNIST/
Each… See the full description on the dataset page: https://huggingface.co/datasets/Demoren/nes-surrogate-architectures.new_architecture_test_act_joystick_ticksThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "mcx",
"total_episodes": 10,
"total_frames": 4970,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/antwoor/new_architecture_test_act_joystick_ticks.architecture2022old_architecture_test_act_joystick_ticksThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "mcx",
"total_episodes": 30,
"total_frames": 10290,
"total_tasks": 1,
"total_videos": 30,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/antwoor/old_architecture_test_act_joystick_ticks.Text-Gen-Architecturesoftware-architecture-instructionsrepro-conservation-laws-for-modern-neural-architectures-traces
Agent traces
Agent sessions published from a Trackio Logbook.
architecture-corpus
Architecture Image Text Data Notes
Dataset summary
Preparation notes and schema examples for Architecture tasks using Image Text data. Full source material is intentionally not bundled, so provenance and licensing remain explicit.
Included material
dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data… See the full description on the dataset page: https://huggingface.co/datasets/Leechaeyoungvoj/architecture-corpus.photo-architecture
Photo Architecture Dataset
Pulled from Pexels in 2023.
Images contain a majority of images of buildings and unique architecture. Some buildings may be copyrighted, though training is currently understood to fall under fair-use.
Image filenames may be used as captions, or, the parquet table contains the same values.
This dataset contains the full images.
Captions were created with CogVLM.
csc-wireless-latency-synthetic-100k
CSC Wireless Latency Synthetic Dataset (100k)
This synthetic dataset provides 100,000 prompt-completion pairs designed for training and evaluating PHY/MAC cross-layer optimization models in hybrid Li-Fi/RF wireless networks.
Official Core Implementation & Runtime
To parse, simulate, or process this dataset according to the official protocol specifications, please utilize the official runtime library:
Core Protocol Library (npm):… See the full description on the dataset page: https://huggingface.co/datasets/csc-architecture/csc-wireless-latency-synthetic-100k.architecture-audio-text29
Architecture Audio Text Data Notes
Dataset summary
This data card accompanies a lightweight Architecture loader for Audio Text metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
build_dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md —… See the full description on the dataset page: https://huggingface.co/datasets/Nicholassmith1220/architecture-audio-text29.Results_main_architecture_triple_audio_tie_breakersResults_main_architecture_triple_video_tie_breakersarchitecture-corpus-2024
Architecture Image Text Data Notes
Dataset summary
A documented Architecture data-preparation workflow for Image Text records. The bundled rows demonstrate the schema and validation path rather than pretending to be a full training corpus.
Included material
dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.… See the full description on the dataset page: https://huggingface.co/datasets/Takumiabe0213/architecture-corpus-2024.
