datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WN-ComponentsPCB_COMPONENTS_LABELLED
PCB Components Labelled
YOLO-format object detection datasets and trained weights for detecting electronic
components on PCB (printed circuit board) images. Two dataset/model pairs are
included, covering a coarse 5-class label set and a fine-grained 17-class label set.
Contents
pcb_final_v4/ 17-class dataset (train/valid, YOLO format)
pcb_5class_tiled/ 5-class dataset, tiled images (train/valid, YOLO format)… See the full description on the dataset page: https://huggingface.co/datasets/tutitata/PCB_COMPONENTS_LABELLED.Mechanical-Components
Mechanical Components Vibration Dataset
Comprehensive multi-source mechanical vibration dataset for training cross-component fault diagnosis and prognostics models. Designed for the Mechanical-JEPA project.
Total: ~12,000+ samples | 9.5 GB | 16 sources | 5 component types
Quick Start
from datasets import load_dataset
bearings = load_dataset("Forgis/Mechanical-Components", "bearings", split="train")
gearboxes = load_dataset("Forgis/Mechanical-Components", "gearboxes"… See the full description on the dataset page: https://huggingface.co/datasets/Forgis/Mechanical-Components.stride-architecture-components-v1
STRIDE Architecture Threat Modeling Dataset (AWS & Azure)
📌 Overview
This dataset was created to enable automatic STRIDE threat modeling from cloud architecture diagrams (AWS and Azure).
The goal is to detect architectural components in diagrams and support automated threat identification based on data flows and trust boundaries.
Annotations were created using Label Studio in YOLO format.
Total images: 4190Total classes: 32
🎯 Purpose
Detect cloud… See the full description on the dataset page: https://huggingface.co/datasets/guillherms/stride-architecture-components-v1.camel-componentselectronic-components-supply-chain
Electronic Components Supply Chain Dataset
Dataset Description
This dataset contains 791 electronic components collected from the Nexar API, enriched with supply chain risk analysis and supplier information. The dataset combines general electronic components and telecom-specific components, designed for supply chain transparency research, risk assessment, and component sourcing analysis.
Dataset Summary
Total Components: 791
Source: Nexar API (via Octopart)… See the full description on the dataset page: https://huggingface.co/datasets/mdnh/electronic-components-supply-chain.component-static-buildsk8s-kubectl-35k
Dataset Card for "k8s-kubectl-35k"
More Information needed
pc-components-reviews
Dataset Card for pc-components-reviews
This dataset has been created with distilabel.
The pipeline script was uploaded to easily reproduce the dataset:
generate_pc_components.py.
It can be run directly using the CLI:
distilabel pipeline run --script "https://huggingface.co/datasets/argilla/pc-components-reviews/raw/main/generate_pc_components.py"
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that… See the full description on the dataset page: https://huggingface.co/datasets/argilla/pc-components-reviews.powerline-components-and-faults
Powerline Components and Faults Dataset
Overview
The Powerline Components and Faults Dataset is a dataset designed for object detection tasks involving powerline components and associated faults. It provides images of powerline infrastructure along with annotated bounding boxes for various components and faults.
This dataset is augmented with mosaic augmentation useful for training and evaluating models on powerline inspection, maintenance, and safety applications.… See the full description on the dataset page: https://huggingface.co/datasets/docmhvr/powerline-components-and-faults.cdg-tech-components-integration-dataset
Character Pool Dataset: 81 Characters - Generated by Conversation Dataset Generator
This dataset was generated using the Conversation Dataset Generator script available at https://cahlen.github.io/conversation-dataset-generator/.
Generation Parameters
Number of Conversations Requested: 1000
Number of Conversations Successfully Generated: 1000
Total Turns: 5132
Model ID: meta-llama/Meta-Llama-3-8B-Instruct
Generation Mode:
Topic & Scenario
Initial Topic:… See the full description on the dataset page: https://huggingface.co/datasets/cahlen/cdg-tech-components-integration-dataset.k8s-kubectl-cot-20k
Dataset Card for "k8s-kubectl-cot-20k"
More Information needed
asia-who-financial-hardship-in-health-and-components-population
Financial hardship in health and components (SDG 3.8.2, 2025 definition): population (%) | Asia (WHO GHO)
🌏 4,884 observations · 39 Asia countries · 1992–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 4,884 observations of Financial hardship in health and components (SDG 3.8.2, 2025 definition): population (%) data across 39 Asia countries, spanning 1992–2024, covering 1 distinct indicators.
About the source
Source: WHO… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-who-financial-hardship-in-health-and-components-population.shadcn-components-v0Mechanical-Components
Mechanical Components Vibration Dataset
Comprehensive multi-source mechanical vibration dataset for training cross-component fault diagnosis and prognostics models. Designed for the Mechanical-JEPA project.
Total: ~12,000+ samples | 9.5 GB | 16 sources | 5 component types
Quick Start
from datasets import load_dataset
bearings = load_dataset("Forgis/Mechanical-Components", "bearings", split="train")
gearboxes = load_dataset("Forgis/Mechanical-Components"… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/Mechanical-Components.powerline-components-and-faults
Powerline Components and Faults Dataset
Overview
The Powerline Components and Faults Dataset is a dataset designed for object detection tasks involving powerline components and associated faults. It provides images of powerline infrastructure along with annotated bounding boxes for various components and faults.
This dataset is augmented with mosaic augmentation useful for training and evaluating models on powerline inspection, maintenance, and safety… See the full description on the dataset page: https://huggingface.co/datasets/meekL/powerline-components-and-faults.electronic_componentsconstruct-v1-componentsUMA-Dataset-LLM-Components-v2
UMA-IA/VELA-Components-v2
Authors
Youri LALAIN, Engineering student at French Engineering School ECE
Lilian RAGE, Engineering student at French Engineering School ECE
Dataset Summary
The UMA-IA/VELA-Components-v2 is a specialized dataset designed for training language models in the field of aerospace and aeronautical engineering components. It consists of structured question-answer pairs related to aerospace propulsion components, particularly focusing on… See the full description on the dataset page: https://huggingface.co/datasets/Alyon-AI/UMA-Dataset-LLM-Components-v2.mirrai-sd-training-open-components
MirrAI SD Training Open Components
This dataset repo contains the redistributable subset that was actually used in the MirrAI Stable Diffusion hair inpainting training pipeline.
Included In This Release
face_sketches_refined_generation/
images_512/: 139 processed 512px training images
masks/: hair, face protect, and cloth protect masks
controls/canny/: canny control inputs
face_crops/: identity-preserving face crops
manifests/recon_hair_external.jsonl: training manifest… See the full description on the dataset page: https://huggingface.co/datasets/siik/mirrai-sd-training-open-components.electrical-componentsgradio-componentscomponent_sort_v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 30,
"total_frames": 33452,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sadhana1818/component_sort_v2.shadcn-componentscomponent_sortThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 27,
"total_frames": 25678,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:27"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sadhana1818/component_sort.tailwindcss_componentsflowchat_components基于
galirage/FC-Detection
shreyanshu09/Block_Diagram
的重新标注。
使用makesense.ai对数据集的重新标注,标注分为以下几类
0: text
1: arrow
2: terminator
3: data
4: process
5: decision
6: connection
7: arrow_start
8: arrow_end
此分类参考论文 Arrow-Guided VLM: Enhancing Flowchart Understanding via Arrow Direction Encoding,也对论文所提到的数据集galirage/FC-Detection进行了重新标注,并加入了手动标注的几十张shreyanshu09/Block_Diagram流程图
示例:
africa-uganda-unemployment-rate-14-64-years-components-2012-13-2016-17-2-422dd326
Unemployment Rate 14 64 Years Components 2012 13 2016 17 2 | Africa (Uganda Bureau of Statistics)
18 rows - 1 Africa country/area - 2012-2022 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 18 rows from Uganda Bureau of Statistics, covering Unemployment Rate 14 64 Years Components 2012 13 2016 17 2. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-uganda-unemployment-rate-14-64-years-components-2012-13-2016-17-2-422dd326.Cell_Components_Fine-Tuning_DatasetA PubMed-based dataset, used for the fine-tuning of the BiomedNLP-PubMedBERT-base-uncased-abstract-fulltext model for the context-based classification of the names of molecular pathways.
HuggingFace card of the fine-tuned model.
GitHub link with a notebooks, for the fine-tuning and application of the model.
Citing
If you found the developed datasets to be useful in your research, please cite the following articles:
Ivanisenko, T.V., Saik, O.V., Demenkov, P.S. et al. ANDDigest:… See the full description on the dataset page: https://huggingface.co/datasets/Timofey/Cell_Components_Fine-Tuning_Dataset.k8s-kubectl
Dataset Card for "k8s-kubectl"
More Information needed
