datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.Wake-Vision
Dataset Card for Wake Vision
Dataset Description
"Wake Vision" is a large, high-quality dataset featuring over 6 million images, significantly exceeding the scale and diversity of
current tinyML datasets (100x). This dataset includes images with annotations of whether each image contains a person. Additionally,
it incorporates a comprehensive fine-grained benchmark to assess fairness and robustness, covering perceived gender, perceived age,
subject distance, lighting… See the full description on the dataset page: https://huggingface.co/datasets/Harvard-Edge/Wake-Vision.Prostate-Anatomical-Edge-Cases
Prostate-Anatomical-Edge-Cases
Stress-Testing Pelvic Autosegmentation Algorithms Using Anatomical Edge Cases —
a TCIA collection of pelvic radiotherapy planning CT with manually contoured
organs at risk, curated so that most cases contain anatomy known to break
autosegmentation algorithms (Kanwar et al., Phys Imaging Radiat Oncol 2023).
Read before using — the name is misleading in two ways:
This is CT, not MRI. Despite "Prostate" in the name it is not a prostate
mpMRI/zonal… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/Prostate-Anatomical-Edge-Cases.v1Wake-Vision-Train-LargeEdge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/BlueIsGreen/Edge-Agent-Reasoning-WebSearch-260K.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/DEMIRUNC/Edge-Agent-Reasoning-WebSearch-260K.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/Torenn/Edge-Agent-Reasoning-WebSearch-260K.edge-agent-reasoning-websearch-260k
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/ppenner/edge-agent-reasoning-websearch-260k.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/JACKYS999/Edge-Agent-Reasoning-WebSearch-260K.edge-ML-metabolomics-graph
edge_ML metabolomics co-response graph
An undirected graph of 18,494 nodes and 2,709,209 edges built from pairwise
metabolite co-response statistics across 83 MetaboLights
studies, together with the node properties, the PyTorch Geometric graph object, and the
full pipeline that produces them.
A node is one differential comparison within one study assay (MTBLS1405_0002_00003332
= study MTBLS1405, assay 002, feature 00003332). An edge carries the association
between two… See the full description on the dataset page: https://huggingface.co/datasets/kozo2/edge-ML-metabolomics-graph.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/svryn/Edge-Agent-Reasoning-WebSearch-260K.edges2handbagscoffee-lamp
Synthetic Image-Classification Dataset
Synthetic image-classification dataset generated with stable diffusion
(zerogpu_sdxl_turbo) using text-to-image from class names + short descriptions.
Classes
Label
Images
background
20
coffee-mug
20
lamp
20
Layout
train/<label>/<label>.<id>.jpg
test/<label>/<label>.<id>.jpg
metadata.csv
Loading
from datasets import load_dataset
ds = load_dataset("imagefolder"… See the full description on the dataset page: https://huggingface.co/datasets/edgeimpulse/coffee-lamp.PhilEOBench-road_density_regression
Simulated PhiSat Bench Dataset - Roads
This dataset comprises simulated PhiSat2 data derived from Sentinel-2, tailored for pixel-wise regression tasks aimed at estimating road coverage.
Dataset Overview
Each sample in the dataset includes a single-channel label.
The labels are stored as floating-point values that represent the estimated percentage of roads area within each pixel.
For a pixel with a 10-meter resolution (representing 100 square meters), the label… See the full description on the dataset page: https://huggingface.co/datasets/ESA-PhiLab-Edge/PhilEOBench-road_density_regression.mapillary_vistas_semantic_edges_and_segmentationedges2shoesPhilEOBench-building_density_regression
Simulated PhiSat Bench Dataset - Buildings
This repository contains a simulated dataset derived from Sentinel-2 data for building analysis.
Specifically, the dataset simulates outputs from the PhiSat2 satellite.
Label Description
Each sample in the dataset includes a single-channel label.
The labels are stored as floating-point values that represent the estimated percentage of building coverage within each pixel.
For a pixel with a 10-meter resolution… See the full description on the dataset page: https://huggingface.co/datasets/ESA-PhiLab-Edge/PhilEOBench-building_density_regression.celeba_with_llava_captions_and_edges
Dataset Card for "celeba_with_llava_captions_and_edges"
More Information needed
scanqa_images_16_keyframes_120_non_keyframes_min_532_long_edgeCyberpunk-Edgerunnerscubes-on-conveyor-beltThis dataset has been collected by Edge Impulse and used extensively to design the FOMO (Faster Objects, More Objects) object detection architecture. See FOMO documentation or the announcement blog post.
The dataset is composed of 70 images including:
32 blue cubes,
32 green cubes,
30 red cubes
28 yellow cubes
Download link: cubes on a conveyor belt dataset in Edge Impulse Object Detection format.
You can also retrieve this dataset from this Edge Impulse public project.
Data exported from… See the full description on the dataset page: https://huggingface.co/datasets/edgeimpulse/cubes-on-conveyor-belt.test_upload6_mini_LAION_ART_with_hard_comped_edgeedgetear_ssEdgeMMEval
EdgeMMEval
Minimal multimodal evaluation dataset for on-device inference testing.
Covers functional correctness, accuracy, latency stress, and memory
pressure across image, audio, text, multi-turn, combination, structured
output, and tool-calling cases.
Dataset summary
The test split is defined in data/test/metadata.jsonl (200 rows). Each
row has a test_id (for example IMG-001, STO-020) and a modality.
Modality
Samples
Focus
Image
34
VQA, OCR, description… See the full description on the dataset page: https://huggingface.co/datasets/CortexSwarm/EdgeMMEval.edge-vision-modelsemantic_edges_and_segmentation_placepulseScratches_roller_edgeEdge_Mapping_train_subsetedges-controlnet-dataset-main
