datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EdgeBench
Overview
EdgeBench is a benchmark of 134 real-world tasks for evaluating how autonomous AI agents learn from real-world environments. Instead of measuring one-shot performance, EdgeBench places agents in executable task environments with realistic, multi-level feedback and lets them iterate for 12+ hours per task — tracking the full trajectory of improvement, not just the final score. We publicly release 51 tasks… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/EdgeBench.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.Wake-Vision
Dataset Card for Wake Vision
Dataset Description
"Wake Vision" is a large, high-quality dataset featuring over 6 million images, significantly exceeding the scale and diversity of
current tinyML datasets (100x). This dataset includes images with annotations of whether each image contains a person. Additionally,
it incorporates a comprehensive fine-grained benchmark to assess fairness and robustness, covering perceived gender, perceived age,
subject distance, lighting… See the full description on the dataset page: https://huggingface.co/datasets/Harvard-Edge/Wake-Vision.IDA_Edgesankaku-webp-256shortest-edgev1voice-acting-edge-top3
Edge-case reward-Top-3 — voice annotations
1,494,503 synthetic expressive-speech utterances (the reward-Top-3 selection of the
edge-case corpus), annotated with:
a CrisperWhisper-format transcript with corrected vocal bursts — each surviving burst
carries a class name from laion/vocal-burst-detector-v2
at its original timestamp;
raw voice scores at two granularities — whole utterance and per sentence — from
laion/Empathic-Insight-Voice-Plus
(all 40 emotions), the 57 VoiceNet… See the full description on the dataset page: https://huggingface.co/datasets/laion/voice-acting-edge-top3.Prostate-Anatomical-Edge-Cases
Prostate-Anatomical-Edge-Cases
Stress-Testing Pelvic Autosegmentation Algorithms Using Anatomical Edge Cases —
a TCIA collection of pelvic radiotherapy planning CT with manually contoured
organs at risk, curated so that most cases contain anatomy known to break
autosegmentation algorithms (Kanwar et al., Phys Imaging Radiat Oncol 2023).
Read before using — the name is misleading in two ways:
This is CT, not MRI. Despite "Prostate" in the name it is not a prostate
mpMRI/zonal… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/Prostate-Anatomical-Edge-Cases.bleeding-edge-gameplay-sampleThis dataset contains 1024 60 second video clips of Bleeding Edge gameplay (75GB). The data has already been processed into the following format: 300x180 videos sampled at 10 fps.
Dataset Structure
Data Files
testing_dataset_part1.zip & testing_dataset_part2.zip – Contains all 1024 60 second trajectories used for our evaluation.
4 examples from the dataset:
FB[…].npz – .npz file (described below)
FB[…].mp4 – 60 seconds .mp4 video of the images from the .npz file.… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/bleeding-edge-gameplay-sample.gpa-demosEdge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/BlueIsGreen/Edge-Agent-Reasoning-WebSearch-260K.Wake-Vision-Train-LargeEdge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/DEMIRUNC/Edge-Agent-Reasoning-WebSearch-260K.Hey-Edge
hey_edge — Wake Word Synthetic Speech Dataset
Synthetic, augmented audio for training a small wake-word / keyword-spotting
model. Generated with piper_tts and local audio augmentation.
Classes
Label
Samples
background_noise
200
hey_edge
378
unknown
1071
hey_edge — the target wake phrase and close variants.
unknown — near-miss and unrelated short phrases.
background_noise — synthetic background noise.
Audio Specification… See the full description on the dataset page: https://huggingface.co/datasets/edgeimpulse/Hey-Edge.tiger_layer_edgesAn unofficial re-packaged parquet files of TIGER/Line® Edges data provided by the US Census Bureau.
See LICENSE.pdf for more details.
Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/Torenn/Edge-Agent-Reasoning-WebSearch-260K.scanqa_images_16_keyframes_120_non_keyframes_min_532_long_edgeoprefill-hicache-2p1d-glm52-edgebench-ic1
Optimistic Prefill × HiCache 2P1D Ablation — full experiment artifacts
This dataset contains the complete artifacts (client traces, router receipts, engine
telemetry/logs, per-request exports, analysis scripts, and results) of a matched-pair
ablation of SGLang optimistic prefill in a PD-disaggregated deployment, run on
2026-08-01 on Together's research-b200-ic1 cluster. It is published gated so results
can be re-analyzed later without cluster access.
What is NOT here (by… See the full description on the dataset page: https://huggingface.co/datasets/weili-0234/oprefill-hicache-2p1d-glm52-edgebench-ic1.edge-agent-reasoning-websearch-260k
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/ppenner/edge-agent-reasoning-websearch-260k.edge-aui-framework-data
Dataset Card: Edge-Native Adaptive UI Behavioral Logs
This repository stores aggregated behavioral interaction logs in raw and processed forms, primarily in the .parquet data storage format.
These logs support research on edge-native adaptive user interfaces.
1. Hosted Data
The repository hosts microtensor parquet files derived from five public research datasets:
AdSERP Search and Interaction Logs (Arapakis et al., 2025).
High-Volume Trajectories (Mendeley Mouse… See the full description on the dataset page: https://huggingface.co/datasets/T40/edge-aui-framework-data.kg-edges
Dataset Card for Every Cure Integrated Knowledge Graph (Edges)
Dataset Summary
The Every Cure KG is currently (as of February 2026) essentially an integrated, simplified and filtered merged KG comprising ROBOKOP and RTX-KG2.
The "Edges" dataset contains the records for all edges in the graph, including metadata.
See nodes dataset for the corresponding set of nodes.
Source Data
Attribution
First-level knowledge sources
Primary knowledge sources… See the full description on the dataset page: https://huggingface.co/datasets/everycure/kg-edges.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/JACKYS999/Edge-Agent-Reasoning-WebSearch-260K.domain-resurrect-edges
Dataset Card for domain-resurrect-edges
This dataset contains link counts between domains on the Internet. The data is based on CommonCrawl.
See mbrt/domain-resurrect for the companion dataset with scored domains based on Page Rank, and
the Blog post on how this was computed.
Dataset Details
Dataset Description
This dataset is a processed version of the CommonCrawl September crawl.
Each row is the count of how many hyperlinks exist between the source… See the full description on the dataset page: https://huggingface.co/datasets/mbrt/domain-resurrect-edges.OEOBench-CloudSEN12demo_grab_pcb_from_edge_igor1_pcbcrop1_arecord-EdgeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 60,
"total_frames": 20252,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/U-RIL/record-Edge.ark-asr-open-asr-leaderboard-results
ARK-ASR Open ASR Leaderboard Results
This dataset contains JSONL prediction manifests for AutoArk-AI/ARK-ASR-0.6B on hf-audio/open-asr-leaderboard public English short-form splits.
These files are intended for Open ASR Leaderboard maintainer verification.
Scoring summary from normalizer.eval_utils.score_results:
Split
WER
RTFx
ami/test
10.02
352.12
earnings22/test
9.77
331.88
gigaspeech/test
8.00
217.72
librispeech/test.clean
1.53
412.12
librispeech/test.other… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-open-asr-leaderboard-results.demo_grab_pcb_from_edge1_center_annotEdge-Compute-Reasoning-Distillation-Dataset
Edge-Compute Reasoning Distillation 🧠⚡
This repository contains a pipeline for Reasoning Distillation—fine-tuning a tiny 1.5B Small Language Model (Qwen 2.5) that runs entirely offline on Edge hardware (e.g., a 4GB VRAM RTX 3050).
Note: The original dataset generation scripts have been removed. To use this repository, you must provide your own dataset in data/train.jsonl formatted with <think> tags.
Architecture
The project focuses on the training and deployment… See the full description on the dataset page: https://huggingface.co/datasets/athul020/Edge-Compute-Reasoning-Distillation-Dataset.ark-asr-3b-open-asr-leaderboard-results
ARK-ASR-3B Open ASR Leaderboard Results
Raw JSONL manifests for AutoArk-AI/ARK-ASR-3B on the public English
short-form hf-audio/open-asr-leaderboard splits.
These manifests were generated on a local 8x RTX 4090 machine and scored with
the shared Open ASR Leaderboard scorer:
PYTHONPATH=. python - <<'PY'
from normalizer.eval_utils import score_results
score_results(
'ark_asr/results.AutoArk-AI-ARK-ASR-3B_20260622_official',
'AutoArk-AI/ARK-ASR-3B',
)
PY
Important:… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-3b-open-asr-leaderboard-results.
