datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EdgeBench
Overview
EdgeBench is a benchmark of 134 real-world tasks for evaluating how autonomous AI agents learn from real-world environments. Instead of measuring one-shot performance, EdgeBench places agents in executable task environments with realistic, multi-level feedback and lets them iterate for 12+ hours per task — tracking the full trajectory of improvement, not just the final score. We publicly release 51 tasks… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/EdgeBench.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.Wake-Vision
Dataset Card for Wake Vision
Dataset Description
"Wake Vision" is a large, high-quality dataset featuring over 6 million images, significantly exceeding the scale and diversity of
current tinyML datasets (100x). This dataset includes images with annotations of whether each image contains a person. Additionally,
it incorporates a comprehensive fine-grained benchmark to assess fairness and robustness, covering perceived gender, perceived age,
subject distance, lighting… See the full description on the dataset page: https://huggingface.co/datasets/Harvard-Edge/Wake-Vision.Prostate-Anatomical-Edge-Cases
Prostate-Anatomical-Edge-Cases
Stress-Testing Pelvic Autosegmentation Algorithms Using Anatomical Edge Cases —
a TCIA collection of pelvic radiotherapy planning CT with manually contoured
organs at risk, curated so that most cases contain anatomy known to break
autosegmentation algorithms (Kanwar et al., Phys Imaging Radiat Oncol 2023).
Read before using — the name is misleading in two ways:
This is CT, not MRI. Despite "Prostate" in the name it is not a prostate
mpMRI/zonal… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/Prostate-Anatomical-Edge-Cases.v1Wake-Vision-Train-LargeEdge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/BlueIsGreen/Edge-Agent-Reasoning-WebSearch-260K.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/DEMIRUNC/Edge-Agent-Reasoning-WebSearch-260K.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/Torenn/Edge-Agent-Reasoning-WebSearch-260K.tiger_layer_edgesAn unofficial re-packaged parquet files of TIGER/Line® Edges data provided by the US Census Bureau.
See LICENSE.pdf for more details.
edge-agent-reasoning-websearch-260k
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/ppenner/edge-agent-reasoning-websearch-260k.domain-resurrect-edges
Dataset Card for domain-resurrect-edges
This dataset contains link counts between domains on the Internet. The data is based on CommonCrawl.
See mbrt/domain-resurrect for the companion dataset with scored domains based on Page Rank, and
the Blog post on how this was computed.
Dataset Details
Dataset Description
This dataset is a processed version of the CommonCrawl September crawl.
Each row is the count of how many hyperlinks exist between the source… See the full description on the dataset page: https://huggingface.co/datasets/mbrt/domain-resurrect-edges.kg-edges
Dataset Card for Every Cure Integrated Knowledge Graph (Edges)
Dataset Summary
The Every Cure KG is currently (as of February 2026) essentially an integrated, simplified and filtered merged KG comprising ROBOKOP and RTX-KG2.
The "Edges" dataset contains the records for all edges in the graph, including metadata.
See nodes dataset for the corresponding set of nodes.
Source Data
Attribution
First-level knowledge sources
Primary knowledge sources… See the full description on the dataset page: https://huggingface.co/datasets/everycure/kg-edges.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/JACKYS999/Edge-Agent-Reasoning-WebSearch-260K.ark-asr-open-asr-leaderboard-results
ARK-ASR Open ASR Leaderboard Results
This dataset contains JSONL prediction manifests for AutoArk-AI/ARK-ASR-0.6B on hf-audio/open-asr-leaderboard public English short-form splits.
These files are intended for Open ASR Leaderboard maintainer verification.
Scoring summary from normalizer.eval_utils.score_results:
Split
WER
RTFx
ami/test
10.02
352.12
earnings22/test
9.77
331.88
gigaspeech/test
8.00
217.72
librispeech/test.clean
1.53
412.12
librispeech/test.other… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-open-asr-leaderboard-results.edges2shoes
Citation
@article{pix2pix2017,
title={Image-to-Image Translation with Conditional Adversarial Networks},
author={Isola, Phillip and Zhu, Jun-Yan and Zhou, Tinghui and Efros, Alexei A},
journal={CVPR},
year={2017}
}
ark-asr-3b-open-asr-leaderboard-results
ARK-ASR-3B Open ASR Leaderboard Results
Raw JSONL manifests for AutoArk-AI/ARK-ASR-3B on the public English
short-form hf-audio/open-asr-leaderboard splits.
These manifests were generated on a local 8x RTX 4090 machine and scored with
the shared Open ASR Leaderboard scorer:
PYTHONPATH=. python - <<'PY'
from normalizer.eval_utils import score_results
score_results(
'ark_asr/results.AutoArk-AI-ARK-ASR-3B_20260622_official',
'AutoArk-AI/ARK-ASR-3B',
)
PY
Important:… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-3b-open-asr-leaderboard-results.edge-ai-kg
Edge AI Deployment Knowledge Graph
25,152 nodes. 76,306 edges. Boards, kernels and neural networks in one graph — so you can
ask what actually runs on your silicon.
Built with Samyama Graph.
Loader and generator: samyama-ai/edge-ai-kg.
Part real, part synthetic — and every node says which
Every node carries a provenance property ("real" or "synthetic") and a source. No
node is unstamped:
provenance
Nodes
synthetic
23,910
real
1,242
Do not… See the full description on the dataset page: https://huggingface.co/datasets/VaidhyaMegha/edge-ai-kg.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/svryn/Edge-Agent-Reasoning-WebSearch-260K.Changelog-Nightly-Repositoriesarchive of all the repositories incl. metadata of: https://changelog.com/nightly
will be used to train a spam classifier with spacy; hence the "text" column, but kept submeta in case this is useful for anyone else to re-format.
The Dataset is provided ""AS IS"" and ""AS AVAILABLE"" without warranty of any kind, express or implied, including but not limited to warranties of merchantability, fitness for a particular purpose, title, or non-infringement.
The Provider disclaims all liability for… See the full description on the dataset page: https://huggingface.co/datasets/Edgerunners/Changelog-Nightly-Repositories.Edge-Computing-JEV
EdgeIntent v1
EdgeIntent v1 is a benchmark of natural-language requests to edge services, each paired with the typed intent
contract it expresses. It was built for the paper
Replacing Large Language Models with Jev Decision Models for Low-Latency Edge Service Orchestration
Delong Li, Xu Wang, Haochen Gong, Rui Lang, and Guangsheng Yu. University of Technology Sydney.
Code, evaluation harness, and reproduction instructions: https://github.com/OniReimu/Edge-Computing-JEV
This… See the full description on the dataset page: https://huggingface.co/datasets/OniReimu/Edge-Computing-JEV.rithmic-tick-dataedgevane-trainset-1Dataset with wikipedia fitst paragraph and public literature
friendship-graph-modular-edge-irregularity-proof
Defect Conservation and Exact Modular Edge-Irregularity Strength of Friendship Graphs
Public AI-friendly research release · candidate proof · independently verifiable artifacts
This repository contains a complete candidate resolution of Open Problem 3.3 from Koam, Ahmad, Bača, and Semaničová-Feňovčíková, AIMS Mathematics 8(1), 2023, concerning the modular edge irregularity strength of friendship graphs.
Main candidate theorem
For the friendship graph (F_n=K_1\vee… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/friendship-graph-modular-edge-irregularity-proof.MAICA_ds_basisEdgeReason
EdgeReason
EdgeReason is a compact verifier-backed dataset for improving small and
edge-deployable language models on tool use, structured JSON outputs,
state/table/unit/date reasoning, compact Mathlib-derived SFT, and routing
between direct answer, tool use, retrieval, clarification, and escalation.
The dataset is designed for teams training small models with SFT, DPO, RLVR,
GRPO, rejection sampling, and internal evaluation loops. It is not tied to any
model vendor or… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/EdgeReason.edge-ids-threats
HookProbe Edge IDS Threat Telemetry
Real-world, anonymised threat verdicts from the HookProbe production
edge intrusion-detection system. Unlike synthetic lab datasets
(CICIDS2017, UNSW-NB15, Kitsune) this is what an actual edge sensor
mesh observes on the open internet, labelled by the SENTINEL ensemble
(isolation forest + calibrated naive-Bayes) that ships with HookProbe.
Sensor: Raspberry Pi edge node + NAPSE AI-native flow classifier
Enrichment: RDAP country + ASN lookups… See the full description on the dataset page: https://huggingface.co/datasets/hookprobe/edge-ids-threats.jamjuri-edge-v4-stage2-datasets
Jamjuri-Edge V4 — Stage 2 Datasets (E1 / E2 / E3)
Cleaned, deduplicated, non-thinking SFT datasets for the Jamjuri-Edge V4 Stage-2 experts.
Series: part of JamjuriEDGE (4B) — collection · series card
✦ English
Overview
Three SFT datasets used to train the Stage-2 experts of
Cheva123/Jamjuri-EDGE-Preview-100,
plus the Stage-1 curriculum mixture (data/stage1/train.parquet) that trained the shared
Stage-1 parent. Every row is pre-rendered with… See the full description on the dataset page: https://huggingface.co/datasets/Cheva123/jamjuri-edge-v4-stage2-datasets.metabolomics-edges-expected-ge5
Cross-study metabolomics co-response edges (expected frequency ≥ 5)
620,265 edges over 34,378 nodes, drawn from pairwise metabolite co-response
statistics across MetaboLights and
Metabolomics Workbench studies, together with
the node properties, two PyTorch Geometric graphs, and the full pipeline that produces
them.
A node is one differential comparison within one study assay — MTBLS1285_0001_00000028
is study MTBLS1285, assay 0001, feature 00000028; ST002832_AN004625_00002191… See the full description on the dataset page: https://huggingface.co/datasets/kozo2/metabolomics-edges-expected-ge5.synthengine-cot-edge-case-v1
SynthEngine CoT Edge Case Dataset v1.0
Premium synthetic Chain-of-Thought reasoning data for autonomous driving, robotics, and embodied AI edge cases.
🔗 Full dataset (1000 records) available on Gumroad
This HuggingFace repo contains a free sample (10 records) under CC BY-NC-SA 4.0.
🎯 Why This Dataset?
In 2025, NVIDIA Alpamayo-R1 proved that Chain-of-Causation reasoning improves autonomous driving planning accuracy by +12% and reduces close encounters by -35%.… See the full description on the dataset page: https://huggingface.co/datasets/NeroSeungSan/synthengine-cot-edge-case-v1.
