datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EdgeBench
Overview
EdgeBench is a benchmark of 134 real-world tasks for evaluating how autonomous AI agents learn from real-world environments. Instead of measuring one-shot performance, EdgeBench places agents in executable task environments with realistic, multi-level feedback and lets them iterate for 12+ hours per task — tracking the full trajectory of improvement, not just the final score. We publicly release 51 tasks… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/EdgeBench.ark-asr-open-asr-leaderboard-results
ARK-ASR Open ASR Leaderboard Results
This dataset contains JSONL prediction manifests for AutoArk-AI/ARK-ASR-0.6B on hf-audio/open-asr-leaderboard public English short-form splits.
These files are intended for Open ASR Leaderboard maintainer verification.
Scoring summary from normalizer.eval_utils.score_results:
Split
WER
RTFx
ami/test
10.02
352.12
earnings22/test
9.77
331.88
gigaspeech/test
8.00
217.72
librispeech/test.clean
1.53
412.12
librispeech/test.other… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-open-asr-leaderboard-results.ark-asr-3b-open-asr-leaderboard-results
ARK-ASR-3B Open ASR Leaderboard Results
Raw JSONL manifests for AutoArk-AI/ARK-ASR-3B on the public English
short-form hf-audio/open-asr-leaderboard splits.
These manifests were generated on a local 8x RTX 4090 machine and scored with
the shared Open ASR Leaderboard scorer:
PYTHONPATH=. python - <<'PY'
from normalizer.eval_utils import score_results
score_results(
'ark_asr/results.AutoArk-AI-ARK-ASR-3B_20260622_official',
'AutoArk-AI/ARK-ASR-3B',
)
PY
Important:… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-3b-open-asr-leaderboard-results.Changelog-Nightly-Repositoriesarchive of all the repositories incl. metadata of: https://changelog.com/nightly
will be used to train a spam classifier with spacy; hence the "text" column, but kept submeta in case this is useful for anyone else to re-format.
The Dataset is provided ""AS IS"" and ""AS AVAILABLE"" without warranty of any kind, express or implied, including but not limited to warranties of merchantability, fitness for a particular purpose, title, or non-infringement.
The Provider disclaims all liability for… See the full description on the dataset page: https://huggingface.co/datasets/Edgerunners/Changelog-Nightly-Repositories.Edge-Computing-JEV
EdgeIntent v1
EdgeIntent v1 is a benchmark of natural-language requests to edge services, each paired with the typed intent
contract it expresses. It was built for the paper
Replacing Large Language Models with Jev Decision Models for Low-Latency Edge Service Orchestration
Delong Li, Xu Wang, Haochen Gong, Rui Lang, and Guangsheng Yu. University of Technology Sydney.
Code, evaluation harness, and reproduction instructions: https://github.com/OniReimu/Edge-Computing-JEV
This… See the full description on the dataset page: https://huggingface.co/datasets/OniReimu/Edge-Computing-JEV.EdgeReason
EdgeReason
EdgeReason is a compact verifier-backed dataset for improving small and
edge-deployable language models on tool use, structured JSON outputs,
state/table/unit/date reasoning, compact Mathlib-derived SFT, and routing
between direct answer, tool use, retrieval, clarification, and escalation.
The dataset is designed for teams training small models with SFT, DPO, RLVR,
GRPO, rejection sampling, and internal evaluation loops. It is not tied to any
model vendor or… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/EdgeReason.synthengine-cot-edge-case-v1
SynthEngine CoT Edge Case Dataset v1.0
Premium synthetic Chain-of-Thought reasoning data for autonomous driving, robotics, and embodied AI edge cases.
🔗 Full dataset (1000 records) available on Gumroad
This HuggingFace repo contains a free sample (10 records) under CC BY-NC-SA 4.0.
🎯 Why This Dataset?
In 2025, NVIDIA Alpamayo-R1 proved that Chain-of-Causation reasoning improves autonomous driving planning accuracy by +12% and reduces close encounters by -35%.… See the full description on the dataset page: https://huggingface.co/datasets/NeroSeungSan/synthengine-cot-edge-case-v1.slicenet-edge-slicing-scenarios
SliceNet Edge Slicing Scenarios
Synthetic edge computing and 5G network slicing placement scenarios for telecom, cloud, smart factory, connected vehicles, gaming, video, digital health, IoT, and smart city workloads.
Samples
File
Industry
Policy
sample_smart_city.json
Smart City / IoT
Min Latency
sample_connected_vehicle.json
Automotive V2X
Min Latency
sample_online_gaming.json
Cloud Gaming
Min Latency
sample_smart_factory.json
Industry 4.0
Min… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/slicenet-edge-slicing-scenarios.Edge-Industrial-Anomaly-Phi3
Edge-Industrial-Anomaly-Phi3: A Curated Dataset for SLMs
This dataset is a curated collection of industrial sensor data formatted specifically for Small Language Models (SLMs) like Phi-3. It merges three high-value industrial domains into a unified "Natural Language Reasoning" format to move beyond simple binary classification.
🚀 Purpose
Standard anomaly detection uses CSVs and Scikit-Learn. This dataset enables Generative Anomaly Detection, where a model like Phi-3 can… See the full description on the dataset page: https://huggingface.co/datasets/ssam17/Edge-Industrial-Anomaly-Phi3.douvras-network-edge-telemetry
Douvras Network and Edge Telemetry v0.1
Synthetic TCP/UDP telemetry episodes with latency, packet loss, throughput,
CPU and queue depth. Labels cover normal operation, latency, loss, congestion
and resource saturation. It contains 60 records (40/10/10) across 12 episodes,
split by episode. No real network traffic is included.
This is a diagnostic benchmark only; it never remediates or changes a network.
EdgeWisePersona
Dataset Card for EdgeWisePersona
Dataset Summary
The core component of the dataset consists of natural language sessions between users and their smart home systems. These dialogues simulate realistic, free-form interactions in which users express commands, preferences, or queries. The sessions are grounded in underlying formalized behavioral routines. These routines, along with the user profiles they compose, are also included in the dataset as ground truth… See the full description on the dataset page: https://huggingface.co/datasets/TCLResearchEurope/EdgeWisePersona.github-issuesNobodyExistsOnTheInternet_ToxicDPOqa_llama_factoryreformat of: NobodyExistsOnTheInternet/ToxicDPOqa for llama-factory DPO format
usage example:
"toxic_dpo_reformat": {
"hf_hub_url": "lucyknada/NobodyExistsOnTheInternet_ToxicDPOqa_llama_factory",
"ranking": true,
"columns": {
"prompt": "prompt",
"response": "response",
"system": "system"
}
},
Phoebus-86kraw human created erotica stories
very aggresively cleaned version of Phoebus-127k to try to remove:
product spambots, the website was being spammed a few times (usually with html tags)
warning only pages ("Warning" and nothing else)
edits (authors adding editorial history footnotes)
patreon and alike callouts (author asking for donations)
author notes, summaries, tagging
some still got through, classifier used: Edgerunners/Phoebus-Spam-Classifier-v2
The Dataset is provided ""AS IS"" and… See the full description on the dataset page: https://huggingface.co/datasets/Edgerunners/Phoebus-86k.EDGE-DatasetThis is the dataset repository of paper EDGE: Enhanced Grounded GUI Understanding with Enriched Multi-Granularity Synthetic Data.
Considering the huge number of images, the all_items.jsonl provided here contains the final QA pairs for training, but does not contain images for the time being. We will release all images as soon as possible.
You can also follow the mark_webpages and dataset.py scripts provided in the code repository to generate your own webpage image-question-answering dataset.… See the full description on the dataset page: https://huggingface.co/datasets/EDGEwww25/EDGE-Dataset.screenplay-format-edge-cases
Screenplay Format Edge Cases
48 original Fountain specimens in 24 contrast pairs — two near-identical inputs
per pair, at the points where the Fountain syntax leaves a choice. In 17 pairs
the one difference changes how the lines are classified. In the other 7 it
changes the surface and the labels hold: a lowercase scene prefix, a cue
extension, a non-Latin cue, escaped characters, a dual-dialogue caret, an inline
note and centered-text markers.
Version: 1.0.0 · Maintainer:… See the full description on the dataset page: https://huggingface.co/datasets/Inkwell-Software/screenplay-format-edge-cases.glaive-first-human-only-dedupedglaive function calling dataset filtered just for the original human prompt and deduplicated afterwards
The Dataset is provided ""AS IS"" and ""AS AVAILABLE"" without warranty of any kind, express or implied, including but not limited to warranties of merchantability, fitness for a particular purpose, title, or non-infringement.
The Provider disclaims all liability for any damages or losses resulting from the use or misuse of the Dataset, including but not limited to any damages or losses… See the full description on the dataset page: https://huggingface.co/datasets/Edgerunners/glaive-first-human-only-deduped.token-counting-edge-cases
token-counting-edge-cases
20 short strings with approximate token counts across three tokenizer families: Claude, GPT (cl100k_base), and Llama (SentencePiece). Built for sanity-checking token counters / chunkers / context-window fitters.
The numbers are approximate — exact counts depend on tokenizer version, BOS/EOS handling, and surrounding context. Expect ±1–2 token jitter. Use these to catch order-of-magnitude bugs (e.g. "your counter says 200 tokens for one emoji"), not as… See the full description on the dataset page: https://huggingface.co/datasets/mukunda1729/token-counting-edge-cases.EdgeMMEval
EdgeMMEval
Minimal multimodal evaluation dataset for on-device inference testing.
Covers functional correctness, accuracy, latency stress, and memory
pressure across image, audio, text, multi-turn, combination, structured
output, and tool-calling cases.
Dataset summary
The test split is defined in data/test/metadata.jsonl (200 rows). Each
row has a test_id (for example IMG-001, STO-020) and a modality.
Modality
Samples
Focus
Image
34
VQA, OCR, description… See the full description on the dataset page: https://huggingface.co/datasets/CortexSwarm/EdgeMMEval.Phoebus-127k-labelsPhoebus-127k but with labels added so users can connect chapters together.
raw human created erotica stories, needs filtering
things that need to be filtered:
product spambots, the website was being spammed a few times (usually with html tags)
warning only pages ("Warning" and nothing else)
edits (authors adding editorial history footnotes)
patreon and alike callouts (author asking for donations)
author notes, summaries, tagging
The Dataset is provided ""AS IS"" and ""AS AVAILABLE""… See the full description on the dataset page: https://huggingface.co/datasets/Edgerunners/Phoebus-127k-labels.Phoebus-127kraw human created erotica stories, needs filtering
things that need to be filtered:
product spambots, the website was being spammed a few times (usually with html tags)
warning only pages ("Warning" and nothing else)
edits (authors adding editorial history footnotes)
patreon and alike callouts (author asking for donations)
author notes, summaries, tagging
The Dataset is provided ""AS IS"" and ""AS AVAILABLE"" without warranty of any kind, express or implied, including but not limited to… See the full description on the dataset page: https://huggingface.co/datasets/Edgerunners/Phoebus-127k.nemo-grpo-from083-full-edge-curation
Nemotron 0.83 Edge-Prompt Curation
This private dataset contains edge-prompt curation rollouts for the DGXChen/Tong CoT dataset.
Seed edge prompts: 134
New rollout rows after seed exclusion: 7668
New edge prompts: 1648
Full edge prompts, seed plus rollout: 1782
Full dataset rows: 7830
Edge rate over full dataset: 0.2276
The Hugging Face dataset viewer is configured to load only data/full_edge_prompts_seed_plus_rollout.jsonl.
The larger rollout and metadata files remain… See the full description on the dataset page: https://huggingface.co/datasets/dvyomkesh/nemo-grpo-from083-full-edge-curation.wind-edge-1.6-sft
Wind Lite SFT
Custom supervised fine-tuning dataset for Wind Lite 1.6 by North AI.
Dataset Summary
20,000 high-quality instruction-response pairs covering identity grounding, math reasoning, coding, general knowledge, and multi-turn conversations.
Data Composition
Category
Count
Description
Math & Reasoning
~7,000
Arithmetic, algebra, percentages, unit conversions — with step-by-step working
Coding
~4,000
Python, JavaScript, SQL, systems — with… See the full description on the dataset page: https://huggingface.co/datasets/North-ML1/wind-edge-1.6-sft.Edgerunners__meta-llama-3-8b-instruct-hf-ortho-baukit-34fail-3000total-bf16-details
Dataset Card for Evaluation run of Edgerunners/meta-llama-3-8b-instruct-hf-ortho-baukit-34fail-3000total-bf16
Dataset automatically created during the evaluation run of model Edgerunners/meta-llama-3-8b-instruct-hf-ortho-baukit-34fail-3000total-bf16
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Edgerunners__meta-llama-3-8b-instruct-hf-ortho-baukit-34fail-3000total-bf16-details.DreadPoor__Blunt_Edge-8B-SLERP-details
Dataset Card for Evaluation run of DreadPoor/Blunt_Edge-8B-SLERP
Dataset automatically created during the evaluation run of model DreadPoor/Blunt_Edge-8B-SLERP
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Blunt_Edge-8B-SLERP-details.Thalia-Functions-13kBetter approach to Edgerunners/glaive-first-human-only-deduped, containing general questions a human might ask an AI assistant.
Best performers
I went through all OpenRouter offered models with theme-based filtering and then landed on a few top performers:
koboldai/psyfighter-13b-2 (72)
intel/neural-chat-7b (54)
mistralai/mistral-small (42)
openchat/openchat-7b (37)
gryphe/mythomax-l2-13b:nitro (31)
mistralai/mistral-tiny (31)
qwen/qwen-7b-chat (29)
the number representing how often it… See the full description on the dataset page: https://huggingface.co/datasets/Edgerunners/Thalia-Functions-13k.Thalia-Functions-OC3.6-33kSame idea as Edgerunners/Thalia-Functions-13k except instead of running through all OpenRouter models and then also generating locally, I have used the newly released OpenChat 3.6 which performed the best on its own, semantically deduped.
Topics covered:
General Trivia and Entertainment Questions
Lifestyle, Self-Improvement, and Productivity Questions
Music and Media Requests
Financial and Economic Questions
Technology and Computer-Related Questions
Food, Cooking, and Grocery Shopping… See the full description on the dataset page: https://huggingface.co/datasets/Edgerunners/Thalia-Functions-OC3.6-33k.simon-arc-solve-edge-v3
Version 1
ARC-AGI Tasks where the job is to identify where the top/bottom/left/right edge of the object is located.
example count: 4-5.
test count: 1-2.
image size: 3-5.
Version 2
image size: 3-10.
Version 3
image size: 3-5.
Focus on identifying diagonal edges.
simon-arc-solve-edge-v4
Version 1
ARC-AGI Tasks where the job is to identify where the top/bottom/left/right edge of the object is located.
example count: 4-5.
test count: 1-2.
image size: 3-5.
Version 2
image size: 3-10.
Version 3
image size: 3-5.
Focus on identifying diagonal edges.
Version 4
image size: 3-10.
Focus on identifying diagonal edges.
simon-arc-solve-edge-v5
Version 1
ARC-AGI Tasks where the job is to identify where the top/bottom/left/right edge of the object is located.
example count: 4-5.
test count: 1-2.
image size: 3-5.
Version 2
image size: 3-10.
Version 3
image size: 3-5.
Focus on identifying diagonal edges.
Version 4
image size: 3-10.
Focus on identifying diagonal edges.
Version 5
image size: 3-10.
Enabled all edge_names: top_left, top, top_right, left, right, bottom_left, bottom… See the full description on the dataset page: https://huggingface.co/datasets/neoneye/simon-arc-solve-edge-v5.
