datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ark-asr-open-asr-leaderboard-results
ARK-ASR Open ASR Leaderboard Results
This dataset contains JSONL prediction manifests for AutoArk-AI/ARK-ASR-0.6B on hf-audio/open-asr-leaderboard public English short-form splits.
These files are intended for Open ASR Leaderboard maintainer verification.
Scoring summary from normalizer.eval_utils.score_results:
Split
WER
RTFx
ami/test
10.02
352.12
earnings22/test
9.77
331.88
gigaspeech/test
8.00
217.72
librispeech/test.clean
1.53
412.12
librispeech/test.other… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-open-asr-leaderboard-results.ark-asr-3b-open-asr-leaderboard-results
ARK-ASR-3B Open ASR Leaderboard Results
Raw JSONL manifests for AutoArk-AI/ARK-ASR-3B on the public English
short-form hf-audio/open-asr-leaderboard splits.
These manifests were generated on a local 8x RTX 4090 machine and scored with
the shared Open ASR Leaderboard scorer:
PYTHONPATH=. python - <<'PY'
from normalizer.eval_utils import score_results
score_results(
'ark_asr/results.AutoArk-AI-ARK-ASR-3B_20260622_official',
'AutoArk-AI/ARK-ASR-3B',
)
PY
Important:… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-3b-open-asr-leaderboard-results.github-issuesEdgeMMEval
EdgeMMEval
Minimal multimodal evaluation dataset for on-device inference testing.
Covers functional correctness, accuracy, latency stress, and memory
pressure across image, audio, text, multi-turn, combination, structured
output, and tool-calling cases.
Dataset summary
The test split is defined in data/test/metadata.jsonl (200 rows). Each
row has a test_id (for example IMG-001, STO-020) and a modality.
Modality
Samples
Focus
Image
34
VQA, OCR, description… See the full description on the dataset page: https://huggingface.co/datasets/CortexSwarm/EdgeMMEval.Phoebus-127k-labelsPhoebus-127k but with labels added so users can connect chapters together.
raw human created erotica stories, needs filtering
things that need to be filtered:
product spambots, the website was being spammed a few times (usually with html tags)
warning only pages ("Warning" and nothing else)
edits (authors adding editorial history footnotes)
patreon and alike callouts (author asking for donations)
author notes, summaries, tagging
The Dataset is provided ""AS IS"" and ""AS AVAILABLE""… See the full description on the dataset page: https://huggingface.co/datasets/Edgerunners/Phoebus-127k-labels.nemo-grpo-from083-full-edge-curation
Nemotron 0.83 Edge-Prompt Curation
This private dataset contains edge-prompt curation rollouts for the DGXChen/Tong CoT dataset.
Seed edge prompts: 134
New rollout rows after seed exclusion: 7668
New edge prompts: 1648
Full edge prompts, seed plus rollout: 1782
Full dataset rows: 7830
Edge rate over full dataset: 0.2276
The Hugging Face dataset viewer is configured to load only data/full_edge_prompts_seed_plus_rollout.jsonl.
The larger rollout and metadata files remain… See the full description on the dataset page: https://huggingface.co/datasets/dvyomkesh/nemo-grpo-from083-full-edge-curation.Edgerunners__meta-llama-3-8b-instruct-hf-ortho-baukit-34fail-3000total-bf16-details
Dataset Card for Evaluation run of Edgerunners/meta-llama-3-8b-instruct-hf-ortho-baukit-34fail-3000total-bf16
Dataset automatically created during the evaluation run of model Edgerunners/meta-llama-3-8b-instruct-hf-ortho-baukit-34fail-3000total-bf16
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Edgerunners__meta-llama-3-8b-instruct-hf-ortho-baukit-34fail-3000total-bf16-details.DreadPoor__Blunt_Edge-8B-SLERP-details
Dataset Card for Evaluation run of DreadPoor/Blunt_Edge-8B-SLERP
Dataset automatically created during the evaluation run of model DreadPoor/Blunt_Edge-8B-SLERP
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Blunt_Edge-8B-SLERP-details.
