datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
c4-en-html-with-metadata-ppl-cleanFile list:
"c4-en-html_cc-main-2019-18_pq00-000.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-001.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-002.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-003.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-004.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-005.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-006.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-007.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-008.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-009.jsonl.gz"… See the full description on the dataset page: https://huggingface.co/datasets/masoudjs/c4-en-html-with-metadata-ppl-clean.swe-agent-tool-rubrics-860
SWE Agent 逐 turn 工具调用评判数据集(860 个决策点)
本数据集来自 2026-08-06 的一次实验:**从真实 SWE agent 轨迹中归纳"怎么判断一次工具调用的好坏"**。
包含两个文件:
文件
行数
大小
内容
cases.jsonl
860
5.0 MB
决策点原始数据(题目、历史、两个候选命令、执行结果、现役判官打分)
map_io.jsonl
860
9.6 MB
每个决策点喂给 GPT-5.6 的完整 prompt 原文与完整回复
两个文件通过 case_id 一一对应。
背景:为什么是"按动作分类"而不是"按工具分类"
轨迹来自 slime 的 minimal harness,该 harness 只暴露一个工具 bash
(slime/agent/harness/minimal.py 里的 BASH_TOOL),全部 328,270 次调用的工具名都是 bash。
所以"不同工具用不同 rubric"无法按工具名实现,只能按命令在干什么分类。… See the full description on the dataset page: https://huggingface.co/datasets/MasterVito/swe-agent-tool-rubrics-860.ovos-stt-bench-speech-massive-fr-FR
OVOS stt bench — speech-massive-fr-FR
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
FBK-MT/Speech-MASSIVE-test.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-speech-massive-fr-FR.ovos-stt-bench-speech-massive-de-DE
OVOS stt bench — speech-massive-de-DE
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
FBK-MT/Speech-MASSIVE-test.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-speech-massive-de-DE.ovos-stt-bench-speech-massive-nl-NL
OVOS stt bench — speech-massive-nl-NL
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
FBK-MT/Speech-MASSIVE-test.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-speech-massive-nl-NL.mASNQ
Dataset Description
mASNQ is a translated version of ASNQ which is an AS2 dataset created by adapting the Natural Question corpus from Machine Reading (MR) to the AS2 task.
The dataset has been translated into five European languages: French, German, Italian, Portuguese, and Spanish, as described in this paper: Datasets for Multilingual Answer Sentence Selection.
Splits:
For each language (English, French, German, Italian, Portuguese, and Spanish), we provide:… See the full description on the dataset page: https://huggingface.co/datasets/matteogabburo/mASNQ.ovos-stt-bench-speech-massive-ru-RU
OVOS stt bench — speech-massive-ru-RU
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
FBK-MT/Speech-MASSIVE-test.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-speech-massive-ru-RU.massive-yt-edu-queue
Massive YouTube Educational Video Queue
Full metadata and content classification for 4,489,228 YouTube educational videos totaling 3,975,157 hours.
Description
This dataset contains metadata, content categorization, and license risk assessment for ~4.5M YouTube videos identified as potentially educational. It serves as the discovery and processing queue for the massive-yt-edu-transcriptions project, which aims to create the world's largest open educational transcript… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/massive-yt-edu-queue.ovos-stt-bench-speech-massive-hu-HU
OVOS stt bench — speech-massive-hu-HU
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
FBK-MT/Speech-MASSIVE-test.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-speech-massive-hu-HU.ovos-stt-bench-speech-massive-tr-TR
OVOS stt bench — speech-massive-tr-TR
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
FBK-MT/Speech-MASSIVE-test.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-speech-massive-tr-TR.ovos-stt-bench-speech-massive-pl-PL
OVOS stt bench — speech-massive-pl-PL
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
FBK-MT/Speech-MASSIVE-test.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-speech-massive-pl-PL.ariel-2025-cube-masked-cacheovos-stt-bench-speech-massive-vi-VN
OVOS stt bench — speech-massive-vi-VN
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
FBK-MT/Speech-MASSIVE-test.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-speech-massive-vi-VN.bilibili-masterpieces
Dataset Card for Bilibili Masterpieces
Dataset Summary
The bilibili-masterpieces dataset is a curated collection of representative works from some of the early well-known content creators (up 主) on the Bilibili platform. This dataset captures key metadata from these videos, providing a snapshot of the creative output that has significantly influenced the Bilibili community.
Supported Tasks and Leaderboards
The dataset can be used for various tasks such as video… See the full description on the dataset page: https://huggingface.co/datasets/wencan2024/bilibili-masterpieces.repro-causal-jepa-learning-world-models-through-object-level-latent-masking-traces
Agent traces
Agent sessions published from a Trackio Logbook.
dataset-A-routing-eval
Dataset A - Routing Evaluation
Total rows: 5339 | Gated (masked) rows: 0
Stratified dataset to evaluate 3 LLM models and 3 routing systems across 6 capabilities.
Multi-config layout
from datasets import load_dataset
# Full dataset
ds = load_dataset("massaindustries/dataset-A-routing-eval", "all")
# Per dimension
ds_math = load_dataset("massaindustries/dataset-A-routing-eval", "math_reasoning")
ds_code = load_dataset("massaindustries/dataset-A-routing-eval", "coding")
#… See the full description on the dataset page: https://huggingface.co/datasets/massaindustries/dataset-A-routing-eval.MaskGroups-HQovos-intent-bench-speech-massive-nl-NL
ovos-intent-bench-speech-massive-nl-NL
An OVOS Plugin Arena benchmark repository. The arena's prediction runner
publishes each plugin's raw output on the speech-massive-nl-NL dataset here, one JSON-lines
file per plugin under predictions/<lang>/, and where a sample-set manifest
governs scoring it lives under sample_sets/. The public leaderboard at
https://openvoiceos.github.io/ovos-plugin-arena/ is computed from these rows,
and anyone can recompute an entry from them without… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-speech-massive-nl-NL.ovos-intent-bench-speech-massive-pt-PT
ovos-intent-bench-speech-massive-pt-PT
An OVOS Plugin Arena benchmark repository. The arena's prediction runner
publishes each plugin's raw output on the speech-massive-pt-PT dataset here, one JSON-lines
file per plugin under predictions/<lang>/, and where a sample-set manifest
governs scoring it lives under sample_sets/. The public leaderboard at
https://openvoiceos.github.io/ovos-plugin-arena/ is computed from these rows,
and anyone can recompute an entry from them without… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-speech-massive-pt-PT.ovos-intent-bench-speech-massive-ru-RU
ovos-intent-bench-speech-massive-ru-RU
An OVOS Plugin Arena benchmark repository. The arena's prediction runner
publishes each plugin's raw output on the speech-massive-ru-RU dataset here, one JSON-lines
file per plugin under predictions/<lang>/, and where a sample-set manifest
governs scoring it lives under sample_sets/. The public leaderboard at
https://openvoiceos.github.io/ovos-plugin-arena/ is computed from these rows,
and anyone can recompute an entry from them without… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-speech-massive-ru-RU.ovos-intent-bench-speech-massive-pl-PL
ovos-intent-bench-speech-massive-pl-PL
An OVOS Plugin Arena benchmark repository. The arena's prediction runner
publishes each plugin's raw output on the speech-massive-pl-PL dataset here, one JSON-lines
file per plugin under predictions/<lang>/, and where a sample-set manifest
governs scoring it lives under sample_sets/. The public leaderboard at
https://openvoiceos.github.io/ovos-plugin-arena/ is computed from these rows,
and anyone can recompute an entry from them without… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-speech-massive-pl-PL.ovos-intent-bench-speech-massive-ar-SA
ovos-intent-bench-speech-massive-ar-SA
An OVOS Plugin Arena benchmark repository. The arena's prediction runner
publishes each plugin's raw output on the speech-massive-ar-SA dataset here, one JSON-lines
file per plugin under predictions/<lang>/, and where a sample-set manifest
governs scoring it lives under sample_sets/. The public leaderboard at
https://openvoiceos.github.io/ovos-plugin-arena/ is computed from these rows,
and anyone can recompute an entry from them without… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-speech-massive-ar-SA.ovos-intent-bench-speech-massive-de-DE
ovos-intent-bench-speech-massive-de-DE
An OVOS Plugin Arena benchmark repository. The arena's prediction runner
publishes each plugin's raw output on the speech-massive-de-DE dataset here, one JSON-lines
file per plugin under predictions/<lang>/, and where a sample-set manifest
governs scoring it lives under sample_sets/. The public leaderboard at
https://openvoiceos.github.io/ovos-plugin-arena/ is computed from these rows,
and anyone can recompute an entry from them without… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-speech-massive-de-DE.ovos-intent-bench-speech-massive-fr-FR
ovos-intent-bench-speech-massive-fr-FR
An OVOS Plugin Arena benchmark repository. The arena's prediction runner
publishes each plugin's raw output on the speech-massive-fr-FR dataset here, one JSON-lines
file per plugin under predictions/<lang>/, and where a sample-set manifest
governs scoring it lives under sample_sets/. The public leaderboard at
https://openvoiceos.github.io/ovos-plugin-arena/ is computed from these rows,
and anyone can recompute an entry from them without… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-speech-massive-fr-FR.ovos-intent-bench-speech-massive-hu-HU
ovos-intent-bench-speech-massive-hu-HU
An OVOS Plugin Arena benchmark repository. The arena's prediction runner
publishes each plugin's raw output on the speech-massive-hu-HU dataset here, one JSON-lines
file per plugin under predictions/<lang>/, and where a sample-set manifest
governs scoring it lives under sample_sets/. The public leaderboard at
https://openvoiceos.github.io/ovos-plugin-arena/ is computed from these rows,
and anyone can recompute an entry from them without… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-speech-massive-hu-HU.ovos-intent-bench-speech-massive-vi-VN
ovos-intent-bench-speech-massive-vi-VN
An OVOS Plugin Arena benchmark repository. The arena's prediction runner
publishes each plugin's raw output on the speech-massive-vi-VN dataset here, one JSON-lines
file per plugin under predictions/<lang>/, and where a sample-set manifest
governs scoring it lives under sample_sets/. The public leaderboard at
https://openvoiceos.github.io/ovos-plugin-arena/ is computed from these rows,
and anyone can recompute an entry from them without… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-speech-massive-vi-VN.ovos-intent-bench-speech-massive-ko-KR
ovos-intent-bench-speech-massive-ko-KR
An OVOS Plugin Arena benchmark repository. The arena's prediction runner
publishes each plugin's raw output on the speech-massive-ko-KR dataset here, one JSON-lines
file per plugin under predictions/<lang>/, and where a sample-set manifest
governs scoring it lives under sample_sets/. The public leaderboard at
https://openvoiceos.github.io/ovos-plugin-arena/ is computed from these rows,
and anyone can recompute an entry from them without… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-speech-massive-ko-KR.massive-pt-br
MASSIVE pt-BR — localização brasileira
Localização para português do Brasil do split pt-PT do
MASSIVE (Amazon; CC-BY-4.0) — o benchmark
original não tem locale pt-BR. 15.852 frases (train 10.948 / validation 1.929 /
test 2.930 aceitas), com anotação de slots preservada ([campo : valor]):
a frase limpa é derivada da anotada por construção, então os valores de slot são
literais por construção.
Método: professor LLM recebe apenas a frase ANOTADA e devolve a versão pt-BR anotada;… See the full description on the dataset page: https://huggingface.co/datasets/Magurofg/massive-pt-br.ovos-intent-bench-speech-massive-es-ES
ovos-intent-bench-speech-massive-es-ES
An OVOS Plugin Arena benchmark repository. The arena's prediction runner
publishes each plugin's raw output on the speech-massive-es-ES dataset here, one JSON-lines
file per plugin under predictions/<lang>/, and where a sample-set manifest
governs scoring it lives under sample_sets/. The public leaderboard at
https://openvoiceos.github.io/ovos-plugin-arena/ is computed from these rows,
and anyone can recompute an entry from them without… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-speech-massive-es-ES.dataset-A-routing
Dataset A - Routing (3 modelli, verdict-level)
Totale query: 5504 | Gated (masked) query: 0
Dataset di valutazione per LLM routing systems su 6 capability. Ogni query è stata eseguita su 3 modelli (qwen3.5-9b, deepseek-v4-flash, kimi2.6) e giudicata con grader deterministici (math/coding/ifeval) o LLM judge panel 2-of-3 (planning_agentic) / single judge (creative_synthesis, world_knowledge).
Configs
from datasets import load_dataset
# Pivot verdict per query (default… See the full description on the dataset page: https://huggingface.co/datasets/massaindustries/dataset-A-routing.
