datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
slue-phase-2
Licensing Information
SLUE-HVB
SLUE-HVB dataset contains a subset of the Gridspace-Stanford Harper Valley speech dataset and the copyright of this subset remains the same with the original license, CC-BY-4.0. See also original license notice (https://github.com/cricketclub/gridspace-stanford-harper-valley/blob/master/LICENSE)
Additionally, we provide dialog act classification annotation and it is covered with the same license as CC-BY-4.0.
SLUE-SQA-5… See the full description on the dataset page: https://huggingface.co/datasets/asapp/slue-phase-2.slurp
Dataset Card for "slurp"
More Information needed
qrecc-passages
QReCC Passages (54M Web Crawl)
This repository hosts the QReCC passage collection—a raw web-crawl dataset of 54 million passages. It includes only "id" and "contents" per record, stored in compressed Parquet format for efficient loading and streaming.
Source & Context
This dataset complements the QReCC retrieval setup outlined in the Apple ML-QReCC GitHub repository. Use this passage collection as the retrieval corpus for query rewriting and conversational… See the full description on the dataset page: https://huggingface.co/datasets/slupart/qrecc-passages.glove.6B.100d.txtglove.6B.100d.txt for practice
MAC_SLU
MAC-SLU: A Benchmark for Multi-Intent Spoken Language Understanding in Automotive Cabins
Paper | Code
This repository hosts the MAC-SLU dataset, a novel Multi-Intent Automotive Cabin Spoken Language Understanding Benchmark. MAC-SLU is designed to evaluate Spoken Language Understanding (SLU) systems on complex, multi-intent user commands within an automotive environment, addressing the limitations of existing SLU datasets in terms of diversity and complexity. It features authentic… See the full description on the dataset page: https://huggingface.co/datasets/Gatsby1984/MAC_SLU.SLUE-processedslurp_synthetic_barkslurp_slu_intentslugvision
slugvision
~32,000 image → URL-slug pairs for training small vision-language models to
generate 3–5 word kebab-case permalink slugs, e.g. three-brown-teddy-bears,
nursing-workforce-growth-scaling. Built as the distillation corpus for
slugvision-500m and
slugvision-2.2b.
The distribution mimics editorial-site image uploads (article cards): a mix
of everyday photos, conceptual editorial illustrations, product shots,
clinical stock imagery, and web-page screenshots. All slugs were… See the full description on the dataset page: https://huggingface.co/datasets/kelnei/slugvision.slurpdetails_chargoddard__duplicitous-slurpbeast-13b
Dataset Card for Evaluation run of chargoddard/duplicitous-slurpbeast-13b
Dataset Summary
Dataset automatically created during the evaluation run of model chargoddard/duplicitous-slurpbeast-13b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_chargoddard__duplicitous-slurpbeast-13b.asr-slu_whisper
Dataset Card for "asr-slu_whisper"
More Information needed
project-gutenberg-alice-wonderlandaliaksandr-serzhputouski-kazki-i-apaviadanni-belarusau-slutskaga-pavetu-iury-zhy
Казкі і апавяданні беларусаў Слуцкага павету
Metadata
Author: Аляксандр Сержпутоўскі
Title: Казкі і апавяданні беларусаў Слуцкага павету
Narrator: Юры Жыгамонт
Source Group: Аўдыёкнігі
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/aliaksandr-serzhputouski-kazki-i-apaviadanni-belarusau-slutskaga-pavetu-iury-zhy.M3-SLU-Task2slue_p2_sqa5_test@article{shon2022slue,
title={SLUE phase-2: A benchmark suite of diverse spoken language understanding tasks},
author={Shon, Suwon and Arora, Siddhant and Lin, Chyi-Jiunn and Pasad, Ankita and Wu, Felix and Sharma, Roshan and Wu, Wei-Lun and Lee, Hung-yi and Livescu, Karen and Watanabe, Shinji},
journal={arXiv preprint arXiv:2212.10525},
year={2022}
}
@article{wang2024audiobench,
title={AudioBench: A Universal Benchmark for Audio Large Language Models},
author={Wang, Bin and Zou… See the full description on the dataset page: https://huggingface.co/datasets/AudioLLMs/slue_p2_sqa5_test.slurp-mask-v2Auto-SLURP
Auto-SLURP: A Benchmark Dataset for Evaluating Multi-Agent Frameworks in Smart Personal Assistant
Repository for the paper Auto-SLURP: A Benchmark Dataset for Evaluating Multi-Agent Frameworks in Smart Personal Assistant
requirements
To test the multi-agent frameworks, you need to first install the framework according to the instruction of the framework.
We have tested CamelAI, Langgraph, AgentLite, and AutoGEN.
1. start simulated servers
cd server
sh run.sh… See the full description on the dataset page: https://huggingface.co/datasets/lorashen/Auto-SLURP.slurpslue(Jan. 8 2024) Test set labels are released
Dataset Card for SLUE
Dataset Summary
We introduce the Spoken Language Understanding Evaluation (SLUE) benchmark. The goals of our work are to
Track research progress on multiple SLU tasks
Facilitate the development of pre-trained representations by providing fine-tuning and eval sets for a variety of SLU tasks
Foster the open exchange of research by focusing on freely available datasets that all academic and industrial groups… See the full description on the dataset page: https://huggingface.co/datasets/asapp/slue.slurp_slu_intent_with_transcriptionopen-bottleneck-ranklong27b-slurm-364982-rollouts
Open Bottleneck RankLong 27B — Slurm array 364982
Compact rollout evidence archived from completed Slurm array 364982.
Config
Files / steps
Records
JSONL bytes
Note
rank_a40
60 (1–60)
15,360
66,445,670
Complete local rollout evidence
rank_a80
54 (1–54)
13,824
59,386,451
Includes the cancelled arm's final dumped step (54.jsonl)
Each JSONL record contains input, output, gts, score, acc,
response_length, grouprel_reward, and step.
Only rollout evidence is archived… See the full description on the dataset page: https://huggingface.co/datasets/ryankim17920/open-bottleneck-ranklong27b-slurm-364982-rollouts.SLURP-TN
SLURP-TN : Resource for Tunisian Dialect Spoken Language Understanding
Contact person : fethi.bougares@elyadata.com
Spoken Language Understanding (SLU) aims to extract the semantic information from the speech utterance of user
queries. It is a core component in a task-oriented dialogue system. With the spectacular progress of deep neural
network models and the evolution of pre-trained language models, SLU has obtained significant breakthroughs.
However, only a few… See the full description on the dataset page: https://huggingface.co/datasets/Elyadata/SLURP-TN.slurp-bab-Qwen2.5-Omni-3B-asr1-v1Libri-Trans-Spatialized_SLURP-Spatialized_datasetSpatialized Libri-Trans and Spatialized SLURP (LT-S and SLURP-S), Enhancement for Translation and Understanding dataset
slue-voxceleb
Dataset Card for "slue-voxceleb"
More Information needed
slurp-asr-bab-v1slurp
Dataset Card for "slurp"
More Information needed
mask-slurpslue-sqa-code-l22-c500
SLUE-SQA-5 HuBERT Layer-22 K=500 Discrete Units
Packed discrete-unit files for SpeechGR experiments on SLUE-SQA-5.
The units were produced with HuBERT layer 22 and a K=500 k-means model, then deduplicated with consecutive counts retained. The packed format avoids one .code and .cnt file per utterance.
Files
documents.npz: packed document/passage units
train.npz: packed train question units
validation.npz: packed validation question units
test.npz: packed test question… See the full description on the dataset page: https://huggingface.co/datasets/dodofk/slue-sqa-code-l22-c500.
