slu
Datasets
All datasets matching “slu”slue-phase-2
Licensing Information
SLUE-HVB
SLUE-HVB dataset contains a subset of the Gridspace-Stanford Harper Valley speech dataset and the copyright of this subset remains the same with the original license, CC-BY-4.0. See also original license notice (https://github.com/cricketclub/gridspace-stanford-harper-valley/blob/master/LICENSE)
Additionally, we provide dialog act classification annotation and it is covered with the same license as CC-BY-4.0.
SLUE-SQA-5… See the full description on the dataset page: https://huggingface.co/datasets/asapp/slue-phase-2.slurp
Dataset Card for "slurp"
More Information needed
qrecc-passages
QReCC Passages (54M Web Crawl)
This repository hosts the QReCC passage collection—a raw web-crawl dataset of 54 million passages. It includes only "id" and "contents" per record, stored in compressed Parquet format for efficient loading and streaming.
Source & Context
This dataset complements the QReCC retrieval setup outlined in the Apple ML-QReCC GitHub repository. Use this passage collection as the retrieval corpus for query rewriting and conversational… See the full description on the dataset page: https://huggingface.co/datasets/slupart/qrecc-passages.glove.6B.100d.txtglove.6B.100d.txt for practice
MAC_SLU
MAC-SLU: A Benchmark for Multi-Intent Spoken Language Understanding in Automotive Cabins
Paper | Code
This repository hosts the MAC-SLU dataset, a novel Multi-Intent Automotive Cabin Spoken Language Understanding Benchmark. MAC-SLU is designed to evaluate Spoken Language Understanding (SLU) systems on complex, multi-intent user commands within an automotive environment, addressing the limitations of existing SLU datasets in terms of diversity and complexity. It features authentic… See the full description on the dataset page: https://huggingface.co/datasets/Gatsby1984/MAC_SLU.SLUE-processed
