datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
event-bench
Event Bench
29-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as an event planning assistant.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as an event planning assistant managing venue bookings, catering, and guest logistics. The conversation features cascading changes — a venue switch triggers catering… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/event-bench.even_speech_biblicalThis dataset consists of audiofiles with a speech in Even language.
The correspondence between text and audio is in the table metadata.csv.
The data was collected from religious texts written down by Institute for Bible Translation
Һөвки Дукундукун укчэнэкэл. Institute for Bible Translation, Moscow, 2018.
Притчал. Institute for Bible Translation, Moscow, 2019.
Sourse
Dialect
Total length (min)
Religious texts
Lamunkhin
67.96
Another dataset of Even speech:
field records of… See the full description on the dataset page: https://huggingface.co/datasets/tbkazakova/even_speech_biblical.even_speech_hseThis dataset consists of audiofiles with a speech in Even language.
The correspondence between text and audio is in the table metadata.csv.
The data was collected during field trips of HSE University expedition "Languages and Cultures of Kamchatka"
Sourse
Dialect
Total length (min)
Expedition records
Bystraja
TBA
Another dataset of Even speech:
biblical texts: https://huggingface.co/datasets/tbkazakova/even_speech_biblical
There is also unified data from the project (Aralova… See the full description on the dataset page: https://huggingface.co/datasets/tbkazakova/even_speech_hse.even_speech_pakendorfThis dataset consists of audiofiles with a speech in Even language.
The correspondence between text and audio is in the table metadata.csv.
The data was collected from the project (Aralova et al. 2007-2023)[1] and then brought to a unified format.
These are conversations about life, folklore and personal narratives, explanatory and procedural texts, individual words and phrases recorded in three districts: Bystraja District (Esso and Anavgaj villages), Sebyan-Kyuyol, Topolinoe.
This is field… See the full description on the dataset page: https://huggingface.co/datasets/tbkazakova/even_speech_pakendorf.
