datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SwitchLingua_audio
Dataset Card for SwitchLingua_text
🚀 News
[19/09/2025] SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset is accepted by NeurIPS 2025!
[30/05/2024] The manuscript can be found on arXiv.
Dataset Summary
SwitchLingua is a comprehensive multilingual and multicultural code-switching dataset designed to advance research in automatic speech recognition, natural language processing, and conversational AI. The… See the full description on the dataset page: https://huggingface.co/datasets/Shelton1013/SwitchLingua_audio.Ficbook-Audio-Instruct-10K
Ficbook Audio Instruct 10K
Synthetic audio instruction dataset for training Russian audio-language models.
Contains ~10K samples of fiction text voiced with OpenAI TTS and paired with diverse instruction tasks.
Dataset Description
This dataset was created for training and evaluating audio-language models on Russian fiction content.
Each sample contains:
Audio: Fiction text voiced using OpenAI's gpt-4o-mini-tts model
Text: Original text from ficbook stories
Question:… See the full description on the dataset page: https://huggingface.co/datasets/Vikhrmodels/Ficbook-Audio-Instruct-10K.
