datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EgoIT-99KCheckout the paper EgoLife (https://arxiv.org/abs/2503.03803) for more information.
vibravox
Dataset Card for VibraVox
👀 While waiting for the TooBigContentError issue to be resolved by the HuggingFace team, you can explore the dataset viewer of vibravox-test
which has exactly the same architecture.
DATASET SUMMARY
The VibraVox dataset is a general purpose audio dataset of french speech captured with body-conduction transducers.
This dataset can be used for various audio machine learning tasks :
Automatic Speech Recognition (ASR) (Speech-to-Text… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/vibravox.mmaulibrispeechvoicebench
License
The dataset is available under the Apache 2.0 license.
Citation
If you use the VoiceBench dataset in your research, please cite the following paper:
@article{chen2024voicebench,
title={VoiceBench: Benchmarking LLM-Based Voice Assistants},
author={Chen, Yiming and Yue, Xianghu and Zhang, Chen and Gao, Xiaoxue and Tan, Robby T. and Li, Haizhou},
journal={arXiv preprint arXiv:2410.17196},
year={2024}
}
vibravox-test
Dataset Card for Vibravox-test
Important Note
This dataset contains a very small proportion (1.2 %) of the original Vibravox Dataset.
vibravox-test is a only a dummy dataset for use with test pipelines in the Vibravox project. It is therefore not intended for training or testing models.
For full access to the complete dataset and documentation suitable for training and testing various audio and speech-related tasks, please visit the Vibravox Dataset page on Hugging Face.… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/vibravox-test.WenetSpeech_tempPerceptionTest_ValWorldSensegigaspeechAIR_benchClothoAQAOmni_Bench_fixcommon_voice_15PerceptionTestOpenS2S_Datasets
How to Use?
Download, merge the files, and extract
You can run the following command to merge the compressed file parts after downloading.
cat en_response_wav.tar.gz.* > en_response_wav.tar.gz
cat zh_response_wav.tar.gz.* > zh_response_wav.tar.gz
covost2_en-zhnon_curated_vibravox
Dataset Card for non-curated VibraVox
👀 This is the non-curated version of the VibraVox dataset. For a full documentation and dataset usage, please refer to https://huggingface.co/datasets/Cnam-LMSSC/vibravox
DATASET SUMMARY
The VibraVox dataset is a general purpose audio dataset of french speech captured with body-conduction transducers.
This dataset can be used for various audio machine learning tasks :
Automatic Speech Recognition (ASR) (Speech-to-Text… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/non_curated_vibravox.muchomusicDataset Summary
MuChoMusic is a benchmark designed to evaluate music understanding in multimodal audio-language models (Audio LLMs). The dataset comprises 1,187 multiple-choice questions created from 644 music tracks, sourced from two publicly available music datasets: MusicCaps and the Song Describer Dataset (SDD). The questions test knowledge and reasoning abilities across dimensions such as music theory, cultural context, and functional applications. All questions and answers have been… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-audio/muchomusic.common_voice_13_french_phoneme
Common Voice 13 French Phoneme
Dataset Summary
This dataset is a curated version of the French subset of Common Voice 13.0, enriched with a phonetic transcription column (phoneme).
It was created by the Laboratoire de Mécanique des Structures et des Systèmes Couplés (Cnam-LMSSC) to support research in speech processing, specifically for tasks requiring phonetic alignment, phoneme recognition, and robust speech-to-text applications in French.
The dataset retains the… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/common_voice_13_french_phoneme.shoebox_rir_updatedvocalsoundLMD-AI-Detection
LMD AI-Generated Music Detection Benchmark
(Note: The corresponding research paper will be released later.)
Dataset Description
The rapid advancement of AI music generation has raised growing concerns about the authenticity of digital music. While deepfake detection has been extensively studied in the audio domain, symbolic music (MIDI) remains largely unexplored.
This dataset presents a comprehensive benchmark for AI-generated symbolic music detection, examining… See the full description on the dataset page: https://huggingface.co/datasets/dhlee3000/LMD-AI-Detection.french_librispeech_vibravoxed"Vibravoxed" version on the french split of facebook/multilingual_librispeech
Deteriorated clean speech with reverse EBEN models:
Cnam-LMSSC/EBEN_reverse_forehead_accelerometer
Cnam-LMSSC/EBEN_reverse_rigid_in_ear_microphone
Cnam-LMSSC/EBEN_reverse_soft_in_ear_microphone
Cnam-LMSSC/EBEN_reverse_throat_microphone
Cnam-LMSSC/EBEN_reverse_temple_vibration_pickuppeoples_speechsyntheticlarge_shoebox_rirfleursmultilingual_librispeech_french_phoneme
Multilingual LibriSpeech French Phoneme
Dataset Summary
This dataset is a curated version of the French subset of Multilingual LibriSpeech (MLS), enriched with a phonetic transcription column (phoneme).
The Laboratoire de Mécanique des Structures et des Systèmes Couplés (Cnam-LMSSC) created this version to facilitate research into French acoustic modeling, phoneme recognition, and speech synthesis. It builds upon the high-quality audio derived from LibriVox audiobooks… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/multilingual_librispeech_french_phoneme.AV_Odyssey_Bench_LMMs_Eval
