datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hindi_audio_dataset_testsmart-turn-data-v3.2-testTesting dataset for Smart Turn v3.2.
Thank you to the following contributors whose audio samples are included in this dataset:
The Pipecat team
Liva AI: https://www.theliva.ai/
Midcentury: https://www.midcentury.xyz/
MundoAI: https://mundoai.world/
Also, thank you to the following people for the CC-0 background noise sample data which has been used in this dataset:
https://freesound.org/people/4team/sounds/214995/
https://freesound.org/people/tomhannen/sounds/698090/… See the full description on the dataset page: https://huggingface.co/datasets/pipecat-ai/smart-turn-data-v3.2-test.audio_test_dataset
Dataset Card for "audio_test_dataset"
This dataset consists of the first 5 samples of mozilla-foundation/common_voice_13_0 and is only used for unit testing.
smart-turn-data-v3.1-testTesting dataset for Smart Turn v3.1.
Thank you to the following contributors whose audio samples are included in this dataset:
The Pipecat team
Liva AI: https://www.theliva.ai/
Midcentury: https://www.midcentury.xyz/
MundoAI: https://mundoai.world/
wav2vec2-test-datasetsmart-turn-data-v3-testtest_audio_datasetIMDA-NSC-datasets-testaudio_data_kaggle_test_taskb_synthetic-data-indonesia_testaudio_data_kaggle_test_taskcTestDataset5149test-datasetqaloon_dataset_testnormalized_test_ATC_datasetSartify_ITU_Zindi_Testdataset
Sartify Your Voice, Your Device, Your Language Challenge
Dataset Description
This dataset contains audio files designed for automatic speech recognition (ASR) and transcription tasks, created as part of an Sartify competition. The dataset is structured to facilitate machine learning model evaluation on audio-to-text transcription.
Dataset Summary
The Test Dataset is an audio collection formatted for speech recognition tasks. Each sample contains an audio file… See the full description on the dataset page: https://huggingface.co/datasets/sartifyllc/Sartify_ITU_Zindi_Testdataset.anv-data-ke-somali-testsmart-turn-data-v3.2-testTesting dataset for Smart Turn v3.2.
Thank you to the following contributors whose audio samples are included in this dataset:
The Pipecat team
Liva AI: https://www.theliva.ai/
Midcentury: https://www.midcentury.xyz/
MundoAI: https://mundoai.world/
Also, thank you to the following people for the CC-0 background noise sample data which has been used in this dataset:
https://freesound.org/people/4team/sounds/214995/
https://freesound.org/people/tomhannen/sounds/698090/… See the full description on the dataset page: https://huggingface.co/datasets/0x3/smart-turn-data-v3.2-test.Test_Data_filtered_samplesdataset-farma-test
Dataset Card for "dataset-farma-test"
More Information needed
test_data_subsettest_datatest-asr-datasynthetic-data-indonesia_2_4_testaudio-text-pair_train-test-dataset-hindi
dataset contains pairs of audio-text where text are hindi transcriptions and audio are corresponding audio file
structure:
`
DatasetDict({
train: Dataset({
features: ['audio', 'text'],
num_rows: 200
})
test: Dataset({
features: ['audio', 'text'],
num_rows: 60
})
sample row
{'audio': {'path': '/content/drive/MyDrive/sarvam.ai/1.mp3', 'array': array([ 5.85489014e-13, -6.08550428e-13, 5.14475181e-13, ..., -3.15216061e-13, -1.82061695e-13, 0.00000000e+00])… See the full description on the dataset page: https://huggingface.co/datasets/PYD4320/audio-text-pair_train-test-dataset-hindi.Test_Audio_Generate_Dataset
Hinglish Audio Dataset
Generated by Sarvam AI.
dataset1_testtest_dataset_2_metadatasample-dataset-test-neural-nopack
sample-dataset-test-neural-nopack
Speech dataset prepared with Trelis Studio.
Statistics
Metric
Value
Source files
6
Train samples
184
Validation samples
20
Total duration
62.6 minutes
Columns
Column
Type
Description
audio
Audio
Audio segment (16kHz) - speech only, silence stripped via VAD
text
string
Plain transcription (no timestamps) - backwards compatible
text_ts
string
Transcription WITH Whisper timestamp tokens (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/sample-dataset-test-neural-nopack.test_dataset
Dataset Card for "test_dataset"
More Information needed
