CoolFace
Datasetpublic

FBK-MT/fama-data

Dataset Description, Collection, and Source The FAMA training data is the collection of English and Italian datasets for automatic speech recognition (ASR) and speech translation (ST) used to train the FAMA models family. The ASR section of FAMA is derived from the MOSEL data collection, including the automatic transcripts obtained with Whisper and available in the HuggingFace MOSEL Dataset. The ASR is further augmented with automatically transcribed speech from the… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/fama-data.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
2likes196downloads
13 commits on main
ff0a66f1y ago

Update README.md

spapi
3316dff1y ago

Rename train_mls-it-en.tsv to train_mls_it-en.tsv

spapi
6a01b111y ago

Add FAMA data

spapi
9df35ed1y ago

Add MT inference code to the main README

spapi
26048571y ago

Correct YouTube-Commons README with the correct call to the scripts

spapi
e5933f11y ago

Add segment-ytc.py script

spapi
d2b34bb1y ago

Add speech_only.py script

spapi
bbf4f691y ago

Add YouTube-Commons ids and Silero logs

spapi
2d86fb71y ago

Create SplitAudioUsingSileroLog.pl

spapi
7c755741y ago

Add main FAMA data README

spapi
213fb231y ago

Add ratio filtering script for ST

spapi
779a0411y ago

Add YouTube-Commons README

spapi
8e9f10e1y ago

initial commit

spapi