CoolFace
Datasetpublic

Zeldeo/transatlantic-voice-archive_distille

Distillation brucemacd/transatlantic-voice-archive Dataset ASR distillé via Cohere Transcribe. Source : brucemacd/transatlantic-voice-archive Modèle ASR : cohere-transcribe Langue ASR : en Exemples : 1427 (dataset source intégral) Colonnes : audio — clip audio (16 kHz) transcription_base — référence brute du dataset source transcription_cohere — hypothèse Cohere brute langue_accent — langue / accent détecté wer, cer — métriques item (normalisation training_v3, textes stockés… See the full description on the dataset page: https://huggingface.co/datasets/Zeldeo/transatlantic-voice-archive_distille.

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes19downloads
Dataset Card

Distillation brucemacd/transatlantic-voice-archive

Dataset ASR distillé via Cohere Transcribe.

  • —Source : brucemacd/transatlantic-voice-archive
  • —Modèle ASR : cohere-transcribe
  • —Langue ASR : en
  • —Exemples : 1427 (dataset source intégral)

Colonnes :

  • —audio — clip audio (16 kHz)
  • —transcription_base — référence brute du dataset source
  • —transcription_cohere — hypothèse Cohere brute
  • —langue_accent — langue / accent détecté
  • —wer, cer — métriques item (normalisation training_v3, textes stockés bruts)
  • —source — identifiant dataset source

Les transcriptions sont brutes ; WER/CER sont calculés sur texte normalisé.