datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stt-cefc-fr-test
CEFC-Orfeo FR — long-form oral test mirror
Mirror non-officiel du Corpus d'Études du Français Contemporain (CEFC)
agrégé par le projet Orfeo (ANR), tel que distribué sur le portail
projet-orfeo.fr (release 13).
Long-form : 1 row = 1 fichier audio entier (30-60 min en moyenne).
12 sous-corpus oraux du français contemporain, 303 heures au total,
901 fichiers. Idéal pour bench Whisper / Canary en conditions réelles
(chunked decoding, dérive temporelle, multi-locuteurs).… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-cefc-fr-test.orfeo-cefc-frautoeval-staging-eval-project-a0b7f8d6-f4e4-45b3-a9ae-cefcb10962b0-128122
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Summarization
Model: autoevaluate/summarization-not-evaluated
Dataset: autoevaluate/xsum-sample
Config: autoevaluate--xsum-sample
Split: test
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @lewtun for evaluating this model.
