CoolFace
Datasetpublic

thibault-baneras-roux/HATS-fr

🗃️ HATS Dataset HATS (Human Assessed Transcription Side-by-Side) is a data set for French 🇫🇷 which consists of 1,000 triplets (reference, hypothesis A, hypothesis B) and 7,150 human choice annotated by 143 subjects 🫂 Their objective was to select, given a textual reference, which of two erroneous hypotheses is the best. Curated by: Thibault Bañeras-Roux, Richard Dufour, Jane Wottawa, Mickael Rouvier, Teva Merlin Funded by: Agence Nationale de la Recherche - DIETS… See the full description on the dataset page: https://huggingface.co/datasets/thibault-baneras-roux/HATS-fr.

sourceHugging Facemitupdated 3mo agoView on Hugging Face
1likes11downloads
Dataset Card

🗃️ HATS Dataset

HATS (Human Assessed Transcription Side-by-Side) is a data set for French 🇫🇷 which consists of 1,000 triplets (reference, hypothesis A, hypothesis B) and 7,150 human choice annotated by 143 subjects 🫂 Their objective was to select, given a textual reference, which of two erroneous hypotheses is the best.

<!-- Provide a longer summary of what this dataset is. -->

  • —Curated by: Thibault Bañeras-Roux, Richard Dufour, Jane Wottawa, Mickael Rouvier, Teva Merlin
  • —Funded by: Agence Nationale de la Recherche - DIETS project (contract ANR-20-CE23-0005)
  • —Language: French 🇫🇷
  • —License: MIT

🧑‍🏫 Metric-Evaluator

This toolkit calculates the percentage of time a metric agrees with human judgments. Recognizing that human judgments can vary, instances arise where no consensus exists, and choices may be influenced by randomness 🎲 To filter the dataset based on consensus cases, utilize the certitude argument. This parameter represents the percentage of humans who selected the same hypothesis (set it to 1 when 100% of subjects make the same choice, and 0.7 when 70% of subjects choose the same hypothesis)."

📊 Results

Metrics100%70%Full
Word Error Rate63%53%49%
Character Error Rate77%64%60%
BERTScore CamemBERT-large80%68%65%
SemDist CamemBERT-large80%71%67%
SemDist Sentence CamemBERT-large90%78%73%
Phoneme Error Rate80%69%64%

To add the results of your metric, contact me at thibault [le dot] roux [le at] idiap.ch ✉️

📜 Citation

<!-- If there is a paper or blog post introducing the dataset, the APA and Bibtex information for that should go in this section. -->

@inproceedings{baneras2023hats,
  title={HATS: An Open data set Integrating Human Perception Applied to the Evaluation of Automatic Speech Recognition Metrics},
  author={Ba{\~n}eras-Roux, Thibault and Wottawa, Jane and Rouvier, Mickael and Merlin, Teva and Dufour, Richard},
  booktitle={Text, Speech and Dialogue 2023 - Interspeech Satellite},
  year={2023}
}