datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
red_ace_asr_error_detection_and_correction
RED-ACE
Dataset Summary
This dataset can be used to train and evaluate ASR Error Detection or Correction models. It was introduced in the RED-ACE paper (Gekhman et al, 2022).
The dataset contains ASR outputs on the LibriSpeech corpus (Panayotov et al., 2015) with annotated transcription errors.
Dataset Details
The LibriSpeech corpus was decoded using Google Cloud Speech-to-Text API, with the default and video models.
The word-level confidence was enabled… See the full description on the dataset page: https://huggingface.co/datasets/google/red_ace_asr_error_detection_and_correction.instruction-error-detection-en-id
instruction-error-detection-en-id
Description
instruction-error-detection-en-id is a bilingual benchmark dataset for detecting, explaining, and correcting flawed or ambiguous instructions.
The dataset focuses on instruction robustness by introducing graded difficulty levels and partially incorrect instructions. It is designed to evaluate how well models can reason about contradictions, ambiguities, and incomplete constraints before responding.
Dataset Fields… See the full description on the dataset page: https://huggingface.co/datasets/YosepMulia/instruction-error-detection-en-id.
