datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SLT-Task1-Post-ASR-Text-Correction
Dataset Name: Pilot dataset for Multi-domain ASR corrections
Description
This dataset is a pilot version of a larger dataset for automatic speech recognition (ASR) corrections across multiple domains.
It contains paired hypotheses and corrected transcriptions for various ASR tasks consolidated from PeacefulData/HyPoradise-v0
Structure
Data Split
The dataset is divided into training and test splits:
Training Data: 281,082 entries
Approximately 6,255… See the full description on the dataset page: https://huggingface.co/datasets/GenSEC-LLM/SLT-Task1-Post-ASR-Text-Correction.vn-text-correction-0001
Dataset Card for Vietnamese Text Correction Dataset
Dataset Description
This dataset contains Vietnamese text pairs for training and evaluating text correction models. Each example consists of an erroneous text and its corrected version, making it ideal for:
Grammar correction
Spelling correction
Text normalization
Language model fine-tuning
Dataset Summary
Language: Vietnamese (vi)
Format: Text correction pairs
Size: ~4.0M examples across… See the full description on the dataset page: https://huggingface.co/datasets/LakoreAI/vn-text-correction-0001.text-correction-enwiki_text_correctiontext-correctiontext-correction-base
