dataset-processing
latin-asr-post-processing-dataset
Latin ASR Post-Processing Dataset
A sequence-labeling dataset built for fine-tuning BERT-style models (e.g., latin-bert) on Inverse Text Normalization (ITN) — restoring capitalization and punctuation on raw, lowercased Latin text (such as njand/wav2vec2-xls-r-latin ASR outputs).
Compiled from 2,141 files in the CLTK Latin Library and augmented with transcripts from the njand/llpsi-speech-dataset (currently private). Cleaned and transformed through a specialized classical Latin… See the full description on the dataset page: https://huggingface.co/datasets/njand/latin-asr-post-processing-dataset.rubber_duck_dataset_smoothed_20fps_256x256_xbox_NO_PROCESSINGThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "UR5",
"total_episodes": 1,
"total_frames": 161,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 20,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nsimonato25/rubber_duck_dataset_smoothed_20fps_256x256_xbox_NO_PROCESSING.string-matching-and-processing-dataset-alpacadataset_processingASR_Post-processing-datasetDigital-Signal-Processing-QA-Dataset
