datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ocr-test-challenging-2multilingual_ocrmultilingual_ocr_2tokenizer_traintachiwin_voice_raw
Tachiwin Indigenous Languages of Mexico Voice Pretrain Dataset (RAW)
This dataset collects a large amount of speech hours in indigenous languages of Mexico recorded from 23 public radio stations that broadcast to the indigenous communities.
How it was obtained
After the raw recording of all transmissions, the music and inaudible sound were removed and then the samples of pure speech were classified by language. However, as there are no existing speech language classifiers… See the full description on the dataset page: https://huggingface.co/datasets/tachiwin/tachiwin_voice_raw.multilingual_ocr_llm_2
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/tachiwin/multilingual_ocr_llm_2.ocr-test-challenging-3tiyat-ground-pretrain-m1024tiyat-ground-pretraintachiwin_multilingual_audiomultilingual_ocr_llmmultilingual_ocr_erniesdklanguage_utilsocr-test-challenging-rescanned
ocr-test-challenging-rescanned
Re-scanned pages from the original tachiwin/ocr-test-challenging-3 dataset.
200 pages were printed on physical paper and scanned back to digital images,
providing a real-world OCR degradation benchmark.
Columns are identical to the original dataset:
page_id — unique page identifier
pdf_hash — hash of the original PDF document
page_number — page number within the original document
text — original ground-truth text
uncommon_char_score — score… See the full description on the dataset page: https://huggingface.co/datasets/tachiwin/ocr-test-challenging-rescanned.tachiwin_translatemultilingual_parallel_corpustachiwin_multilingual_audio_curricularmultilingual_language_identificatortachiwin_biblespretraintachiwin_constituciones
Dataset de la Constitución Política de los Estados Unidos Mexicanos
Traducida a varias lenguas originarias de México
Como parte del proyecto Tachiwin
multilingual_translatortachiwin_pretrainpretrain_rawmultilingual_finetune
