datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
maltese-ocr-corpus
Real Maltese OCR Paragraph Pairs
Paragraph-level (image, text) pairs from real, digitally-typeset
Maltese PDFs. Built to close the synthetic-only gap in Maltese OCR
training data - see the accompanying paper (LV-ROVER-MLT, DocEng 2026)
for context: the paper and the corpus-building scripts
(package_for_hf.py, align_pdf_paragraphs.py, under
experiments/neural_resume/corpus/) are at
github.com/adamd1985/doceng2026.
The frozen competition submission (the Tesseract LV-ROVER-MLT… See the full description on the dataset page: https://huggingface.co/datasets/radmada/maltese-ocr-corpus.demo-maltesersThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha-stationary",
"total_episodes": 90,
"total_frames": 32336,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:90"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ccop/demo-maltesers.imnet1k_Maltese_dog_Maltese_terrier_Maltese
