datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
oercommons-v1-optimized
OERCommons v1 Optimized
Authors: Junjie Wang and Yuhan SunHosted by: PIN TeamDataset: pin-team/oercommons-v1-optimized
OERCommons v1 Optimized is a provenance-preserving multimodal pretraining corpus built on the OERCommons subset of The Common Pile v0.1, which serves as its upstream data and licensing baseline. We extend it with full-page recovery, canonical Markdown, ordered image/PDF/link metadata, conservative corrections, and integrity evidence.
At a glance… See the full description on the dataset page: https://huggingface.co/datasets/pin-team/oercommons-v1-optimized.spro-optimized-prompts-fullOptimized-MoveToRedBallexams-ocr-optimized-testbhl-240p-60v-optimized
bhl-240p-60v-optimized
OCR-model comparison published with ocrscout. Each row pairs one source page with one model's normalized output.
Summary
Pages: 224
Models: 3
Rows: 718
Pages with errors: 0
Mean disagreement: 0.163
Median disagreement: 0.092
Source images embedded: yes
Generated: 2026-06-27 10:27 UTC by ocrscout v0.1.0
Per-model metrics
Model
Format
Pages OK
Pages errored
Total tokens
Mean tokens/pg
Mean s/pg
Mean prepare s
Mean chars… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/bhl-240p-60v-optimized.VAEimages_optimized
