datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
faroese-dynaword
🧨 Faroese Dynaword
Version
0.0.7 (Changelog)
Language
Faroese (fo, fao)
License
Openly Licensed, See the respective dataset
Models
Currently there are no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 405.81K
Number of tokens (Llama 3): 45.40M
Average document length in tokens (min, max): 111.87 (2, 109.50K)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/faroese-dynaword.Faroese-Handwritten-OCR
Faroese Handwritten OCR
Public draft, version 0.1: an alignment pilot with 16 line-image/text pairs from one
historical Faroese manuscript page. No rows are verified benchmark ground truth.
The page is image 3 of D IV – Ániasar táttur, held by Landsbókasavnið
(National Library of the Faroe Islands) and digitized on HandRit. The manuscript
is associated with the scribe Jóhan Hendrik Schrøter (1842–1911). Proposed
reference text is aligned from Eivind Weyhe's scholarly edition of… See the full description on the dataset page: https://huggingface.co/datasets/V4ldeLund/Faroese-Handwritten-OCR.
