datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
icelandic-dynaword
🧨 Icelandic Dynaword
Version
0.0.15 (Changelog)
Language
Icelandic (is, isl)
License
Openly Licensed, See the respective dataset
Models
Currently there is no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 39.85M
Number of tokens (Llama 3): 2.67B
Average document length in tokens (min, max): 66.98 (3, 1.03M)
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dynaword.icelandic_asr
Icelandic ASR Collection
This repository collects six Icelandic speech corpora in directly loadable
Parquet form. Audio is embedded as 16 kHz mono FLAC bytes. The repository is a
convenience repackaging: the linked CLARIN-IS records and original dataset
repositories remain the canonical sources and should be cited when using the
data.
No configuration is selected by default. Choose a corpus configuration and,
for this large collection, normally choose a split explicitly.… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/icelandic_asr.Icelandic-Flan
Icelandic FLAN
Icelandic instruction-following data, built by pairing licensed, human-written Icelandic
texts with deterministic instruction templates.
Status
16 sources · 46 tasks · 602,057 rows · 45.6M response characters.
Source
Register
Licence
Rows
Response chars
Share
umbodsmadur
administrative law — Ombudsman
art-9
3,914
9,265,216
20.3%
igc_news
journalism
CC BY 4.0
27,711
8,984,257
19.7%
rafbokavefur
literary — diacritic restoration over… See the full description on the dataset page: https://huggingface.co/datasets/Frejams/Icelandic-Flan.icelandic-common-crawl-corpus-IC3This is the Icelandic Common Crawl Corpus (IC3).
icelandic-dyna-instruct
🧨 Icelandic dyna-instruct
Version
0.1.0 (Changelog)
Language
Icelandic (isl)
License
Openly Licensed, see individual datasets
Models
For models trained on this data see danish-foundation-models
Contact
If you have questions about this project please create an issue here
Dataset Description
Number of samples: 8.11K
Number of tokens (Llama 3): 7.09M
Average conversation length in tokens (min, max): 874.89 (182, 1.39K)
Average number… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dyna-instruct.tsdae-icelandic-iceberticelandic_qa_scandevalA question answering dataset for evaluating LLMs' ability to answer Icelandic questions on Icelandic culture and history.
The dataset contains 2,000 pairs of questions and answers in Icelandic on the topic of Icelandic culture and history. All pairs were automatically created using GPT-4-turbo and then manually reviewed and augmented. 1,900 pairs were created from Icelandic Wikipedia articles and 100 pairs were created from Icelandic online news, the RÚV subcorpus of the Icelandic Gigaword… See the full description on the dataset page: https://huggingface.co/datasets/mideind/icelandic_qa_scandeval.icelandic-blimp-single-erroricelandic_wiki_qaThe dataset is intended as test data and can be used freely as such. If any other use is intended, see OpenAI's terms and conditions on using its output. The dataset contains questions and answers created by GPT-4-turbo, which have been manually reviewed and corrected.
icelandic-ocr-benchmark
Dataset Card for Icelandic OCR Benchmark
Dataset Details
Dataset Description
Icelandic OCR Benchmark is a ground-truth dataset for evaluating OCR accuracy on
Icelandic-language documents. It consists of manually transcribed page images with
matching layout annotations (text regions, line polygons, baselines) in both ALTO
and PAGE XML.
Curated by: Sigurdur Haukur Birgisson
Language(s): Icelandic (is)
License: CC BY-SA 4.0
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Sigurdur/icelandic-ocr-benchmark.icelandic-arc-challenge
Dataset Card for Icelandic ARC-Challenge
This dataset is an Icelandic machine-translated version of the original English ARC-Challenge.
icelandic-winogrande
Icelandic WinoGrande dataset
This is the Icelandic WinoGrande dataset described in the IceBERT paper https://aclanthology.org/2022.lrec-1.464.pdf .
Translation and localization
The records were manually translated and localized (skipped if localization was not possible) from English.
For the examples which were singlets instead of sentence pairs we added a corresponding sentence.
The "translations per se" are not exact since accurately preserving the original semantics is… See the full description on the dataset page: https://huggingface.co/datasets/mideind/icelandic-winogrande.Icelandic-Instructdpo-icelandic-interpreted
Dataset Card for DPO Icelandic Interpreted
Dataset Description
This dataset contains Direct Preference Optimization (DPO) pairs translated and culturally adapted into Icelandic from the original argilla/ultrafeedback-binarized-preferences-cleaned dataset.
Origin and Methodology
Source Dataset: argilla/ultrafeedback-binarized-preferences-cleaned
Model Used for Synthesis: kimi-k3 (via local endpoint)
Translation Strategy: The dataset was generated… See the full description on the dataset page: https://huggingface.co/datasets/jonasaise/dpo-icelandic-interpreted.19th-century-icelandic-letters
19th Century Icelandic Letters OCR Benchmark
Handwritten letter dataset from Bréfasafn Árnastofnunar.
Dataset
Source: Bréfasafn 19. aldar — Árni Magnússon Institute for Icelandic Studies
Size: ~1,640 handwritten letters from ~350 writers
Images: Full-resolution color scans (3264×2176 px JPG), 0–12 images per letter
Text: Diplomatic transcriptions with TEI-like markup, plus cleaned plain text
License: CC BY 4.0
Features
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/Sigurdur/19th-century-icelandic-letters.smol-smoltalk-icelandicicelandic-error-corpus-IceECThe Icelandic Error Corpus (IceEC) is a collection of texts in modern Icelandic annotated for mistakes related to spelling, grammar, and other issues. The texts are organized by genre. The current version includes sentences from student essays, online news texts and Wikipedia articles.
Sentences within texts in the student essays had to be shuffled due to the license which they were originally published under, but neither the online news texts nor the Wikipedia articles needed to be shuffled.icelandic-parallel-abstracts-corpus-IPACSee https://arxiv.org/abs/2108.05289
goldfish-Dp-icelandic-10mbIcelandic-flanicelandic-ner-MIM-GOLD-NERThis Icelandic named entity (NE) corpus, MIM-GOLD-NER, is a version of the MIM-GOLD corpus tagged for NEs. Over 48 thousand NEs are tagged in this corpus of one million tokens, which can be used for training named entity recognizers for Icelandic.
The MIM-GOLD-NER corpus was developed at Reykjavik University in 2018–2020, funded by the Strategic Research and Development Programme for Language Technology (LT). Two LT students were in charge of the corpus annotation and of training named entity recognizers using machine learning methods.
A semi-automatic approach was used for annotating the corpus. Lists of Icelandic person names, location names, and company names were compiled and used for extracting and classifying as many named entities as possible. Regular expressions were then used to find certain numerical entities in the corpus. After this automatic pre-processing step, the whole corpus was reviewed manually to correct any errors. The corpus is tagged for eight named entity types:
PERSON – names of humans, animals and other beings, real or fictional.
LOCATION – names of locations, real or fictional, i.e. buildings, street and place names, both real and fictional. All geographical and geopolitical entities such as cities, countries, counties and regions, as well as planet names and other outer space entities.
ORGANIZATION – companies and other organizations, public or private, real or fictional. Schools, churches, swimming pools, community centers, musical groups, other affiliations.
MISCELLANEOUS – proper nouns that don’t belong to the previous three categories, such as products, books and movie titles, events, such as wars, sports tournaments, festivals, concerts, etc.
DATE – absolute temporal units of a full day or longer, such as days, months, years, centuries, both written numerically and alphabetically.
TIME – absolute temporal units shorter than a full day, such as seconds, minutes, or hours, both written numerically and alphabetically.
MONEY – exact monetary amounts in any currency, both written numerically and alphabetically.
PERCENT – percentages, both written numerically and alphabetically
MIM-GOLD-NER is intended for training of named entity recognizers for Icelandic. It is in the CoNLL format, and the position of each token within the NE is marked using the BIO tagging format. The corpus can be used in its entirety or by training on subsets of the text types that best fit the intended domain.
The Named Entity Corpus corpus is distributed with the same special user license as MIM-GOLD, which is based on the MIM license, since the texts in MIM-GOLD were sampled from the MIM corpus.icelandic-qa-hugi
Icelandic question-answering dataset
The same dataset as in https://huggingface.co/datasets/Sigurdur/hugi_korkar but the first response has been saved, the rest have been thrown out.
The dataset is still not cleaned and may contain question answer pair that is not for all audiences.
Author: Sigurdur Haukur Birgisson
icelandic-english-translationicelandic-inflection-mediumalpaca_icelandic_tacoThis repository contains the dataset used for the TaCo paper.
The dataset follows the style outlined in the TaCo paper, as follows:
{
"instruction": "instruction in xx",
"input": "input in xx",
"output": "Instruction in English: instruction in en ,
Response in English: response in en ,
Response in xx: response in xx "
}
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_icelandic_taco.icelandic-qa-NQiI\grade-school-math-icelandic
GSM8K-Icelandic
GSM8K-Icelandic is an Icelandic translation of the GSM8K dataset created by OpenAI. The translation was performed using the Google Translate API by Sigurdur Haukur Birgisson.
Dataset Description
Dataset Summary
GSM8K (Grade School Math 8K) consists of 8,500 high-quality grade school math word problems. This Icelandic version maintains the same structure as the original dataset but provides all content in Icelandic, making it accessible for… See the full description on the dataset page: https://huggingface.co/datasets/Sigurdur/grade-school-math-icelandic.icelandic-speech-datasetart-lang-uniform-icelandic-beforeicelandic-blimp-nouns-def-to-indf-experimental
