datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
icelandic-dynaword
🧨 Icelandic Dynaword
Version
0.0.15 (Changelog)
Language
Icelandic (is, isl)
License
Openly Licensed, See the respective dataset
Models
Currently there is no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 39.85M
Number of tokens (Llama 3): 2.67B
Average document length in tokens (min, max): 66.98 (3, 1.03M)
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dynaword.Icelandic-Flan
Icelandic FLAN
Icelandic instruction-following data, built by pairing licensed, human-written Icelandic
texts with deterministic instruction templates.
Status
16 sources · 46 tasks · 602,057 rows · 45.6M response characters.
Source
Register
Licence
Rows
Response chars
Share
umbodsmadur
administrative law — Ombudsman
art-9
3,914
9,265,216
20.3%
igc_news
journalism
CC BY 4.0
27,711
8,984,257
19.7%
rafbokavefur
literary — diacritic restoration over… See the full description on the dataset page: https://huggingface.co/datasets/Frejams/Icelandic-Flan.icelandic-common-crawl-corpus-IC3This is the Icelandic Common Crawl Corpus (IC3).
icelandic-dyna-instruct
🧨 Icelandic dyna-instruct
Version
0.1.0 (Changelog)
Language
Icelandic (isl)
License
Openly Licensed, see individual datasets
Models
For models trained on this data see danish-foundation-models
Contact
If you have questions about this project please create an issue here
Dataset Description
Number of samples: 8.11K
Number of tokens (Llama 3): 7.09M
Average conversation length in tokens (min, max): 874.89 (182, 1.39K)
Average number… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dyna-instruct.Roleplay-Icelandic
RolePlay-Icelandic
Roleplay-Icelandic Dataset is a dataset for roleplaying in the Icelandic language for Large Language Model.
The base dataset is GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API.
For more information and other language datasets for roleplay, see this github repo.
For… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Icelandic.
