CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01danish-foundation-models /icelandic-dynaword 🧨 Icelandic Dynaword Version 0.0.15 (Changelog) Language Icelandic (is, isl) License Openly Licensed, See the respective dataset Models Currently there is no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 39.85M Number of tokens (Llama 3): 2.67B Average document length in tokens (min, max): 66.98 (3, 1.03M) Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dynaword.imagetext-generation100M<n<1B4 likes860 downloads18d agoHugging Face02Aalto-Speech-Synthesis /icelandic_asr Icelandic ASR Collection This repository collects six Icelandic speech corpora in directly loadable Parquet form. Audio is embedded as 16 kHz mono FLAC bytes. The repository is a convenience repackaging: the linked CLARIN-IS records and original dataset repositories remain the canonical sources and should be cited when using the data. No configuration is selected by default. Choose a corpus configuration and, for this large collection, normally choose a split explicitly.… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/icelandic_asr.audioautomatic-speech-recognition1M<n<10M0 likes592 downloads21d agoHugging Face03Frejams /Icelandic-Flan Icelandic FLAN Icelandic instruction-following data, built by pairing licensed, human-written Icelandic texts with deterministic instruction templates. Status 16 sources · 46 tasks · 602,057 rows · 45.6M response characters. Source Register Licence Rows Response chars Share umbodsmadur administrative law — Ombudsman art-9 3,914 9,265,216 20.3% igc_news journalism CC BY 4.0 27,711 8,984,257 19.7% rafbokavefur literary — diacritic restoration over… See the full description on the dataset page: https://huggingface.co/datasets/Frejams/Icelandic-Flan.texttext-generation1M<n<10M0 likes271 downloads28d agoHugging Face04mideind /icelandic-common-crawl-corpus-IC3This is the Icelandic Common Crawl Corpus (IC3). texttext-generation1M<n<10M1 likes233 downloads4y agoHugging Face05danish-foundation-models /icelandic-dyna-instruct 🧨 Icelandic dyna-instruct Version 0.1.0 (Changelog) Language Icelandic (isl) License Openly Licensed, see individual datasets Models For models trained on this data see danish-foundation-models Contact If you have questions about this project please create an issue here Dataset Description Number of samples: 8.11K Number of tokens (Llama 3): 7.09M Average conversation length in tokens (min, max): 874.89 (182, 1.39K) Average number… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dyna-instruct.texttext-generation10K<n<100K1 likes203 downloads24d agoHugging Face06Sigurdur /tsdae-icelandic-iceberttext1M<n<10M0 likes158 downloads1y agoHugging Face07mideind /icelandic_qa_scandevalA question answering dataset for evaluating LLMs' ability to answer Icelandic questions on Icelandic culture and history. The dataset contains 2,000 pairs of questions and answers in Icelandic on the topic of Icelandic culture and history. All pairs were automatically created using GPT-4-turbo and then manually reviewed and augmented. 1,900 pairs were created from Icelandic Wikipedia articles and 100 pairs were created from Icelandic online news, the RÚV subcorpus of the Icelandic Gigaword… See the full description on the dataset page: https://huggingface.co/datasets/mideind/icelandic_qa_scandeval.text1K<n<10K2 likes157 downloads2y agoHugging Face08liu-nlp /icelandic-blimp-single-errortext10K<n<100K0 likes157 downloads1y agoHugging Face09mideind /icelandic_wiki_qaThe dataset is intended as test data and can be used freely as such. If any other use is intended, see OpenAI's terms and conditions on using its output. The dataset contains questions and answers created by GPT-4-turbo, which have been manually reviewed and corrected. text1K<n<10K1 likes107 downloads8mo agoHugging Face10Sigurdur /icelandic-ocr-benchmark Dataset Card for Icelandic OCR Benchmark Dataset Details Dataset Description Icelandic OCR Benchmark is a ground-truth dataset for evaluating OCR accuracy on Icelandic-language documents. It consists of manually transcribed page images with matching layout annotations (text regions, line polygons, baselines) in both ALTO and PAGE XML. Curated by: Sigurdur Haukur Birgisson Language(s): Icelandic (is) License: CC BY-SA 4.0 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Sigurdur/icelandic-ocr-benchmark.imageimage-to-textn<1K1 likes87 downloads10d agoHugging Face11mideind /icelandic-arc-challenge Dataset Card for Icelandic ARC-Challenge This dataset is an Icelandic machine-translated version of the original English ARC-Challenge. textquestion-answering1K<n<10K1 likes82 downloads2y agoHugging Face12mideind /icelandic-winogrande Icelandic WinoGrande dataset This is the Icelandic WinoGrande dataset described in the IceBERT paper https://aclanthology.org/2022.lrec-1.464.pdf . Translation and localization The records were manually translated and localized (skipped if localization was not possible) from English. For the examples which were singlets instead of sentence pairs we added a corresponding sentence. The "translations per se" are not exact since accurately preserving the original semantics is… See the full description on the dataset page: https://huggingface.co/datasets/mideind/icelandic-winogrande.text1K<n<10K2 likes73 downloads2y agoHugging Face13V4ldeLund /Icelandic-Instructtext1K<n<10K0 likes66 downloads29d agoHugging Face14jonasaise /dpo-icelandic-interpreted Dataset Card for DPO Icelandic Interpreted Dataset Description This dataset contains Direct Preference Optimization (DPO) pairs translated and culturally adapted into Icelandic from the original argilla/ultrafeedback-binarized-preferences-cleaned dataset. Origin and Methodology Source Dataset: argilla/ultrafeedback-binarized-preferences-cleaned Model Used for Synthesis: kimi-k3 (via local endpoint) Translation Strategy: The dataset was generated… See the full description on the dataset page: https://huggingface.co/datasets/jonasaise/dpo-icelandic-interpreted.textn<1K0 likes60 downloads1d agoHugging Face15Sigurdur /19th-century-icelandic-letters 19th Century Icelandic Letters OCR Benchmark Handwritten letter dataset from Bréfasafn Árnastofnunar. Dataset Source: Bréfasafn 19. aldar — Árni Magnússon Institute for Icelandic Studies Size: ~1,640 handwritten letters from ~350 writers Images: Full-resolution color scans (3264×2176 px JPG), 0–12 images per letter Text: Diplomatic transcriptions with TEI-like markup, plus cleaned plain text License: CC BY 4.0 Features Column Type… See the full description on the dataset page: https://huggingface.co/datasets/Sigurdur/19th-century-icelandic-letters.image1K<n<10K0 likes57 downloads3mo agoHugging Face16liu-nlp /smol-smoltalk-icelandictext100K<n<1M0 likes56 downloads11mo agoHugging Face17mideind /icelandic-error-corpus-IceECThe Icelandic Error Corpus (IceEC) is a collection of texts in modern Icelandic annotated for mistakes related to spelling, grammar, and other issues. The texts are organized by genre. The current version includes sentences from student essays, online news texts and Wikipedia articles. Sentences within texts in the student essays had to be shuffled due to the license which they were originally published under, but neither the online news texts nor the Wikipedia articles needed to be shuffled.text100K<n<1M1 likes49 downloads4y agoHugging Face18vesteinn /icelandic-parallel-abstracts-corpus-IPACSee https://arxiv.org/abs/2108.05289 text10K<n<100K0 likes44 downloads4y agoHugging Face19fpadovani /goldfish-Dp-icelandic-10mbtext100K<n<1M0 likes42 downloads3mo agoHugging Face20Sigurdur /icelandic-qa-hugi Icelandic question-answering dataset The same dataset as in https://huggingface.co/datasets/Sigurdur/hugi_korkar but the first response has been saved, the rest have been thrown out. The dataset is still not cleaned and may contain question answer pair that is not for all audiences. Author: Sigurdur Haukur Birgisson textquestion-answering100K<n<1M0 likes26 downloads3y agoHugging Face21mideind /icelandic-english-translationtextn<1K0 likes24 downloads3y agoHugging Face22mideind /icelandic-inflection-mediumtextn<1K0 likes23 downloads3y agoHugging Face23saillab /alpaca_icelandic_tacoThis repository contains the dataset used for the TaCo paper. The dataset follows the style outlined in the TaCo paper, as follows: { "instruction": "instruction in xx", "input": "input in xx", "output": "Instruction in English: instruction in en , Response in English: response in en , Response in xx: response in xx " } Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_icelandic_taco.text10K<n<100K0 likes23 downloads2y agoHugging Face24Sigurdur /grade-school-math-icelandic GSM8K-Icelandic GSM8K-Icelandic is an Icelandic translation of the GSM8K dataset created by OpenAI. The translation was performed using the Google Translate API by Sigurdur Haukur Birgisson. Dataset Description Dataset Summary GSM8K (Grade School Math 8K) consists of 8,500 high-quality grade school math word problems. This Icelandic version maintains the same structure as the original dataset but provides all content in Icelandic, making it accessible for… See the full description on the dataset page: https://huggingface.co/datasets/Sigurdur/grade-school-math-icelandic.text1K<n<10K0 likes22 downloads1y agoHugging Face25shunyalabs /icelandic-speech-datasetaudio1K<n<10K0 likes21 downloads1y agoHugging Face26liu-nlp /icelandic-blimp-nouns-def-to-indf-experimentaltext1K<n<10K0 likes19 downloads1y agoHugging Face27liu-nlp /icelandic-blimp-verbs-ind-to-sbjv-experimentaltext1K<n<10K0 likes18 downloads1y agoHugging Face28mbruton /icelandic_encrypted_HistCiph Dataset Card for HistCiph — Icelandic Dataset Description Dataset Summary The Icelandic subset of HistCiph is part of the first publicly available multilingual collection of historically grounded plaintext–ciphertext pairs for classical homophonic substitution ciphers. It pairs diachronically balanced historical Icelandic plaintext with independently generated homophonic substitution keys and controlled transcription noise, producing four distinct ciphertext… See the full description on the dataset page: https://huggingface.co/datasets/mbruton/icelandic_encrypted_HistCiph.tabular100K<n<1M0 likes18 downloads5mo agoHugging Face29liu-nlp /icelandic-blimp-verbs-3rd-to-2nd-pers-experimentaltext1K<n<10K0 likes17 downloads1y agoHugging Face30oddadmix /synthetic_pages_icelandic_ltr_v1image10K<n<100K1 likes17 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.