Frejams/faroese-flan
Faroese FLAN Faroese instruction-following data, built by pairing licensed or public-domain, human-written Faroese texts with deterministic instruction templates. The sibling of the Icelandic collection, with the same row schema and the same release checks. Status 13 sources · 31 tasks · 1,000,543 rows · 34.0M response characters. Source Register Licence Rows Response chars Share logir consolidated law public-domain-fo-p9 29,178 14,798,010 43.6%… See the full description on the dataset page: https://huggingface.co/datasets/Frejams/faroese-flan.
Faroese FLAN
Faroese instruction-following data, built by pairing licensed or public-domain, human-written Faroese texts with deterministic instruction templates. The sibling of the Icelandic collection, with the same row schema and the same release checks.
Status
13 sources · 31 tasks · 1,000,543 rows · 34.0M response characters.
Rows and response characters rank the collection almost inversely. The three ravnlex tasks are 68% of all rows and 23% of the response characters; logir's diacritic restoration is 2% of rows and 43% of the response characters. A user who samples uniformly by row trains almost entirely on one-word phonetic and word-class answers.
Much of what the model learns to produce is spelling, not composition. 43% of the response characters come from two diacritic-restoration tasks, where the response is the prompt's own text with its accented letters restored. The model chooses how to spell, not what to say. Legal Faroese is therefore the largest register by characters and still mostly a spelling exercise. The Faroese a model must produce from a different input is sprotin's translated sentences (14.2%) and the composed prose of lum, logting_spurningar, fo_wikipedia, fpsc and the gazette (7.2%).
Registers with nothing in them: journalism, dialogue, academic prose, popular science, statistics, finance, technical and instructional text. sprotin is everyday conversational Faroese in sentence form, translated, not dialogue. Every Faroese news outlet is under ordinary copyright, and the two transcribed conversation corpora are non-commercial or restricted.
Loading
from datasets import load_dataset
ds = load_dataset("<repo>") # everything
ds = load_dataset("<repo>", "lum") # one source at a timeOne parquet file per source, so a source can be reviewed, filtered or replaced on its own. The default config concatenates them.
Split on `source_id`, not randomly — see the field notes below.
Dataset Structure
Each row is a single-turn conversation:
{
"id": "lum__opinion_to_title__26-09428_p1141",
"messages": [
{"role": "user", "content": "Tímar til námsfrøðilig tiltøk í frítíðarskúla Innleiðandi skal viðmerkjast, at sambært § 6, stk. 1 í umboðsmanslógini (...)"},
{"role": "assistant", "content": "Álit um avgerð hjá Tórshavnar kommunu um tímar til námsfrøðilig tiltøk í frítíðarskúla og noktað innlit (...)"}
],
"source": "lum",
"subsource": "nidurstoda",
"task_name": "opinion_to_title",
"template_id": "7",
"license": "public-domain-fo-p9",
"source_id": "26-09428_p1141",
"source_url": null
}Fields
id— unique within this dataset.messages— the conversation. Always exactly oneuserturn and oneassistantturn.source— which underlying dataset the material came from.subsource— the division within it: the document type, the paradigm class, the source language of a translation. What you filter or audit on. Info_wiktionaryit is the language of the prompt word — 91 languages — so a prompt there is not necessarily Faroese.task_name— what kind of question the row poses.template_id— which of the task's instruction phrasings produced this row. Recorded so that the evenness of template use is checkable from the data itself, and so phrasings can be held out at evaluation time.license— the licence of the underlying text, per row, so a user with narrower licence requirements than ours can filter rather than take the whole thing.source_id— the identifier of the source document. Rows sharing asource_idderive from the same document, so split by `source_id`, not randomly, or your evaluation set will contain text the model saw in training.source_url— the row's own public page, where the source has one;nullotherwise. Filled forlogir,fpsc,fo_wikipedia,fo_wikisource,fo_wiktionary,kunngerdaportalurandlogting_spurningar.
The instruction templates live in `templates/` and are part of this repo.
Dataset Construction Process
The same five steps for every source.
- Read the source at a pinned revision. Openly licensed or public domain by statute, already structured — no scraping of copyrighted sites and no crawl-derived text.
- Recover the pair. Every source already contains the (input, target) pair; the build finds it rather than inventing it. A title and the act it names, a headword and its inflected forms, a stanza and the same stanza with its accents removed.
- Filter. Drop what is not a task: pairs too short to be one, targets recoverable by copying the input, and targets that could not be written from the input at all. See Validity.
- Template. Attach one of ten hand-written Faroese instructions, chosen by hashing the source identifier so the assignment is stable and uncorrelated with corpus order.
- Validate and write parquet.
No machine translation and no language model is involved anywhere in the construction.
Sources
logir — Faroese consolidated law
Lógasavn (logir.fo), the official database of Faroese law maintained by the Ministry of Justice: statutes, regulations, circulars and guidance in consolidated form, with numbered sections, chapters and named headings. The 5,262 rules written in Faroese; Danish instruments extended to the Faroes are not included.
The largest task restores the accented letters to a section of statute; the two heading tasks ask for the heading the publisher gave a paragraph or a chapter. The text is the version standing in the database, not as enacted, and roughly half the rules are no longer in force.
Details in `docs/logir.md`
ravnlex — Faroese pronunciation and word class
RAVNlex, the full-form Faroese lexicon of the Ravnur Project at the University of the Faroe Islands: each word form with a phonetic transcription in SAMPA and a PAROLE part-of-speech tag. The only pronunciation data in the collection.
`word_to_wordclass` has nine answers over 224,028 rows and about 65% of them are `navnorð` — a true fact about the lexicon, shipped in full so that you can subset it rather than have it sampled for you.
Details in `docs/ravnlex.md`
sprotin — English into Faroese, everyday sentences
Sprotin's English–Faroese sentence bank: 126,500 short sentences translated from English into Faroese by human translators and published by Sprotin, the Faroese dictionary publisher, in 2021. The English side is largely drawn from Tatoeba, in its didactic house style — Tom appears in about a quarter of the prompts. The only English→Faroese task, and the only everyday conversational Faroese in the response position. The Faroese is translated, not natively authored; median 35 characters.
Details in `docs/sprotin.md`
islex_fo — Icelandic into Faroese
The Faroese side of ISLEX, the Icelandic–Scandinavian multilingual dictionary of the Árni Magnússon Institute: Icelandic headwords, phrases and example sentences with the Faroese equivalents written by the Faroese editorial team. Every pair is by a lexicographer; the response is always Faroese.
Icelandic and Faroese are cognate, so much of the term task is a respelling (silfur → silvur); the sentence task is where the two languages diverge.
Details in `docs/islex_fo.md`
fmd — Faroese inflection
Føroyski bendingargrunnurin, the Faroese Morphological Database (bendingar.fo), a joint publication of the Árni Magnússon Institute and Fróðskaparsetur Føroya. Nouns through four cases in two numbers with and without the suffixed article, verb principal parts, adjective comparison and the mediopassive. Every response is one or more inflected forms in the order the database's own paradigm tables print them.
Details in `docs/fmd.md`
lum — opinions of the Faroese Ombudsman
The published case archive of Løgtingsins umboðsmaður, the Ombudsman of the Faroese parliament (lum.fo/savn). Every response is text the office wrote itself: reasoned conclusions, headnotes, case titles, and the subject and administrative-law categories it files cases under. The largest body of composed Faroese prose in the collection. Conclusions that quote commercially published Danish legal literature are excluded, because the Ombudsman's statutory public-domain status does not extend to third-party works quoted inside an opinion.
Details in `docs/lum.md`
logting_spurningar — written questions to Faroese ministers, and their answers
Members of the Løgting put questions to ministers in writing and the minister answers in writing (www.logting.fo). The prompt is the question document with the member's commentary; the response is the ministry's answer — dated chronologies, tonnages, budget figures, statute citations. The longest composed institutional Faroese in the collection, median 2,446 characters. This release covers one instrument (§ 52a), 213 cases from 2008–2013; the parliament's index holds 3,021 cases back to 1992.
Details in `docs/logting_spurningar.md`
fpsc — Løgtingið parliamentary speech, paired with the agenda
The Faroese Parliament Speech Corpus, segmented speeches from the floor of Løgtingið published by the University of the Faroe Islands. Each row gives a speech transcript and asks which item on the day's order paper it belongs to; the response is the Løgting's own agenda heading, a legislative noun phrase rather than a class label — 602 distinct answers.
The transcripts in the prompt are automatic speech recognition output, as the corpus publishes them. No speaker is identified.
Details in `docs/fpsc.md`
fo_wikipedia — article to lead paragraph
Faroese Wikipedia. Each row gives an article without its opening paragraph, with its infobox rendered into the prompt, and asks for that paragraph. The only summarisation data in Faroese with an open licence. Articles without an infobox are excluded, because a Wikipedia lead routinely states facts that live only there.
Details in `docs/fo_wikipedia.md`
kunngerdaportalur — Kunngerðablaðið, the official gazette
Kunngerðablaðið, the official gazette of the Faroe Islands, published by Løgmansskrivstovan. Every response is a field the gazette publishes alongside the instrument: the act's official title, and — for an amending act — the publisher's statement of what it changes. As-published text, where logir is consolidated.
Details in `docs/kunngerdaportalur.md`
fo_wikisource — restoring the accents to a stanza of verse
Faroese Wikisource, a volunteer-transcribed library of 19th- and early-20th-century Faroese poetry, hymns and ballads — Nólsoyar Páll, the Djurhuus brothers, Effersøe, Patursson. Each row gives a stanza with its accented letters flattened to bare Latin and asks for the stanza as printed. The only literary Faroese in the collection; some texts use an older orthography, reproduced as published.
Details in `docs/fo_wikisource.md`
fo_wiktionary — Faroese Wiktionary
The Faroese-language edition of the volunteer dictionary, from the Wikimedia dump of 1 August 2026. The prompt is a word in one of 91 other languages, or a Faroese headword; the response is always Faroese. 40.7% of the translation rows answer with one of 54 Bible book names, because that is what the wiki's contributors wrote; nothing is removed to hide it, and subsource carries the prompt language so you can subset.
Details in `docs/fo_wiktionary.md`
gerdabokur — Løgtingið minute books, bill → committee
The official minute books of the Løgting, 1992–2009. A bill's title paired with the standing committee the minute records it as referred to — the parliament's own act, not an annotation. Eighteen committees, several of them of their period and no longer existing; subsource is the committee.
Details in `docs/gerdabokur.md`
Not in this release
Waitlist — licence-clear and waiting to be built
- The rest of the Løgtingið questions and minute books —
logting_spurningarships one instrument and 213 of 3,021 indexed cases, andgerdabokurthe 1992–2009 sittings; the remainder of both is public domain under § 9 and not yet harvested. - Faroese Pronunciation Dictionaries (Ravnur, CC BY 4.0) — Central and East Faroese as separate word-level dictionaries. The dialect axis only; general pronunciation is already
ravnlex. - Ravnur BLARK text grids (CC BY 4.0) — the spoken-form ↔ written-form pairs for numbers, times, place names and licence plates, which Ravnur wrote itself. Not the read-aloud news text.
Held back — and what each is waiting on
Built or buildable, and deliberately not here.
Validity
Automated checks, enforced at build time by src/scripts/validate.py and shared with the Icelandic collection:
- Template diversity — no single phrasing may dominate a task.
- Template resolution — every
template_idpresent in the data resolves to a template in `templates/`. - No trivial pairs — the response must not be recoverable by copying a span of the prompt, and the input must be substantially longer than the target, except in tasks declared as transformations, where the response is the prompt's own text corrected.
- Answerability — the target must be writable from the input, measured as the share of the target's content words with no counterpart in the input, against a ceiling declared per task. This is why
fo_wikipediaexcludes articles without an infobox: the lead states facts the body does not. - No duplicate responses — the same response text may not appear under more than one prompt, within a source or across sources, except in tasks declared as lexical mappings.
- Source URLs — every
source_urlpresent is well formed. - Licence values — every row's
licenseis one of the four values listed under License.
Every (sub-source × task) cell was read by hand before shipping: four rows per cell, ten if any of the four failed, and the cell was fixed or excluded at two failures in ten.
The Faroese instruction phrasings have not been reviewed by a native speaker. They are written in Faroese rather than translated, using each publisher's own vocabulary for its material, but whether they read naturally is unverified. The responses are unaffected: every response is text a Faroese institution or writer published.
Limitations
- 43% of the response characters are determined form. In
logir.diacritic_restorationandfo_wikisource.verse_diacritic_restorationthe response is the prompt with its accented letters restored, so a model learns Faroese orthography from them and not composition. The composed prose islum,logting_spurningar,fo_wikipedia,fpscandkunngerdaportalur;sprotinis translation. - Rows are not a guide to content.
ravnlexis 68% of rows with one-word answers; filtersource != "ravnlex"and 322,557 rows remain. - `sprotin`'s Faroese is translated from English, and the English is Tatoeba-style didactic sentences. It teaches everyday vocabulary and sentence patterns; it is not natively composed Faroese, and 2,020 Faroese sentences answer more than one English prompt.
- `ravnlex.word_to_wordclass` has nine answers and about 65% are one label.
fo_wiktionary.foreign_word_to_fo_wordanswers 40.7% of the time with one of 54 Bible book names. Both are shipped in full and stated here rather than resampled. - `logir` is consolidated, not as enacted, and about half its rules are no longer in force. No prompt claims a text is current law.
kunngerdaportaluris the as-published text, so the same act can appear in both in different wording. - `logting_spurningar` and `gerdabokur` are partial harvests — one question instrument, 213 cases from 2008–2013, and the older half of the minute books. Roughly one answer in eight in the question source is unavailable as text (scans without a text layer).
- `gerdabokur` answers with a committee name over 18 labels, several of them historical; two rows carry the minute clerk's spelling of a committee as published (
Skttanevndin,Mentanevndin). - `fpsc` prompts are automatic transcripts as the corpus publishes them, with recognition errors; the response is the parliament's own agenda heading.
- `islex_fo`'s term task is often a respelling. About 15% of its pairs are 80% or more character-identical to the Icelandic; this is a property of two cognate languages.
- Older orthography survives in `fo_wikisource`. 76 rows (7.8%) use
öwhere modern Faroese hasø; both sides of each row keep the spelling the publisher printed, so no row asks the model to choose between the two. - Invisible and control characters survive from the source pages. Zero-width spaces and joiners in
lum,fo_wikipediaandfo_wiktionaryprompts and in onefo_wikipediaresponse — under twenty rows; and form feeds (page breaks) in 218logting_spurningarrows, 155 in the response, where a PDF page ended. Strip\fbefore use. - Personal data. Five sources are public-authority material outside copyright (
logir,kunngerdaportalur,lum,logting_spurningar,gerdabokur). The Ombudsman anonymises complainants at source (ein borgari,A,B); regulations reproducing EU sanctions lists that name individuals are excluded fromlogirandkunngerdaportalur; no shipped row carries a personal identifier, and every street address found is an institution's. MPs, ministers and signing officials are named as the parliament and the gazette name them. A person identified by circumstance alone is not detectable by any check run here. - No journalism, no dialogue, no academic or technical prose. Every Faroese news outlet is under ordinary copyright and the transcribed conversation corpora are non-commercial or restricted, so these registers are absent rather than thin.
License
This collection grants no single licence over its contents. Each row carries its own, in the `license` column, and that is the licence that governs that row.
Attribution is due per source, not to this collection. The wording each rights holder requires is in that source's document under `docs/`, linked from its entry above. Rows are template-wrapped, which is a modification of the source text.
The repository metadata declares cc-by-sa-4.0 because that is the strictest licence present, so a user who takes the whole file and complies with it complies with everything. If share-alike is a problem, filter it out:
ds = ds.filter(lambda r: r["license"] != "cc-by-sa-4.0") # 843,934 rowsCreators and Funders
Created at the Alexandra Institute as part of the EU Horizon project TrustLLM (grant agreement number 101135671).
