V4ldeLund/faroese-taks-fo
TAKS (taks_fo) Status: Standalone website collection. It is not part of Faroese Dynaword. About this data Tax and customs laws, circulars, guidance, forms, HTML pages, PDFs, DOCX/CSV files and linked Kærustovnur decisions. Source: TAKS Size Measure Value Documents 1,178 Characters 13,463,572 Tokens 5,127,975 Possible official documents 205 Tokens in possible official documents 765,561 File type Documents HTML 852… See the full description on the dataset page: https://huggingface.co/datasets/V4ldeLund/faroese-taks-fo.
TAKS (taks_fo)
Status: Standalone website collection. It is not part of Faroese Dynaword.
About this data
Tax and customs laws, circulars, guidance, forms, HTML pages, PDFs, DOCX/CSV files and linked Kærustovnur decisions.
Source: TAKS
Size
The token count uses the AI-Sweden-Models/Llama-3-8B-instruct tokenizer. It counts all collected text, including other languages and repeated website text.
Files
rows.parquetcontains the collected documents.parser/contains the crawler, site settings and Python requirements.
Crawl details
- Crawl started: 2026-08-16
- Crawl completed: 2026-08-16
The crawler used public sitemaps and links from the website. It stayed within the websites and file hosts listed in parser/sites.csv and followed robots.txt.
Text was taken from web pages, text-based PDFs, Word files and CSV files. Scanned PDFs and images were not read with OCR. Very short pages and exact duplicate texts were left out.
Columns
Rights
Some laws, decisions, notices and similar public documents may be outside copyright under section 9 of the Faroese Copyright Act.
likely_official_document only means that the title or URL matched a list of words. It is not a legal classification.
Note for this site
Laws, circulars, official tax guidance and final appeal decisions are the clearest candidates. Use the Logir version of a law when available. Service pages, forms and outside attachments may still be copyrighted.
Known limits
- The crawl may not cover every page or old document.
- The text has not been checked by hand or filtered by language.
- Menus, headers and other repeated website text may remain.
- Text from PDFs may appear in the wrong order.
- Scanned PDFs and text inside images are missing.
- Some pages may contain personal information.
Parser
The crawler used for this website is included in the parser folder. The copy of sites.csv contains only this website.
cd parser
python -m pip install -r requirements.txt
python scrape.py --site taks_fo