datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scandi-reddit
Dataset Card for ScandiReddit
Dataset Summary
ScandiReddit is a filtered and post-processed corpus consisting of comments from Reddit.
All Reddit comments from December 2005 up until October 2022 were downloaded through PushShift, after which these were filtered based on the FastText language detection model. Any comment which was classified as Danish (da), Norwegian (no), Swedish (sv) or Icelandic (is) with a confidence score above 70% was kept.
The resulting comments… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/scandi-reddit.task131_scan_long_text_generation_action_command_long
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task131_scan_long_text_generation_action_command_long
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task131_scan_long_text_generation_action_command_long.task127_scan_long_text_generation_action_command_all
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task127_scan_long_text_generation_action_command_all
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task127_scan_long_text_generation_action_command_all.task128_scan_structured_text_generation_command_action_short
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task128_scan_structured_text_generation_command_action_short
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task128_scan_structured_text_generation_command_action_short.scandinavian-educational-annotations
Scandinavian Educational Annotations
Created using a CommonCrawl dump (April 2024), and annotations with Gemini 1.5 Flash.
task129_scan_long_text_generation_action_command_short
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task129_scan_long_text_generation_action_command_short
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task129_scan_long_text_generation_action_command_short.scandinavian-linguistic-annotations
Scandinavian Educational Annotations
Created using a CommonCrawl dump (April 2024), and annotations with Gemini 1.5 Flash.
scandi-wikiScandiWiki is a parsed and deduplicated version of the Danish, Norwegian Bokmål,
Norwegian Nynorsk, Swedish, Icelandic and Faroese Wikipedia corpora, as of January
2023.task130_scan_structured_text_generation_command_action_long
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task130_scan_structured_text_generation_command_action_long
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task130_scan_structured_text_generation_command_action_long.FinDoc-Pro-Scanned
FinDoc-Pro-Scanned Dataset
Dataset Description
This dataset contains OCR and document parsing results for scanned financial documents using four different models:
PaddleOCR-VL - 6 tasks per image (ocr, table, formula, chart, spotting, seal)
Qwen3-VL - 3 prompts per image (ocr_json, qwenvl_html, qwenvl_markdown)
Bonsai-2-27B - 2 active prompts per image (spotting_json, qwenvl_markdown) with thinking traces; partial ocr_markdown/qwenvl_html
Docling-Heron-101 -… See the full description on the dataset page: https://huggingface.co/datasets/dhruvilp/FinDoc-Pro-Scanned.linical_inference_debt_scanner_v0.1Clinical Inference Debt Scanner
PurposeDetect when a clinical plan relies on stacked assumptions rather than evidence.
You receive:
evidence_signals
a narrative_chain
a planned_action
You output:
inference_debt_level0 to 3
debt_itemthe single most dangerous leap
paydown_stepthe corrective step that restores evidence grounding
Debt scale0 none1 minor2 moderate3 severe
Scoring
debt_level_scoregraded by distance from gold
debt_item_similaritytoken overlap similarity… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/linical_inference_debt_scanner_v0.1.scandi-human-instruct
scandi-human-instruct
Human–LLM instruction tuning dataset for the four main Scandinavian languages. Rows are normalized into a chat-friendly messages format and tagged with the detected language using NbAiLab/nb-nordic-lid.
Dataset size
Total rows: 9,642
Per-language counts:
Norwegian Bokmål (nb): 4,833
Danish (da): 2,633
Swedish (sv): 2,042
Norwegian Nynorsk (nn): 134
Sources
allenai/WildChat-4.8M
ministere-culture/comparia-conversations (both… See the full description on the dataset page: https://huggingface.co/datasets/V4ldeLund/scandi-human-instruct.task126_scan_structured_text_generation_command_action_all
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task126_scan_structured_text_generation_command_action_all
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task126_scan_structured_text_generation_command_action_all.scandi-translated-instruct
scandi-translate
A Scandinavian instruction‑tuning dataset built from machine‑translated instruction/response pairs. It unifies Danish, Swedish, Norwegian Bokmål and Norwegian Nynorsk data into a chat-friendly messages schema.
Dataset size
Total rows: 1,252,683
Per-language counts measured during build:
Danish (da): 377,413
Swedish (sv): 332,558
Norwegian Nynorsk (nn): 270,772
Norwegian Bokmål (nb): 268,875
By source:
V4ldeLund/da-translated-instruct… See the full description on the dataset page: https://huggingface.co/datasets/V4ldeLund/scandi-translated-instruct.scandi-reddit-filtered
Dataset Card for ScandiRedditFiltered
Dataset Summary
ScandiRedditFiltered is manually filtered and post-processed corpus consisting of comments from ScandiReddit.
The intended use of the filtered sentences is for Text-To-Speech (TTS) models.
Supported Tasks and Leaderboards
Training language models is the intended task for this dataset. No leaderboard is active at this point.
Languages
The dataset is available in Danish (da).
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/scandi-reddit-filtered.ganjoor-ipa-scansion
Ganjoor Persian Classical Poetry — Meter & Phonemic Transliteration
A corpus of 124,404 classical Persian poems collected via the Ganjoor API,
enriched with two things every poem now has:
Prosodic meter (ʿarūż / vazn) — the metrical feet and a binary scansion for every poem,
including the ~21% that Ganjoor left unlabeled (reconstructed here from the Persian feet).
Phonemic transliteration — Latin and IPA for every hemistich, produced by the
Homo-GE2PE grapheme-to-phoneme model.… See the full description on the dataset page: https://huggingface.co/datasets/nafisehNik/ganjoor-ipa-scansion.
