datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kb-books
open-rdl-books
Dataset Description
Language
dan, dansk, Danish
License
Public Domain, cc0-1.0
Dataset Summary
Documents from the Royal Danish Library published between 1750 and 1930.
The dataset has each page of each document in image and text format. The text was extracted with OCR.
The documents (books of various genres) were obtained from the library. The dataset was assembled to make these public domain Danish texts more accessible.… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/kb-books.eno-newspapers-enriched
Danish Historical Newspaper Articles Dataset (enriched)
This dataset contains approximately 4.9 million Danish historical newspaper articles (1666–1850) with document embeddings and assigned fictionality tags, providing a comprehensive resource for studying Danish language, culture, and history through primary journalistic sources.
Dataset Details
Dataset Description
This dataset comprises digitized newspaper articles from Danish newspapers… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/eno-newspapers-enriched.CHC-Bench
Dataset Card for "CHC-Bench"
🌐 Homepage | 🤗 MAP-CC | 🤗 CHC-Bench | 🤗 CT-LLM | 📖 arXiv | GitHub
Introduction
We propose a well-chosen multidisciplinary Chinese Hard Case Benchmark (CHC-Bench). We collect the problems from various sources e.g. ziya, gaokao, and CIF-Bench to form hard-case Chinese instructions understanding and following evaluation benchmark (CHC-Bench in short) The categories of problems in CHC-Bench include writing, humanity and history, science… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/CHC-Bench.eno-embs-old-news
Danish Historical Newspaper Articles Dataset
This dataset contains approximately 4.9 million Danish historical newspaper articles (1666–1850) with document embeddings, providing a comprehensive resource for studying Danish language, culture, and history through primary journalistic sources.
Dataset Details
Dataset Description
This dataset comprises digitized newspaper articles from Danish newspapers, featuring full-text content along with metadata including… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/eno-embs-old-news.periphery-aviser-e5grundtvigs-works
Grundtvig's Works (Grundtvigs Værker)
Grundtvig's Works is a comprehensive digital humanities dataset containing the complete collected writings of
Nicolai Frederik Severin Grundtvig (1783-1872) was one of Denmark’s most influential cultural and intellectual figures.
As a critical edition, it includes editorial commentary by philologists and is continually updated.
The project is scheduled for completion in 2030 and will comprise 1,000 individual works spanning 35,000 pages. The… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/grundtvigs-works.dacy-data
Combined CDT, DDT and DaNE dataset
This dataset merges the Danish UD treebank (DDT), Danish Dependency Treebank (DaNE) and Copenhagen Dependency Treebank (CDT). The DDT contains part-of-speech, dependency and morphology tags and has been further annotated for entities by Alexandra Institute in DaNE. DDT is based on CDT to assign tags consistent with the universal dependencies project (UD). However, this process split the data in DDT into singular sentences, therefore models… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/dacy-data.StreamVLN-ScanQA-SQA3D-Datawikidata_benchmarkingThis dataset was compiled for the purpose of finetuning models in the context of benchmarking for art historical research.
Images scraped from Wikimedia Commons via Wikidata; metadata scraped from
Wikidata (CC0). Image licenses vary per file (predominantly public domain,
some CC-BY-SA) see the license_short_name / license_url columns in the
parquet files for the exact terms of each individual image, and the
commons file page for full details.
northern-landscape-painting
License
This dataset was compiled for the purpose of art historical research on Scandinavian landscape paintings.
Images scraped from Wikimedia Commons via Wikidata; metadata scraped from
Wikidata (CC0). Image licenses vary per file (predominantly public domain,
some CC-BY-SA). See the license_short_name / license_url columns in
the parquet files for the exact terms of each individual image, and the
Commons file page for full details.
dansk-ner
Dataset Summary
DANSK: Danish Annotations for NLP Specific TasKs is a dataset consisting of texts from multiple domains, sampled from the Danish GigaWord Corpus (DAGW).
The dataset was created to fill in the gap of Danish NLP datasets from different domains, that are required for training models that generalize across domains. The Named-Entity annotations are moreover fine-grained and have a similar form to that of OntoNotes v5, which significantly broadens the use cases of the… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/dansk-ner.africa-worldbank-main-cooking-fuel-charcoal-of-households-sg-cok-chco-zs
Main cooking fuel: charcoal (% of households) | Africa (World Bank — Gender Statistics) | Africa (World Bank)
Size category: n<1K - Formats: parquet - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-worldbank-main-cooking-fuel-charcoal-of-households-sg-cok-chco-zs.CHR_detection_trees
Motif Segmentation in Paintings
Segmentation masks and metadata for trees detected in a corpus of paintings, produced with SAM3 as part of research in the golden matrix project at the Center for Humanities Computing Aarhus (chcaa).
Dataset Description
This dataset supports a studies of landscape and tree motifs in 19th-century Northern European paintings. It covers 4,727 paintings drawn from museum and Wikidata sources, with tree instances detected using [SAM3]… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/CHR_detection_trees.finetuning-landscape-paintingThis dataset was compiled for the purpose of finetuning models in the context of benchmarking for art historical research.
Images scraped from Wikimedia Commons via Wikidata; metadata scraped from
Wikidata (CC0). Image licenses vary per file (predominantly public domain,
some CC-BY-SA) see the license_short_name / license_url columns in the
parquet files for the exact terms of each individual image, and the
commons file page for full details.
dagw-word-frequencies
Dataset Card for DAGW Word Frequencies
Paper: Derczynski, L., Ciosici, M. R., Baglini, R., Christiansen, M. H., Dalsgaard, J. A., Fusaroli, R., ... & Varab, D. (2021). The Danish Gigaword Corpus. In Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa) (pp. 413-421).
Point of Contact: Kenneth Enevoldsen (Kennethcenevoldsen (at) gmail (dot) com )
This is a list of word frequencies derived from the Danish Gigaword (collected before 2022-22-01).
These… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/dagw-word-frequencies.dagw-word-frequencies-normalized-by-domain
Dataset Card for DAGW Word Frequencies (normalized)
Paper: Derczynski, L., Ciosici, M. R., Baglini, R., Christiansen, M. H., Dalsgaard, J. A., Fusaroli, R., ... & Varab, D. (2021). The Danish Gigaword Corpus. In Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa) (pp. 413-421).
Point of Contact: Kenneth Enevoldsen (Kennethcenevoldsen (at) gmail (dot) com )
This is a list of word frequencies derived from the Danish Gigaword (collected before… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/dagw-word-frequencies-normalized-by-domain.mlrbench-tasksThis repository contains the benchmark dataset of MLR-Bench.
We collect 201 tasks from ICLR/NeurIPS/ICML workshops over the past three years. The followings record the metadata of our collection.
Workshop without an official website or deleted
icml2024_fminwild
neurips2024_attrib_late
neurips2024_gsai
neurips2024_rlfm
iclr2023_ai4abm
iclr2023_ml4iot
iclr2023_mldd
iclr2023_NeSy_GeMs
neurips2023_new_in_ml
neurips2023_ai4mat
icml2023_esfomo
Non-general workshops… See the full description on the dataset page: https://huggingface.co/datasets/chchenhui/mlrbench-tasks.dagw-word-frequencies-by-domain
Dataset Card for DAGW Word Frequencies (by domain)
Paper: Derczynski, L., Ciosici, M. R., Baglini, R., Christiansen, M. H., Dalsgaard, J. A., Fusaroli, R., ... & Varab, D. (2021). The Danish Gigaword Corpus. In Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa) (pp. 413-421).
Point of Contact: Kenneth Enevoldsen (Kennethcenevoldsen (at) gmail (dot) com )
This is a list of word frequencies derived from the Danish Gigaword (collected before… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/dagw-word-frequencies-by-domain.africa-worldbank-wbl-supportive-framework-childcare-score-scale-0-100-gd-wbl-chc-sfr-t
WBL: Supportive Framework, Childcare, Score (scale 0-100) | Africa (World Bank — Gender Statistics) | Africa (World Bank)
Size category: n<1K - Formats: parquet - Sector: demographics_social - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-worldbank-wbl-supportive-framework-childcare-score-scale-0-100-gd-wbl-chc-sfr-t.dagw-word-frequencies-by-domain-with-pos-tags
Dataset Card for DAGW Word Frequencies (with pos tags)
Paper: Derczynski, L., Ciosici, M. R., Baglini, R., Christiansen, M. H., Dalsgaard, J. A., Fusaroli, R., ... & Varab, D. (2021). The Danish Gigaword Corpus. In Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa) (pp. 413-421).
Point of Contact: Kenneth Enevoldsen (Kennethcenevoldsen (at) gmail (dot) com )
This is a list of word frequencies derived from the Danish Gigaword (collected before… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/dagw-word-frequencies-by-domain-with-pos-tags.africa-worldbank-wbl-enforcement-perceptions-childcare-score-scale-0-100-gd-wbl-chc-enf-t
WBL: Enforcement Perceptions, Childcare, Score (scale 0-100) | Africa (World Bank — Gender Statistics) | Africa (World Bank)
Size category: n<1K - Formats: parquet - Sector: demographics_social - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-worldbank-wbl-enforcement-perceptions-childcare-score-scale-0-100-gd-wbl-chc-enf-t.TACK_Tunnel_Data
TACK Tunnel Data (TTD): A Benchmark Dataset for Deep Learning-Based Defect Detection in Tunnels
Tunnels are essential elements of transportation infrastructure, but are increasingly affected by ageing and deterioration mechanisms such as cracking. Regular inspections are required to ensure their safety, yet traditional manual procedures are time-consuming, subjective, and costly. Recent advances in mobile mapping systems and Deep Learning (DL) enable automated visual inspections.… See the full description on the dataset page: https://huggingface.co/datasets/chckch/TACK_Tunnel_Data.Press-and-Plot
Press&Plot: Curated Danish 19th-Century Stories & Serial Fiction (v1.0)
Short description:A curated collection of 29 Danish newspaper stories (1816–1832), including single-part and multi-part fiction, manually inspected, cleaned, and categorized for research use. The dataset is a growing resource.
Dowloading the dataset
# using python
from datasets import load_dataset
ds = load_dataset("chcaa/press-and-plot", split="train")
# if you want it as a pandas DataFrame:
df =… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/Press-and-Plot.fiction4sentiment
Dataset description
A dataset of literary sentences human-annotated for valence (0-10) used for developing multilingual SA
🔬 Data
No. texts
No. annotations
No. words
Period
Fairy tales
3
772
18,597
1837-1847
Hymns
65
2,026
12,798
1798-1873
Prose
1
1,923
30,279
1952
Poetry
40
1,579
11,576
1965
This is the Fiction4 dataset of literary texts, spanning 109 individual texts across 4 genres and two languages (English and Danish) in the 19th and 20th… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/fiction4sentiment.danish-book-ads
Books and Journals in Danish Newspaper Advertisements (1800–1850)
A dataset of extracted and cleaned book titles and author names from Danish newspaper advertisements published between 1800 and 1850. The records were automatically extracted from digitized newspapers using a combination of rule-based methods, named-entity recognition, and a trained category classifier.
Dataset Description
Summary
The dataset contains 80,938 advertisement records drawn from nine… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/danish-book-ads.africa-worldbank-wbl-legal-framework-childcare-score-scale-0-100-gd-wbl-chc-law-t
WBL: Legal Framework, Childcare, Score (scale 0-100) | Africa (World Bank — Gender Statistics) | Africa (World Bank)
Size category: n<1K - Formats: parquet - Sector: governance_security - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-worldbank-wbl-legal-framework-childcare-score-scale-0-100-gd-wbl-chc-law-t.memo-canonical-novels
Dataset description
source_datasets:
The corpus was created and made available by Jens Bjerring-Hansen and Philip Diderichsen, Dorte Haltrup Hansen, June 2023, see: https://huggingface.co/datasets/MiMe-MeMo/Corpus-v1.1
- Here, we make a more accessible, annotated version available.
Additional tags:
CE Canon: Cultural/Educational Canon, referring to novels whose titles are included in the Cultural Canon, or whose author is included in the Educational Canon.
LEX Canon:… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/memo-canonical-novels.shinchan-chat-datasetnaturalistic_social_norms_alignment
Naturalistic Social Norms Alignment
A dataset of 3,023 real-world social dilemmas in Danish, extracted from the popular radio show Sara og Monopolet. Each dilemma comes with reference solutions derived from a panel of three guests, making the dataset suitable for evaluating social norm alignment of LLMs and humans in naturalistic, open-ended conversations.
Paper: Naturalistic measure of social norms alignment
Code & Framework: GitHub repository
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/naturalistic_social_norms_alignment.smk_canon_paintings
