CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01chcaa /kb-books open-rdl-books Dataset Description Language dan, dansk, Danish License Public Domain, cc0-1.0 Dataset Summary Documents from the Royal Danish Library published between 1750 and 1930. The dataset has each page of each document in image and text format. The text was extracted with OCR. The documents (books of various genres) were obtained from the library. The dataset was assembled to make these public domain Danish texts more accessible.… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/kb-books.image1M<n<10M4 likes20k downloads10mo agoHugging Face02chcaa /eno-newspapers-enriched Danish Historical Newspaper Articles Dataset (enriched) This dataset contains approximately 4.9 million Danish historical newspaper articles (1666–1850) with document embeddings and assigned fictionality tags, providing a comprehensive resource for studying Danish language, culture, and history through primary journalistic sources. Dataset Details Dataset Description This dataset comprises digitized newspaper articles from Danish newspapers… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/eno-newspapers-enriched.tabular1M<n<10M1 likes283 downloads3mo agoHugging Face03m-a-p /CHC-Bench Dataset Card for "CHC-Bench" 🌐 Homepage | 🤗 MAP-CC | 🤗 CHC-Bench | 🤗 CT-LLM | 📖 arXiv | GitHub Introduction We propose a well-chosen multidisciplinary Chinese Hard Case Benchmark (CHC-Bench). We collect the problems from various sources e.g. ziya, gaokao, and CIF-Bench to form hard-case Chinese instructions understanding and following evaluation benchmark (CHC-Bench in short) The categories of problems in CHC-Bench include writing, humanity and history, science… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/CHC-Bench.textn<1K9 likes243 downloads2y agoHugging Face04chcaa /eno-embs-old-news Danish Historical Newspaper Articles Dataset This dataset contains approximately 4.9 million Danish historical newspaper articles (1666–1850) with document embeddings, providing a comprehensive resource for studying Danish language, culture, and history through primary journalistic sources. Dataset Details Dataset Description This dataset comprises digitized newspaper articles from Danish newspapers, featuring full-text content along with metadata including… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/eno-embs-old-news.texttext-classification1M<n<10M3 likes190 downloads4mo agoHugging Face05chcaa /periphery-aviser-e5tabular1M<n<10M0 likes152 downloads1y agoHugging Face06chcaa /grundtvigs-works Grundtvig's Works (Grundtvigs Værker) Grundtvig's Works is a comprehensive digital humanities dataset containing the complete collected writings of Nicolai Frederik Severin Grundtvig (1783-1872) was one of Denmark’s most influential cultural and intellectual figures. As a critical edition, it includes editorial commentary by philologists and is continually updated. The project is scheduled for completion in 2030 and will comprise 1,000 individual works spanning 35,000 pages. The… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/grundtvigs-works.textn<1K4 likes150 downloads1y agoHugging Face07chcaa /dacy-data Combined CDT, DDT and DaNE dataset This dataset merges the Danish UD treebank (DDT), Danish Dependency Treebank (DaNE) and Copenhagen Dependency Treebank (CDT). The DDT contains part-of-speech, dependency and morphology tags and has been further annotated for entities by Alexandra Institute in DaNE. DDT is based on CDT to assign tags consistent with the universal dependencies project (UD). However, this process split the data in DDT into singular sentences, therefore models… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/dacy-data.texttoken-classification1K<n<10K3 likes136 downloads9d agoHugging Face08chchnii /StreamVLN-ScanQA-SQA3D-Datatext10K<n<100K2 likes126 downloads1y agoHugging Face09chcaa /wikidata_benchmarkingThis dataset was compiled for the purpose of finetuning models in the context of benchmarking for art historical research. Images scraped from Wikimedia Commons via Wikidata; metadata scraped from Wikidata (CC0). Image licenses vary per file (predominantly public domain, some CC-BY-SA) see the license_short_name / license_url columns in the parquet files for the exact terms of each individual image, and the commons file page for full details. image1K<n<10K0 likes121 downloads29d agoHugging Face10chcaa /northern-landscape-painting License This dataset was compiled for the purpose of art historical research on Scandinavian landscape paintings. Images scraped from Wikimedia Commons via Wikidata; metadata scraped from Wikidata (CC0). Image licenses vary per file (predominantly public domain, some CC-BY-SA). See the license_short_name / license_url columns in the parquet files for the exact terms of each individual image, and the Commons file page for full details. image1K<n<10K0 likes119 downloads1mo agoHugging Face11chcaa /dansk-ner Dataset Summary DANSK: Danish Annotations for NLP Specific TasKs is a dataset consisting of texts from multiple domains, sampled from the Danish GigaWord Corpus (DAGW). The dataset was created to fill in the gap of Danish NLP datasets from different domains, that are required for training models that generalize across domains. The Named-Entity annotations are moreover fine-grained and have a similar form to that of OntoNotes v5, which significantly broadens the use cases of the… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/dansk-ner.texttoken-classification10K<n<100K3 likes98 downloads2y agoHugging Face12electricsheepafrica /africa-worldbank-main-cooking-fuel-charcoal-of-households-sg-cok-chco-zs Main cooking fuel: charcoal (% of households) | Africa (World Bank — Gender Statistics) | Africa (World Bank) Size category: n<1K - Formats: parquet - Sector: energy - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-worldbank-main-cooking-fuel-charcoal-of-households-sg-cok-chco-zs.tabulartabular-classificationn<1K0 likes95 downloads1mo agoHugging Face13chcaa /CHR_detection_trees Motif Segmentation in Paintings Segmentation masks and metadata for trees detected in a corpus of paintings, produced with SAM3 as part of research in the golden matrix project at the Center for Humanities Computing Aarhus (chcaa). Dataset Description This dataset supports a studies of landscape and tree motifs in 19th-century Northern European paintings. It covers 4,727 paintings drawn from museum and Wikidata sources, with tree instances detected using [SAM3]… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/CHR_detection_trees.image1K<n<10K0 likes74 downloads2mo agoHugging Face14chcaa /finetuning-landscape-paintingThis dataset was compiled for the purpose of finetuning models in the context of benchmarking for art historical research. Images scraped from Wikimedia Commons via Wikidata; metadata scraped from Wikidata (CC0). Image licenses vary per file (predominantly public domain, some CC-BY-SA) see the license_short_name / license_url columns in the parquet files for the exact terms of each individual image, and the commons file page for full details. tabularn<1K0 likes56 downloads2mo agoHugging Face15chcaa /dagw-word-frequencies Dataset Card for DAGW Word Frequencies Paper: Derczynski, L., Ciosici, M. R., Baglini, R., Christiansen, M. H., Dalsgaard, J. A., Fusaroli, R., ... & Varab, D. (2021). The Danish Gigaword Corpus. In Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa) (pp. 413-421). Point of Contact: Kenneth Enevoldsen (Kennethcenevoldsen (at) gmail (dot) com ) This is a list of word frequencies derived from the Danish Gigaword (collected before 2022-22-01). These… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/dagw-word-frequencies.tabular10M<n<100M0 likes48 downloads4y agoHugging Face16chcaa /dagw-word-frequencies-normalized-by-domain Dataset Card for DAGW Word Frequencies (normalized) Paper: Derczynski, L., Ciosici, M. R., Baglini, R., Christiansen, M. H., Dalsgaard, J. A., Fusaroli, R., ... & Varab, D. (2021). The Danish Gigaword Corpus. In Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa) (pp. 413-421). Point of Contact: Kenneth Enevoldsen (Kennethcenevoldsen (at) gmail (dot) com ) This is a list of word frequencies derived from the Danish Gigaword (collected before… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/dagw-word-frequencies-normalized-by-domain.tabular10M<n<100M0 likes48 downloads4y agoHugging Face17chchenhui /mlrbench-tasksThis repository contains the benchmark dataset of MLR-Bench. We collect 201 tasks from ICLR/NeurIPS/ICML workshops over the past three years. The followings record the metadata of our collection. Workshop without an official website or deleted icml2024_fminwild neurips2024_attrib_late neurips2024_gsai neurips2024_rlfm iclr2023_ai4abm iclr2023_ml4iot iclr2023_mldd iclr2023_NeSy_GeMs neurips2023_new_in_ml neurips2023_ai4mat icml2023_esfomo Non-general workshops… See the full description on the dataset page: https://huggingface.co/datasets/chchenhui/mlrbench-tasks.textn<1K1 likes44 downloads1y agoHugging Face18chcaa /dagw-word-frequencies-by-domain Dataset Card for DAGW Word Frequencies (by domain) Paper: Derczynski, L., Ciosici, M. R., Baglini, R., Christiansen, M. H., Dalsgaard, J. A., Fusaroli, R., ... & Varab, D. (2021). The Danish Gigaword Corpus. In Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa) (pp. 413-421). Point of Contact: Kenneth Enevoldsen (Kennethcenevoldsen (at) gmail (dot) com ) This is a list of word frequencies derived from the Danish Gigaword (collected before… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/dagw-word-frequencies-by-domain.tabular10M<n<100M0 likes35 downloads4y agoHugging Face19electricsheepafrica /africa-worldbank-wbl-supportive-framework-childcare-score-scale-0-100-gd-wbl-chc-sfr-t WBL: Supportive Framework, Childcare, Score (scale 0-100) | Africa (World Bank — Gender Statistics) | Africa (World Bank) Size category: n<1K - Formats: parquet - Sector: demographics_social - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-worldbank-wbl-supportive-framework-childcare-score-scale-0-100-gd-wbl-chc-sfr-t.tabulartabular-classificationn<1K0 likes23 downloads1mo agoHugging Face20chcaa /dagw-word-frequencies-by-domain-with-pos-tags Dataset Card for DAGW Word Frequencies (with pos tags) Paper: Derczynski, L., Ciosici, M. R., Baglini, R., Christiansen, M. H., Dalsgaard, J. A., Fusaroli, R., ... & Varab, D. (2021). The Danish Gigaword Corpus. In Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa) (pp. 413-421). Point of Contact: Kenneth Enevoldsen (Kennethcenevoldsen (at) gmail (dot) com ) This is a list of word frequencies derived from the Danish Gigaword (collected before… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/dagw-word-frequencies-by-domain-with-pos-tags.tabular10M<n<100M0 likes22 downloads4y agoHugging Face21electricsheepafrica /africa-worldbank-wbl-enforcement-perceptions-childcare-score-scale-0-100-gd-wbl-chc-enf-t WBL: Enforcement Perceptions, Childcare, Score (scale 0-100) | Africa (World Bank — Gender Statistics) | Africa (World Bank) Size category: n<1K - Formats: parquet - Sector: demographics_social - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-worldbank-wbl-enforcement-perceptions-childcare-score-scale-0-100-gd-wbl-chc-enf-t.tabulartabular-classificationn<1K0 likes22 downloads1mo agoHugging Face22chckch /TACK_Tunnel_Data TACK Tunnel Data (TTD): A Benchmark Dataset for Deep Learning-Based Defect Detection in Tunnels Tunnels are essential elements of transportation infrastructure, but are increasingly affected by ageing and deterioration mechanisms such as cracking. Regular inspections are required to ensure their safety, yet traditional manual procedures are time-consuming, subjective, and costly. Recent advances in mobile mapping systems and Deep Learning (DL) enable automated visual inspections.… See the full description on the dataset page: https://huggingface.co/datasets/chckch/TACK_Tunnel_Data.image1K<n<10K0 likes20 downloads7mo agoHugging Face23chcaa /Press-and-Plot Press&Plot: Curated Danish 19th-Century Stories & Serial Fiction (v1.0) Short description:A curated collection of 29 Danish newspaper stories (1816–1832), including single-part and multi-part fiction, manually inspected, cleaned, and categorized for research use. The dataset is a growing resource. Dowloading the dataset # using python from datasets import load_dataset ds = load_dataset("chcaa/press-and-plot", split="train") # if you want it as a pandas DataFrame: df =… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/Press-and-Plot.tabularn<1K1 likes19 downloads11mo agoHugging Face24chcaa /fiction4sentiment Dataset description A dataset of literary sentences human-annotated for valence (0-10) used for developing multilingual SA 🔬 Data No. texts No. annotations No. words Period Fairy tales 3 772 18,597 1837-1847 Hymns 65 2,026 12,798 1798-1873 Prose 1 1,923 30,279 1952 Poetry 40 1,579 11,576 1965 This is the Fiction4 dataset of literary texts, spanning 109 individual texts across 4 genres and two languages (English and Danish) in the 19th and 20th… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/fiction4sentiment.tabular1K<n<10K1 likes17 downloads1y agoHugging Face25chcaa /danish-book-ads Books and Journals in Danish Newspaper Advertisements (1800–1850) A dataset of extracted and cleaned book titles and author names from Danish newspaper advertisements published between 1800 and 1850. The records were automatically extracted from digitized newspapers using a combination of rule-based methods, named-entity recognition, and a trained category classifier. Dataset Description Summary The dataset contains 80,938 advertisement records drawn from nine… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/danish-book-ads.texttext-classification10K<n<100K0 likes17 downloads4mo agoHugging Face26electricsheepafrica /africa-worldbank-wbl-legal-framework-childcare-score-scale-0-100-gd-wbl-chc-law-t WBL: Legal Framework, Childcare, Score (scale 0-100) | Africa (World Bank — Gender Statistics) | Africa (World Bank) Size category: n<1K - Formats: parquet - Sector: governance_security - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-worldbank-wbl-legal-framework-childcare-score-scale-0-100-gd-wbl-chc-law-t.tabulartabular-classificationn<1K0 likes17 downloads1mo agoHugging Face27chcaa /memo-canonical-novels Dataset description source_datasets: The corpus was created and made available by Jens Bjerring-Hansen and Philip Diderichsen, Dorte Haltrup Hansen, June 2023, see: https://huggingface.co/datasets/MiMe-MeMo/Corpus-v1.1 - Here, we make a more accessible, annotated version available. Additional tags: CE Canon: Cultural/Educational Canon, referring to novels whose titles are included in the Cultural Canon, or whose author is included in the Educational Canon. LEX Canon:… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/memo-canonical-novels.tabularn<1K1 likes16 downloads1y agoHugging Face28chcho /shinchan-chat-datasettextn<1K0 likes12 downloads1y agoHugging Face29chcaa /naturalistic_social_norms_alignmentgated Naturalistic Social Norms Alignment A dataset of 3,023 real-world social dilemmas in Danish, extracted from the popular radio show Sara og Monopolet. Each dilemma comes with reference solutions derived from a panel of three guests, making the dataset suitable for evaluating social norm alignment of LLMs and humans in naturalistic, open-ended conversations. Paper: Naturalistic measure of social norms alignment Code & Framework: GitHub repository Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/naturalistic_social_norms_alignment.texttext-generation1K<n<10K1 likes12 downloads4mo agoHugging Face30chcaa /smk_canon_paintingsimage1K<n<10K0 likes11 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.