CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01DDSC /lcc Dataset Card for LCC Dataset Summary This dataset consists of Danish data from the Leipzig Collection that has been annotated for sentiment analysis by Finn Årup Nielsen. Supported Tasks and Leaderboards This dataset is suitable for sentiment analysis. Languages This dataset is in Danish. Dataset Structure Data Instances Every entry in the dataset has a document and an associated label. Data Fields An entry in the… See the full description on the dataset page: https://huggingface.co/datasets/DDSC/lcc.texttext-classificationn<1K3 likes8.8k downloads3y agoHugging Face02DDSC /angry-tweets Dataset Card for AngryTweets Dataset Summary This dataset consists of anonymised Danish Twitter data that has been annotated for sentiment analysis through crowd-sourcing. All credits go to the authors of the following paper, who created the dataset: Pauli, Amalie Brogaard, et al. "DaNLP: An open-source toolkit for Danish Natural Language Processing." Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa). 2021 Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/DDSC/angry-tweets.texttext-classification1K<n<10K4 likes2.5k downloads3y agoHugging Face03DDSC /dkhategated Dataset Card for DKHate Dataset Summary This dataset consists of anonymised Danish Twitter data that has been annotated for hate speech. All credits go to the authors of the following paper, who created the dataset: Offensive Language and Hate Speech Detection for Danish (Sigurbergsson & Derczynski, LREC 2020) Supported Tasks and Leaderboards This dataset is suitable for hate speech detection. PwC leaderboard for Task A: Hate Speech Detection on DKhate… See the full description on the dataset page: https://huggingface.co/datasets/DDSC/dkhate.texttext-classification1K<n<10K5 likes607 downloads3y agoHugging Face04DDSC /reddit-da Dataset Card for SQuAD-da Dataset Summary This dataset consists of 1,908,887 Danish posts from Reddit. These are from this Reddit dump and have been filtered using this script, which uses FastText to detect the Danish posts. Supported Tasks and Leaderboards This dataset is suitable for language modelling. Languages This dataset is in Danish. Dataset Structure Data Instances Every entry in the dataset contains short Reddit… See the full description on the dataset page: https://huggingface.co/datasets/DDSC/reddit-da.texttext-generation1M<n<10M2 likes273 downloads4y agoHugging Face05Rosalia1212 /cbis-ddsm-r CBIS-DDSM-R: A Curated Radiomic Feature Dataset for Breast Cancer Classification Dataset Summary CBIS-DDSM-R is an open-source, radiomics-ready extension of the Curated Breast Imaging Subset of the Digital Database for Screening Mammography (CBIS-DDSM). It is designed to facilitate reproducible radiomics and quantitative imaging research in breast cancer analysis. The dataset provides a standardized preprocessing pipeline for mammograms and includes IBSI-compliant… See the full description on the dataset page: https://huggingface.co/datasets/Rosalia1212/cbis-ddsm-r.tabularimage-classification1K<n<10K0 likes175 downloads4mo agoHugging Face06DDSC /nordic-embedding-training-data Thanks to Arrow Denmark and Nvidia for sponsoring the compute used to generate this dataset The purpose of this dataset is to pre- or post-train embedding models for Danish on text similarity tasks. The dataset is structured for training using InfoNCE loss (also known as SimCSE loss, Cross-Entropy Loss with in-batch negatives, or simply in-batch negatives loss), with hard-negative samples for the tasks of retrieval and unit-triplet. Beware that if fine-tuning the unit-triplets for… See the full description on the dataset page: https://huggingface.co/datasets/DDSC/nordic-embedding-training-data.text100K<n<1M4 likes155 downloads9mo agoHugging Face07DDSC /reddit-da-asr-preprocessedtext1M<n<10M0 likes154 downloads5y agoHugging Face08DDSC /europarl Dataset Card for DKHate Dataset Summary This dataset consists of Danish data from the European Parliament that has been annotated for sentiment analysis by the Alexandra Institute - all credits go to them. Supported Tasks and Leaderboards This dataset is suitable for sentiment analysis. Languages This dataset is in Danish. Dataset Structure Data Instances Every entry in the dataset has a document and an associated label.… See the full description on the dataset page: https://huggingface.co/datasets/DDSC/europarl.texttext-classificationn<1K2 likes150 downloads4y agoHugging Face09sotetsuk /dds_datasetSee http://github.com/sotetsuk/pgx dds_results_2.5M.npy : for training dds_results_10M.npy : for larger training dds_results_100M.npy : for larger training dds_results_500K.npy : for test If you use this dataset in your research, please cite our paper. To make your own dataset, use sotetsuk/make-dds-dataset. @inproceedings{koyamada2023pgx, title={Pgx: Hardware-Accelerated Parallel Game Simulators for Reinforcement Learning}, author={Koyamada, Sotetsu and Okano, Shinri and Nishimori… See the full description on the dataset page: https://huggingface.co/datasets/sotetsuk/dds_dataset.tabular100M<n<1B2 likes79 downloads2y agoHugging Face10helloerikaaa /cbis-ddsm-r CBIS-DDSM-R: A Curated Radiomic Feature Dataset for Breast Cancer Classification Dataset Summary CBIS-DDSM-R is an open-source, radiomics-ready extension of the Curated Breast Imaging Subset of the Digital Database for Screening Mammography (CBIS-DDSM). It is designed to facilitate reproducible radiomics and quantitative imaging research in breast cancer analysis. The dataset provides a standardized preprocessing pipeline for mammograms and includes IBSI-compliant… See the full description on the dataset page: https://huggingface.co/datasets/helloerikaaa/cbis-ddsm-r.tabularimage-classification1K<n<10K0 likes58 downloads8mo agoHugging Face11ZhiyuanQiu /DD_sans_symboletext10K<n<100K0 likes46 downloads4y agoHugging Face12wopdevries /bridge-dds-datasetDouble dummy solved data textothern<1K0 likes38 downloads7mo agoHugging Face13DDSC /da-wikipedia-queries Danish dataset for training embedding models for retrieval - sponsored by Arrow Denmark and Nvidia The purpose of this dataset is to train embedding models for retrieval in Danish. This dataset was made by showing ~30k Wikipedia paragraphs to LLMs and asking the LLMs to generate queries that would return the paragraph. For each of the 30k paragraphs in the original Wikipedia dataset, we used 3 different LLMs to generate queries: ThatsGroes/Llama-3-8b-instruct-SkoleGPT… See the full description on the dataset page: https://huggingface.co/datasets/DDSC/da-wikipedia-queries.tabular10K<n<100K5 likes35 downloads2y agoHugging Face14edi45 /gearxai-dds-seugated GearXAI DDS-SEU PGB Release This repository contains the public dataset and participant devkit for GearXAI: An Explainable Neuro-Symbolic Gearbox Fault Diagnosis Challenge, accepted to the IJCAI-ECAI 2026 Competitions and Challenges Track. The release provides processed planetary gearbox (PGB) vibration data from the DDS-SEU drivetrain setup, packaged for direct use with Hugging Face Datasets and the GearXAI evaluator. The competition task is 9-class gearbox fault diagnosis from… See the full description on the dataset page: https://huggingface.co/datasets/edi45/gearxai-dds-seu.tabular1M<n<10M1 likes34 downloads4mo agoHugging Face15LexiconShiftInnovations /DDSLU_Benchmarktext1K<n<10K0 likes32 downloads2y agoHugging Face16DDSC /angry-tweets-binary Dataset Card for "angry-tweets-binary" More Information needed text1K<n<10K0 likes28 downloads3y agoHugging Face17FLARE25-Agent-Xray /CBIS-DDSM-SEGimage1K<n<10K1 likes25 downloads1y agoHugging Face18korir8 /dd-seedstabular100K<n<1M0 likes21 downloads8mo agoHugging Face19Tokerss /ddsadadatext1K<n<10K0 likes20 downloads2y agoHugging Face20DDSC /partial-danish-gigaword-small-test-sample Dataset Card for "Danish Gigaword Test Sample" This is a small sample of the dataset DDSC/partial-danish-gigaword-no-twitter. It is meant as a small dataset for testing code. It is constructed using the following code: from datasets import concatenate_datasets, load_dataset # download dataset from huggingface dataset = load_dataset("DDSC/partial-danish-gigaword-no-twitter") # All of the dataset is available in the train split - we can simply: dataset = dataset["train"] #… See the full description on the dataset page: https://huggingface.co/datasets/DDSC/partial-danish-gigaword-small-test-sample.text1K<n<10K0 likes19 downloads4y agoHugging Face21martbern /DDSCtext1K<n<10K0 likes18 downloads3y agoHugging Face22DDSC /da-wikipedia-queries-gemma Dataset generated with compute sponsored by Arrow Electronics and NVIDIA This is a subset of https://huggingface.co/datasets/DDSC/da-wikipedia-queries/ tabular10K<n<100K0 likes15 downloads2y agoHugging Face23DDSC /da-wikipedia-queries-gemma-processedThis is a processed version of https://huggingface.co/datasets/DDSC/da-wikipedia-queries-gemma The dataset was created using this script: https://github.com/Dansk-Data-Science-Community/embedding_model/blob/main/create_processed_data.py text10K<n<100K1 likes14 downloads2y agoHugging Face24DDSC /ddiscotext1K<n<10K0 likes13 downloads3y agoHugging Face25DDSameera /items_raw_litetabular10K<n<100K0 likes10 downloads4mo agoHugging Face26fathyshalaby /ddstextn<1K0 likes9 downloads3y agoHugging Face27DDSameera /items_raw_fulltabular100K<n<1M0 likes5 downloads4mo agoHugging Face28failproof /ddshitstextn<1K0 likes4 downloads7mo agoHugging Face29dsrestrepo /cbis-ddsm-datathongated CBIS-DDSM Dataset (448px resolution) / Dataset CBIS-DDSM (resolución 448px) English Dataset Description This folder contains a simplified CBIS-DDSM dataset prepared for the Medical AI Datathon. It includes resized full mammograms and one compact labels.csv file. Original dataset: https://www.cancerimagingarchive.net/collection/cbis-ddsm/ Structure CBIS-DDSM-clean/ ├── images/ ├── labels.csv └── README.md The image column contains… See the full description on the dataset page: https://huggingface.co/datasets/dsrestrepo/cbis-ddsm-datathon.tabular1K<n<10K0 likes4 downloads3mo agoHugging Face30failproof /ddstextn<1K1 likes2 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.