dds
Datasets
All datasets matching “dds”lcc
Dataset Card for LCC
Dataset Summary
This dataset consists of Danish data from the Leipzig Collection that has been annotated for sentiment analysis by Finn Årup Nielsen.
Supported Tasks and Leaderboards
This dataset is suitable for sentiment analysis.
Languages
This dataset is in Danish.
Dataset Structure
Data Instances
Every entry in the dataset has a document and an associated label.
Data Fields
An entry in the… See the full description on the dataset page: https://huggingface.co/datasets/DDSC/lcc.angry-tweets
Dataset Card for AngryTweets
Dataset Summary
This dataset consists of anonymised Danish Twitter data that has been annotated for sentiment analysis through crowd-sourcing. All credits go to the authors of the following paper, who created the dataset:
Pauli, Amalie Brogaard, et al. "DaNLP: An open-source toolkit for Danish Natural Language Processing." Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa). 2021
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/DDSC/angry-tweets.dkhate
Dataset Card for DKHate
Dataset Summary
This dataset consists of anonymised Danish Twitter data that has been annotated for hate speech. All credits go to the authors of the following paper, who created the dataset:
Offensive Language and Hate Speech Detection for Danish (Sigurbergsson & Derczynski, LREC 2020)
Supported Tasks and Leaderboards
This dataset is suitable for hate speech detection.
PwC leaderboard for Task A: Hate Speech Detection on DKhate… See the full description on the dataset page: https://huggingface.co/datasets/DDSC/dkhate.reservoir-neural-bench
ReservoirNeuralBench — Norne dataset and OOD suite
Data companion of DDSirota/reservoir-neural-bench —
a controlled benchmark of 20+ neural surrogates for 3-D reservoir simulation on the real
Norne field geometry (46×112×22 corner-point grid, 44 431 active cells, OPM Flow ground truth).
No trained checkpoints are distributed here — train the released architectures yourself
with the code in the GitHub repository (see Reproduce below).
Contents
path
size
what… See the full description on the dataset page: https://huggingface.co/datasets/DDSirota/reservoir-neural-bench.CBIS-DDSMreddit-da
Dataset Card for SQuAD-da
Dataset Summary
This dataset consists of 1,908,887 Danish posts from Reddit. These are from this Reddit dump and have been filtered using this script, which uses FastText to detect the Danish posts.
Supported Tasks and Leaderboards
This dataset is suitable for language modelling.
Languages
This dataset is in Danish.
Dataset Structure
Data Instances
Every entry in the dataset contains short Reddit… See the full description on the dataset page: https://huggingface.co/datasets/DDSC/reddit-da.
