schneiderkamplab/dfm10-danish-lexical-sentiment-sft
dfm10-danish-lexical-sentiment-sft Gold lexical-polarity supervision derived from the Danish Sentiment Lexicon. Contents Format: gzip-compressed JSON Lines under data/train-*.jsonl.gz Schema: chat messages, optional condition and tools, plus provenance Shards: 1 Rows: 13,698 Category: Danish lexical sentiment Upstream material dsldk/danish-sentiment-lexicon Processing Gold batched mappings are supplemented by separately generated… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm10-danish-lexical-sentiment-sft.
dfm10-danish-lexical-sentiment-sft
Gold lexical-polarity supervision derived from the Danish Sentiment Lexicon.
Contents
- Format: gzip-compressed JSON Lines under
data/train-*.jsonl.gz - Schema: chat
messages, optionalconditionandtools, plus provenance - Shards: 1
- Rows: 13,698
- Category: Danish lexical sentiment
Upstream material
dsldk/danish-sentiment-lexicon
Processing
Gold batched mappings are supplemented by separately generated and audited natural Danish questions, without inventing sentence context.
Selection policy: all rows in the packaged source artifact.
Every packaged row is taken from the accepted source tree identified in the package manifest. Tokenized arrays and epoch sampling indices are not included; export staging alone does not imply inclusion in a sampled training union.
License and release review
The source and this transformed package are distributed under CC BY-SA 4.0. Attribute DSL and CST and preserve the ShareAlike notice. Attribute DSL and CST and preserve ShareAlike terms.
Validate
python recreate_dataset.py