CoolFace
Datasetpublic

schneiderkamplab/dfm10-danish-lexical-sentiment-sft

dfm10-danish-lexical-sentiment-sft Gold lexical-polarity supervision derived from the Danish Sentiment Lexicon. Contents Format: gzip-compressed JSON Lines under data/train-*.jsonl.gz Schema: chat messages, optional condition and tools, plus provenance Shards: 1 Rows: 13,698 Category: Danish lexical sentiment Upstream material dsldk/danish-sentiment-lexicon Processing Gold batched mappings are supplemented by separately generated… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm10-danish-lexical-sentiment-sft.

sourceHugging Facecc-by-sa-4.0updated 28d agoView on Hugging Face
0likes67downloads
Dataset Card

dfm10-danish-lexical-sentiment-sft

Gold lexical-polarity supervision derived from the Danish Sentiment Lexicon.

Contents

  • —Format: gzip-compressed JSON Lines under data/train-*.jsonl.gz
  • —Schema: chat messages, optional condition and tools, plus provenance
  • —Shards: 1
  • —Rows: 13,698
  • —Category: Danish lexical sentiment

Upstream material

  • —dsldk/danish-sentiment-lexicon

Processing

Gold batched mappings are supplemented by separately generated and audited natural Danish questions, without inventing sentence context.

Selection policy: all rows in the packaged source artifact.

Every packaged row is taken from the accepted source tree identified in the package manifest. Tokenized arrays and epoch sampling indices are not included; export staging alone does not imply inclusion in a sampled training union.

License and release review

The source and this transformed package are distributed under CC BY-SA 4.0. Attribute DSL and CST and preserve the ShareAlike notice. Attribute DSL and CST and preserve ShareAlike terms.

Validate

bash
python recreate_dataset.py