datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
courtlistener_opinionsMultilingual-Opinion-Target-ExtractionThis repository contains the English 'SemEval-2014 Task 4: Aspect Based Sentiment Analysis'. translated with DeepL into Spanish, French, Russian, and Turkish. The labels have been manually projected. For more details, read this paper: Model and Data Transfer for Cross-Lingual Sequence Labelling in Zero-Resource Settings.
Intended Usage: Since the datasets are parallel across languages, they are ideal for evaluating annotation projection algorithms, such as T-Projection.
Label… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Multilingual-Opinion-Target-Extraction.OpinionQAThis is the OpinionQA dataset from the BERDS benchmark. (Paper link: )The purpose of this dataset is evaluating diversity of retrieval systems given subjective questions.
Each instance consists of a question and a list of valid perspectives (opinions) for the question."Type" indicates the number of perspectives a question has. "Binary" questions come with two perspectives, while "Multi" questions come with more than two.
We repurpose the OpinionQA dataset into the desired setting.We first… See the full description on the dataset page: https://huggingface.co/datasets/timchen0618/OpinionQA.northwind_opinion_mining_corpus
Opinion Mining Text Corpus
A labeled text corpus for opinion mining and sentiment analysis tasks, compiled from an open product review text corpus dataset publicly hosted on this Hub. The source corpus was assembled by a university research center.
This card does not yet list the source dataset or the applicable usage terms.
weibo-opinion-dynamic-single-dim
Weibo Sentiment Evolution Dataset
This dataset contains Weibo posts and their associated comment threads used for studying sentiment evolution and opinion dynamics in social media discussions.
The dataset is distributed as a single JSON Lines file:
weibo_dataset.jsonl
Each line is one Weibo post record. Comments for that post are embedded in the comments field.
Dataset Details
Number of post records: 1,379
Number of embedded comments: 93,569
Number of Weibo… See the full description on the dataset page: https://huggingface.co/datasets/hreyulog/weibo-opinion-dynamic-single-dim.sapient-synth-flan-niv2-fsopt-data-task902-deceptive-opinion-spam-classification
sapient-synth-flan-niv2-fsopt-data-task902-deceptive-opinion-spam-classification
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 2935
Task: synthetic anonymous instruction replacement
Generation… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-niv2-fsopt-data-task902-deceptive-opinion-spam-classification.tos_court_opinions_filteredsame_opinion_tripletsfact-or-opinionsapient-synth-flan-niv2-fsopt-data-task903-deceptive-opinion-spam-classification
sapient-synth-flan-niv2-fsopt-data-task903-deceptive-opinion-spam-classification
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 1703
Task: synthetic anonymous instruction replacement
Generation… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-niv2-fsopt-data-task903-deceptive-opinion-spam-classification.sapient-synth-flan-flan-fsnoopt-data-opinion-abstracts-rotten-tomatoes
sapient-synth-flan-flan-fsnoopt-data-opinion-abstracts-rotten-tomatoes
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 2404
Task: synthetic anonymous instruction replacement
Generation
Rows were… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-flan-fsnoopt-data-opinion-abstracts-rotten-tomatoes.sapient-synth-flan-flan-zsopt-data-opinion-abstracts-idebate
sapient-synth-flan-flan-zsopt-data-opinion-abstracts-idebate
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 1307
Task: synthetic anonymous instruction replacement
Generation
Rows were generated… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-flan-zsopt-data-opinion-abstracts-idebate.sapient-synth-flan-niv2-zsopt-data-task902-deceptive-opinion-spam-classification
sapient-synth-flan-niv2-zsopt-data-task902-deceptive-opinion-spam-classification
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 1234
Task: synthetic anonymous instruction replacement
Generation… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-niv2-zsopt-data-task902-deceptive-opinion-spam-classification.sapient-synth-flan-flan-fsopt-data-opinion-abstracts-rotten-tomatoes
sapient-synth-flan-flan-fsopt-data-opinion-abstracts-rotten-tomatoes
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 4649
Task: synthetic anonymous instruction replacement
Generation
Rows were… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-flan-fsopt-data-opinion-abstracts-rotten-tomatoes.sapient-synth-flan-flan-zsnoopt-data-opinion-abstracts-rotten-tomatoes
sapient-synth-flan-flan-zsnoopt-data-opinion-abstracts-rotten-tomatoes
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 2895
Task: synthetic anonymous instruction replacement
Generation
Rows were… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-flan-zsnoopt-data-opinion-abstracts-rotten-tomatoes.sapient-synth-flan-niv2-zsopt-data-task903-deceptive-opinion-spam-classification
sapient-synth-flan-niv2-zsopt-data-task903-deceptive-opinion-spam-classification
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 865
Task: synthetic anonymous instruction replacement
Generation… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-niv2-zsopt-data-task903-deceptive-opinion-spam-classification.sapient-synth-flan-flan-zsnoopt-data-opinion-abstracts-idebate
sapient-synth-flan-flan-zsnoopt-data-opinion-abstracts-idebate
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 1302
Task: synthetic anonymous instruction replacement
Generation
Rows were generated… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-flan-zsnoopt-data-opinion-abstracts-idebate.sapient-synth-flan-flan-fsnoopt-data-opinion-abstracts-idebate
sapient-synth-flan-flan-fsnoopt-data-opinion-abstracts-idebate
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 831
Task: synthetic anonymous instruction replacement
Generation
Rows were generated… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-flan-fsnoopt-data-opinion-abstracts-idebate.sapient-synth-flan-flan-zsopt-data-opinion-abstracts-rotten-tomatoes
sapient-synth-flan-flan-zsopt-data-opinion-abstracts-rotten-tomatoes
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 2919
Task: synthetic anonymous instruction replacement
Generation
Rows were… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-flan-zsopt-data-opinion-abstracts-rotten-tomatoes.weibo-opinion-dynamic-multi-dimopinion-extract
