datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BnSentMix
BnSentMix: A Diverse Bengali-English Code-Mixed Dataset for Sentiment Analysis
Dataset Overview
Column Title
Description
Data Sources
Facebook, YouTube, E-commerce Sites
#Samples
20000
Sentiment Labels
1:Positive, 2:Negative, 3:Neutral, 4:Mixed
Filtering Method
Automated using mBERT
#Annotators
64
Annotation/Sample
2 or 3 (if tie)
Dataset Statistics
Statistic
Value
Mean Character Length
62.77
Max Character Length
1985… See the full description on the dataset page: https://huggingface.co/datasets/aplycaebous/BnSentMix.aave-bns-data
Aave-BNS Multidimensional Protocol-Network Evidence
Hugging Face release status: PRIVATE RELEASE CANDIDATE. The dataset has been uploaded to the Hugging Face Hub for post-upload validation. Local validation covers 14 configurations, Croissant 1.1 core metadata, substantive Responsible AI metadata, checksums, and deterministic handoff auditing. Dataset Viewer validation, clean Hub loading, platform-generated Croissant reconciliation, and immutable release tagging are the… See the full description on the dataset page: https://huggingface.co/datasets/zlysunshine/aave-bns-data.IPC_and_BNS_transformationbn_sentiment_noisy_dataset
Dataset Card for "SentNoB"
Dataset Summary
Social Media User Comments' Sentiment Analysis Dataset. Each user comments are labeled with either positive (1), negative (2), or neutral (0).
Citation Information
@inproceedings{islam2021sentnob,
title={SentNoB: A Dataset for Analysing Sentiment on Noisy Bangla Texts},
author={Islam, Khondoker Ittehadul and Kar, Sudipta and Islam, Md Saiful and Amin, Mohammad Ruhul},
booktitle={Findings of the Association for… See the full description on the dataset page: https://huggingface.co/datasets/sustcsenlp/bn_sentiment_noisy_dataset.bnsbns2bnsBnSentMix
BnSentMix: A Diverse Bengali-English Code-Mixed Dataset for Sentiment Analysis
Dataset Overview
Column Title
Description
Data Sources
Facebook, YouTube, E-commerce Sites
#Samples
20000
Sentiment Labels
1:Positive, 2:Negative, 3:Neutral, 4:Mixed
Filtering Method
Automated using mBERT
#Annotators
64
Annotation/Sample
2 or 3 (if tie)
Dataset Statistics
Statistic
Value
Mean Character Length
62.77
Max Character Length
1985… See the full description on the dataset page: https://huggingface.co/datasets/rifat101/BnSentMix.IPC_and_BNS_transformationbnsaddcsvIPC_and_BNS_transformationieee-dataIPC_and_BNS_transformation
