datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
go_emotions
Dataset Card for GoEmotions
Dataset Summary
The GoEmotions dataset contains 58k carefully curated Reddit comments labeled for 27 emotion categories or Neutral.
The raw data is included as well as the smaller, simplified version of the dataset with predefined train/val/test
splits.
Supported Tasks and Leaderboards
This dataset is intended for multi-class, multi-label emotion classification.
Languages
The data is in English.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/go_emotions.discofuse
Dataset Card for "discofuse"
Dataset Summary
DiscoFuse is a large scale dataset for discourse-based sentence fusion.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
discofuse-sport
Size of downloaded dataset files: 4.33 GB
Size of the generated dataset: 15.04 GB
Total amount of disk used: 19.36 GB
An example of 'train' looks as follows.
{… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/discofuse.Copy_Dakshina_Google_research_dataset
Copy_Dakshina_Google_research_dataset
This repository is a structured, processed version of the Dakshina Dataset, originally released by Google Research. It has been reorganized into a unified Hugging Face format to support NLP research in South Asian languages, specifically focusing on sentence-level and word-level transliteration tasks.
Dataset Overview
The original Dakshina dataset is a collection of text in both Latin and native scripts for 12 South Asian… See the full description on the dataset page: https://huggingface.co/datasets/Anvesh-Lankala/Copy_Dakshina_Google_research_dataset.
