datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sst2
Dataset Card for [Dataset Name]
Dataset Summary
The Stanford Sentiment Treebank is a corpus with fully labeled parse trees that allows for a complete analysis of the
compositional effects of sentiment in language. The corpus is based on the dataset introduced by Pang and Lee (2005)
and consists of 11,855 single sentences extracted from movie reviews. It was parsed with the Stanford parser and
includes a total of 215,154 unique phrases from those parse trees, each… See the full description on the dataset page: https://huggingface.co/datasets/stanfordnlp/sst2.sst5
Stanford Sentiment Treebank - Fine-Grained
Stanford Sentiment Treebank with 5 labels: very positive, positive, neutral, negative, very negative
Splits are from:
https://github.com/AcademiaSinicaNLPLab/sentiment_dataset/tree/master/data
Training data is on sentence level, not on phrase level!
sst2
Stanford Sentiment Treebank - Binary
Stanford Sentiment Treebank with 2 labels: negative, positive
Splits are from:
https://github.com/AcademiaSinicaNLPLab/sentiment_dataset/tree/master/data
Training data is on sentence level, not on phrase level!
rendered-sst2
Rendered SST-2
The Rendered SST-2 Dataset from Open AI.
Rendered SST2 is an image classification dataset used to evaluate the models capability on optical character recognition. This dataset was generated by rendering sentences in the Standford Sentiment Treebank v2 dataset.
This dataset contains two classes (positive and negative) and is divided in three splits: a train split containing 6920 images (3610 positive and 3310 negative), a validation split containing 872 images (444… See the full description on the dataset page: https://huggingface.co/datasets/nateraw/rendered-sst2.sstThe Stanford Sentiment Treebank, the first corpus with fully labeled parse trees that allows for a
complete analysis of the compositional effects of sentiment in language.MICO-SST2
MICO SST-2 challenge dataset
Mico Argentatus (Silvery Marmoset) - William Warby/Flickr
For the accompanying code, visit the GitHub repository of the competition: https://github.com/microsoft/MICO/.
Getting Started
The starting kit notebook for this task is available at: https://github.com/microsoft/MICO/tree/main/starting-kit.
In the starting kit notebook you will find a walk-through of how to load the data and make your first submission.
We also provide a library… See the full description on the dataset page: https://huggingface.co/datasets/szanella/MICO-SST2.earth-sst-dailymur-sst-ml-benchmark-Zarr
MUR SST ML Benchmark (Pacific, Zarr)
Machine-learning friendly Zarr subset of NASA/JPL GHRSST MUR SST.
Upstream source (public, no auth): s3://mur-sst/zarr
Subset:
Region: Pacific (20–50°N, 180–240°E); longitude is stored as 0–360°E
Time: 2018-01-01 → 2019-12-30 (729 daily frames; upstream coverage for this slice ends on 2019-12-30)
Variable: analysed_sst only (float32, °C)
Chunking for ML: (time, lat, lon) = (7, 256, 256) (weekly windows)
Why mur-sst/zarr-v1 during extraction?… See the full description on the dataset page: https://huggingface.co/datasets/KokosDev/mur-sst-ml-benchmark-Zarr.viirs-sst-daily-nonNRT
VIIRS/SNPP Daily Sea Surface Temperature — 4 km, Science Quality
Daily global sea surface temperature from the VIIRS instrument aboard Suomi-NPP,
Level-3 Standard Mapped Image at 4 km, as distributed by the NASA Ocean Biology
Processing Group. This is the science-quality (refined) feed — fully calibrated
and reprocessed.
A companion repository, PranavKonijeti/viirs-sst-daily-nrt,
holds the near-real-time (NRT) feed of the same product. See
Which feed should I use? below.… See the full description on the dataset page: https://huggingface.co/datasets/PranavKonijeti/viirs-sst-daily-nonNRT.sst2viirs-sst-daily-nrt
VIIRS/SNPP Daily Sea Surface Temperature — 4 km, Near-Real-Time
Daily global sea surface temperature from the VIIRS instrument aboard Suomi-NPP,
Level-3 Standard Mapped Image at 4 km, as distributed by the NASA Ocean Biology
Processing Group. This is the near-real-time (NRT) feed — produced within hours
of acquisition, with preliminary calibration.
If you are training a model or computing a trend, use the science-quality feed
instead: PranavKonijeti/viirs-sst-daily-nonNRT.
See… See the full description on the dataset page: https://huggingface.co/datasets/PranavKonijeti/viirs-sst-daily-nrt.Lipi-Ghor-bn-882-SSTT
🗣️ Lipi-Ghor | লিপিঘর — Bengali Speech Dataset (bn-882-SSTT)
Lipi-Ghor (লিপিঘর, meaning "House of Scripts") is a large-scale Bengali speech dataset designed for automatic speech recognition (ASR), speaker diarization, and spoken language research. It is one of the largest open Bengali speech corpora with aligned speaker, transcription, and timestamp annotations.
Built by Team_Villagers as part of DL Sprint 4.0.
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/Sanjidh090/Lipi-Ghor-bn-882-SSTT.sst2_ptThe Stanford Sentiment Treebank consists of sentences from movie reviews and
human annotations of their sentiment. The task is to predict the sentiment of a
given sentence. We use the two-way (positive/negative) class split, and use only
sentence-level labels.SS_tiercommon_voice_enopencryptoassetpricing
Open Crypto Asset Pricing
Cryptocurrency risk factors (CMKT, CSMB, CMOM) constructed following
Liu, Tsyvinski & Wu (2022),
"Common Risk Factors in Cryptocurrency," Journal of Finance, 77(2), 1133-1177.
Website: sstoeckl.github.io/crypto_data
Author: Sebastian Stoeckl,
Professor of Financial Economics, University of Liechtenstein.
Dataset Structure
data/
spec_grid.parquet # Specification grid (parameter lookup)
spec_grid.csv # Same, for easy… See the full description on the dataset page: https://huggingface.co/datasets/sstoeckl/opencryptoassetpricing.coco640-sstvae
coco320-sstvae
COCO 2017 images resized (cover) and center-cropped to 320x240, JPEG
quality 92. Built for training the SSTVAE radio autoencoder
(image-over-HF-radio). Splits follow the original COCO train2017 /
val2017 membership. Images smaller than 320x240 were dropped.
Original images: https://cocodataset.org — see COCO terms of use;
image copyrights belong to their Flickr owners.
sst-v2-step1500-5casessst_migaretThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so101_follower",
"total_episodes": 61,
"total_frames": 53650,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:61"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/daxum34/sst_migaret.sstirSST2Perturbed
Dataset Card for "SST2Perturbed"
More Information needed
sst1Dataset used in the paper:
A thorough benchmark of automatic text classification
From traditional approaches to large language models
https://github.com/waashk/atcBench
To guarantee the reproducibility of the obtained results, the dataset and its respective CV train-test partitions is available here.
Each dataset contains the following files:
data.parquet: pandas DataFrame with texts and associated encoded labels for each document.
split_<k>.pkl: pandas DataFrame with k-cross validation… See the full description on the dataset page: https://huggingface.co/datasets/waashk/sst1.task363_sst2_polarity_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task363_sst2_polarity_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task363_sst2_polarity_classification.Our_ILSVRCaugmented-glue-sst2
Dataset Card for Augmented-GLUE-SST2
Automatically augmented data from train split of SST-2 dataset using conditional text generation approach.
Code used to generate this file will be soon available at https://github.com/IntelLabs/nlp-architect.
SST2_train67k_test1.8k_valid0.8k
Dataset Card for "SST2_train67k_test1.8k_valid0.8k"
More Information needed
sst2-textbugger
Stanford Sentiment Treebank - Binary
hssd-sstksst2-pwws
Stanford Sentiment Treebank - Binary
SSTQA
