Set
Datasets
All datasets matching “Set”sst5
Stanford Sentiment Treebank - Fine-Grained
Stanford Sentiment Treebank with 5 labels: very positive, positive, neutral, negative, very negative
Splits are from:
https://github.com/AcademiaSinicaNLPLab/sentiment_dataset/tree/master/data
Training data is on sentence level, not on phrase level!
meccog-final-set
MecCog final set
This challenge is closed (ended 2026-08-28). No new registrations or PRs
are accepted, and the merge-bot no longer merges — see CONTRIBUTING.md. The
final set below reflects the state at closure.
The curated final set of papers for five APOE4 / Alzheimer's disease
mechanism hypotheses, built by autonomous agents through native Hugging Face
Pull Requests on this dataset.
Agents do not search for papers here. They curate: every paper below was
picked out of an… See the full description on the dataset page: https://huggingface.co/datasets/MecCogAgenticChallenge/meccog-final-set.emotion** Attention: There appears an overlap in train / test. I trained a model on the train set and achieved 100% acc on test set. With the original emotion dataset this is not the case (92.4% acc)**
SID_Set
Dataset Card for SID_Set
Dataset Summary
We provide Social media Image Detection dataSet (SID-Set), which offers three key advantages:
Extensive volume: Featuring 300K AI-generated/tampered and authentic images with comprehensive annotations.
Broad diversity: Encompassing fully synthetic and tampered images across various classes.
Elevated realism: Including images that are predominantly indistinguishable from genuine ones through mere visual inspection.
Please check… See the full description on the dataset page: https://huggingface.co/datasets/saberzl/SID_Set.Ezaris-Training-Sets
Ezaris-Training-Sets
Complete training data, checkpoints, code and provenance for the Ezaris-1B program (formerly LUNA-1B) — a 1.2B-parameter Llama-style model trained from scratch by ASTERIZER.
Layout
Path
Contents
pretrain_raw_datasets/
Raw pretraining sources (SmolLM-corpus, fineweb-edu, OpenWebMath, Finemath, Wiki, code)
pretrain_cleaned_datasets/
Cleaned / deduplicated / tokenized pretrain sets + phase-2 exact & near-dedup outputs… See the full description on the dataset page: https://huggingface.co/datasets/ASTERIZER/Ezaris-Training-Sets.20_newsgroupsThis is a version of the 20 newsgroups dataset that is provided in Scikit-learn. From the Scikit-learn docs:
The 20 newsgroups dataset comprises around 18000 newsgroups posts on 20 topics split in two subsets: one for training (or development) and the other one for testing (or for performance evaluation). The split between the train and test set is based upon a messages posted before and after a specific date.
We followed the recommended practice to remove headers, signature blocks, and… See the full description on the dataset page: https://huggingface.co/datasets/SetFit/20_newsgroups.
