datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
maze-30x30-hard-1ksnappfood-sentiment-refined
Refined SnappFood Persian Sentiment Dataset
This dataset contains 13,000 Persian food-delivery reviews from the publicly available SnappFood sentiment dataset, with labels refined using an LLM-assisted annotation procedure described in:
Improving Persian Sentiment Classification through LLM-Assisted Label Refinement and Small Language Models
Dataset Summary
The released dataset contains:
Split
Positive
Negative
Total
Train
5,000
5,000
10,000… See the full description on the dataset page: https://huggingface.co/datasets/maziyaramini/snappfood-sentiment-refined.btc-15min-6monthsThere-Are-No-Games-Dataset-Encodings
Source and License
This dataset is derived from the applications.csv file of the Steam Dataset 2025: Multi-Modal Gaming Analytics, originally published on Kaggle.
This dataset (csv and encodings) were modified to be used with our There-Are-No-Games fine-tuned SPLADEV3 model in the scope of our AIR course project.
Original dataset:
https://www.kaggle.com/datasets/crainbramp/steam-dataset-2025-multi-modal-gaming-analytics/data
The original dataset is licensed under the Creative… See the full description on the dataset page: https://huggingface.co/datasets/mazombieme/There-Are-No-Games-Dataset-Encodings.snappfood-reviews-relabeledmaze_40x40maze_30x30maze_20x20_dead_end_near_oodmaze-100x100-wilson-hard-1ksteam_recommendation_systemmaze_20x20_very_far_ood_path_difficultymaze_50x50maze_20x20_dead_end_oodmaze_20x20_path80_100_deadend_oodmaze2d-hard-110-100kmaze_10x10_ood_difficultyaymara-english
Cite
@inproceedings{tan-2023-shot,
title = "Few-shot {S}panish-{A}ymara Machine Translation Using {E}nglish-{A}ymara Lexicon",
author = "Gillin, Nat and
Gummibaerhausen, Brian",
editor = "Mager, Manuel and
Ebrahimi, Abteen and
Oncevay, Arturo and
Rice, Enora and
Rijhwani, Shruti and
Palmer, Alexis and
Kann, Katharina",
booktitle = "Proceedings of the Workshop on Natural Language Processing for Indigenous Languages of… See the full description on the dataset page: https://huggingface.co/datasets/Mazen130203/aymara-english.maze_20x20_far_ood_path_difficultymaze_10x10maze_30x30_near_ood_difficultymaze_20x20_near_ood_path_difficultymaze_20x20_ood_difficultymaze_25x25maze_20x20
