datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ScentSet
ScentSet: A Synthetic Dataset for Smell Description and Classification
ScentSet is a synthetic dataset containing 572,293 entries and approximately 15 million tokens. Each entry is a short natural language description in simple english of a smell, often followed by a hint or guess about its source. The dataset is designed to support machine learning research in scent recognition, classification, and multimodal representation learning.
Format
{"text": "There's a bright… See the full description on the dataset page: https://huggingface.co/datasets/sixf0ur/ScentSet.ScentSet
ScentSet: A Synthetic Dataset for Smell Description and Classification
ScentSet is a synthetic dataset containing 572,293 entries and approximately 15 million tokens. Each entry is a short natural language description in simple english of a smell, often followed by a hint or guess about its source. The dataset is designed to support machine learning research in scent recognition, classification, and multimodal representation learning.
Format
{"text": "There's a bright… See the full description on the dataset page: https://huggingface.co/datasets/Strangefiction/ScentSet.
