datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
si_for_sdsift-audio
SIFT Audio Dataset
Self-Instruction Fine-Tuning (SIFT) dataset for training audio understanding models.
Dataset Description
This dataset contains audio samples paired with LLM-generated responses following the
AZeroS multi-mode approach. Each audio sample is processed in three different modes
to train models that can both respond conversationally AND describe/analyze audio.
SIFT Modes
Each audio sample generates three training samples with different behaviors:… See the full description on the dataset page: https://huggingface.co/datasets/mazesmazes/sift-audio.ptb-sifosift-128-euclidean
Dataset Overview
dataset: sift-128-euclidean
Metadata
Creation Time: 2025-01-07 11:37:52+0000
Update Time: 2025-01-07 11:38:07+0000
Source: https://github.com/erikbern/ann-benchmarks
Task: N/A
Train Samples: N/A
Test Samples: N/A
License: DISCLAIMER AND LICENSE NOTICE:
This dataset is intended for benchmarking and research purposes only.
The source data used in this dataset retains its original license and copyright. Users must comply with the respective licenses of… See the full description on the dataset page: https://huggingface.co/datasets/open-vdb/sift-128-euclidean.banglish_bench
BanglishBench
A smoke test for Banglish models. It answers one question: did this build break?
700 prompts, 7 categories, a floor per category, and an exit code. The same job
pytest does before you demo a feature.
It does not rank models and it does not measure quality. It tells you whether a
build is worth the time it takes to read its answers. Whether the answers are
any good still takes a person who reads Banglish.
Run it
pip install huggingface_hub
hf… See the full description on the dataset page: https://huggingface.co/datasets/sifat-febo/banglish_bench.smoleval
SmolEval
Pick the right base model before you fine-tune
Which small base model is worth your training time? There are dozens under
2B, and fine-tuning the wrong one costs hours. This scores one in a few
minutes, on what base models actually do: continue text.
90 prompts, 3-run average
coherence
relevance
diversity
SmolLM2-135M
███░░░░░░░ 34%
█░░░░░░░░░ 7%
█░░░░░░░░░ 8%
SmolLM2-360M
██░░░░░░░░ 21%
░░░░░░░░░░ 0%
█░░░░░░░░░ 14%
SmolLM2-1.7B
████░░░░░░… See the full description on the dataset page: https://huggingface.co/datasets/sifat-febo/smoleval.SIF-VLM-Fingerprint-Triggers
SIF and AGDI VLM Fingerprint Triggers
This public repository contains 12,000 model-specific visual fingerprint
trigger images for research on fingerprint transfer and robustness in Large
Vision-Language Models:
9,000 SIF baseline triggers generated with Ordinary, RNA, and PLA;
3,000 AGDI triggers generated for the same three base models.
Dataset configs
Config
Training model
Method
Rows
qwen2.5-vl-7b
Qwen/Qwen2.5-VL-7B-Instruct
Ordinary, RNA, PLA
3… See the full description on the dataset page: https://huggingface.co/datasets/autoRiver/SIF-VLM-Fingerprint-Triggers.sifta-document-forgery-dataseteq-esconv-sifted
EQ-ESConv-Sifted: Elo-Ranked Emotional Support Conversations
The ESConv dataset (Liu et al., ACL 2021) ranked by empathetic quality via Swiss-style Elo tournament. All 1,300 conversations scored and sorted.
Why this exists
ESConv is a widely-used emotional support dataset but quality varies significantly — some conversations have excellent empathetic support, others are low-effort or off-topic. This dataset adds Elo rankings so you can filter by quality.
For… See the full description on the dataset page: https://huggingface.co/datasets/nivvis/eq-esconv-sifted.DMSD-ood
Debiasing Multimodal Sarcasm Detection with Contrastive Learning
This is a replication of the DMSD-ood dataset for easier access.
Reference
Jia, M., Xie, C., & Jing, L. (2024). Debiasing Multimodal Sarcasm Detection with Contrastive Learning. Proceedings of the AAAI Conference on Artificial Intelligence, 38(16), 18354-18362.
electricity-productionThis dataset is used for demo purposes to illustrate using the time series forecasting models present in the Transformers library.
Source: https://www.kaggle.com/datasets/shenba/time-series-datasets
sifahane-turkish-medical-complaintsRedEval-ood
Leveraging Generative Large Language Models with Visual Instruction and Demonstration Retrieval for Multimodal Sarcasm Detection
This is a replication of the RedEval-ood dataset for easier access.
Reference
Binghao Tang, Boda Lin, Haolong Yan, and Si Li. 2024. Leveraging Generative Large Language Models with Visual Instruction and Demonstration Retrieval for Multimodal Sarcasm Detection. In Proceedings of the 2024 Conference of the North American Chapter of the… See the full description on the dataset page: https://huggingface.co/datasets/sifan077/RedEval-ood.si-follow-dummy
