CoolFace
Datasetpublic

Rakancorle1/hans-sft-4k

Hans-SFT-4K · SFT recipe for the audio-visual Clever Hans Supervised fine-tuning (SFT) data accompanying the paper When Vision Speaks for Sound. Like the original Clever Hans — the horse that looked like he could do arithmetic but was actually reading his trainer's body language — video-capable MLLMs often look like they can hear: they answer audio questions by reading visual cues and never verifying the audio stream. Hans-SFT-4K is the 3,834-sample SFT mix that teaches models… See the full description on the dataset page: https://huggingface.co/datasets/Rakancorle1/hans-sft-4k.

sourceHugging Facecc-by-nc-4.0updated 4mo agoView on Hugging Face
1likes27downloads
filemedia.zip7.39 GBdownload
filesft_train_media.zip7.39 GBdownload

Rakancorle1/hans-sft-4k · main · files are served by the source, never re-hosted here