CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01humanify /si_for_sdaudio10K<n<100K0 likes4.1k downloads2mo agoHugging Face02amazon-agi /SIFT-50M Dataset Card for SIFT-50M SIFT-50M (Speech Instruction Fine-Tuning) is a 50-million-example dataset designed for instruction fine-tuning and pre-training of speech-text large language models (LLMs). It is built from publicly available speech corpora containing a total of 14K hours of speech and leverages LLMs and off-the-shelf expert models. The dataset spans five languages, covering diverse aspects of speech understanding and controllable speech generation instructions. SIFT-50M… See the full description on the dataset page: https://huggingface.co/datasets/amazon-agi/SIFT-50M.textaudio-text-to-text10M<n<100M39 likes1.4k downloads1y agoHugging Face03edmundluan /SWE-V-SIFtextn<1K0 likes579 downloads27d agoHugging Face04mazesmazes /sift-audio SIFT Audio Dataset Self-Instruction Fine-Tuning (SIFT) dataset for training audio understanding models. Dataset Description This dataset contains audio samples paired with LLM-generated responses following the AZeroS multi-mode approach. Each audio sample is processed in three different modes to train models that can both respond conversationally AND describe/analyze audio. SIFT Modes Each audio sample generates three training samples with different behaviors:… See the full description on the dataset page: https://huggingface.co/datasets/mazesmazes/sift-audio.audioautomatic-speech-recognition100K<n<1M0 likes469 downloads8mo agoHugging Face05HaimingW /ptb-sifotabularn<1K0 likes191 downloads3mo agoHugging Face06open-vdb /sift-128-euclidean Dataset Overview dataset: sift-128-euclidean Metadata Creation Time: 2025-01-07 11:37:52+0000 Update Time: 2025-01-07 11:38:07+0000 Source: https://github.com/erikbern/ann-benchmarks Task: N/A Train Samples: N/A Test Samples: N/A License: DISCLAIMER AND LICENSE NOTICE: This dataset is intended for benchmarking and research purposes only. The source data used in this dataset retains its original license and copyright. Users must comply with the respective licenses of… See the full description on the dataset page: https://huggingface.co/datasets/open-vdb/sift-128-euclidean.tabular1M<n<10M0 likes178 downloads2y agoHugging Face07sifat-febo /banglish_bench BanglishBench A smoke test for Banglish models. It answers one question: did this build break? 700 prompts, 7 categories, a floor per category, and an exit code. The same job pytest does before you demo a feature. It does not rank models and it does not measure quality. It tells you whether a build is worth the time it takes to read its answers. Whether the answers are any good still takes a person who reads Banglish. &nbsp; Run it pip install huggingface_hub hf… See the full description on the dataset page: https://huggingface.co/datasets/sifat-febo/banglish_bench.texttext-generationn<1K0 likes152 downloads15d agoHugging Face08sifat-febo /smoleval SmolEval Pick the right base model before you fine-tune Which small base model is worth your training time? There are dozens under 2B, and fine-tuning the wrong one costs hours. This scores one in a few minutes, on what base models actually do: continue text. 90 prompts, 3-run average coherence relevance diversity SmolLM2-135M ███░░░░░░░ 34% █░░░░░░░░░ 7% █░░░░░░░░░ 8% SmolLM2-360M ██░░░░░░░░ 21% ░░░░░░░░░░ 0% █░░░░░░░░░ 14% SmolLM2-1.7B ████░░░░░░… See the full description on the dataset page: https://huggingface.co/datasets/sifat-febo/smoleval.texttext-generationn<1K1 likes149 downloads20d agoHugging Face09yil384 /sift-archive Sift — 研究数据归档 Sift 是一个 CPU/DDR-primary + GPU-assisted 的分层内存 MoE + 长上下文推理系统研究项目 (用便宜的大容量 DDR/CXL 承载放不进 HBM 的大型稀疏 MoE + 长上下文;头条指标是 tokens-per-dollar / tokens-per-Joule)。 本仓是该项目自产实验数据的归档,用于把数据从本地磁盘卸下来。 这里没有模型权重 —— 模型是上游公开 GGUF,见 MODELS.manifest.json + restore_models.sh。 取数据 hf download yil384/sift-archive --repo-type dataset fetch_archive.sh --local-dir . bash fetch_archive.sh # 列出仓里有什么 bash fetch_archive.sh ssd2/traces/v2lite #… See the full description on the dataset page: https://huggingface.co/datasets/yil384/sift-archive.text10B<n<100B0 likes63 downloads4d agoHugging Face10autoRiver /SIF-VLM-Fingerprint-Triggers SIF and AGDI VLM Fingerprint Triggers This public repository contains 12,000 model-specific visual fingerprint trigger images for research on fingerprint transfer and robustness in Large Vision-Language Models: 9,000 SIF baseline triggers generated with Ordinary, RNA, and PLA; 3,000 AGDI triggers generated for the same three base models. Dataset configs Config Training model Method Rows qwen2.5-vl-7b Qwen/Qwen2.5-VL-7B-Instruct Ordinary, RNA, PLA 3… See the full description on the dataset page: https://huggingface.co/datasets/autoRiver/SIF-VLM-Fingerprint-Triggers.imageimage-to-text10K<n<100K0 likes52 downloads3mo agoHugging Face11Sifal /Kabyle-French French - Kabyle (Tatoeba) This dataset contains translation pairs for French (fr) and Kabyle (kab). The data was collected from the Tatoeba Project, a free collaborative online database of example sentences. ⚠️ Important Note on Quality Disclaimer: This dataset has been exported automatically and has not been manually verified. While Tatoeba relies on community contributions, errors or inconsistencies in translation pairs may exist. Use with appropriate caution. tabular100K<n<1M5 likes42 downloads10mo agoHugging Face12zodumair /sifta-document-forgery-datasetimagen<1K0 likes34 downloads3mo agoHugging Face13Sifi-world /DetectAIRevDetectAIRev, an AI-generated review detection dataset curated from human- written and LLM-generated reviews across diverse domains and diverse- LLMs text100K<n<1M0 likes28 downloads11mo agoHugging Face14nivvis /eq-esconv-sifted EQ-ESConv-Sifted: Elo-Ranked Emotional Support Conversations The ESConv dataset (Liu et al., ACL 2021) ranked by empathetic quality via Swiss-style Elo tournament. All 1,300 conversations scored and sorted. Why this exists ESConv is a widely-used emotional support dataset but quality varies significantly — some conversations have excellent empathetic support, others are low-effort or off-topic. This dataset adds Elo rankings so you can filter by quality. For… See the full description on the dataset page: https://huggingface.co/datasets/nivvis/eq-esconv-sifted.tabulartext-generation1K<n<10K0 likes25 downloads6mo agoHugging Face15fzliu /sift1btextn<1K1 likes24 downloads3y agoHugging Face16Sifal /KabyleWikipediatext1K<n<10K3 likes21 downloads3y agoHugging Face17sifan077 /DMSD-ood Debiasing Multimodal Sarcasm Detection with Contrastive Learning This is a replication of the DMSD-ood dataset for easier access. Reference Jia, M., Xie, C., & Jing, L. (2024). Debiasing Multimodal Sarcasm Detection with Contrastive Learning. Proceedings of the AAAI Conference on Artificial Intelligence, 38(16), 18354-18362. imagen<1K0 likes17 downloads2y agoHugging Face18sifan2026 /electricity-productionThis dataset is used for demo purposes to illustrate using the time series forecasting models present in the Transformers library. Source: https://www.kaggle.com/datasets/shenba/time-series-datasets texttime-series-forecastingn<1K0 likes17 downloads2mo agoHugging Face19open-llm-leaderboard /Sourjayon__DeepSeek-R1-8b-Sify-detailsgated Dataset Card for Evaluation run of Sourjayon/DeepSeek-R1-8b-Sify Dataset automatically created during the evaluation run of model Sourjayon/DeepSeek-R1-8b-Sify The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sourjayon__DeepSeek-R1-8b-Sify-details.tabular10K<n<100K0 likes11 downloads2y agoHugging Face20cihatyldz /sifahane-turkish-medical-complaintstabular1K<n<10K0 likes9 downloads5mo agoHugging Face21sifan077 /RedEval-ood Leveraging Generative Large Language Models with Visual Instruction and Demonstration Retrieval for Multimodal Sarcasm Detection This is a replication of the RedEval-ood dataset for easier access. Reference Binghao Tang, Boda Lin, Haolong Yan, and Si Li. 2024. Leveraging Generative Large Language Models with Visual Instruction and Demonstration Retrieval for Multimodal Sarcasm Detection. In Proceedings of the 2024 Conference of the North American Chapter of the… See the full description on the dataset page: https://huggingface.co/datasets/sifan077/RedEval-ood.image1K<n<10K0 likes7 downloads2y agoHugging Face22sifat1221 /english_voice_512audion<1K0 likes4 downloads2y agoHugging Face23sifat1221 /english_voice_256audion<1K0 likes4 downloads2y agoHugging Face24Sourjayon /Sify_CoT_Datasettext1K<n<10K0 likes4 downloads2y agoHugging Face25sifar0 /book_datatext10K<n<100K0 likes4 downloads4mo agoHugging Face26sifat1221 /english_voice_dummyaudion<1K0 likes3 downloads2y agoHugging Face27sifujohn /testtext10K<n<100K0 likes3 downloads1y agoHugging Face28luvox-ai /sift_vivoice700k_qwen27Bgatedtext100K<n<1M0 likes2 downloads5mo agoHugging Face29OneBottleKick /si-follow-dummygatedtextn<1K0 likes1 downloads2y agoHugging Face30jnvqc /sifitotabular100K<n<1M0 likes1 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.