CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01amazon-agi /SIFT-50M Dataset Card for SIFT-50M SIFT-50M (Speech Instruction Fine-Tuning) is a 50-million-example dataset designed for instruction fine-tuning and pre-training of speech-text large language models (LLMs). It is built from publicly available speech corpora containing a total of 14K hours of speech and leverages LLMs and off-the-shelf expert models. The dataset spans five languages, covering diverse aspects of speech understanding and controllable speech generation instructions. SIFT-50M… See the full description on the dataset page: https://huggingface.co/datasets/amazon-agi/SIFT-50M.textaudio-text-to-text10M<n<100M39 likes1.4k downloads1y agoHugging Face02mazesmazes /sift-audio SIFT Audio Dataset Self-Instruction Fine-Tuning (SIFT) dataset for training audio understanding models. Dataset Description This dataset contains audio samples paired with LLM-generated responses following the AZeroS multi-mode approach. Each audio sample is processed in three different modes to train models that can both respond conversationally AND describe/analyze audio. SIFT Modes Each audio sample generates three training samples with different behaviors:… See the full description on the dataset page: https://huggingface.co/datasets/mazesmazes/sift-audio.audioautomatic-speech-recognition100K<n<1M0 likes468 downloads8mo agoHugging Face03open-vdb /sift-128-euclidean Dataset Overview dataset: sift-128-euclidean Metadata Creation Time: 2025-01-07 11:37:52+0000 Update Time: 2025-01-07 11:38:07+0000 Source: https://github.com/erikbern/ann-benchmarks Task: N/A Train Samples: N/A Test Samples: N/A License: DISCLAIMER AND LICENSE NOTICE: This dataset is intended for benchmarking and research purposes only. The source data used in this dataset retains its original license and copyright. Users must comply with the respective licenses of… See the full description on the dataset page: https://huggingface.co/datasets/open-vdb/sift-128-euclidean.tabular1M<n<10M0 likes178 downloads2y agoHugging Face04yil384 /sift-archive Sift — 研究数据归档 Sift 是一个 CPU/DDR-primary + GPU-assisted 的分层内存 MoE + 长上下文推理系统研究项目 (用便宜的大容量 DDR/CXL 承载放不进 HBM 的大型稀疏 MoE + 长上下文;头条指标是 tokens-per-dollar / tokens-per-Joule)。 本仓是该项目自产实验数据的归档,用于把数据从本地磁盘卸下来。 这里没有模型权重 —— 模型是上游公开 GGUF,见 MODELS.manifest.json + restore_models.sh。 取数据 hf download yil384/sift-archive --repo-type dataset fetch_archive.sh --local-dir . bash fetch_archive.sh # 列出仓里有什么 bash fetch_archive.sh ssd2/traces/v2lite #… See the full description on the dataset page: https://huggingface.co/datasets/yil384/sift-archive.text10B<n<100B0 likes60 downloads3d agoHugging Face05zodumair /sifta-document-forgery-datasetimagen<1K0 likes34 downloads3mo agoHugging Face06fzliu /sift1btextn<1K1 likes29 downloads3y agoHugging Face07nivvis /eq-esconv-sifted EQ-ESConv-Sifted: Elo-Ranked Emotional Support Conversations The ESConv dataset (Liu et al., ACL 2021) ranked by empathetic quality via Swiss-style Elo tournament. All 1,300 conversations scored and sorted. Why this exists ESConv is a widely-used emotional support dataset but quality varies significantly — some conversations have excellent empathetic support, others are low-effort or off-topic. This dataset adds Elo rankings so you can filter by quality. For… See the full description on the dataset page: https://huggingface.co/datasets/nivvis/eq-esconv-sifted.tabulartext-generation1K<n<10K0 likes28 downloads6mo agoHugging Face08luvox-ai /sift_vivoice700k_qwen27Bgatedtext100K<n<1M0 likes2 downloads5mo agoHugging Face09redamancyguy /sift1btextn<1K0 likes1 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.