sifta
Datasets
All datasets matching “sifta”sift-audio
SIFT Audio Dataset
Self-Instruction Fine-Tuning (SIFT) dataset for training audio understanding models.
Dataset Description
This dataset contains audio samples paired with LLM-generated responses following the
AZeroS multi-mode approach. Each audio sample is processed in three different modes
to train models that can both respond conversationally AND describe/analyze audio.
SIFT Modes
Each audio sample generates three training samples with different behaviors:… See the full description on the dataset page: https://huggingface.co/datasets/mazesmazes/sift-audio.sift-archive
Sift — 研究数据归档
Sift 是一个 CPU/DDR-primary + GPU-assisted 的分层内存 MoE + 长上下文推理系统研究项目
(用便宜的大容量 DDR/CXL 承载放不进 HBM 的大型稀疏 MoE + 长上下文;头条指标是 tokens-per-dollar / tokens-per-Joule)。
本仓是该项目自产实验数据的归档,用于把数据从本地磁盘卸下来。
这里没有模型权重 —— 模型是上游公开 GGUF,见 MODELS.manifest.json + restore_models.sh。
取数据
hf download yil384/sift-archive --repo-type dataset fetch_archive.sh --local-dir .
bash fetch_archive.sh # 列出仓里有什么
bash fetch_archive.sh ssd2/traces/v2lite #… See the full description on the dataset page: https://huggingface.co/datasets/yil384/sift-archive.sifta-document-forgery-datasetANTON-SIFTA
