CoolFace
Agents
Live
Leaderboard
Models
Community
Search
Create
Alerts
Menu
9 results
sm_86
sm_86
Search
in
all
models
datasets
apps
agents
people
projects
Models
All models matching “sm_86”
gbueno86 /
Meta-LLama-3-Cat-Smaug-LLama-70b
text-generation
transformers
en
1 likes
8.6k downloads
2y ago
Hugging Face
wei231 /
ternary-bonsai-3080-sm86-ninfer
ninfer
zh
0 likes
1.3k downloads
2d ago
Hugging Face
featherless-ai-quants /
gbueno86-Meta-LLama-3-Cat-Smaug-LLama-70b-GGUF
text-generation
0 likes
457 downloads
1y ago
Hugging Face
nith3n /
qwen3-4b-instruct-2507-qairt-sm8650-w4a16
0 likes
373 downloads
1mo ago
Hugging Face
zLLM-Lab /
zllm-minicpm5-2b-qnn-sm8635
0 likes
113 downloads
14d ago
Hugging Face
leslie721007 /
babylm-strict-small-coherent86
fill-mask
transformers
en
1 likes
33 downloads
22d ago
Hugging Face
leslie721007 /
babylm-strict-small-coherent86-alpha075
fill-mask
transformers
en
1 likes
29 downloads
21d ago
Hugging Face
satish860 /
sms_detection_algorithm
text-classification
transformers
0 likes
20 downloads
4y ago
Hugging Face
Datasets
All datasets matching “sm_86”
Relativ3pa1n /
dsv4-flash-sm86-8x3090
DeepSeek-V4-Flash on 8x RTX 3090 (SM86): 262K context, 120 tok/s aggregate Serving recipes, launch wrappers, and the measured throughput/context ladder for running a W4A16 DeepSeek-V4-Flash-class model on 8x RTX 3090 (SM 8.6, 24 GiB each) with CUDA graph decode, FlashInfer sparse MLA, Marlin MoE, and compressed hybrid KV. The ladder Concurrent sequences amortize the TP8 allreduce that dominates each decode step, so aggregate throughput scales near-linear while… See the full description on the dataset page: https://huggingface.co/datasets/Relativ3pa1n/dsv4-flash-sm86-8x3090.
0 likes
90 downloads
28d ago
Hugging Face