datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
apple-silicon-llm-benchmarks
Apple Silicon Local LLM Benchmarks — M2 Max 32GB
Measurements taken while trying to get Qwen3.8-27B usable locally on a 32GB M2 Max. Most of
the popular speedup advice did not transfer from CUDA, so these are mostly negative results.
Everything here was measured on one machine. Treat it as a datapoint, not a law.
Hardware and software
Chip
Apple M2 Max
Unified memory
32 GB (~21.8 GB wireable to the GPU)
macOS
26.5.2
llama.cpp
build c1d0e7a00… See the full description on the dataset page: https://huggingface.co/datasets/RaynarDM/apple-silicon-llm-benchmarks.humaneval-apple-silicon
Mac Coding Bench Results v1 — Speed + Code Quality Benchmarks on Apple Silicon
Speed and code quality benchmarks for quantized LLMs running locally on Apple Silicon Macs. The dataset pairs inference speed measurements (tokens/sec) with HumanEval+ functional correctness scores for 21 models, across three hardware configurations (M1, M2 Max, M5) totaling 123 benchmark results.
Key Highlights
Qwen 3.6 35B-A3B achieves 89.6% HumanEval+ pass@1 at 16.7 tok/s — best quality… See the full description on the dataset page: https://huggingface.co/datasets/enescingoz/humaneval-apple-silicon.
