datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hemmingway-1-omlx-quantization-benchmark-v1
Hemmingway-1 oMLX Quantization Benchmark
This is the public-safe benchmark package for the Hemmingway-1 oMLX
quantization study on Apple Silicon.
The release contains the authored task prompts, selected local execution
metadata, aggregate blind-judge results, reliability metadata, and the policy
used to select records when a condition was run more than once.
What is in the dataset
File
Rows
Purpose
data/train.jsonl
184
Mixed rows. Filter record_type for… See the full description on the dataset page: https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-benchmark-v1.splash-omlx-ollama-benchmark
Qwen3.8 27B: oMLX vs Splash vs Ollama
Reproducible local benchmark on a MacBook Pro M4 Max with 64 GB unified memory.
Result in one sentence
For this workload, Ollama is the best overall backend: it wins most TTFT/decode comparisons and the concurrency tests. Splash is interesting specifically for cold long-context prefill, where it is faster than oMLX and slightly faster than Ollama at 32K tokens.
Scope
Prompt lengths: 1K, 4K, 8K, 16K, 32K tokens… See the full description on the dataset page: https://huggingface.co/datasets/fparrav/splash-omlx-ollama-benchmark.
