omlx
Datasets
All datasets matching “omlx”hemmingway-1-omlx-quantization-evidence-v2
Hemmingway-1 Quantization Evidence v2
This package records two local evidence lanes for the Hemmingway-1 oQ4e build: teacher-forced numerical fidelity against a BF16 reference, and controlled runtime telemetry on Apple Silicon. It complements the frozen blind-preference study in Hemmingway-1 oMLX Quantization Benchmark v1.
This dataset is sixstringzen/hemmingway-1-omlx-quantization-evidence-v2. The quality dataset remains unchanged because blind preference, distribution fidelity… See the full description on the dataset page: https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-evidence-v2.hemmingway-1-omlx-quantization-benchmark-v1
Hemmingway-1 oMLX Quantization Benchmark
This is the public-safe benchmark package for the Hemmingway-1 oMLX
quantization study on Apple Silicon.
Altworld developed and published
Hemmingway-1. Bobby Pierce
published these quantizations and the evaluation package. The
collection
links the upstream model and all six builds.
Analysis revision 2, corrected on 2026-09-22, fixes A/B attribution and matching
across reversed packets. Read CORRECTION.md before using the
aggregate… See the full description on the dataset page: https://huggingface.co/datasets/sixstringzen/hemmingway-1-omlx-quantization-benchmark-v1.splash-omlx-ollama-benchmark
Qwen3.8 27B: oMLX vs Splash vs Ollama
Reproducible local benchmark on a MacBook Pro M4 Max with 64 GB unified memory.
Result in one sentence
For this workload, Ollama is the best overall backend: it wins most TTFT/decode comparisons and the concurrency tests. Splash is interesting specifically for cold long-context prefill, where it is faster than oMLX and slightly faster than Ollama at 32K tokens.
Scope
Prompt lengths: 1K, 4K, 8K, 16K, 32K tokens… See the full description on the dataset page: https://huggingface.co/datasets/fparrav/splash-omlx-ollama-benchmark.
