concurrency
nemotron-terminal-8b-eval-terminal-bench-lite-concurrency-100nemotron-terminal-8b-5pct-rand-skill-based-eval-terminal-bench-lite-concurrency-25nemotron-terminal-8b-eval-terminal-bench-lite-concurrency-25qwen3.8-27b-apple-silicon-concurrency
Qwen3.8-27B concurrent serving on Apple Silicon — benchmarks, patches & recipes
Research artifacts from making 16 concurrent Qwen3.8-27B requests work on a
single Apple Silicon Mac (M-series, 128 GB unified memory) — across three
serving engines, including the patches that make SGLang's native MLX
backend serve this model for the first time.
Code / full history: https://github.com/bluehawana/Qwen3.827B-SGLang-mpbm5max
Run it with Ollama: ollama run bluehawana/qwen3.8-27b-q8
Fast… See the full description on the dataset page: https://huggingface.co/datasets/bluehawana/qwen3.8-27b-apple-silicon-concurrency.fastapi_pydantic_v2_migration_concurrency_teaser
🚀 Python Backend - FastAPI & Pydantic v2 Migration & Concurrency Suite (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (498 Samples) & Commercial EULA on Gumroad:👉 Python Backend - FastAPI & Pydantic v2 Migration & Concurrency Suite on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout!
📦 What is Inside the Full Production Package:
498 Verified FAANG v2.0 Scenarios (100%… See the full description on the dataset page: https://huggingface.co/datasets/emgena/fastapi_pydantic_v2_migration_concurrency_teaser.pg19-concurrency-bench
PG-19 Concurrency Evaluation Prompts
Long-context prompts in chat-message format at six bucket sizes
(1K, 2K, 4K, 8K, 16K, 32K user-message tokens), designed for benchmarking
LLM inference-server throughput across context-length regimes.
Each prompt is a 2-turn conversation (system + user) ready to send to any
OpenAI-compatible /v1/chat/completions endpoint. The system message
instructs the model to summarise the passage in exactly five words, so the
generated output is bounded… See the full description on the dataset page: https://huggingface.co/datasets/nnilayy/pg19-concurrency-bench.
