cross-platform
qwen36-cross-platform-benchmark
Qwen3.6-35B-A3B cross-platform benchmark dataset
Complete sanitized evidence for the Qwen3.6 GX10 versus M2 Max benchmark.
Headline results
At 128K, median cold-prompt TTFT was 45.70 s on GX10 and 552.79 s on M2 Max, a 12.10x difference.
Near 256K, GX10 completed and strictly passed 12/12 requests. M2 Max completed 9/12 and strictly passed 7/9 completed requests.
On the same Q4_K_M coding control, MTP improved median decode 26.87% on GX10 CUDA and regressed… See the full description on the dataset page: https://huggingface.co/datasets/tekosML/qwen36-cross-platform-benchmark.cross-platform-reviews-booking-tripadvisor-gmapscross-platform_datasetchatbench-cross-platform-transfer
Cross-Platform Transfer (slack)
Part of ChatBench: a benchmark for evaluating embedding models on chat/conversational retrieval tasks.
Task Description
Thread retrieval evaluated on a held-out chat platform.
Dataset Statistics
Split
Queries
Corpus Documents
test
113
1232
Usage
from datasets import load_dataset
# Load corpus
corpus = load_dataset("GabeA/chatbench-cross-platform-transfer", "corpus", split="test")
# Load queries… See the full description on the dataset page: https://huggingface.co/datasets/GabeA/chatbench-cross-platform-transfer.arcanum-cross-platform-queries-syntheticarcanumsearch-cross-platform
