YuvrajSingh9886/bonsai-jetson-benchmark-15w
Bonsai Jetson Benchmark — 15W Platform: NVIDIA Jetson Orin Nano Super 8GB · Power mode: 15W Backend: llama.cpp build-jetson · CUDA · -ngl 99 Sweep: prompt ∈ {256, 512, 1024, 2048} tok × gen ∈ {128, 256, 512} tok · 20 reqs/combo Status: Complete — 57 combos (5 models × 12 prompt/gen configs) Key metric: tok/J = output tok/s ÷ VDD_CPU_GPU_CV (W) Models Model Quant Size Bonsai-1.7B Q1_0 (1-bit) ~237 MB Bonsai-4B Q1_0 (1-bit) ~540 MB Bonsai-8B Q1_0… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/bonsai-jetson-benchmark-15w.
Bonsai Jetson Benchmark — 15W
Platform: NVIDIA Jetson Orin Nano Super 8GB · Power mode: 15W Backend: llama.cpp build-jetson · CUDA · -ngl 99 Sweep: prompt ∈ {256, 512, 1024, 2048} tok × gen ∈ {128, 256, 512} tok · 20 reqs/combo Status: Complete — 57 combos (5 models × 12 prompt/gen configs) Key metric: tok/J = output tok/s ÷ VDDCPUGPU_CV (W)
Models
Dataset viewer columns
model · quant · prompt_tokens · gen_tokens · ttft_avg_ms · ttft_p50/p90/p99_ms · itl_avg_ms · itl_p50/p90/p99_ms · tok_s · prefill_tok_s · req_latency_avg/p99_ms · power_w · tok_j · avg_temp_c
Bonsai All-Model Benchmark — Jetson Orin Nano Super 8GB
Date: 2026-05-27 22:13 Backend: CUDA (-ngl 99) Context: 2560 tokens Concurrency: 1 Sweep: prompt in {256,512,1024,2048} gen in {128,256,512} Artifacts: /home/yuvrajsingh/Desktop/benchmark/artifacts/bonsai-all-20260527-0200
Skipped Models (smoke test failed or file not found)
- Bonsai-8B (server failed to start)
- Ternary-Bonsai-4B (server failed to start)
- Ternary-Bonsai-8B (server failed to start)
Full Results
Per-Model Best Tok/J
Thermal Summary
Generated by `benchmark_all_bonsai.sh` on 2026-05-27 22:13:35
License & citation
- Benchmark results and artifacts (aiperf exports, server logs,
tegrastatslogs, generated reports): CC BY 4.0. Reuse and adaptation allowed, including commercially, provided you credit Yuvraj Singh, link the license, and indicate changes. - Harness code that produced these results: Apache-2.0.
If you use these results, please cite:
@misc{singh2026smolperfbenchmark,
title={smolperfbenchmark: On-Device LLM Leaderboard},
author={Yuvraj Singh},
year={2026},
howpublished={\url{https://github.com/YuvrajSingh-mist/smolperfbenchmark}},
}<!-- ad-footer --> ---
Released freely — support more experiments like it:
 
More: <https://huggingface.co/YuvrajSingh9886>
