YuvrajSingh9886/bonsai-jetson-benchmark-7w
Bonsai Jetson Benchmark: 7W Platform: NVIDIA Jetson Orin Nano Super 8GB · Power mode: 7W Backend: llama.cpp build-jetson · CUDA · -ngl 99 Sweep: prompt in {256, 512, 1024, 2048} tok x gen in {128, 256, 512} tok · 20 reqs/combo Context: 2560 tok · Concurrency: 1 Part of smolperfbenchmark, a public on-device LLM benchmark leaderboard. Headline metric is output tok/J (tokens per joule), computed over the decode phase. Files Bonsai-*, Ternary-Bonsai-*: per-combo… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/bonsai-jetson-benchmark-7w.
Bonsai Jetson Benchmark: 7W
Platform: NVIDIA Jetson Orin Nano Super 8GB · Power mode: 7W Backend: llama.cpp build-jetson · CUDA · -ngl 99 Sweep: prompt in {256, 512, 1024, 2048} tok x gen in {128, 256, 512} tok · 20 reqs/combo Context: 2560 tok · Concurrency: 1
Part of smolperfbenchmark, a public on-device LLM benchmark leaderboard. Headline metric is output tok/J (tokens per joule), computed over the decode phase.
Files
Bonsai-*,Ternary-Bonsai-*: per-combo aiperf exports.*-server.log: llama.cpp server logs.tegrastats.log: 1 Hz power and thermal samples.model_timing.log: per-combo timing.report.md: generated summary of the tables below.
Results
Date: 2026-06-04 02:15 Backend: llamacpp Context: 2560 tokens Concurrency: 1 Sweep: prompt in {256,512,1024,2048} gen in {128,256,512} Artifacts: artifacts/llamacpp/bonsai-all-20260528-0328-7W
Full Results
Per-Model Best Tok/J
Thermal Summary
Generated by `benchmark_all_bonsai.sh` (llamacpp) on 2026-06-04 02:15:00
Notes
- Results cover 5 model/quant configs: Bonsai 1.7B/4B/8B at Q10, and Ternary-Bonsai 1.7B/4B at Q20.
- A
Ternary-Bonsai-8B-server.logis included, but that model produced no timing rows in this run. - Power is the
VDD_CPU_GPU_CVrail (CPU + GPU + CV) fromtegrastats, averaged over each aiperf run window.
License & citation
- Benchmark results and artifacts (aiperf exports, server logs,
tegrastatslogs, generated reports): CC BY 4.0. Reuse and adaptation allowed, including commercially, provided you credit Yuvraj Singh, link the license, and indicate changes. - Harness code that produced these results: Apache-2.0.
If you use these results, please cite:
@misc{singh2026smolperfbenchmark,
title={smolperfbenchmark: On-Device LLM Leaderboard},
author={Yuvraj Singh},
year={2026},
howpublished={\url{https://github.com/YuvrajSingh-mist/smolperfbenchmark}},
}<!-- ad-footer --> ---
Released freely — support more experiments like it:
 
More: <https://huggingface.co/YuvrajSingh9886>
