Skip to content
Benchmarks / LLM /Frontier / BF16

LLM · Frontier · BF16

A 4B-class model at BF16 on 12 GB. The universal baseline: every supported NVIDIA GPU from an RTX 3060 12 GB upward runs this identically, so it is the one leaderboard that spans the whole hardware range.

Best Decode
tok/s
Ranked results
0
0 distinct systems
GPU models
0
Minimum VRAM
64 GB
~14 min run

Best result per GPU

Decode · tok/s
No results yet for this view.
Show the values behind this chart (0 rows)
GPUGPUsSamplesBest (tok/s)Mean (tok/s)

Download these values as CSV· Full open dataset

No results match these filters yet.

What is pinned

Everything below is fixed by the profile and checked on every submission. A result whose runtime flags differ from these is published, but never ranked.

Model
Repository
Qwen/Qwen3-32B
Revision
pending freeze
Precision
bf16
Parameters
32 B
Licence
Apache-2.0
Runtime
Engine
vllm 0.26.0
Harness
aiperf
dtype
bfloat16
max-model-len
32768
gpu-memory-utilization
0.9
max-num-seqs
16
enforce-eager
false
disable-log-requests
true
tensor-parallel-size
1
swap-space
0
Workloads
interactive
input_tokens=1024 output_tokens=256 concurrency=1 input_source=fixed_synthetic input_seed=20260801
concurrency
input_tokens=1024 output_tokens=256 concurrency_sweep=1/2/4/8 input_source=fixed_synthetic input_seed=20260801
longcontext
input_tokens=8192 output_tokens=1024 concurrency=1 input_source=fixed_synthetic input_seed=20260801
Ranking
Ranked on
Decode (tok/s, higher is better)
Gate
TTFT p95 ≤ 3000
Gate
Inter-token p95 ≤ 200
Gate
Error rate ≤ 0
Also shown
Prefill, TTFT p50, TTFT p95, Inter-token p95, Peak throughput @c8, Energy / token, Model load
Design notes
  • Release candidate. The runtime and harness versions are pinned; the model revision is not frozen until the bake-off confirms fit on the tier's minimum-VRAM reference system, so this leaderboard is provisional.
  • A 32B at BF16 is roughly 64 GB, which fits both a DGX Spark and a 2x96 GB workstation. A 70B BF16 baseline was rejected because it would exclude DGX Spark from its own tier.
  • Both 1 and 2 GPU results are accepted, ranked separately. Tensor parallelism is recorded as a comparability key.
  • RTX PRO 6000 Blackwell has no NVLink, so a 2-GPU result there communicates over PCIe. PCIe generation and width are recorded and shown on the result page.