Benchmarks / Vision /Frontier / BF16
Vision · Frontier · BF16
A 4B-class model at BF16 on 12 GB. The universal baseline: every supported NVIDIA GPU from an RTX 3060 12 GB upward runs this identically, so it is the one leaderboard that spans the whole hardware range.
Best Images / min
—img/min
Ranked results
0
0 distinct systems
GPU models
0
Minimum VRAM
64 GB
~14 min run
Best result per GPU
Images / min · img/minNo results yet for this view.
Show the values behind this chart (0 rows)
| GPU | GPUs | Samples | Best (img/min) | Mean (img/min) |
|---|
Leaderboard
Including unranked results →No results match these filters yet.
What is pinned
Everything below is fixed by the profile and checked on every submission. A result whose runtime flags differ from these is published, but never ranked.
Model
- Repository
- Qwen/Qwen3-VL-32B-Instruct
- Revision
- pending freeze
- Precision
- bf16
- Parameters
- 32 B
- Licence
- Apache-2.0
Runtime
- Engine
- vllm 0.26.0
- Harness
- aiperf
- dtype
- bfloat16
- max-model-len
- 32768
- gpu-memory-utilization
- 0.9
- max-num-seqs
- 16
- enforce-eager
- false
- disable-log-requests
- true
- tensor-parallel-size
- 1
- swap-space
- 0
Workloads
- vision_ocr
- task=ocr image_long_edge_px=1080 max_output_tokens=256 concurrency=1 temperature=0
- vision_chart
- task=chart image_long_edge_px=1080 max_output_tokens=256 concurrency=1 temperature=0
- vision_document
- task=document image_long_edge_px=1080 max_output_tokens=256 concurrency=1 temperature=0
- vision_description
- task=description image_long_edge_px=1080 max_output_tokens=256 concurrency=1 temperature=0
- vision_reasoning
- task=reasoning image_long_edge_px=1080 max_output_tokens=256 concurrency=1 temperature=0
- vision_throughput
- task=description image_long_edge_px=1080 max_output_tokens=128 concurrency=2 temperature=0
Ranking
- Ranked on
- Images / min (img/min, higher is better)
- Gate
- Answer accuracy ≥ 90
- Gate
- Error rate ≤ 0
- Also shown
- TTFT p50, TTFT p95, End to end p50, Decode, Image encode p50, Energy / image
Design notes
- Release candidate. The runtime and harness versions are pinned; the model revision is not frozen until the bake-off confirms fit on the tier's minimum-VRAM reference system, so this leaderboard is provisional.
- Headline is images per minute at concurrency 2, which is what a reader actually wants to know.
- Quality is a gate, not a score. OCR, chart and document tasks have deterministic expected answers; description and reasoning are collected but never scored until a stable evaluator exists.
- Image encode time is separated from prompt prefill so the vision tower cost is visible.