Benchmarks
OpenNN benchmarks
These benchmarks compare specific OpenNN configurations with named alternatives on defined hardware and datasets. Open a test to review its methodology, versions, precision, number of runs and raw results.
Build
Tests covering dependency requirements and memory footprint.
Baseline RAM and GPU-ready VRAM
OpenNN vs PyTorch & TensorFlow · baseline footprint
Dependencies & install friction
OpenNN vs PyTorch & TensorFlow · install requirements
Learn
Tests covering source size and example API usage.
Source lines of code
OpenNN vs PyTorch & TensorFlow · native source
Iris API lines of code
OpenNN vs PyTorch & TensorFlow · same Iris model
Load
Tests covering data handling and memory capacity.
Train
Training tests on the named workloads and hardware.
GPU HIGGS dense training
OpenNN vs PyTorch & TensorFlow · HIGGS
CPU HIGGS dense training
OpenNN vs PyTorch & TensorFlow · 3-run median, 0.7% dispersion
GPU Transformer training
OpenNN vs PyTorch & TensorFlow · GPU training
GPU ResNet-50 training
OpenNN vs PyTorch & TensorFlow · CIFAR-10
GPU on Windows
OpenNN vs PyTorch & TensorFlow · native CUDA path
Optimize
Precision tests on the named GPU workloads.
GPU fp32 vs bf16 precision sweep
OpenNN bf16 vs fp32 · reported results across four workloads
Validate
Accuracy tests using the stated held-out data and methodology.
Deploy
Tests covering package size, startup and export.
Deployment size on GPU (CNN)
OpenNN vs PyTorch & TensorFlow · CNN CUDA build
Deployment size on CPU
OpenNN vs PyTorch & TensorFlow · CPU package
Startup latency
OpenNN vs PyTorch & TensorFlow · first prediction
Model export to standalone code
OpenNN vs PyTorch & TensorFlow · standalone artifact
Operate
Inference tests on the named workloads and hardware.
GPU Transformer inference
OpenNN vs PyTorch & TensorFlow · up to 680.4k tok/s
GPU HIGGS dense inference
OpenNN vs PyTorch & TensorFlow · 15.34M/34.84M samples/s
CPU HIGGS dense inference
OpenNN vs PyTorch & TensorFlow · 441.5k samples/s
GPU ResNet-50 inference
OpenNN vs PyTorch & TensorFlow · fp32/bf16, 5 runs