laya-cuda-bench

by bhushankinge

Benchmarks Laya inference throughput, latency and cost across NVIDIA GPUs and serving backends

How many decisions/s can one NVIDIA GPU serve under a p99 SLO, and what does a million cost? Reproducible benchmark of the Laya 421M decision models on RTX 2000 Ada, RTX PRO 5000/6000 Blackwell and H100 (+MIG), with TensorRT, torch.compile, ONNX Runtime, and LLM/API baselines.

Related projects