laya-serve

by c4bbage

Serves Laya inference with Go, dynamic batching, and TensorRT or CUDA execution

High-throughput Go inference server and quantization toolkit for Laya (unofficial): TensorRT fp16/bf16/fp32, dynamic batching, parity-checked against the Python laya package

Related projects