GPU Speech Inference on AWS
Two ways of serving an open-source speech recognition model on AWS GPUs, compared on cost and reliability.
Architecture
A FastAPI transcription service on ECS with the EC2 GPU launch type behind an Application Load Balancer, and a vLLM container on a multi-GPU instance with tensor parallelism. Both are defined in CDK with ECR, SSM and Deep Learning AMIs. The write-up explains why long audio failed the load balancer health checks and how to fix it.
Highlights
- Tensor-parallel vLLM serving
- Documented failure analysis
- Full test run for about $13