Skip to content
Daniel Fragoso

GPU Speech Inference on AWS

Two ways of serving an open-source speech recognition model on AWS GPUs, compared on cost and reliability.

Architecture

A FastAPI transcription service on ECS with the EC2 GPU launch type behind an Application Load Balancer, and a vLLM container on a multi-GPU instance with tensor parallelism. Both are defined in CDK with ECR, SSM and Deep Learning AMIs. The write-up explains why long audio failed the load balancer health checks and how to fix it.

Highlights

  • Tensor-parallel vLLM serving
  • Documented failure analysis
  • Full test run for about $13