Amazon Web Services announced Amazon SageMaker HyperPod Inference Gateway on September 18, 2026, a Kubernetes-native, GPU-aware routing system for large language model inference that deploys as a single managed add-on for Amazon EKS on existing HyperPod infrastructure. AWS said the gateway can reduce first-token latency by up to 82%. The Routing Problem Behind the Gateway According to AWS, default Kubernetes load-balancing algorithms such as round-robin and least-connections have no visibility…