STN
USA Jobs Board

Software Engineer, ECS

Amazon.com Services LLC · Seattle, WA
Salary $144K – $194K
Category Engineering
Type FULL TIME
Posted 2mo ago
View and Apply

Opens the original job posting in a new tab.

Job description

Are you familiar with Service Meshes, Container Networking, and Ingress Traffic Management? They're the dedicated infrastructure layers that facilitate seamless communication between services and microservices — making it easy to connect, monitor, secure, and control services at scale. Here's your opportunity to be part of the AWS Container Networking offerings that power millions of customer workloads: ECS Service Connect, ECS Service Gateway, and AWS Cloud Map. Our networking stack is built on three foundational pillars: - Envoy Proxy & xDS Control Plane (Service Connect) — At the core of our service mesh lies Envoy, a versatile, open-source proxy deployed as a sidecar to intercept and manage east-west traffic. Our team builds the Envoy Management Server (xDS) control plane that configures thousands of Envoy proxies in real-time — enabling zero-code load balancing, retries, circuit-breaking, observability, and traffic routing across microservices. - Elastic Load Balancing Integration (Service Gateway) — For north-south ingress traffic, we orchestrate AWS Elastic Load Balancers (ALB & NLB) on behalf of customers. Service Gateway automates the provisioning and lifecycle management of load balancers, listener rules, target groups, TLS certificates, and security groups — abstracting away complex networking infrastructure so customers can expose their ECS services to the internet with a single API call. - AWS Cloud Map (Service Discovery) — The foundational registry that powers service discovery across the stack. Cloud Map provides DNS-based and API-based service registration and resolution, enabling ECS services to find and communicate with each other without hard-coded endpoints. It serves as the backbone for both Service Connect's internal routing and Service Discovery's DNS resolution. Key job responsibilities -

  • Design, develop, and operate highly available, scalable distributed systems for container networking and service mesh infrastructure -
  • Build and extend the Envoy Management Server (xDS) control plane — implementing CDS, EDS, LDS, and RDS APIs that configure thousands of Envoy proxies in real-time for east-west service mesh traffic -
  • Design and build ELB orchestration systems for Service Gateway — automating ALB/NLB provisioning, listener rule management, target group bindings, TLS certificate association, and multi-service bin-packing to optimize north-south ingress traffic and customer costs -
  • Develop and scale Cloud Map as the central service registry — building DNS-based and API-based discovery, namespace management, health checking integration, and cross-account service resolution -
  • Design and implement APIs for ingress traffic management (Service Gateway), service-to-service communication (Service Connect), and service discovery (Cloud Map) -
  • Drive operational excellence — on-call, COE investigations, deployment safety, and production health -
  • Build AI-powered operational tooling — develop skills, agents, and automation that improve diagnostics, accelerate root-cause analysis, and enhance operational metrics and alarming -
  • Design and implement customer self-service capabilities — enabling customers to diagnose, troubleshoot, and resolve networking issues independently through better APIs, dashboards, and guided workflows -
  • Collaborate with cross-functional teams including Product Managers, partner AWS service teams (ELB, Route 53, ECS Scheduler), and open-source communities - Influence technical direction through design documents, code reviews, and architectural decisions A day in the life Your morning might start with a code review for a teammate's control plane change, followed by a quick check on the deployment pipeline — yesterday's release is baking in gamma and the canary metrics look healthy. Mid-morning, you dive into a design document for a new feature. You sketch out the system interactions, identify edge cases, and post your proposal for team review. After lunch, you pair with a teammate to debug a subtle production issue. You trace the request flow end to end, reproduce it locally, and draft a fix. Later in the afternoon, you work on an AI-powered diagnostic skill you've been building — it correlates metrics, events, and logs to automatically surface root causes for common customer issues. You test it against a recent ticket and it correctly identifies a misconfigured timeout. Before wrapping up, you join a quick sync with a partner team to align on an API contract, then review the on-call dashboard — all green. Some days you're deep in internals tracing an HTTP/2 issue; other days you're designing a customer-facing API or building automation that makes the next on-call shift smoother. No two days are the same, but the thread that ties them together is making container networking invisible — so customers can focus on their applications, not their infrastructure. About the team
  • Work/Life Balance Our team puts a high value on work-life balance. It isn’t about how many hours you spend at home or at work; it’s about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives. This position involves on-call responsibilities, typically for one week every two months. We don’t like getting paged in the middle of the night or on the weekend, so we work to ensure that our systems are fault tolerant. When we do get paged, we work together to resolve the root cause so that we don’t get paged for the same issue twice. Mentorship & Career Growth Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we’re building an environment that celebrates knowledge sharing and mentorship. Our senior members enjoy one-on-one mentoring and thorough, but kind, code reviews. We care about your career growth and strive to assign projects based on what will help each team member develop into a better-rounded engineer and enable them to take on more complex tasks in the future.