Swoon · 1 month ago
Site Reliability Engineer
Swoon is actively hiring a Site Reliability Engineer (SRE) to join the team. The role involves ensuring the reliability and performance of a real-time AWS-based call platform while leading incident response efforts and enhancing automation processes.
Responsibilities
Own the reliability and performance of a real-time AWS-based call platform by monitoring system health, troubleshooting production issues, and ensuring high availability and low-latency service delivery
Lead incident response efforts for outages and critical production issues, performing root cause analysis and driving permanent fixes to prevent recurrence
Design, build, and maintain cloud infrastructure in AWS and Kubernetes (EKS), ensuring the platform remains scalable, resilient, and capable of supporting growing call volumes
Develop and enhance automation using Infrastructure as Code (Terraform, Ansible) and improve operational processes to reduce manual work and increase platform stability
Manage and optimize observability tools, dashboards, alerts, and monitoring strategies to proactively identify risks, performance bottlenecks, and service degradation
Partner with engineering teams on deployments, CI/CD improvements, system design reviews, and long-term reliability initiatives while mentoring junior engineers on SRE best practices
Qualification
Required
5+ years of experience in Site Reliability Engineering (SRE), DevOps, Cloud Engineering, or Infrastructure Engineering within fast-paced or startup-style environments
3+ years of Expert-level AWS experience supporting and designing highly available, scalable production environments
3+ years of Strong Kubernetes (EKS) administration and troubleshooting experience in enterprise production environments
Proven incident response and production support expertise, including root cause analysis and resolution of critical outages
Hands-on Infrastructure as Code experience using Terraform and/or Ansible to automate infrastructure and operational processes
Experience with monitoring, observability, and alerting tools to proactively identify and resolve performance or reliability issues
Strong automation, CI/CD, and scripting skills with a focus on reducing manual work and improving platform reliability and deployment processes
Work Authorization - US Citizen or Permanent Resident Only
Preferred
Telecommunications or real-time systems experience preferred, including SIP, SBCs, carrier integrations, VoIP platforms, high-availability architectures, and 24x7 on-call production support