Own and scale resilient, security-first cloud infrastructure on AWS, with Kubernetes, Terraform, Datadog, and modern CI/CD at the core.
Remote – United StatesFull-Time4+ years experience
Job Summary
We're looking for an experienced Site Reliability Engineer with deep expertise in cloud infrastructure, platform reliability, and DevOps. You'll work across AWS (EKS, IAM, VPC), Terraform, Kubernetes, Datadog, and GitHub Actions CI/CD pipelines. Experience with Infrastructure as Code, Python automation, and compliance frameworks (HIPAA, HITRUST, SOC 2) is highly desirable.
Responsibilities
Design, scale, and maintain resilient cloud-native infrastructure on AWS, emphasizing Amazon EKS, IAM, RBAC, and security-first architecture.
Build, enhance, and manage CI/CD pipelines using GitHub Actions and GitHub Advanced Security.
Own and improve platform observability through Datadog (metrics, logging, tracing, dashboards, alerting).
Develop and maintain Infrastructure as Code solutions using Terraform and Terragrunt.
Create internal tools and automation scripts in Python to streamline operations.
Maintain comprehensive technical documentation including runbooks and standards.
Collaborate within Agile teams using Jira for planning and tracking.
Participate in on-call rotations, incident response, and post-incident reviews.
Required Qualifications
4+ years in a Senior SRE or DevOps role supporting large-scale production cloud environments.
U.S. Citizens only (no visa sponsorship; no Green Card holders).
Strong expertise in AWS services: IAM, EKS, VPC, EC2, Secrets Manager, and serverless technologies.
Hands-on experience with Terraform, Terragrunt, Helm, and Kubernetes.
CI/CD pipeline design and maintenance using GitHub Actions with GitHub Advanced Security.
Deep Datadog knowledge (dashboards, monitors, alerts, telemetry analysis).
Python proficiency for automation scripts and internal tooling.
Strong commitment to comprehensive documentation.
Agile/Scrum team experience and Jira proficiency.
Infrastructure optimization, capacity planning, and resource utilization expertise.
Preferred Skills
AWS Certified DevOps Engineer – Professional certification.
AWS Lambda, AWS Fargate, and serverless architecture experience.
Familiarity with multi-tenant platforms and customer-isolated deployment models.
Knowledge of HIPAA, HITRUST, and SOC 2 compliance frameworks.
Eligibility: Open exclusively to U.S. Citizens only. No visa sponsorship available. No Green Card holders.