Varo Bank•5h ago
Career Pages
Site Reliability Engineer
United States
Contract
Mid Level
$40/hr - $60/hr
Full Job Description
About the Role
The Site Reliability Engineer will design, build, and operate scalable, distributed infrastructure while managing Kubernetes and AWS environments, infrastructure as code, data platforms, observability, automation, and incident response.
Key Responsibilities
- Manage, upgrade, and autoscale EKS clusters across multiple environments (SIT, UAT, Prod) and AWS accounts
- Write Terraform modules and Helm charts to support GitOps workflows using ArgoCD and GitLab CI/CD pipelines
- Maintain and troubleshoot Kafka (MSK) clusters, including broker health, connectors, and CDC pipelines
- Improve observability using Prometheus, Thanos, Grafana, and ELK while proactively identifying cloud cost-optimization opportunities
- Automate operational tasks with Python and leverage AI/ML techniques for predictive alerting and intelligent runbooks
- Handle Platform Service Desk requests, including Terraform merge request reviews, access management, and deployment support
- Participate in the production on-call rotation, support incident response, and contribute to blameless post-mortems
Requirements
Must Have:
- 3+ years of experience in an SRE, DevOps, or Infrastructure Engineering role, with the ability to work independently and manage multiple workstreams
- Strong hands-on experience with core AWS services, including EKS, EC2, RDS Aurora, MSK, S3, IAM, VPC, and Direct Connect
- Deep production experience with Kubernetes (upgrades, networking, RBAC) alongside Helm and GitOps tools like ArgoCD
- Advanced proficiency with Terraform, including writing modules and managing multi-account/multi-environment states
- Experience supporting and maintaining data platforms such as Airflow, Databricks, EMR, Kafka/MSK, or CDC pipelines
- Solid understanding of networking (VPCs, security groups, Istio, DNS) paired with strong Python scripting skills for tooling and automation
- Experience managing observability stacks (Prometheus, Grafana, ELK) and effectively leveraging AI/LLM tools for automation and incident analysis
Preferred:
- Experience with Karpenter and KEDA, GitLab CI/CD pipeline experience, Hashicorp Vault for secrets management
Company
Varo Bank
Varo is a digital bank that offers innovative, premium banking services wrapped in inclusive design.
United States
Posted on Career Pages