Infrastructure / Site Reliability Engineer (SRE)
Mercor
- Work arrangement & location
- Remote
Track this job
- Engagement
- Part Time
Employer description and requirements
Mercor connects exceptional technical talent with leading organisations working on ambitious technology and AI initiatives. We are looking for experienced Infrastructure / Site Reliability Engineers (SREs) to join a full-time engagement focused on building and operating complex, enterprise-grade infrastructure. We are seeking engineers with strong hands-on experience building, operating, debugging, and scaling sophisticated production systems. The ideal candidate has worked extensively with Kubernetes, AWS, observability platforms such as Datadog, and modern infrastructure tooling.
This is a full-time opportunity, and candidates must be able to commit to full-time engagement.
What You'll Do
- Build, operate, and improve highly available and scalable production infrastructure.
- Manage and optimise Kubernetes-based production environments.
- Design and maintain cloud infrastructure, primarily across AWS.
- Improve system reliability, availability, scalability, and operational efficiency.
- Build and maintain observability across infrastructure and applications using Datadog or similar platforms.
- Investigate production incidents, perform root-cause analysis, and implement durable fixes.
- Improve monitoring, alerting, logging, tracing, and overall production visibility.
- Develop automation and internal tooling to reduce manual operational work.
- Partner closely with software engineering teams on deployments, infrastructure, and production reliability.
- Contribute to infrastructure architecture and technical decisions for complex distributed systems.
Ideal Background
- Professional experience in Infrastructure Engineering, Site Reliability Engineering (SRE), Platform Engineering, DevOps, or Production Engineering.
- Hands-on experience operating complex, enterprise-grade production systems.
- Strong production experience with Kubernetes.
- Strong experience with AWS and cloud-native infrastructure.
- Experience with Datadog, Prometheus, Grafana, or comparable observability platforms.
- Experience with Infrastructure as Code using Terraform, Pulumi, or equivalent technologies.
- Strong understanding of distributed systems, networking, containers, Linux, and cloud architecture.
- Experience building or maintaining CI/CD and production deployment infrastructure.
- Strong debugging, troubleshooting, and incident-response capabilities.
- Proficiency in at least one programming or scripting language, such as Python, Go, or Bash.
Strong Signals
- Experience operating Kubernetes and cloud infrastructure at significant production scale.
- Experience supporting high-traffic or mission-critical applications.
- Experience building infrastructure or platform tooling used by large engineering organisations.
- Ownership of production reliability, on-call operations, incident response, or capacity planning.
- Experience working within sophisticated, large-scale distributed systems.
- Demonstrated improvements to SLOs/SLIs, observability, deployment reliability, infrastructure performance, or operational efficiency.
Why Join
- Solve challenging reliability, scalability, and performance problems across enterprise-grade production systems.
- Work extensively with technologies such as Kubernetes, AWS, Datadog, Terraform/Pulumi, and modern cloud-native tooling.
- Take meaningful ownership of production reliability, observability, infrastructure architecture, and operational improvements.
- Competitive hourly compensation reflecting your experience and technical expertise.
- Join a network of highly skilled engineers working on ambitious projects with leading technology and AI organisations.