← Back to AI Training Jobs

Infrastructure / Site Reliability Engineer (SRE)

Mercor

Work arrangement & location
Remote

Track this job

My Jobs
  • Software Engineering
Engagement
Part Time

Employer description and requirements

Mercor connects exceptional technical talent with leading organisations working on ambitious technology and AI initiatives. We are looking for experienced Infrastructure / Site Reliability Engineers (SREs) to join a full-time engagement focused on building and operating complex, enterprise-grade infrastructure. We are seeking engineers with strong hands-on experience building, operating, debugging, and scaling sophisticated production systems. The ideal candidate has worked extensively with Kubernetes, AWS, observability platforms such as Datadog, and modern infrastructure tooling.

This is a full-time opportunity, and candidates must be able to commit to full-time engagement.

What You'll Do

  • Build, operate, and improve highly available and scalable production infrastructure.
  • Manage and optimise Kubernetes-based production environments.
  • Design and maintain cloud infrastructure, primarily across AWS.
  • Improve system reliability, availability, scalability, and operational efficiency.
  • Build and maintain observability across infrastructure and applications using Datadog or similar platforms.
  • Investigate production incidents, perform root-cause analysis, and implement durable fixes.
  • Improve monitoring, alerting, logging, tracing, and overall production visibility.
  • Develop automation and internal tooling to reduce manual operational work.
  • Partner closely with software engineering teams on deployments, infrastructure, and production reliability.
  • Contribute to infrastructure architecture and technical decisions for complex distributed systems.

Ideal Background

  • Professional experience in Infrastructure Engineering, Site Reliability Engineering (SRE), Platform Engineering, DevOps, or Production Engineering.
  • Hands-on experience operating complex, enterprise-grade production systems.
  • Strong production experience with Kubernetes.
  • Strong experience with AWS and cloud-native infrastructure.
  • Experience with Datadog, Prometheus, Grafana, or comparable observability platforms.
  • Experience with Infrastructure as Code using Terraform, Pulumi, or equivalent technologies.
  • Strong understanding of distributed systems, networking, containers, Linux, and cloud architecture.
  • Experience building or maintaining CI/CD and production deployment infrastructure.
  • Strong debugging, troubleshooting, and incident-response capabilities.
  • Proficiency in at least one programming or scripting language, such as Python, Go, or Bash.

Strong Signals

  • Experience operating Kubernetes and cloud infrastructure at significant production scale.
  • Experience supporting high-traffic or mission-critical applications.
  • Experience building infrastructure or platform tooling used by large engineering organisations.
  • Ownership of production reliability, on-call operations, incident response, or capacity planning.
  • Experience working within sophisticated, large-scale distributed systems.
  • Demonstrated improvements to SLOs/SLIs, observability, deployment reliability, infrastructure performance, or operational efficiency.

Why Join

  • Solve challenging reliability, scalability, and performance problems across enterprise-grade production systems.
  • Work extensively with technologies such as Kubernetes, AWS, Datadog, Terraform/Pulumi, and modern cloud-native tooling.
  • Take meaningful ownership of production reliability, observability, infrastructure architecture, and operational improvements.
  • Competitive hourly compensation reflecting your experience and technical expertise.
  • Join a network of highly skilled engineers working on ambitious projects with leading technology and AI organisations.