We are looking for Infrastructure / Site Reliability Engineer (SRE) candidates for a project delivered through the hiring partner.
What you'll do
- Build, operate, and improve highly available and scalable production infrastructure.
- Manage and optimise Kubernetes-based production environments.
- Design and maintain cloud infrastructure, primarily across AWS.
- Improve system reliability, availability, scalability, and operational efficiency.
- Build and maintain observability across infrastructure and applications using Datadog or similar platforms.
- Investigate production incidents, perform root-cause analysis, and implement durable fixes.
- Improve monitoring, alerting, logging, tracing, and overall production visibility.
- Develop automation and internal tooling to reduce manual operational work.
- Partner closely with software engineering teams on deployments, infrastructure, and production reliability.
- Contribute to infrastructure architecture and technical decisions for complex distributed systems.
What you need
- Hands-on experience operating complex, enterprise-grade production systems.
- Strong production experience with Kubernetes.
- Strong experience with AWS and cloud-native infrastructure.
- Experience with Datadog, Prometheus, Grafana, or comparable observability platforms.
- Experience with Infrastructure as Code using Terraform, Pulumi, or equivalent technologies.
- Strong understanding of distributed systems, networking, containers, Linux, and cloud architecture.
- Experience building or maintaining CI/CD and production deployment infrastructure.
- Strong debugging, troubleshooting, and incident-response capabilities.
- Proficiency in at least one programming or scripting language, such as Python, Go, or Bash.
Nice to have
- Experience operating Kubernetes and cloud infrastructure at significant production scale.
- Experience supporting high-traffic or mission-critical applications.
- Experience building infrastructure or platform tooling used by large engineering organizations.
- Ownership of production reliability, on-call operations, incident response, or capacity planning.
- Experience working within sophisticated, large-scale distributed systems.
- Demonstrated improvements to SLOs/SLIs, observability, deployment reliability, infrastructure performance, or operational efficiency.
Expertise
Who you work with
Project and contracting process: the hiring partner. Applications continue on the provider's website.
Pay and hours
This role pays $200/hr, for 40 h/week. The work is remote. That is about 150% above the $80 an hour median for software roles open now.
Who can apply from where
The provider lists this role as open worldwide. With a profile, Tier1 checks your country and languages against this role and every other before you apply.

