Senior Site Reliability Engineer - 12 months - London - £425/day - Inside IR35
We are seeking a highly experienced and hands-on Senior Site Reliability Engineer to join a global technology services organisation on a 12-month hybrid contract based in London City. There are 2 positions available. The successful candidate will bring a strong background in observability, automation, cloud infrastructure, and DevOps practices on AWS, driving reliability and performance across complex distributed systems whilst designing and optimising a best-in-class monitoring infrastructure using Datadog and Geneos. Experience in financial services, capital markets, or fintech environments is highly advantageous.
Key Responsibilities:
- Architect, implement, and maintain observability platforms using Datadog and Geneos to ensure comprehensive monitoring and alerting
- Design and manage scalable infrastructure using Terraform and Infrastructure as Code (IaC) principles
- Champion GitOps methodologies using GitLab for CI/CD, configuration management, and deployment automation
- Optimise alerting strategies to reduce noise and improve actionable insights
- Oversee and continuously optimise cloud cost management strategies for observability infrastructure in line with Cloud FinOps principles
- Lead incident response, perform root cause analysis, and drive continuous improvement through blameless post-mortems
- Collaborate with software developers across multiple geographies and cross-functional teams to deliver systems within agreed timelines
- Collaborate with executive leadership and cross-functional stakeholders to align infrastructure strategy with long-term business objectives
- Mentor junior engineers and contribute to the evolution of SRE best practices
- Stay updated on industry best practices and emerging technologies in observability, DevOps, and cloud
What You Will Ideally Bring:
- 7+ years of experience in Site Reliability Engineering, DevOps, or related roles (essential)
- Strong hands-on experience with AWS cloud services including EC2, S3, RDS, Lambda, VPC, IAM, CloudWatch, EKS, and ECS (essential)
- Strong hands-on expertise in Infrastructure as Code using Terraform (essential)
- Deep expertise in Datadog for metrics, logs, traces, and dashboards (essential)
- Proficiency with Geneos for Real Time monitoring and alerting (essential)
- Deep knowledge of CI/CD tools including GitLab CI and Jenkins for Java and Python-based microservices architectures
- Solid understanding of GitLab and GitOps workflows
- Strong Scripting and automation skills in Python and Bash
- Experience with containerisation and orchestration using Docker and Kubernetes
- Solid understanding of Linux systems, networking, and distributed systems
- Desirable: AWS certifications, familiarity with SLOs/SLIs and error budgets, experience in multi-cloud or hybrid cloud environments, knowledge of ITIL or SRE frameworks, and experience working in regulated or high-availability environments such as fintech or capital markets
- Bonus: experience in Equity or Fixed Income and working knowledge of Benchmarks and Indices
Contract Details:
- Duration: 12 months
- Rate: £425/day - Inside IR35
- Location: London City (Hybrid)
- Positions: 2 available
- Start Date: ASAP