Senior Cloud Infrastructure Engineer – Open LMS (Remote, UK)
ltg
Job description
About the role
We are looking for a Senior Cloud Infrastructure Engineer to join our team and help build, scale, and evolve our multi‑tenant SaaS hosting platform on AWS. The platform dynamically provisions, manages, and scales hundreds of Moodle LMS instances for education clients, using custom orchestration tooling, distributed service discovery, and infrastructure‑as‑code.
Key responsibilities
- Design, build, and maintain AWS infrastructure (EC2, RDS, S3, SQS, Lambda, ALB, ElastiCache, Route 53, VPC networking) using Terraform.
- Write and maintain Puppet modules to configure and manage fleets of EC2 instances across auto‑scaling groups.
- Develop and extend Python‑based automation and tooling that supports platform operations.
- Operate and improve distributed service discovery and configuration management with etcd.
- Manage and tune a multi‑tier caching strategy (Varnish, Redis/Valkey, PHP OPcache).
- Run and scale the observability stack (Prometheus, Grafana, Loki, Fluentd, PagerDuty) and participate in on‑call rotations.
- Evaluate and implement distributed storage solutions as the platform evolves.
- Improve deployment workflows, release processes, and zero‑downtime strategies for VM‑based environments.
- Collaborate with internal teams on API contracts, integration patterns, and operational tooling.
- Participate in incident response, root‑cause analysis, and platform reliability improvements.
Required profile
- Strong production experience with AWS services, especially EC2, RDS, S3, SQS, Lambda, ALB, ElastiCache, Route 53, IAM, and VPC networking.
- Proficiency in authoring and maintaining Terraform modules for production infrastructure.
- Proficiency in authoring and maintaining Puppet (or equivalent) modules for fleet management.
- Solid Python development skills for building production daemons.
- Deep Linux systems knowledge (Ubuntu) including Apache/Nginx, PHP‑FPM, Varnish, systemd, filesystem mounts, and networking fundamentals.
- Understanding of distributed systems concepts such as consensus, leader election, distributed locking, and eventual consistency.
- Experience building and maintaining observability pipelines (Prometheus, Grafana, Loki or equivalents) in production.
- Comfortable working in a GitLab‑based CI/CD workflow and communicating technical decisions to both technical and non‑technical stakeholders.
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in the United Kingdom.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 8 hours ago
Expires 1 month from now
1 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
ltg