Infrastructure Site Reliability Engineer
radiant · Gloucestershire
Job description
About the role
Radiant is building purpose‑built AI infrastructure and needs an experienced Infrastructure Site Reliability Engineer to run and evolve its bare‑metal, virtualization and orchestration stack. You will ensure 24/7 stability, security and performance for AI/HPC workloads while mentoring teammates and improving automation.
Key responsibilities
- Deploy and operate resilient, scalable infrastructure supporting AI/HPC workloads.
- Optimize Linux system configuration, BIOS/firmware, kernel and disk subsystems for performance.
- Configure, monitor and manage bare‑metal infrastructure using IPMI, Redfish, PXE, etc.
- Build and maintain automation scripts and infrastructure‑as‑code (Bash, Python, Ansible).
- Maintain observability stack (Prometheus, Grafana) and handle incident, major incident and change management processes.
- Participate in on‑call rotation, incident post‑mortems and automate remediations.
Required profile
- 5+ years experience in globally scaled, performance‑intensive environments with 24/7 support.
- Expert‑level Linux administration (Ubuntu) and system‑tuning expertise.
- Strong networking fundamentals (TCP/IP, DNS, DHCP, VLANs, routing, switching).
- Experience with orchestration platforms such as Kubernetes, MAAS and Tinkerbell.
- Excellent communication and mentorship skills.
Required skills
- Linux (Ubuntu)
- Bash, Python, Ansible
- IPMI, Redfish, PXE
- Prometheus, Grafana
- Kubernetes, MAAS, Tinkerbell
- TCP/IP, DNS, DHCP, VLANs, routing, switching
- ITSM, Incident Management, Change Management
What we offer
- 25 days of annual leave.
- Private medical insurance via Bupa.
- Learning time for personal development.
- Cycle‑to‑Work scheme and Gympass subscription.
- Company shares programme and enhanced parental pay.
- Inclusive, results‑focused culture with open communication.
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in the United Kingdom.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 1 day ago
Expires 1 month from now
2 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
radiant
Gloucestershire
Related job offers
-
HPC Infrastructure Site Reliability Engineer
radiant Gloucestershire -
Platform Site Reliability Engineer
radiant Gloucestershire -
Senior C++ Engineer
Ncounter Gloucestershire -
Senior Mobile App Developer – Native iOS & Android (Remote, UK)
Round Table DevOps Series #RTDoS Salary Avaratra -
Senior Recovery Specialist – Cyber Incident Response
cypfer London