Infrastructure Tooling & Observability Engineer
radiant · London
Job description
About the role
We are looking for an Infrastructure Tooling & Observability Engineer to join our fast‑growing GPU‑as‑a‑Service company. Working closely with SRE and platform teams, you will design and build internal control‑plane systems that turn high‑volume telemetry into actionable insight, improving reliability, efficiency and performance of our global infrastructure.
Key responsibilities
- Design, develop and evolve internal tooling and observability platforms for large‑scale distributed environments.
- Transform logs, metrics and events into actionable insight, enhancing alerting quality and operational decision‑making.
- Translate SRE reliability requirements into production‑ready software, including automated incident detection and remediation.
- Automate infrastructure operations such as environment provisioning, cluster onboarding, inventory management and lifecycle workflows.
- Build tooling for capacity planning, performance testing, benchmarking and automated result analysis.
- Integrate and extend existing Ruby/Rails and Go services, and maintain automation workflows using Ansible and AWX.
Required profile
- Degree in Computer Science, Software Engineering or equivalent experience.
- 6–8 years of experience in infrastructure engineering, DevOps, SRE or software engineering.
- Proven experience working with large‑scale or hyperscale distributed infrastructure environments.
Required skills
- Ruby (Rails) and/or Go programming.
- Ansible and AWX automation.
- Grafana stack – Prometheus, Loki, Mimir, Grafana Alloy.
- Kubernetes orchestration.
- CI/CD pipelines with GitHub Actions and self‑hosted runners.
- REST API design and integration.
- Telemetry protocols such as SNMP and syslog.
What we offer
- Exposure to large‑scale distributed infrastructure systems.
- Opportunity to shape foundational internal platforms.
- Collaborative, engineering‑led culture with strong ownership.
- High‑impact work spanning observability, automation and reliability.
- Close partnership with SRE and infrastructure engineering teams.
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in the United Kingdom.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 1 hour ago
Expires 1 month from now
3 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
radiant
London