Software Engineer, ChatGPT Infrastructure
openai · London
Job description
About the role
This role is for software engineers who want to build and evolve backend systems operating at significant scale. You’ll write production code, design shared infrastructure, and solve technical challenges involving performance, distributed systems, and system reliability.
Key responsibilities
- Design, build, and maintain backend systems supporting high‑traffic ChatGPT experiences.
- Develop shared services, APIs, and infrastructure that help product teams build and launch new capabilities safely.
- Improve the performance, scalability, and efficiency of production systems as usage and product complexity grow.
- Build and improve systems for asynchronous processing and other large‑scale backend workloads.
- Lead architectural improvements and infrastructure migrations while maintaining correctness, compatibility, and safe rollout and rollback.
- Strengthen monitoring, alerting, and diagnostics to detect problems early and reduce customer impact.
- Participate in on‑call, incident response, and root‑cause analysis, turning operational learnings into lasting engineering improvements.
- Work across product and infrastructure teams to ensure new systems are reliable, secure, and production‑ready.
- Build automation and tooling that reduce repetitive operational work and improve engineering effectiveness.
Required profile
- Have experience building and operating backend systems at scale.
- Understand distributed systems concepts such as concurrency, consistency, asynchronous processing, and failure handling.
- Enjoy writing production code while taking ownership of how systems behave in real‑world conditions.
- Can identify and address bottlenecks affecting latency, throughput, resource usage, or reliability.
- Have experience introducing production changes safely through testing, staged rollouts, monitoring, and rollback plans.
- Are comfortable investigating ambiguous technical problems and collaborating across teams to resolve them.
- Care about clear interfaces, maintainable systems, and practical engineering trade‑offs.
- Take ownership of projects from design and implementation through deployment and ongoing improvement.
Required skills
- Proficiency in at least one general‑purpose programming language.
- Experience designing, building, or improving services, platforms, or shared infrastructure at scale.
- Familiarity with production reliability practices such as monitoring, incident response, and root‑cause analysis.
- Understanding of distributed systems concepts, data storage, concurrency, asynchronous processing, and networking.
- Familiarity with modern deployment, observability, and cloud infrastructure practices.
- Ability to lead complex technical work and collaborate effectively across product and infrastructure teams.
- Experience with cloud infrastructure, containerized environments, or observability tools (useful but not required).
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in the United Kingdom.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 45 minutes ago
Expires 1 month from now
2 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
openai
London