Jobiglo

No results.

Engineering Manager – Production Inference

deepl · London

New
Hybrid Senior 🇬🇧 English
Production ML systems Inference optimisation Model serving at scale LLM inference Speculative decoding Quantisation Serving infrastructure Hardware utilisation

Job description

About the role

DeepL is looking for an Engineering Manager to lead the Production Inference team, responsible for delivering reliable, low‑latency model serving at scale. You will guide a group of research scientists and ML engineers, shaping both technical direction and people development.

Key responsibilities

  • Lead and develop a high‑performing team, creating development plans and fostering a feedback‑rich culture.
  • Own the research and development roadmap for production inference systems, balancing reliability commitments with long‑term efficiency research.
  • Act as the primary technical liaison between Production Inference and adjacent functions such as foundational models, voice research, infrastructure, and product.
  • Drive reliability, efficiency, and cost performance of the model serving stack, influencing load‑balancing, autoscaling, runtime selection and hardware utilisation.
  • Identify, assess, and recruit research and engineering talent as the team grows.

Required profile

  • Proven experience leading researchers or ML engineers, with a track record of talent development and delivery rigour.
  • Strong background in Computer Science, Mathematics, Physics or a comparable quantitative discipline, or equivalent ML/systems expertise.
  • Excellent communication skills, able to translate complex technical direction for both technical and non‑technical stakeholders.
  • Solution‑oriented and decisive, capable of defining direction without waiting for higher‑level guidance.

Required skills

  • Production ML systems
  • Inference optimisation
  • Model serving at scale
  • LLM inference
  • Speculative decoding
  • Quantisation
  • Serving infrastructure (load balancing, autoscaling, runtime selection)
  • GPU inference runtimes and hardware utilisation

What we offer

  • Diverse, internationally distributed team across more than 90 nationalities.
  • Hybrid work schedule with flexible hours and two days per week in the office.
  • Virtual Shares giving employees a stake in DeepL’s growth.
  • Regular in‑person team events and monthly full‑day hack sessions.
  • 30 days of annual leave plus mental‑health resources.

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec deepl.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.
Source : ats:ashby

Why are you reporting this job?

Thank you for your report. We will review this job.

Explore further

Salaries, guides and searches in the United Kingdom.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published 7 hours ago

Expires 1 month from now

5 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

deepl

London