Member of Technical Staff (AI Inference Engineer)
perplexity · London
Job description
About the role
We are looking for an AI Inference Engineer to join our growing team at Perplexity. The role involves building and running the inference engine behind every Perplexity query, deploying dozens of model architectures at scale while meeting tight latency and cost budgets.
Key responsibilities
- Support new model architectures, including transformer‑based retrieval, text‑generation, and multimodal models across the inference stack.
- Migrate in‑house CUDA kernels to NVIDIA’s CuTe DSL for current and future hardware platforms.
- Develop and maintain a Rust‑native serving runtime to replace Python‑based components.
- Profile and eliminate performance bottlenecks from network ingress through continuous batching and GPU kernel interleaving.
- Build dashboards, alerts, and automated remediation to ensure reliability and observability in production.
Required profile
- Deep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar).
- Strong understanding of modern LLM architectures and ability to deploy them reliably.
- Proven track record building and operating production‑grade distributed systems under real load.
- Comfortable working across Rust, Python, and CUDA/CuTe DSL codebases.
- Self‑directed, able to take problems from research paper to production incident resolution.
Required skills
- Rust
- Python
- CUDA
- CuTe DSL
- Triton, CUTLASS
- PyTorch, JAX, TensorFlow
- NCCL, NVLink, InfiniBand, RDMA
- Kubernetes (GPU scheduling, autoscaling)
- Profiling tools such as Nsight Compute, CUDA‑GDB, PTX/SASS analysis
What we offer
- Competitive base salary (determined by experience and expertise).
- Equity participation as part of the total compensation package.
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in the United Kingdom.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
A question about this job?
Ask it here: you will get the full job summary by e-mail, right away.
Published 1 week ago
Expires 1 month from now
8 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
perplexity
London
Related job offers
-
VP of Tech Expansion – Global Technology Growth
Tether Operations Limited London -
IT Network & Security Senior Specialist
Centre People Appointments London -
Data Engineer
Lynx Recruitment Ltd London -
Service Delivery Manager
Think Specialist Recruitment Watford -
AWS Data Engineer – London (Hybrid)
Lynx Recruitment Ltd London