AI Evaluation Engineer
TXP Technology x People · London
Job description
About the role
We are seeking an experienced AI Evaluation Engineer to join a specialist Public Sector Technology team. The role is hands‑on, focusing on building evaluation frameworks, tooling and harnesses for AI systems, especially large language models and agentic AI solutions. You will work directly with government teams to prototype, test and deliver rapid AI evaluations.
Key responsibilities
- Design and develop evaluation frameworks, tooling and repeatable processes.
- Create evaluation harnesses to test model and agentic AI systems.
- Define metrics, methodologies and benchmark AI performance.
- Investigate AI failure modes and present findings with recommendations.
- Collaborate with government stakeholders to prototype and test solutions.
Required profile
- Strong software engineering background with experience building and evaluating AI systems.
- Deep understanding of LLMs, agentic AI and AI evaluation frameworks.
- Ability to work autonomously, solve problems and communicate results effectively.
Required skills
- Python development
- AI/LLM evaluation
- RAG evaluation and Ragas
- Agentic AI evaluation
- Evaluation harness creation
- LLM testing and benchmarking
- Prompt engineering
- Model and agent integration
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in the United Kingdom.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 2 hours ago
Expires 1 day from now
1 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
TXP Technology x People
London