Data Engineer – Financial Instrument Mastering (Python/PySpark)
Queen Square Recruitment · Londres et périphérie
Job description
About the role
You will join a specialist engineering team that builds and optimises end‑to‑end pipelines for mastering financial instruments. Working across Azure and Microsoft Fabric, you will collaborate with data architects, domain experts and quality‑control engineers to deliver high‑performance, reliable data solutions.
Key responsibilities
- Design, develop and maintain PySpark‑based pipelines that ingest, normalise and process financial data from multiple sources.
- Implement bi‑temporal data models (system time + valid time) with Slice, Resolve, Coalesce and Diff logic.
- Create and optimise Azure Cosmos DB data models, including partitioning, indexing and change‑feed processing.
- Integrate external entity‑resolution APIs (e.g., PermID, IAAS) with robust retry and batching.
- Build publication pipelines that convert bi‑temporal data to uni‑temporal outputs for Microsoft Fabric lakehouse architectures.
- Apply Great Expectations to enforce data quality and compliance.
- Write unit and integration tests using PyTest for PySpark and Cosmos DB components.
- Maintain CI/CD pipelines (GitLab CI), manage Python packaging, Artifactory deployment and ARM‑based infrastructure provisioning.
- Configure mastering rules, schemas and environments via YAML files.
- Monitor production pipelines with Eventstream telemetry, KQL and DataDog, and troubleshoot issues.
- Implement data‑governance controls such as masking, role‑based access and compliance policies.
- Continuously tune workloads for performance, cost efficiency and reliability.
Required profile
- Strong experience with Python and PySpark (Spark SQL, DataFrame API, Structured Streaming).
- Hands‑on experience building large‑scale ETL/streaming pipelines.
- Proficiency with Azure Cosmos DB, Azure Data Lake Storage (ADLS/OneLake/ABFS) and related performance tuning.
- Experience designing bi‑temporal or SCD Type 2 data models.
- Knowledge of data‑quality frameworks such as Great Expectations.
- Familiarity with CI/CD tools (GitLab, Azure DevOps) and automated deployments.
- Solid testing discipline using PyTest, mocking and integration testing.
- Experience with YAML/JSON configuration and ARM templates.
- Understanding of distributed data processing and Spark‑based architectures.
- Background working with financial or time‑series datasets is preferred.
- Excellent communication skills for cross‑functional collaboration.
Required skills
- Python
- PySpark
- Spark SQL
- Structured Streaming
- Azure Cosmos DB
- Azure Data Lake Storage (ADLS, OneLake, ABFS)
- Great Expectations
- GitLab CI / Azure DevOps
- PyTest
- YAML
- JSON
- ARM templates
- Microsoft Fabric
- Parquet
- Eventstream telemetry
- KQL
- DataDog
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in the United Kingdom.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
A question about this job?
Ask it here: you will get the full job summary by e-mail, right away.
3 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Queen Square Recruitment
Londres et périphérie
Related job offers
-
Principal Site Reliability Engineer – Up to £250k + Bonus
Hunter Bond Londres et périphérie -
IT Support Coordinator – London
Set2Recruit Londres et périphérie -
Project Manager – MarTech Transformation
Addition Londres et périphérie -
CyberArk SME – Cloud Migration & Privileged Access Management
Hays Specialist Recruitment Limited London -
Technical Dual Run & Cutover Project Manager
Matchtech London