Financial Services | Infrastructure Engineering & Technology Risk
|
Level:
|
Senior Individual Contributor
|
|
Function:
|
Infrastructure Engineering / Technology Resiliency
|
|
|
|
We are a [retail bank / investment firm / payments infrastructure
provider] operating in a regulated financial services environment. We are
looking for a Senior Backup & Restore Engineer to own the design,
implementation, and assurance of our data protection and recovery
capabilities across our hybrid estate.
This is a high-trust, high-accountability role. Our backup infrastructure
is not just an operational concern — it is a regulatory obligation, a
cyber defence layer, and a core component of our operational resilience
framework. You will be the subject matter expert who ensures our ability
to recover is proven, not assumed, and that our posture satisfies the
expectations of regulators including the FCA, PRA, and the requirements of
DORA.
You will work in close partnership with Information Security, Technology
Risk, Business Continuity, and SRE teams, and will interface directly with
audit and regulatory processes.
•
Architect, implement, and operate backup and restore solutions across
on-premises data centers and cloud environments (AWS, Azure, and/or GCP),
covering the full range of financial services workloads — core banking
systems, trading platforms, payment processing, databases, and
unstructured data.
•
Define, own, and continuously validate RTO and RPO targets in alignment
with Important Business Service (IBS) mapping and impact tolerances as
required under the PRA/FCA operational resilience framework.
•
Own backup policy governance — retention schedules, legal hold processes,
data classification alignment, and lifecycle rules — ensuring consistency
with data management and records retention obligations under GDPR and FCA
COBS/SYSC requirements.
•
Lead a structured restore testing programme including full
Business Service recovery
exercises, recovery simulations, and immutable backup integrity validation
— with evidence packages suitable for regulatory review.
•
Own backup failure investigation and resolution end-to-end, conducting
blameless post-incident reviews and feeding findings into the firm’s risk
event process.
•
Design and maintain air-gapped and immutable backup tiers as a core
ransomware and cyber attack recovery capability, aligned to the firm’s
Cyber Recovery Plan and wider BCBS 239 / DORA ICT resilience obligations.
•
Participate in adversarial recovery exercises and cyber simulation
scenarios (e.g. ransomware tabletops,
service
restores under incident conditions) — providing technical leadership on
recoverability.
•
Maintain alignment with NIST CSF, ISO 27001, and DORA ICT risk management
requirements as they apply to backup and recovery controls.
•
Work with the Information Security team to ensure backup systems are
hardened, access-controlled, and monitored for anomalous activity —
treating backup infrastructure as a high-value target requiring its own
threat model.
•
Maintain audit trails and cryptographic integrity verification for backup
data to support forensic and regulatory investigation requirements.
•
Build and maintain all backup infrastructure using Terraform and
supporting IaC tooling (Ansible, CloudFormation, or Pulumi where
applicable), with all changes managed through version-controlled,
peer-reviewed CI/CD pipelines.
•
Automate restore validation workflows so recovery confidence is
continuously measured, evidenced, and reportable — not dependent on manual
spot checks.
•
Manage Terraform state, module libraries, and environment promotion across
non-prod and production estates within a change-controlled, audit-friendly
delivery model.
•
Contribute to shared platform engineering standards and champion IaC best
practices across the wider infrastructure function.
•
Maintain deep expertise across both on-premises (
Cohesity, Rubrik
, NetApp SnapVault, Dell
Data Domain
,
)
IBM Safeguarded Copy)
and cloud-native (AWS Backup, Azure Backup, snapshot-based recovery,
object storage immutability) paradigms.
•
Design and manage data transfer, replication, and egress cost optimisation
for hybrid backup flows, with particular attention to latency and
bandwidth constraints on critical financial data paths.
•
Ensure backup coverage extends to containerised workloads (Velero or
equivalent) as the firm’s Kubernetes adoption grows.
•
Maintain a live and accurate backup estate inventory, coverage map, and
gap register — with regular reporting to Technology Risk and the CTO
organisation.
•
Produce and own the Backup & Recovery Policy, associated standards,
and operational runbooks, ensuring they are reviewed annually and remain
aligned with regulatory expectations.
•
Serve as the primary technical contact for internal audit, external audit
(Big 4), and regulatory examination processes relating to backup,
recovery, and operational resilience.
•
Ensure the firm’s backup posture can satisfy requirements under DORA —
including ICT risk management, resilience testing, and third-party backup
vendor oversight obligations.
•
Support Technology Risk in maintaining and evidencing compliance with
SS2/21 (PRA operational resilience), PS6/21, FCA PS21/3, and relevant
EBA/EIOPA guidelines as they touch recovery capabilities.
•
Provide input to the firm’s ICAAP/ILAAP and recovery plan where ICT
recoverability is a factor.
•
6+ years in infrastructure or platform engineering with a significant
focus on backup, recovery, and data protection, with at least 3 years in a
financial services or other regulated environment.
•
Hands-on production Terraform experience — modules, remote state,
workspaces, policy-as-code, and integration with enterprise CI/CD
pipelines.
•
Demonstrated hands-on experience across both on-premises infrastructure
(VMware,
Dell, Cohesity, Runrik
, Commvault, NetApp, or equivalent) and at least one major cloud provider.
•
Working knowledge of the UK/EU regulatory landscape for operational
resilience and ICT risk — FCA/PRA SS2/21, DORA, BCBS 239, and their
practical implications for backup and recovery.
•
Deep understanding of cyber-resilient backup design: immutability, air-gap
architecture, encryption, integrity verification, and ransomware recovery
patterns.
•
Proven experience designing, executing, and evidencing RTO/RPO validation
and DR testing for audit and regulatory purposes.
•
Strong scripting capability in Python, Bash, or PowerShell for automation,
remediation, and tooling.
•
Direct experience with DORA ICT resilience testing requirements (TLPT) and
how backup systems are scoped within them.
•
Familiarity with cyber recovery frameworks — NIST CSF Recovery function,
NCSC guidance on offline backups, SWIFT CSCF controls.
•
Experience with Kubernetes workload backup (Velero or equivalent) in a
financial services context.
•
Exposure to regulatory examination processes — preparing evidence packs,
responding to findings, and tracking remediation.
•
Experience with data sovereignty, cross-border data transfer controls, and
their implications for cloud backup architecture in a multi-jurisdiction
financial institution.
•
Relevant certifications: AWS/Azure backup and storage certifications,
HashiCorp Terraform Associate or Professional, CISSP, CRISC, or ISO 22301
Lead Implementer.
•
Treats recoverability as a provable property, not a configuration setting
— if it hasn’t been tested under realistic conditions, it doesn’t count.
•
Communicates technical risk clearly and credibly to non-technical
stakeholders including Risk, Compliance, and senior leadership.
•
Comfortable operating in a high-governance environment where changes
require documentation, approval, and audit trails.
•
Raises standards through peer review, runbook quality, and enabling
colleagues — not just personal output.
•
Maintains discretion and professionalism when handling sensitive data,
regulatory findings, and security incidents.
•
Technology Risk & Information Security — backup posture aligned to
firm risk appetite and cyber recovery plan.
•
Business Continuity & Operational Resilience — recovery capabilities
mapped to Important Business Services.
•
Internal & External Audit — evidence provision and finding
remediation.
•
Platform / SRE Teams — IaC standards, incident response, and day-2
operations.
•
Third-party backup vendors — contract oversight, SLA assurance, and DORA
third-party ICT risk obligations.
Apply through whichever channel suits you best.