Director, Production Services Manager
Job Description
hackajob is collaborating with BNY to connect them with exceptional professionals for this role.
\nHead of Production Services Governance, Incident & Problem Management
\nRole Summary
\nThe Head of Production Services Governance, Incident & Problem Management is accountable for the enterprise governance, standards, and performance of Technology Incident Management and Problem Management (including root cause analysis) across BNY’s Platforms. This leader oversees a team that sets the operating model, drives consistent execution, improves quality and speed of restoration, and strengthens auditability and regulatory credibility.
\nThe role is the senior point of accountability for:
\n- \n
- Firm-wide incident/problem governance and ITIL-aligned standards \n
- High-severity incident command and communications frameworks \n
- End-to-end RCA quality and timeliness, including corrective/preventive actions \n
- Regulatory and client-facing incident narratives and responses \n
- Internal oversight engagement with groups such as ORR and ERO \n
- Automation and AI augmentation to modernize and scale incident/problem practices \n
This position partners closely with engineering, SRE/operations, cyber, resiliency, risk, compliance, and business stakeholders to ensure stability, transparency, and continuous improvement of production services.
\nKey Objectives
\n- \n
- Protect service availability and client experience by ensuring rapid restoration and disciplined incident handling. \n
- Improve resiliency and reduce repeat incidents through high-quality problem management, robust RCAs, and effective remediation governance. \n
- Strengthen governance and audit defensibility by ensuring consistent process adherence, evidence capture, and clear accountability. \n
- Modernize production governance through automation, AIOps capabilities, and AI-assisted workflows. \n
- Elevate operational excellence through measurable improvements in MTTR, recurrence, SLA adherence, and control effectiveness. \n
\n
Primary Responsibilities
\n1) Enterprise Incident Management Governance (ITIL)
\n- \n
- Own the Incident Management practice and ensure it is implemented consistently across Platform Production Services and aligned to ITIL principles. \n
- Establish and maintain incident taxonomy, severity models, prioritization rules, escalation paths, and functional/organizational RACI. \n
- Define Major Incident Management (MIM) framework: incident command roles, war-room orchestration, communications cadence, stakeholder engagement, and decision rights. \n
- Ensure end-to-end controls: accurate incident logging, categorization, impact assessment, timeline reconstruction, evidence retention, and closure criteria. \n
- Drive performance through standard KPIs (e.g., MTTA/MTTR, reopen rate, SLA compliance, major incident frequency, customer-impact minutes, incident backlog health). \n
2) Enterprise Problem Management & RCA Excellence (ITIL)
\n- \n
- Own the Problem Management practice including proactive problem identification, trending, and prevention of recurrence. \n
- Establish RCA standards (methodologies such as 5 Whys, fishbone, fault tree, “cause–trigger–control gap” framing) and ensure consistent quality across teams. \n
- Govern Corrective and Preventive Action (CAPA) management: remediation backlog, prioritization, due dates, owner accountability, and validation of effectiveness. \n
- Maintain governance for Known Errors and Workarounds, enabling faster recovery and better knowledge reuse. \n
- Drive systemic improvements by connecting incidents/problems to resiliency risks, architectural weaknesses, control gaps, and engineering quality. \n
3) Regulatory, Client, and Executive Communications & Responses
\n- \n
- Serve as accountable executive for regulatory responses and supervisory requests relating to incidents, outages, recovery actions, RCA findings, and resiliency improvements. \n
- Lead firm readiness for time-sensitive regulatory deliverables—ensuring accuracy, consistency, and defensible evidence. \n
- Coordinate and quality-assure client communications for impactful incidents (internal/external statements, timelines, cause, remediation, and prevention). \n
- Provide clear executive narratives and materials for senior leadership, risk committees, audit committees, and business stakeholders. \n
4) Oversight & Partnership Model (ORR, ERO, Risk, Audit, Compliance)
\n- \n
- Act as the primary interface to internal oversight groups (e.g., ORR, ERO, Operational Risk, Compliance, Internal Audit, and Technology Risk Management). \n
- Ensure incidents/problems are appropriately mapped to relevant governance constructs (e.g., operational risk events where applicable) with clear traceability. \n
- Lead continuous improvement of control coverage and evidence quality to support audits and examinations. \n
- Partner with Resiliency teams to connect operational learning to scenario testing, dependency mapping, recovery planning, and service resiliency metrics. \n
5) Standardization, Quality Assurance, and Continuous Improvement
\n- \n
- Build and run a Quality Management System for incident/problem practices: sampling, assurance reviews, coaching, playbooks, and maturity assessments. \n
- Develop and maintain standard artifacts (runbooks, major incident playbooks, comms templates, RCA templates, PIR guidance). \n
- Run Continual Improvement programs: trend analysis, “top drivers” remediation themes, performance benchmarking, and maturity roadmaps. \n
- Drive adoption of consistent tooling, workflows, and data standards across platforms. \n
6) Automation & AI Enablement (AIOps / Intelligent Operations)
\nThis role is expected to use AI responsibly to improve speed, quality, and scale of incident/problem management while meeting security, privacy, and model-risk expectations.
\nKey AI and automation outcomes include:
\n- \n
- AI-assisted triage: classification, routing, deduplication, and severity recommendation based on history and signals. \n
- Correlation and probable cause insights using telemetry, topology, and change data to identify likely blast radius and suspects. \n
- Automation for repetitive tasks: stakeholder updates, timeline capture, evidence packaging, and post-incident documentation generation. \n
- RCA acceleration: AI-supported timeline reconstruction, log summarization, anomaly explanation, and “similar incident” retrieval. \n
- Knowledge management uplift: automated drafting of knowledge articles/workarounds; improvement suggestions based on recurrence patterns. \n
- Establish governance for AI usage: model transparency, human-in-the-loop controls, data handling, audit logs, and bias/quality monitoring. \n
7) Leadership & Talent Development
\n- \n
- Lead and develop a high-performing team of incident/problem governance professionals (e.g., problem managers, automation analysts). \n
- Establish role clarity, training paths, and ITIL-aligned capability development. \n
- Foster a culture of calm, disciplined execution during crises and a learning culture post-incident—focused on prevention, not blame. \n
\n
Scope & Decision Rights
\n- \n
- Enterprise-level authority to define and enforce incident/problem standards and minimum controls. \n
- Authority to convene major incident response, direct escalations, and require timely executive updates. \n
- Authority to gate incident/problem closure based on quality criteria (documentation, evidence, RCA completeness, CAPA commitments). \n
- Joint governance with engineering/production leaders to prioritize remediation work and measure effectiveness. \n
\n
Key Interfaces
\n- \n
- Platform Production Services leaders, SRE/Operations, Engineering, Architecture \n
- Cybersecurity Operations, Fraud/Financial Crime Technology (as relevant) \n
- Enterprise Resiliency Office (ERO) \n
- Office of Regulatory Relations (ORR) \n
- Operational Risk, Compliance, Legal, Privacy \n
- Internal Audit, Technology Risk Management \n
- Business/Product leadership and client coverage teams \n
\n
Required Qualifications
\n- \n
- 10–15+ years in technology operations, SRE/production services, service management, or resiliency roles in complex enterprises; regulated financial services strongly preferred. \n
- Demonstrated leadership in Major Incident Management and Problem Management/RCA at enterprise scale. \n
- Strong command of ITIL practices (Incident, Problem, Monitoring & Event, Service Level, Change Enablement, Continual Improvement; familiarity with CMDB/Service Configuration is a plus). \n
- Proven experience driving process standardization, operating model change, and measurable performance improvements (e.g., MTTR reduction, recurrence reduction). \n
- Experience leading regulatory/audit-facing responses with strong evidence discipline and executive communication. \n
\n
Preferred Qualifications / Certifications
\n- \n
- ITIL 4 Managing Professional (MP) and/or ITIL Strategic Leader (SL); ITIL Foundation minimum. \n
- Familiarity with ISO/IEC 20000, NIST, and resiliency/operational risk expectations in financial services (helpful but not required). \n
- Experience with AIOps platforms/observability tooling (e.g., event correlation, log analytics, tracing, anomaly detection). \n
- Experience with Agile/DevOps/SRE operating models and integrating incident/problem practices into product/platform delivery. \n
\n
Core Competencies (What “Great” Looks Like)
\n- \n
- Crisis leadership: calm command presence, structured decision-making, clear communications under pressure. \n
- Governance rigor: sets standards that are pragmatic, scalable, and audit-defensible. \n
- Analytical excellence: uses trends and data to drive prevention, not just restoration. \n
- Influence without friction: partners effectively with engineering leaders to get remediation done. \n
- Automation mindset: removes manual steps, improves quality through workflow and tooling. \n
- AI fluency with controls: leverages AI safely with strong human oversight and evidence trails. \n
\n
Success Metrics (Illustrative)
\n- \n
- Reduced major incident frequency and customer-impact minutes (YoY). \n
- Improved MTTR/MTTA and decreased escalations due to better routing/triage. \n
- Increased RCA timeliness and quality scores, fewer incomplete RCAs, higher CAPA completion on time. \n
- Reduced repeat incidents driven by top recurring causes. \n
- Improved audit/regulatory outcomes: fewer findings, faster response cycles, higher evidence quality. \n
- Increased automation coverage: % of incidents with AI-assisted classification/correlation; reduction in manual documentation hours. \n
At BNY, our culture allows us to run our company better and enables employees’ growth and success. As a leading global financial services company at the heart of the global financial system, we influence nearly 20% of the world’s investible assets. Every day, our teams harness cutting-edge AI and breakthrough technologies to collaborate with clients, driving transformative solutions that redefine industries and uplift communities worldwide.
\nRecognized as a top destination for innovators, BNY is where bold ideas meet advanced technology and exceptional talent. Together, we power the future of finance – and this is what #LifeAtBNY is all about. Join us and be part of something extraordinary.
