M&T Bank

M&T Bank

Full-service banking with mortgage, deposits, loans

Manager, Site Reliability Engineering

Full-Time
$139.7k - $232.9k/yr
Expert
Bachelor's, Master's
Buffalo, NY, USA
In Person

About the job

Requirements
  • A combined minimum of 11 years of higher education and/or work experience, including a minimum of 4 years of engineering and/or architecture experience and 5 years of leadership experience, including people management.
  • Complete understanding of the Software Development Life Cycle and modern software delivery practices.
  • Demonstrated leadership delivering enterprise-wide Site Reliability Engineering, reliability engineering, production engineering, or comparable technology capabilities, including operating models and adoption mechanisms.
  • Proven ownership of Software Development Life Cycle reliability and operational readiness controls, enforceable standards, templates, quality gates, and supporting evidence requirements.
  • Experience establishing governance for Service Level Objectives, Service Level Indicators, error budgets, observability, release readiness, deployment safety, resiliency, and production operations.
  • Experience with release and deployment orchestration, including progressive delivery practices such as blue-green deployments, canary releases, rolling deployments, feature flags, automated health validation, and rollback.
  • Strong execution management skills, including prioritization, dependency management, resource planning, and institutionalizing solutions beyond direct involvement.
  • Experience managing complex initiatives, large system enhancements, cloud or platform transformations, application conversions, and production issue resolution.
  • Experience leading multidisciplinary engineering teams and coordinating delivery across application, infrastructure, platform, cloud, operations, risk, and business organizations.
  • Confidence leading multiple teams across geographies and time zones.
  • Prior experience presenting technology strategy, delivery progress, operational performance, investment needs, and material risks to senior management.
  • Strong understanding of Site Reliability Engineering practices, including Service Level Objectives, Service Level Indicators, error budgets, observability, incident and problem management, Root Cause Analysis, operational readiness, resiliency, capacity planning, and toil reduction.
  • Strong understanding of cloud-native and hybrid architectures, distributed systems, application programming interfaces, containers, service dependencies, and Microsoft Azure.
  • Experience leading observability capabilities using Dynatrace, OpenTelemetry, Azure Monitor, Application Insights, Log Analytics, or comparable technologies.
  • Experience governing Infrastructure as Code, preferably using Terraform, and continuous integration/continuous delivery pipeline practices.
  • Understanding of automated regression, performance, resiliency, recovery, and post-deployment testing.
  • Knowledge of release quality gates, traffic management, automated rollback, self-healing, high-availability, and disaster-recovery patterns.
  • Ability to provide credible leadership to senior engineers and evaluate architectural proposals for reliability, release, and operational risks.
  • Excellent verbal and written communication skills, with strong analytical, decision-making, negotiation, and organizational capabilities.
Responsibilities
  • Establish and execute the vision, strategy, operating model, service offerings, and multiyear roadmap for the Site Reliability Engineering Center of Excellence.
  • Define enterprise Site Reliability Engineering standards, engineering practices, governance processes, engagement models, and maturity expectations.
  • Lead the adoption of reliability engineering practices across application development, infrastructure, platform engineering, cloud engineering, and technology operations.
  • Translate enterprise technology, business, resiliency, and risk priorities into an actionable Site Reliability Engineering portfolio and delivery roadmap.
  • Establish a scalable Site Reliability Engineering service model including consulting, enablement, embedded engineering, Forward Deployed Site Reliability Engineering engagements, reusable capabilities, and sustained ownership by application and platform teams.
  • Define intake, prioritization, engagement, transition, and exit criteria for Site Reliability Engineering services.
  • Develop and maintain a Site Reliability Engineering maturity model used to assess service teams, identify reliability gaps, and guide improvement plans.
  • Establish communities of practice, technical forums, training programs, playbooks, reference architectures, and reusable engineering patterns that expand Site Reliability Engineering capabilities across the organization.
  • Represent the Site Reliability Engineering organization in senior leadership forums, architecture reviews, operational governance meetings, and enterprise transformation initiatives.
  • Lead the Forward Deployed Site Reliability Engineering program, placing Site Reliability Engineering professionals into high-priority application and platform teams to address complex reliability challenges and improve operational maturity.
  • Manage program managers, Site Reliability Engineering leaders, and technical leads responsible for coordinating engagements across multiple technology domains.
  • Establish a transparent intake and prioritization process based on customer impact, service criticality, operational risk, incident history, reliability maturity, strategic importance, and anticipated business value.
  • Partner with application and platform leaders to define engagement objectives, scope, deliverables, staffing, success measures, dependencies, and duration.
  • Ensure Forward Deployed Site Reliability Engineering teams deliver sustainable engineering improvements rather than becoming long-term substitutes for application support or production operations.
  • Establish shared accountability for participation, knowledge transfer, remediation activities, and long-term ownership of implemented reliability capabilities.
  • Develop transition and exit plans that enable application and platform teams to sustain Site Reliability Engineering practices after engagements conclude.
  • Evaluate engagement effectiveness using measurable outcomes such as availability, Service Level Objective attainment, incident frequency, restoration time, alert quality, automation adoption, toil reduction, change-failure rate, deployment reliability, and engineering maturity.
  • Convert common findings and lessons learned into reusable standards, automation, tools, training, and engineering patterns.
  • Continuously optimize the Forward Deployed Site Reliability Engineering operating model based on demand, capacity, outcomes, stakeholder feedback, and changes in technology strategy.
  • Lead and develop an organization of employees and contingent resources through direct and indirect management relationships.
  • Manage and develop program managers, Site Reliability Engineering managers, technical leaders, and senior engineering professionals.
  • Establish clear roles, responsibilities, decision rights, performance expectations, and accountability across the Site Reliability Engineering organization.
  • Recruit, retain, coach, and develop engineering and program management talent.
  • Conduct workforce, capacity, and succession planning to ensure the organization has the leadership and technical capabilities required to meet current and future demand.
  • Define career paths and skill-development plans for Site Reliability Engineering professionals in partnership with engineering and human resources leaders.
  • Provide regular feedback, performance management, recognition, coaching, and development opportunities.
  • Manage staffing and sourcing strategies, including the appropriate use of employees, contingent labor, managed services, and specialized partners.
  • Ensure third-party resources and service providers meet applicable engineering, security, risk, performance, financial, and contractual expectations.
  • Establish standards and governance for Service Level Indicators, Service Level Objectives, error budgets, availability targets, and reliability reporting.
  • Partner with service owners and business stakeholders to align reliability objectives with customer expectations, service criticality, business impact, risk tolerance, and cost.
  • Define common reliability metrics and executive-level reporting for service health, operational performance, deployment health, risk, and improvement outcomes.
  • Establish governance for error-budget decisions, including remediation, release-risk evaluation, delivery tradeoffs, escalation, and investment prioritization.
  • Drive adoption of reliability-by-design practices throughout the Software Development Lifecycle.
  • Establish operational readiness requirements for applications and platforms entering production or undergoing material change.
  • Partner with architecture and engineering leaders to incorporate reliability requirements into solution designs, architecture reviews, and engineering standards.
  • Identify systemic reliability risks and sponsor cross-organizational remediation programs.
  • Establish enterprise Site Reliability Engineering standards for safe, repeatable, observable, and recoverable application and infrastructure releases.
  • Partner with application development, platform engineering, cloud engineering, quality engineering, change management, and technology operations teams to improve release and deployment practices.
  • Define approved deployment patterns based on service criticality, architecture, customer impact, technical capability, and risk.
  • Promote progressive delivery practices, including blue-green deployments, canary releases, rolling deployments, ring-based deployments, feature flags, traffic splitting, and controlled production experimentation where appropriate.
  • Establish requirements for automated pre-deployment, in-deployment, and post-deployment validation using technical health signals, service-level indicators, business metrics, and customer-experience measures.
  • Promote deployment orchestration that integrates continuous integration/continuous delivery pipelines, Infrastructure as Code, automated testing, observability, approval controls, and policy enforcement.
  • Establish standards for automated rollback, roll-forward, traffic evacuation, deployment pausing, and recovery when release health thresholds are breached.
  • Define release health criteria and quality gates based on error rates, latency, saturation, availability, dependency health, business transactions, and customer-impact indicators.
  • Partner with platform teams to provide reusable deployment templates, pipeline capabilities, policy-as-code controls, and self-service release patterns.
  • Establish traceability between changes, deployments, configuration updates, incidents, service telemetry, and business outcomes.
  • Improve release observability through deployment markers, version-aware dashboards, automated change correlation, and real-time health analysis.
  • Measure and improve deployment frequency, lead time for changes, change-failure rate, rollback effectiveness, failed deployment recovery time, and release-related customer impact.
  • Ensure deployment practices comply with applicable technology risk, cybersecurity, change-management, segregation-of-duties, and regulatory requirements.
  • Define the enterprise observability strategy for supported services, including standards for telemetry, distributed tracing, metrics, logging, dashboards, alerting, dependency mapping, and customer-experience monitoring.
  • Govern the effective implementation and use of Dynatrace, OpenTelemetry, Azure Monitor, Application Insights, Log Analytics, and other approved enterprise tools.
  • Drive standardization of telemetry and observability patterns across cloud-native, hybrid, distributed, and legacy environments.
  • Establish expectations for actionable alerts, signal quality, service-health visibility, release observability, and real-time operational insight.
  • Reduce alert fatigue and operational noise by improving monitoring coverage, alert thresholds, routing, correlation, suppression, ownership, and automation.
  • Sponsor reusable dashboards, instrumentation patterns, automation libraries, reference architectures, and reliability controls.
  • Develop and execute a strategy to reduce operational toil through automation, self-service capabilities, automated recovery, and self-healing solutions.
  • Establish measurable toil-reduction goals and ensure reclaimed engineering capacity is redirected toward reliability, automation, and product improvements.
  • Promote Infrastructure as Code using Terraform and other approved technologies to improve repeatability, control, recoverability, and environment consistency.
  • Partner with Azure and enterprise platform teams to improve scalability, resiliency, deployment safety, lifecycle management, and operational controls.
  • Provide leadership and executive coordination during significant technology incidents affecting customers, critical business services, or enterprise operations.
  • Ensure appropriate technical leadership, stakeholder communication, escalation, decision-making, and recovery focus during high-severity events.
  • Partner with incident management, technology operations, application teams, infrastructure teams, and business leaders to minimize customer impact and restore services safely.
  • Establish expectations for timely, objective, and technically rigorous Root Cause Analysis.
  • Ensure corrective and preventive actions address underlying technical, process, monitoring, testing, release, deployment, architecture, and organizational causes.
  • Track material reliability actions to completion and escalate overdue or inadequately addressed risks.
  • Analyze incident, change, deployment, and operational telemetry to identify recurring failure patterns and enterprise-level improvement opportunities.
  • Sponsor proactive problem-management efforts that reduce repeat incidents and improve service stability.
  • Establish standards for resiliency testing, fault-tolerance validation, performance testing, capacity planning, disaster recovery, and recovery automation.
  • Promote game days, controlled failure testing, and other approved validation methods to test service behavior and organizational readiness.
  • Partner with technology resiliency and continuity teams to ensure recovery capabilities meet applicable business and regulatory requirements.
  • Ensure lessons from incidents, failed changes, deployment events, and resiliency exercises are incorporated into engineering standards, training, tooling, and future designs.
  • Communicate material reliability risks, customer impacts, remediation strategies, and progress to senior management and governance bodies.
  • Manage a portfolio of Site Reliability Engineering initiatives and Forward Deployed Site Reliability Engineering engagements spanning multiple lines of business, platforms, applications, and technology domains.
  • Establish portfolio governance, delivery cadences, health reporting, dependency management, escalation paths, and outcome tracking.
  • Provide direction and oversight to program managers responsible for planning, sequencing, coordinating, and reporting Site Reliability Engineering initiatives.
  • Define measurable objectives, key results, milestones, and success criteria for the Site Reliability Engineering Center of Excellence and Forward Deployed Site Reliability Engineering program.
  • Monitor execution against approved roadmaps, budgets, commitments, capacity, risk tolerances, and benefit expectations.
  • Resolve delivery barriers that span organizational boundaries and escalate decisions requiring senior leadership involvement.
  • Establish demand-management and capacity-planning processes for Site Reliability Engineering services.
  • Prioritize investments based on customer impact, operational risk, service criticality, strategic alignment, regulatory considerations, and measurable reliability benefit.
  • Develop business cases for staffing, tooling, platform capabilities, release engineering, automation, observability, and engineering transformation.
  • Manage applicable budgets, forecasts, vendor relationships, contracts, licensing requirements, and financial commitments.
  • Demonstrate the value of Site Reliability Engineering investments through measurable improvements in operational performance, customer experience, risk reduction, release quality, and engineering productivity.
  • Build trusted partnerships with senior leaders across application engineering, infrastructure, cloud, platform engineering, architecture, cybersecurity, technology operations, risk, audit, and business-aligned technology organizations.
  • Advise technology and business leaders on reliability risks, service-health trends, engineering priorities, deployment risks, operational readiness, and investment decisions.
  • Communicate complex technical issues, tradeoffs, risks, and recommendations in language appropriate for technical, business, risk, and executive audiences.
  • Present Site Reliability Engineering strategy, operating metrics, program status, investment needs, material risks, and reliability outcomes to leadership and governance forums.
  • Facilitate decisions where reliability, delivery speed, cost, customer impact, and operational risk must be balanced.
  • Influence teams outside the direct reporting structure to adopt enterprise Site Reliability Engineering standards and resolve cross-functional reliability concerns.
  • Establish clear accountability among the Site Reliability Engineering organization, product teams, service owners, platform teams, release teams, and technology operations.
  • Understand and adhere to the Company’s risk and regulatory standards, policies, and controls in accordance with the Company’s Risk Appetite.
  • Ensure Site Reliability Engineering practices, tools, automation, release processes, and operating models comply with applicable technology, cybersecurity, data, privacy, resiliency, change-management, and third-party risk requirements.
  • Identify, assess, document, and escalate material reliability, operational, technology, program, release, and control risks.
  • Ensure appropriate preventive and detective controls are incorporated into Site Reliability Engineering standards, automation, continuous integration/continuous delivery pipelines, deployment processes, observability solutions, and operational workflows.
  • Maintain internal control standards, including timely implementation of internal and external audit findings and regulatory requirements as applicable.
  • Support examinations, audits, risk assessments, control testing, and remediation activities related to the Site Reliability Engineering organization and supported services.
  • Ensure risk acceptances, exceptions, and material reliability or deployment decisions are documented, approved, monitored, and escalated in accordance with Company requirements.
  • Establish governance that provides transparency into unresolved reliability risks and associated remediation commitments.
  • Complete other related duties as assigned.
Desired Qualifications
  • Bachelor’s or Master’s degree in engineering, computer science, information systems, business administration, or a related discipline.
  • Minimum of 12 years of technology management or large program leadership experience.
  • Experience leading a Site Reliability Engineering Center of Excellence, production engineering organization, platform engineering function, or Forward Deployed Site Reliability Engineering operating model.
  • Hands-on familiarity with enterprise observability, incident management, service management, deployment orchestration, and engineering workflow platforms.
  • Experience building reusable Site Reliability Engineering capabilities, automation frameworks, deployment patterns, reference architectures, and self-service engineering solutions with multi-team and vendor coordination.
  • Exposure to artificial intelligence-assisted operations, AIOps, automated incident correlation, predictive analytics, performance engineering, or chaos engineering capabilities.
  • Subject matter expertise in business-critical applications, distributed platforms, and integrated systems.
  • Strong understanding of the Bank’s application framework, business plan, risk environment, and strategic objectives.
  • Proven mentoring, coaching, organizational development, and enterprise leadership capabilities.

About the company

M&T Bank is a full-service financial institution that offers a wide range of banking solutions for individuals, small businesses, and larger enterprises. Its products include personal and business checking accounts, mortgage assistance programs, loans, deposits, investment products, and mobile banking. The bank serves mainly customers in the Northeastern and Mid-Atlantic United States and emphasizes community engagement and customer-focused service. It generates revenue from interest income, fees, and service charges tied to loans, deposits, and other financial products. The company’s recent merger with United Bank, N.A. expands its footprint and enhances its service offerings. The goal is to provide accessible financial services to communities while growing its market presence and deepening local connections.

Company Size

10,001+

Company Stage

IPO

Headquarters

Buffalo, New York

Founded

1993

Get referred to M&T Bank

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • On July 15, 2026, M&T reported record quarterly EPS of $5.32.
  • Average loans rose $3.0 billion in Q2 2026, led by commercial lending and CRE.
  • M&T repurchased $465 million of stock in Q2 2026 while CET1 stayed 10.19%.

What critics are saying

  • Average deposits fell $0.7 billion in Q2 2026, tightening funding and forcing pricier borrowings.
  • Commercial real estate balances reached $24.5 billion; 2027 refinancing stress hits credit costs and earnings.
  • A severe CRE downturn forces reserve spikes, buyback cuts, and capital-trapping regulation through 2027.

What makes M&T Bank unique

  • M&T Bank’s Q2 2026 balance sheet paired $168.9 billion deposits with $141.4 billion loans.
  • Ninety percent of C&I businesses grew quarter-over-quarter in Q2 2026, proving deep relationship banking.
  • The People’s United merger expanded M&T’s Northeast-Atlantic branch and deposit footprint after 2022.

Help us improve and share your feedback! Did you find this helpful?

Benefits

401(k) Company Match

401(k) Retirement Plan

Flexible Work Hours

Hybrid Work Options

Paid Vacation

Paid Holidays

Health Insurance

Dental Insurance

Vision Insurance

Life Insurance

Disability Insurance

Health Savings Account/Flexible Spending Account

Growth & Insights and Company News

Headcount

6 month growth

↑ 10%

1 year growth

↑ 11%

2 year growth

↑ 11%
TipRanks
Sep 18th, 2026
Constellation Brands secures $300M delayed draw term loan for corporate flexibility

Constellation Brands has entered into a new term loan credit agreement worth up to $300 million with Manufacturers and Traders Trust Company and other lenders. The delayed draw facility, established on 18 September 2026, will be used for general corporate purposes, including potential debt repayment. The two-year term loans carry interest linked to Term SOFR or a base rate with rating-based margins. Unused commitments will incur a ticking fee beginning 30 days after closing. The agreement includes covenants that mirror the company's existing revolving credit facility. These impose limits on subsidiary indebtedness, liens, mergers, affiliate transactions, and sale-leasebacks, whilst requiring minimum interest coverage and maximum net leverage ratios. The structure is designed to reinforce Constellation Brands' financial flexibility and balance-sheet management.

Yahoo Finance
Sep 15th, 2026
M&T Bank gains 18% YTD, outpacing Nasdaq amid strong loan growth and rising net interest income

M&T Bank Corporation has outperformed the Nasdaq Composite over recent periods, with shares gaining 18% year-to-date and 19.6% over the past 52 weeks, compared to the Nasdaq's 12.7% and 18.3% returns respectively. Over the past three months, the Buffalo-based bank's shares rose 2.6% against the Nasdaq's 1.2% gain. The outperformance follows strong second-quarter results reported in July, with record earnings per share of $5.32, up 25.5% year-on-year. Taxable-equivalent net interest income increased 4.8% to $1.80 billion, whilst total average loan balances grew 4.4%. Nonaccrual loans fell 23.2% to $1.2 billion, indicating improved asset quality. M&T Bank has a market capitalisation of approximately $34.6 billion.

Business Wire
Sep 7th, 2026
Cari Raises Over $30 Million in First Tranche of Initial Funding Round, Backed Entirely by Banks

Cari, the bank-governed digital money network, today announced it has raised $32.5 million in the first tranche of its initial external funding round, with t...

Yahoo Finance
Aug 28th, 2026
M&T Bank beats Q2 revenue estimates by 1.9% as regional banks show mixed results

Regional banks delivered mixed second-quarter results, with revenues matching analyst expectations on average. M&T Bank reported revenues of $2.52 billion, up 4.9% year on year and exceeding consensus estimates by 1.9%. The bank also beat earnings per share projections and narrowly surpassed net interest income forecasts. Despite the strong quarter, M&T Bank's shares fell 1.3% following the announcement, trading at $238.67. The decline suggests investor expectations may have been higher than published Wall Street projections. Regional banks as a group have struggled, with share prices down an average of 1.5% since their latest earnings releases. These institutions face challenges including fintech competition, deposit outflows, credit deterioration, regulatory costs, and concerns about commercial real estate exposure following recent high-profile bank failures.

Pulse 2.0
Aug 7th, 2026
Charles River Associates expands credit facility to $400M with six-bank syndicate

Charles River Associates has expanded its credit facility to $400 million through a six-bank lending syndicate. The five-year agreement includes a $75 million term loan and a $325 million revolving credit facility, replacing a previous $300 million facility set to mature in August 2027. The revolving facility can be reduced to $250 million between 16 July and 15 January each year when working-capital requirements are lower. CRA will use proceeds to repay outstanding borrowings, support working capital, fund growth investments, and cover general corporate purposes. The lending group includes Bank of America, Citizens Financial Group, Eastern Bank, and Beacon Bank & Trust, with BMO and M&T Bank joining as new lenders. The expanded facility provides additional liquidity as the consulting firm invests in its global economic, financial, and management advisory operations.