Full-Time

Associate Director

Observability and Service Reliability

Kyndryl

Kyndryl

10,001+ employees

Managed IT infrastructure and cloud services

No salary listed

Toronto, ON, Canada

Remote

Bachelor's

Category
Engineering Management (1)
Required Skills
Dynatrace
Data Visualization
ServiceNow
Machine Learning
OpenTelemetry
Docker
Observability
REST APIs
APM
DevOps
Data Analysis

Get referred to Kyndryl

See people who can refer or advise you

Requirements
  • A bachelor's degree in Information Technology, Computer Science, Engineering, or a related discipline, or equivalent professional experience.
  • At least 8 years of experience in observability, application performance, service reliability, infrastructure, cloud, engineering, or enterprise technology operations.
  • At least 3 years of technical leadership or people leadership experience.
  • Experience designing and operating monitoring and observability capabilities in complex enterprise environments.
  • Experience with Dynatrace and familiarity with Nexthink or comparable Digital Employee Experience technologies.
  • Knowledge of observability technologies, monitoring architectures, application performance management, infrastructure monitoring, event management, logging, tracing, synthetic monitoring, and user experience monitoring.
  • Understanding of modern enterprise applications, cloud technologies, containers, APIs, networks, databases, infrastructure, and distributed architectures.
  • Experience defining service health, reliability measures, Service Level Indicators, Service Level Objectives, and operational performance standards.
  • Ability to influence senior technical leaders and translate complex technical concepts into business risk and operational outcomes.
  • Strong leadership, architecture, analytical, problem-solving, communication, and stakeholder management skills.
Responsibilities
  • Own and mature the enterprise observability and service reliability strategy.
  • Define standards for monitoring applications, infrastructure, cloud platforms, networks, endpoints, APIs, databases, middleware, and other critical technology services.
  • Establish expectations for metrics, logs, traces, events, synthetic monitoring, real user monitoring, digital experience, service health, and business transaction visibility.
  • Create a consistent enterprise approach while allowing teams to use monitoring technologies suited to their platforms and services.
  • Identify monitoring gaps, redundant capabilities, excessive alerting, and opportunities to improve visibility.
  • Move the organization from traditional monitoring toward proactive, predictive, and automated operations.
  • Establish and mature the organization’s service reliability framework.
  • Partner with technical service owners to define monitoring requirements for critical applications and services.
  • Ensure monitoring reflects the complete service, including application performance, infrastructure, dependencies, integrations, user experience, business transactions, capacity, and failure conditions.
  • Define minimum observability requirements based on service criticality and business impact.
  • Help teams establish meaningful Service Level Indicators, Service Level Objectives, availability targets, performance thresholds, and health measures.
  • Use reliability data, incidents, problem records, capacity trends, and telemetry to identify systemic weaknesses and prioritize improvements.
  • Serve as the enterprise technical authority for observability, monitoring, and service reliability.
  • Provide architectural guidance to application, infrastructure, cloud, engineering, DevOps, Site Reliability Engineering, cybersecurity, and operations teams.
  • Influence solution design so services are observable, measurable, supportable, and resilient by design.
  • Develop enterprise monitoring patterns, reference architectures, standards, and reusable capabilities.
  • Guide technical teams in selecting appropriate monitoring methods and technologies for specific platforms and use cases.
  • Evaluate emerging observability, AIOps, automation, analytics, and service reliability capabilities for measurable operational value.
  • Provide strategic oversight for the enterprise observability and monitoring tool ecosystem.
  • Lead the strategy, architecture, governance, adoption, and optimization of major platforms, including Dynatrace, Nexthink, and related enterprise monitoring technologies.
  • Ensure monitoring tools operate as an integrated ecosystem rather than isolated platforms.
  • Establish standards for instrumentation, tagging, alerting, dashboards, integrations, service mapping, ownership, and data quality.
  • Partner with technical teams to maximize platform value while reducing tooling duplication and complexity.
  • Manage strategic technology and vendor relationships to ensure observability investments deliver measurable operational value.
  • Provide strategic leadership for enterprise use of Dynatrace across applications, infrastructure, cloud, and digital services.
  • Drive adoption of application performance monitoring, distributed tracing, real user monitoring, synthetic monitoring, infrastructure monitoring, logs, topology, service health, and intelligent problem detection.
  • Partner with technical owners to improve application instrumentation and ensure Dynatrace provides meaningful service visibility.
  • Use Dynatrace capabilities to improve root cause identification, dependency visibility, anomaly detection, and incident response.
  • Lead the enterprise strategy for Nexthink and Digital Employee Experience.
  • Provide visibility into endpoint health, application performance, employee technology experience, device reliability, and technology friction.
  • Partner with Digital Workplace, Service Desk, application teams, and infrastructure teams to identify and remediate issues proactively.
  • Use experience data to reduce support demand, improve employee productivity, and identify systemic technology issues.
  • Expand automation and targeted remediation to resolve employee experience issues at scale.
  • Improve operational signal quality by reducing noise and increasing alert relevance.
  • Drive event correlation, enrichment, anomaly detection, automated diagnostics, and proactive remediation.
  • Integrate observability platforms with IT service management, incident management, automation, collaboration, configuration, and operational data platforms.
  • Develop capabilities that help support teams detect degradation earlier and understand business impact faster.
  • Use artificial intelligence, machine learning, analytics, and automation where they improve outcomes and reduce manual effort.
  • Establish governance for enterprise monitoring standards, tooling, integrations, data quality, licensing, and adoption.
  • Develop executive and operational reporting that provides insight into service health, reliability, performance, and user experience.
  • Use service measures to identify reliability risks, guide investment decisions, and prioritize service improvements.
  • Lead and develop observability, monitoring, reliability, and platform engineering professionals.
  • Create an engineering culture focused on proactive operations, technical excellence, automation, and measurable service outcomes.
  • Build partnerships with application owners, platform owners, infrastructure, cloud, network, cybersecurity, Digital Workplace, DevOps, Site Reliability Engineering, Service Management, and business teams.
  • Influence technical teams without relying solely on direct reporting authority.
  • Create clear accountability for monitoring coverage and service reliability across the enterprise.
  • Partner with Incident, Problem, Change, Major Incident Management, and Operational Resilience teams to improve detection, recovery, and prevention.
Desired Qualifications
  • Advanced Dynatrace experience in a large enterprise environment.
  • Experience implementing or scaling Nexthink.
  • Experience with ServiceNow and enterprise event management platforms.
  • Experience with Site Reliability Engineering, DevOps, AIOps, automation, OpenTelemetry, cloud-native monitoring, and modern observability architectures.
  • Experience establishing enterprise monitoring standards or observability reference architectures.
  • Experience within complex, global, or highly regulated enterprise environments.

Kyndryl provides managed IT infrastructure services for large enterprises, helping them run and modernize their technology environments. Its offerings cover cloud adoption, cybersecurity, security operations, and digital workplace services, all delivered under long-term managed-service contracts. The company works by partnering with major technology providers (including Microsoft) to offer integrated, end-to-end solutions, often complemented by a new security operations center (SOC) and other security capabilities. This focus sets it apart from competitors through scale, enterprise-grade governance, and a history rooted in IBM’s infrastructure services, now independent but aligned with large technology partners. The goal is to support digital transformation by providing reliable, scalable, and secure infrastructure management that enables enterprises to operate efficiently and securely in a cloud-enabled world.

Company Size

10,001+

Company Stage

IPO

Headquarters

New York City, New York

Founded

2021

Get referred to Kyndryl

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • FY2026 hyperscaler-related revenue rose 59% to $1.9 billion, proving alliance monetization.
  • Kyndryl Consult posted double-digit growth for three straight years, expanding higher-margin advisory demand.
  • September 8, 2026 Citi conference keeps capital-markets visibility high after fiscal 2026 margin gains.

What critics are saying

  • SEC enforcement remains open on cash management and internal controls, following February 9, 2026 disclosures.
  • Southern District class action from February 2026 targets Martin Schroeter and former officers.
  • Revenue fell 3% constant-currency in FY2026; IBM-related headwinds and long sales cycles persist.

What makes Kyndryl unique

  • Since spin-off, Kyndryl runs mission-critical infrastructure for 60+ countries and thousands of enterprises.
  • Broadcom partnership on August 27, 2026 hardens VMware Cloud Foundation private clouds for AI.
  • Brazil SOC and global SOC network give local cybersecurity coverage and incident response.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Professional Development Budget

Hybrid Work Options

Growth & Insights and Company News

Headcount

6 month growth

4%

1 year growth

4%

2 year growth

4%
CxOToday
Sep 8th, 2026
Dassault Systèmes collaborates with Kyndryl to accelerate infrastructure and manufacturing transformation in India.

Dassault Systèmes collaborates with Kyndryl to accelerate infrastructure and manufacturing transformation in India. Dassault Systèmes today announced a collaboration with Kyndryl (NYSE:KD), a leading provider of mission-critical enterprise technology services, as a Consulting and Systems Integrator partner in India. This is Dassault Systèmes' first collaboration with Kyndryl in India, reinforcing its commitment to enable large-scale digital transformation across infrastructure, consumer packaged goods (CPG) and retail, and pharmaceutical manufacturing sectors. The collaboration brings together Dassault Systèmes 3DEXPERIENCE platform and expertise in virtual twin experiences with Kyndryl's leadership in mission-critical IT infrastructure services. Together, the companies aim to address the rapidly evolving needs of India's CapEx-intensive industries by combining deep engineering intelligence with scalable, secure and resilient IT operations. By leveraging virtual twin experiences, the collaboration aims to reduce project delays, optimize costs, and enhance stakeholder collaboration. It will also support sustainability goals, particularly in packaging within the CPG and retail sectors, while enabling pharmaceutical manufacturers to align more efficiently with global regulatory standards, including U.S. Food and Drug Administration compliance. Kyndryl and Dassault Systèmes will focus on joint go-to-market initiatives, combine complementary capabilities and unlock new innovation opportunities to deliver measurable business outcomes for customers. Kyndryl's investments in India's AI innovation, talent development, and hybrid IT transformation further extend the partnership to build out resilient, future-ready infrastructure in the country. "Our collaboration with Kyndryl reflects a shared vision to enable smarter, sustainable, and efficient operations across key sectors. By combining virtual twin experiences with deep IT infrastructure expertise, we aim to help organizations transform the way they design, build, and operate complex systems," said Deepak NG, Managing Director, India, Dassault Systèmes. "This collaboration with Dassault Systèmes underscores our commitment to driving innovation and delivering value for customers in India. By integrating advanced virtual twin technologies with our infrastructure and technology services capabilities, we are well-positioned to support enterprises in navigating complexity improving resilience, and accelerating their digital transformation journeys," said Hitesh Shah, VP - Manufacturing Client Unit, Kyndryl India. This integrated "design-to-operate" approach will help governments, engineering, procurement, and construction players, and infrastructure operators reduce project risks, improve asset lifecycle management, and enable data-driven, citizen-centric services. It also aligns with India's continued focus on digital public infrastructure, smart urbanization and sustainability. With India witnessing significant investments in physical and digital infrastructure, including smart cities, transport networks, utilities, airports, ports and railways, the collaboration is uniquely positioned to support the country's next phase of growth. Dassault Systèmes' virtual twin capabilities enable the design, simulation, and optimization of complex infrastructure systems, while Kyndryl's strengths in cloud, artificial intelligence, cybersecurity, and managed services help deploy and operate these systems at scale.

Kyndryl
Sep 3rd, 2026
Pedro Soares: transforming enterprise knowledge into AI value.

Pedro Soares: transforming enterprise knowledge into AI value. Article Sep 3, 2026 Read time: 4 min For organizations to fully realize the promise of AI, they must trust the data it uses. That's why knowledge management, the practice of identifying, capturing, organizing and optimizing critical institutional data, is so important. Newly appointed Kyndryl Distinguished Engineer Pedro Soares leads the firm's efforts to transform unstructured data and undocumented expertise into actionable information. Drawn from documents, repositories, processes and people, this knowledge can become the source of truth for an organization's AI-enabled IT estate. Ultimately, Pedro's work centers on trust. Here, he explains how data-driven transformation provides the foundation for scalable innovation and business growth. Can you describe the importance of Enterprise Knowledge Management (EKM) to the uninitiated? Pedro Soares: EKM is not a new concept. Organizations have always needed to locate, understand and manage their intellectual property and business processes. What has changed in the era of AI is the greater need to make that data accessible, quantifiable, consistent and governable. AI has illuminated the gaps and shortcomings in traditional EKM and has shown Kyndryl Inc. that Kyndryl Inc. need to consolidate disparate pockets of data into taxonomic structures with clearly defined relationships among the information - an ontology. Kyndryl Inc. know that everything Kyndryl Inc. need is in the garage. Now it's time to clean it up, put its tools in the right boxes, label the boxes clearly, and maintain the relationships between all those pieces. Once Kyndryl Inc. has correctly structured and governed its data, Kyndryl Inc. can use AI to generate recommendations, create specialized knowledge packages, establish and maintain information lifecycles, and produce outputs Kyndryl Inc. can trust. Kyndryl Inc. know that AI is only as useful and safe as the data it draws from. EKM is the science of making that data useful. How is Kyndryl helping enterprises infuse AI into their operations, and what are some specific reasons for organizations to undertake these transformations? Soares: Kyndryl Inc. has established that organizations need to prepare their data and infrastructure for AI infusion. But they also need to modernize and, in some cases, redesign their business processes before automating them. Otherwise, they risk automating flawed processes that AI will amplify at speed. That's why Kyndryl begins with the data. Kyndryl Inc. employ its platforms and capabilities - Kyndryl Bridge, Responsible AI at Kyndryl, Kyndryl Agentic AI Framework - to help customers get started on their AI journeys. But technology in a vacuum is meaningless. Its customer engagements begin with consultation and listening. Kyndryl Inc. work to understand each organization's operational challenges and ultimate goals, then apply that knowledge and perspective to the technology solutions Kyndryl Inc. design. Kyndryl's modernization demonstrates that transformation is possible and that Kyndryl Inc. has the expertise to deliver it. Kyndryl Inc. also continue to develop unique capabilities based on its foundational experience, decades of work across every generation of enterprise IT and the cross-industry data Kyndryl Inc. analyze each day. AI is only as useful and safe as the data it draws from. Enterprise Knowledge Management is the science of making that data useful. Pedro Soares Distinguished Engineer, Vice President, Knowledge Management What passions drew you to this line of work? Soares: One might say that I'm a bit obsessed with structure, which led me to study engineering. Although I wouldn't say that coding is my main interest, I've always been fascinated by compilers and the way people can communicate with machines. I began experimenting with neural networks in the 1990s and have been fascinated with building digital structures that enable people to access and reuse knowledge ever since. I also enjoy developing talent and building teams. I've had the privilege of leading architect communities at the country, regional and global levels. Helping people grow energizes me. Those experiences also prompted me to consider structural challenges. What happens when organizations rely too heavily on human memory and scattered pockets of data? They end up with pieces of a puzzle but never the complete picture. Businesses can't scale when their data remains disconnected, or they expend human energy on mundane, repeatable tasks. The goal is to have AI handle the process-intensive work so people can focus on matters requiring discernment and creativity. The key to both outcomes is assembling a useful data ontology that people can trust. What advice would you give to your younger self? Soares: Listen more! Active listening means absorbing what others are sharing and teaching instead of jumping to conclusions or simply waiting for your chance to talk. It's one of the most valuable skills you can have in your professional or personal life. I also would have advised myself to be resilient, be patient and seek feedback. Always keep learning. That doesn't mean accepting everything you hear without evaluating it. As a mentor advised me, it's important to listen and seek feedback, but it's just as important to evaluate what you hear critically. That's where resilience and patience come into play. As you gain experience, you become more skilled at listening, synthesizing information, and working with others to create value. What activities do you enjoy in your spare time? Soares: My family is the most important thing in my life. I have two children, who aren't little anymore, who keep me engaged, challenged and always learning. Kyndryl Inc. live on the western Portuguese coast and enjoy traveling, swimming and water sports. I also enjoy martial arts, scuba diving, making and listening to music and all sorts of reading, including fiction, nonfiction and technical material. But my greatest passion and my therapy is motorcycling. I take short weekend trips when I can and complete at least one long tour with friends every year. Along with offering freedom, scenery and new experiences, responsible motorcycling requires preparation, discipline and concentration. Riding forces me to focus completely on the present and disconnect from daily pressures. Pedro Soares. Distinguished Engineer, Vice President, Knowledge Management

PR Newswire
Sep 1st, 2026
Kyndryl to speak at Citi investor Conference on September 8.

Kyndryl to speak at Citi investor Conference on September 8. Sep 01, 2026, 10:00 ET NEW YORK, Sept. 1, 2026 /PRNewswire/ - Kyndryl (NYSE: KD), a leading provider of mission-critical enterprise technology services, today announced that Chairman and Chief Executive Officer Martin Schroeter and Chief Financial Officer Ellen Johnson will speak at the Citi Global TMT Conference on Tuesday, September 8, 2026 at 8:50 a.m. ET. During the event, they will discuss the company's business and financial performance. To listen to the live webcast, please visit Kyndryl's investor relations website at investors.kyndryl.com. A replay of the webcast will be available approximately 24 hours after the live presentation. About Kyndryl Kyndryl (NYSE: KD) is a leading provider of mission-critical enterprise technology services, offering advisory, implementation and managed service capabilities to thousands of customers in more than 60 countries. As the world's largest IT infrastructure services provider, the company designs, builds, manages and modernizes the complex information systems that the world depends on every day. For more information, visit www.kyndryl.com. SOURCE Kyndryl

Windows Mode
Aug 27th, 2026
Kyndryl and Broadcom strengthen partnership to support private cloud solutions for AI workloads.

Kyndryl and Broadcom strengthen partnership to support private cloud solutions for AI workloads. Key Points Discover more Download Productivity Apps Compare Laptop Specs Track Business News * Broadcom and Kyndryl launched a joint consulting service for VMware Cloud Foundation (VCF) aimed at AI workloads. * Thousands of Kyndryl consultants will be trained on VCF 9.1 and agentic workflow automation. * Analysts praise the skills investment but some note the partnership is not new. What is changing. According to a report from Network World, Kyndryl and Broadcom have introduced new consulting services for VMware Cloud Foundation (VCF) that are positioned as an AI program. The offering is designed to help enterprises build private clouds that can run artificial intelligence workloads. It includes automated deployment templates and policy based controls for data sovereignty. The service is intended to simplify the adoption of AI workloads in regulated environments. The initiative will certify thousands of Kyndryl consultants on VCF 9.1 and related automation tools to support agentic workflows. These consultants will be trained to configure, operate, and secure the platform across hybrid environments. The program also promises to accelerate migration of legacy applications to the private cloud. Why it matters. The change is most relevant to IT leaders who are evaluating private cloud strategies for AI deployment. These professionals are responsible for choosing platforms that can support both current workloads and emerging AI tasks. The new consulting services may help them reduce the learning curve for VCF. Enterprises with existing VCF environments may see a skills gap addressed by the expanded consultant base. Having more certified professionals can shorten migration timelines and improve operational stability. However, the impact is likely limited to organizations already using VMware private cloud solutions. Discover more Compare Cloud Storage Download Productivity Apps Readers, share your experience with VCF deployments in the comments. Post Views: 257 Discover more from Windows mode. Get Windows Licenses Adam LaFoute Adam is a seasoned technology expert and Microsoft specialist with over 15 years of immersion in the company's products and services. From coding and app development to cloud computing and mixed reality, his deep understanding of Microsoft's ecosystem makes him resourceful, unless you catch him in the morning without his coffee!

Beka Publishing
Aug 27th, 2026
Kyndryl expands partnership with Broadcom.

Kyndryl expands partnership with Broadcom. August 27, 2026 Kyndryl has expanded its strategic alliance with Broadcom to deliver end-to-end consulting services for VMware Cloud Foundation (VCF), bringing cloud-like speed, automation and developer experience to private and hybrid cloud environments. Under the expanded collaboration, Kyndryl and Broadcom are helping enterprises modernize, secure and scale mission-critical systems by building sovereign, AI-ready private clouds that reduce operational complexity, are secure by design and strengthen long-term performance. Kyndryl provides end-to-end VCF and VMware Tanzu transformation capabilities across advisory and strategy, architecture and design, upgrade and modernization and secure Day-2 operations, spanning virtualized and containerized workloads, while enabling seamless integration with public cloud landing zones as workloads evolve. Kyndryl is also applying the Kyndryl Agentic AI Framework and policy as code capability to hold agents to approved, deterministic actions informed by business rules and regulatory requirements, so automation stays inside guardrails and drift is contained.