Full-Time

Senior Director Data Center On-Site Operations

Data Center Facilities

DigitalOcean

DigitalOcean

1,001-5,000 employees

Cloud platform enabling rapid app deployment

Compensation Overview

$233.6k - $292k/yr

Seattle, WA, USA

Hybrid

Hybrid role; some on-site presence in Seattle required.

Category
Facilities Operations (1)
Required Skills
Computer Networking
Operating Systems
Linux/Unix

Get referred to DigitalOcean

Find people who can refer or advise you

Requirements
  • Bachelor’s degree in Engineering, Business Administration, Computer Science, Management Information Systems, or a related field; equivalent work experience may be considered.
  • Minimum of 8 years of experience in data center operations, infrastructure operations, colocation operations, cloud infrastructure, or a related technical operations environment.
  • Minimum of 5 to 7 years of senior leadership experience managing operational, technical, or site-support teams in a large enterprise, cloud, or colocation environment.
  • Demonstrated experience managing operations across multiple data center locations or large-scale mission-critical facilities.
  • Proven experience developing and implementing operational improvement roadmaps.
  • Proven experience improving process maturity, operational consistency, team performance, vendor execution, and SLA performance.
  • Experience building scalable processes to support rapid organizational growth, new site launches, and increased operational complexity.
  • Experience managing vendor relationships, service providers, contractors, and third-party data center partners.
  • Experience with budgeting, forecasting, operational planning, performance reporting, and executive communications.
  • Experience with GPU, AI infrastructure, high-density compute, liquid cooling, or large-scale hardware deployment environments is strongly preferred.
  • Strong understanding of data center critical infrastructure, including UPS systems, generators, transfer switches, power distribution, DC power systems, HVAC, cooling systems, liquid cooling, racks, structured cabling, and physical security.
  • Ability to translate business, server, storage, networking, and customer requirements into practical data center operational needs.
  • Strong vendor management and negotiation skills, with the ability to hold partners accountable while maintaining effective business relationships.
  • Excellent communication skills, including the ability to present complex operational issues clearly to executive, technical, and non-technical audiences.
  • Proven ability to develop credibility, influence without authority, and drive change across diverse teams.
  • Strong decision-making, problem-solving, prioritization, and risk-management capabilities.
  • Ability to develop and communicate relevant department metrics, SLA performance reports, operating reviews, and executive-level summaries.
  • Demonstrated ability to develop teams, coach leaders, mentor technical staff, and build organizational capability.
  • Familiarity with industry standards, certifications, and operational best practices for data center environments.
  • Working knowledge of IT infrastructure, including servers, storage, networking, operating systems, virtualization, and cloud infrastructure.
  • Practical knowledge of electronics, electrical systems, energy management, IP networking, Ethernet, Linux, Windows, virtualized environments, and SAN storage.
  • Strong preference for experience with modern infrastructure trends, including AI workloads, GPU clusters, software-defined networking, automation, and high-density data center designs.
Responsibilities
  • Provides senior leadership for data center on-site operations across DigitalOcean’s existing and expanding global footprint.
  • Leads the development and execution of a scalable operating model for data center sites, including standards for staffing, shift coverage, escalation, maintenance, change management, incident response, vendor oversight, inventory control, deployment support, and customer-impact prevention.
  • Drives operational excellence across all data center locations by identifying gaps, standardizing best practices, improving execution discipline, and ensuring consistent adoption of policies, procedures, and operating standards.
  • Leads site operations teams responsible for 24x7 data center support, ensuring availability, reliability, safety, quality, and service delivery meet or exceed business and customer expectations.
  • Improves on-site execution performance by strengthening accountability, defining clear ownership, and implementing measurable performance standards across locations.
  • Build repeatable operational processes that allow DigitalOcean to scale from a regional data center footprint to a larger, more complex, and more globally consistent operating environment.
  • Develops standardized site launch playbooks, operational readiness checklists, escalation models, staffing templates, maintenance routines, training plans, and vendor management practices for new and existing sites.
  • Partners with cross-functional teams to ensure that site operations are prepared for data center expansions, new technology deployments, new customer requirements, higher-density racks, GPU infrastructure, and liquid cooling environments.
  • Identifies operational constraints that limit growth and develops practical solutions to improve throughput, reduce execution friction, and accelerate site readiness.
  • Creates a process-development roadmap that supports organizational growth, improves consistency, and reduces dependency on individual tribal knowledge.
  • Defines and manages key operational KPIs, SLAs, and performance indicators across the data center operations organization.
  • Uses data to identify trends, performance gaps, recurring issues, staffing constraints, vendor deficiencies, and opportunities for automation or process improvement.
  • Develops operational dashboards and review mechanisms for availability, incident response, maintenance completion, change success, inventory accuracy, deployment throughput, backlog management, staffing readiness, vendor performance, and customer-impact events.
  • Leads operational reviews with senior leadership, providing clear visibility into site health, execution risks, improvement plans, and progress against strategic goals.
  • Benchmarks operational performance against industry standards, internal targets, and business requirements.
  • Has responsibility for meeting or exceeding established service levels for data center operations.
  • Oversees site-level execution supporting critical physical infrastructure, including power, cooling, cabling, space, racks, network infrastructure, server deployment, hardware maintenance, and customer environment support.
  • Ensures that work performed in data center environments is completed safely, correctly, and without impact to internal or external customers.
  • Drives stronger change management discipline for work performed in live production environments, including pre-work planning, risk assessment, approvals, execution quality, rollback readiness, and post-change validation.
  • Evaluates and mitigates operational risks across new and existing sites, including staffing gaps, vendor performance issues, maintenance deficiencies, capacity constraints, process weaknesses, and readiness gaps for new deployments.
  • Improves incident response, root cause analysis, corrective action tracking, and recurrence prevention across data center locations.
  • Supports data center expansion efforts by ensuring on-site operational requirements are identified early and integrated into planning, design, construction, commissioning, and turnover processes.
  • Partners with Design, Construction, Capacity Planning, Engineering, Procurement, Networking, and Product teams to translate business demand into operational requirements for space, power, cooling, racks, cabling, staffing, maintenance, security, logistics, and site support.
  • Develops operational acceptance criteria for new sites, new data halls, new pods, and major infrastructure deployments.
  • Ensures new sites are handed over with complete documentation, diagrams, procedures, training, spares, escalation paths, vendor contacts, and support models.
  • Drives readiness for accelerated GPU deployments, high-density rack environments, liquid cooling, and other infrastructure programs tied to business growth.
  • Leads site-level vendor management to ensure service providers, contractors, colocation partners, and maintenance vendors meet operational expectations, contractual commitments, SLAs, and safety requirements.
  • Serves as a senior escalation point for vendor performance issues, site delivery concerns, maintenance execution problems, and operational deficiencies.
  • Develops consistent vendor governance practices, including performance reviews, issue tracking, corrective action plans, escalation paths, and service-quality metrics.
  • Partners with Procurement, Legal, and Finance to identify opportunities to streamline vendor services, optimize costs, and improve accountability.
  • Builds strong working relationships with vendor stakeholders while maintaining clear expectations for execution, quality, safety, and delivery.
  • Leads, coaches, and develops data center operations managers, site leaders, and technical operations teams.
  • Builds a leadership bench capable of supporting rapid site growth, increased operational complexity, and higher customer expectations.
  • Creates training, mentoring, and skill-development programs that improve technical capability, execution discipline, safety awareness, and leadership readiness.
  • Develops staffing models that support service-level requirements, site complexity, customer commitments, GPU concentration, and operational risk.
  • Fosters a culture of ownership, accountability, collaboration, continuous improvement, safety, customer focus, and operational discipline.
  • Removes organizational barriers that slow execution, reduce accountability, or prevent strong cross-functional collaboration.
  • Defines and maintains policies, procedures, standards, and documentation related to data center operations, including performance, capacity, availability, continuity, security, safety, maintenance, and change management.
  • Ensures documentation and diagrams accurately capture critical site information, including physical layout, rack elevations, power paths, cooling configurations, cabling standards, escalation paths, maintenance routines, and operational dependencies.
  • Maintains operational governance routines to ensure standards are being followed across sites and that deviations are identified, reviewed, and corrected.
  • Supports budgeting, forecasting, planning, deployment coordination, incident management, problem management, change management, and operational reporting.
  • Creates and delivers executive-level presentations on site operations performance, operational risks, improvement programs, expansion readiness, and organizational needs.
  • Performs other duties as assigned.
Desired Qualifications
  • Strong preference for experience with modern infrastructure trends, including AI workloads, GPU clusters, software-defined networking, automation, and high-density data center designs

DigitalOcean provides cloud computing infrastructure for developers, startups and SMBs to build, deploy, and scale applications using Droplets, managed databases, Kubernetes, object storage, and networking. It offers simple provisioning via a dashboard and APIs with fully managed services so teams avoid managing underlying infrastructure. It differentiates itself through a focus on simplicity, a strong developer community, open-source alignment, affordable pricing, and responsive support. The goal is to free developers from infrastructure chores so they can focus on coding and growing their business.

Company Size

1,001-5,000

Company Stage

IPO

Headquarters

New York City, New York

Founded

2012

Get referred to DigitalOcean

Find people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • AI customer annual run-rate revenue hit $170M in Q1 2026, up 221% year over year
  • Raised $888M to add 20 new data centers with liquid-cooled NVIDIA B300s in Kansas City
  • Revenue growth expected to accelerate to 29% in Q2 2026, with over 50% growth projected for 2027

What critics are saying

  • Hyperscalers launching SMB inference bundles will erode $170M AI ARR in the $20–$500/month segment within 6–12 months
  • $888M data center expansion amid declining free cash flow forces debt or equity dilution if AI growth drops below 25%
  • Katanemo Labs acquisition may fail to deliver viable inference router within 18 months, losing AI-Native Cloud differentiation

What makes DigitalOcean unique

  • AI-Native Cloud is the first cloud built end-to-end for inference and agentic workloads
  • Flat pricing with no per-query fees unlike hyperscalers, starting at $20/month for Managed Weaviate
  • Open platform with 640,000 customers across 20 data centers and no lock-in agreements

Help us improve and share your feedback! Did you find this helpful?

Benefits

Remote-first

Full health coverage

Wellness coverage

Flexible vacation time

Team-building & social events

401(k) plans

ESPP

Education support

Partner support

Employee giving

Growth & Insights and Company News

Headcount

6 month growth

0%

1 year growth

1%

2 year growth

1%
The Software Scout
Jun 4th, 2026
Railway vs DigitalOcean 2026: which should you deploy on?

Railway vs DigitalOcean 2026: which should you deploy on? By Ben / June 4, 2026 Railway and DigitalOcean both let you get an app online fast, but they sit at different points on the control-versus-convenience line. Railway is a modern platform built around developer experience, where you push code and it just runs. DigitalOcean gives you a broader cloud, from a managed app platform down to raw servers, with more control and more predictable pricing at scale. The right choice depends on whether you want to ship without thinking about infrastructure or keep your hands on the wheel. This comparison breaks it down. Quick verdict Railway is the better pick for most developers who want the fastest path from code to a running app, with a superb developer experience and instant databases. Choose DigitalOcean if you want more control, predictable flat pricing, and room to scale from a managed platform down to raw servers. At a glance. Try Railway Push your code and Railway builds, deploys, and runs it, with instant Postgres, Redis, and more. The fastest developer experience for getting an app live. How The Software Scout compared them. The Software Scout weighed what matters when you are choosing where to deploy: how quickly you get from a git push to a running app, the quality of the developer experience, how pricing behaves as you grow, how much control you have over the underlying infrastructure, the database and add-on story, and how each scales. Railway and DigitalOcean are both excellent, so this is about which trade-off between convenience and control fits your project. Railway. Railway is the modern deployment platform built around making developers productive, and that experience is why it is its default recommendation for getting an app live fast. Developer experience and databases. Railway's whole pitch is that you connect a repo, and it detects your stack, builds it, and deploys it without config files or YAML. Provisioning a Postgres, Redis, or MySQL database is a single click and it wires the connection variables in for you. Preview environments, instant rollbacks, a clean dashboard, and a capable CLI round it out. For solo developers, small teams, and anyone prototyping, the speed from idea to running service is genuinely the best in the category. Scaling and pricing. Railway scales services and databases without you managing servers, and its usage-based pricing means you pay for the compute and resources you actually consume. For small and mid-size projects that is efficient and often cheap. The trade-off is that usage-based billing is less predictable than a flat monthly fee, and at large, steady scale the convenience can cost more than running your own servers. You are also working at a higher level of abstraction, so you trade some low-level control for that simplicity. * Best-in-class developer experience * One-click databases with variables wired in * Git-push deploys, preview environments, rollbacks * Usage-based pricing is efficient for small projects * Usage-based billing is less predictable * Less low-level infrastructure control * Can cost more than raw servers at large scale DigitalOcean. DigitalOcean is a full cloud provider that spans the range from a managed app platform to raw virtual servers, which gives you more control and more predictable economics as you grow. From App Platform to Droplets. DigitalOcean App Platform is its managed PaaS, deploying directly from a repo with builds, scaling, and managed databases, much like Railway but a step less polished on developer experience. Below that sit Droplets, its straightforward virtual machines, plus managed Kubernetes, Spaces object storage, load balancers, and managed databases. That range means you can start on the App Platform and drop down to full server control whenever you need it, all within one provider and one bill. Pricing and control. DigitalOcean's pricing is famously simple and flat, so a Droplet or an App Platform tier costs a predictable amount each month regardless of spikes, which makes budgeting easy and is often cheaper at steady scale. The trade-off is responsibility: the more you move toward Droplets, the more you own patching, scaling, and configuration. The developer experience is good but more hands-on than Railway's, so you are trading some convenience for control and predictable cost. * Predictable, flat monthly pricing * Full range from managed PaaS to raw servers * More control over infrastructure * Often cheaper at steady, larger scale * Developer experience less polished than Railway * More ops responsibility as you move to Droplets * More setup to get to a running app Head to head. Developer experience. Railway wins. Git-push deploys, one-click databases, and zero config get you running faster than anything DigitalOcean offers, including App Platform. Pricing. DigitalOcean wins on predictability. Its flat monthly tiers make costs easy to forecast, while Railway's usage-based model is efficient for small projects but harder to predict and pricier at large steady scale. Control and flexibility. DigitalOcean wins. From App Platform down to Droplets and Kubernetes, you can take as much control as you want. Railway deliberately abstracts the infrastructure away. Scaling. It depends. Railway scales effortlessly with no ops for small to mid-size apps. DigitalOcean scales further and more cheaply at large, steady volumes if you are willing to manage more of the stack. Which should you choose? For most developers, Railway is the smarter choice, with the fastest path from code to a running app, instant databases, and a developer experience nothing else quite matches, which is ideal for solo developers, small teams, and prototypes. Choose DigitalOcean if you want predictable flat pricing, more control over your infrastructure, and the ability to scale from a managed platform down to raw servers within one provider. Both are excellent, so it comes down to convenience versus control. For more options, see its guide to the best hosting platforms for developers, and its Railway vs Render comparison. Get started with Railway Push your code and Railway builds, deploys, and runs it, with instant databases and zero config. The fastest way to get an app live. Frequently asked questions. Is Railway or DigitalOcean easier to use? Railway, clearly. Its git-push deploys, one-click databases, and zero-config workflow get you to a running app faster than DigitalOcean, including the App Platform. Which is cheaper? It depends on scale. Railway's usage-based pricing is efficient for small projects, while DigitalOcean's flat tiers are more predictable and often cheaper at steady, larger scale. Does DigitalOcean have a platform like Railway? Yes, DigitalOcean App Platform is its managed PaaS that deploys from a repo. It is close in concept to Railway but a step less polished on developer experience, and you can drop down to Droplets for more control. Which is better for scaling a large app? DigitalOcean tends to win at large, steady scale thanks to flat pricing and full infrastructure control, if you are willing to manage more of the stack. Railway scales effortlessly with no ops for small to mid-size apps. Can I run databases on both? Yes. Railway offers one-click managed databases with connection variables wired in, and DigitalOcean offers managed databases alongside App Platform and Droplets. The bottom line. Railway and DigitalOcean are both excellent places to deploy, and the right one comes down to how much control you want. For the fastest developer experience and the quickest path from code to running app, Railway is the better choice for most developers. DigitalOcean is the stronger pick if you want predictable flat pricing, full control, and room to scale from a managed platform to raw servers. Decide whether convenience or control matters more, and the right platform becomes clear.

IT Security News
Apr 7th, 2026
Scale smarter: A Practical Guide to Building with Akamai Object Storage.

Scale smarter: A Practical Guide to Building with Akamai Object Storage. 2026-04-07 18:04 Akamai Object Storage provides high-performance, cost-effective Amazon S3-compatible object storage. Here's what it's used for and how to set it up. Read the original article: Hacking & Cracking This post doesn't have text content, please click on the link below to view the original article. This article has been indexed from BlogRead the original article: Scale Faster: A Practical Guide to Building with Akamai Block Storage April 7, 2026 Learn about the early 2026 Terraform update, how the change will affect your workflow, and how to successfully navigate any issues that may arise. This article has been indexed from BlogRead the original article: Akamai Block Storage Makes Block Disk Encryption the Default in Terraform January 23, 2026 DigitalOcean announced Per-Bucket Access Keys for DigitalOcean Spaces, its S3-compatible object storage service. This feature provides customers with identity-based, bucket-level control over access permissions, helping to enhance their data security and simplifying management. Prior to the introduction of Per-Bucket Access Keys, many customers chose to limit the types of applications... January 23, 2025

Business Wire
Apr 3rd, 2026
DigitalOcean Acquires Katanemo Labs to Accelerate the Inference Cloud for the Agentic Era

DigitalOcean (NYSE: DOCN), the Agentic Inference Cloud built for production AI, today announced it has acquired Katanemo Labs, Inc., the models and research ...

Yahoo Finance
Mar 29th, 2026
How much higher can DigitalOcean stock go?

How much higher can DigitalOcean stock go? Anthony Di Pizio, The Motley Fool DigitalOcean (NYSE: DOCN) provides a suite of affordable cloud computing services exclusively to small and medium-sized businesses (SMBs). That was already a lucrative business model, but the company is now also helping its customers deploy artificial intelligence (AI) software, and demand is through the roof. The company's revenue growth accelerated last year, propelling its stock to a 41% gain. But it's up by a further 77% in 2026 already, as investors anticipate a continued surge in businesses' appetite for AI-related computing capacity. In fact, DigitalOcean recently announced plans to raise $800 million from investors to build more data center infrastructure. Will AI create the world's first trillionaire? Its team just released a report on the one little-known company, called an "Indispensable Monopoly" providing the critical technology Nvidia and Intel both need. Continue" Despite its recent gains, the stock is still trading at a relatively attractive valuation. So, how much higher can it go from here? AI for businesses of all sizes. The cloud computing industry is dominated by trillion-dollar hyperscalers like Amazon and Microsoft. Those providers usually compete for large customers with high spending potential, whereas SMB customers don't really move the needle for them in terms of revenue. As a result, SMBs often don't get the level of service they need from the hyperscalers. DigitalOcean exclusively targets those customers by offering cheap and transparent pricing, highly personalized service, and a simple dashboard which makes deploying cloud tools very straightforward. These features are ideal for start-ups and SMBs with limited financial and technical resources. The company is now helping its customers enter the AI era. Its Gradient platform provides them with access to the latest large language models (LLMs) from leading developers like OpenAI and Anthropic, which can be used as a foundation for developing AI software applications. DigitalOcean also operates data centers fitted with thousands of the latest AI chips from suppliers like Nvidia and Advanced Micro Devices, and it rents the computing capacity to SMBs for a fee. While hyperscale cloud providers try to lease thousands of chips to their customers at a time, DigitalOcean allows its customers to start with just one chip and scale as needed, which is plenty for small workloads like running an AI chatbot or deploying a few AI agents. The company says its prices are up to 75% cheaper than the hyperscale cloud providers for the same AI chips, which is a substantial saving for its budget-conscious customers.

Hawkdive
Mar 27th, 2026
GTC 2026 announces arrival of inference era: DigitalOcean confirms.

GTC 2026 announces arrival of inference era: DigitalOcean confirms. News GTC 2026 announces arrival of inference era: DigitalOcean confirms. March 27, 2026 8:1 pm Discover more Machine Learning & Artificial Intelligence At the recent NVIDIA GTC 2026 conference, the focus was on the evolution of AI from the training phase to the production inference phase. This shift signifies a significant development in the AI industry, moving beyond just developing faster chips and more intelligent models to addressing the challenges of running AI at scale in real-world applications. Inference, the stage where AI models are put into production to deliver real products and customer experiences, is now at the forefront of the AI conversation. Factors such as cost efficiency, latency, orchestration, and uptime are becoming just as important as the accuracy of the models themselves. The industry is now looking beyond just hardware advancements to the broader infrastructure needed to support AI-native companies. As AI becomes an integral part of business operations, the focus is on creating a cohesive system that encompasses chips, platforms, models, and applications to meet the demands of customers. DigitalOcean, in collaboration with NVIDIA, announced the launch of the Agentic Inference Cloud, aimed at helping AI developers transition from experimentation to production seamlessly. This initiative includes the introduction of a new data center in Richmond designed specifically for AI inference, equipped with NVIDIA HGX B300 systems and a high-speed RDMA fabric. Additionally, DigitalOcean is integrating NVIDIA Dynamo 1.0 into its Kubernetes platform and expanding model options optimized for various use cases. The momentum towards production-level AI deployment is already evident, with over 43,000 OpenClaw deployments on DigitalOcean, showcasing strong adoption from teams developing always-on assistants and agentic applications. To further explore the practical aspects of running AI inference at scale, leaders from NVIDIA, VAST Data, vLLM, Arcee AI, Character.AI, Workato, and more will be sharing insights at the upcoming DigitalOcean Deploy event in San Francisco on April 28, 2026. This event aims to provide valuable lessons on real-world architecture, performance, economics, and operational efficiency in the field of AI inference. For more Information, Refer to this article. You may also like these: Neil is a highly qualified Technical Writer with an M.Sc(IT) degree and an impressive range of IT and Support certifications including MCSE, CCNA, ACA(Adobe Certified Associates), and PG Dip (IT). With over 10 years of hands-on experience as an IT support engineer across Windows, Mac, iOS, and Linux Server platforms, Neil possesses the expertise to create comprehensive and user-friendly documentation that simplifies complex technical concepts for a wide audience. Discover more Machine Learning & Artificial Intelligence Watch & Subscribe Its YouTube Channel Latest from hawkdive. Neil S - March 27, 2026 10:1 pm Neil S - March 27, 2026 9:1 pm Neil S - March 27, 2026 8:1 pm Neil S - March 27, 2026 7:1 pm Neil S - March 27, 2026 6:1 pm Discover more Machine Learning & Artificial Intelligence