Full-Time

Lead Principal Core Infrastructure Engineer

Network Monitoring, Network Observability Architecture

Oracle

Oracle

10,001+ employees

Enterprise software, cloud infrastructure, databases provider

No salary listed

Bengaluru, Karnataka, India

In Person

Category
DevOps & Infrastructure (1)
Required Skills
Kubernetes
Grafana
Incident Response
Distributed Systems
Data Visualization
Apache Kafka
Java
Data Engineering
Infrastructure as Code (IaC)
Go
Cryptography
Prometheus

Get referred to Oracle

See people who can refer or advise you

Requirements
  • Deep expertise in several of the listed observability, distributed systems, networking, programming, reliability, automation, and security areas, with sufficient breadth to lead architecture across the complete observability stack.
  • Experience with large-scale distributed systems, cloud infrastructure, Kubernetes, service-oriented architectures, distributed state management, high availability, fault tolerance, capacity planning, and performance engineering.
  • Experience with metrics and telemetry architecture, Prometheus, Grafana, time-series data, high-cardinality metrics, alerting, dashboards, service-level objectives, telemetry pipelines, and monitoring large distributed environments.
  • Experience with Kafka, Flink or equivalent technologies, real-time stream processing, event-driven architectures, partitioning, aggregation, backpressure, data pipelines, and large-scale telemetry processing.
  • Strong understanding of Layer 2 and Layer 3 networking, routing and switching, network topology, BGP, LLDP, SNMP, gNMI, network failure modes, and network performance troubleshooting.
  • Experience designing or operating systems that collect and correlate telemetry from large fleets of network devices using streaming telemetry, counters, events, protocol state, and device health information.
  • Strong software engineering expertise in Java, Go, or comparable systems programming languages, with experience building highly concurrent, performance-sensitive production services.
  • Experience with service-level objectives, availability and durability engineering, distributed failure handling, load shedding, throttling, rate limiting, incident response, root-cause analysis, and production readiness.
  • Experience with Infrastructure as Code, automated deployment and configuration management, safe rollout and rollback strategies, and operating large infrastructure fleets with minimal manual intervention.
  • Experience with secure multi-tenant cloud architectures, authentication and authorization, encryption, vulnerability remediation, and security considerations for infrastructure telemetry and management systems.
Responsibilities
  • Define the architecture and technical direction for OCI’s network monitoring and observability platform.
  • Architect distributed systems that collect, process, aggregate, store, query, and visualize network telemetry at cloud scale.
  • Design scalable telemetry ingestion and streaming architectures capable of handling high-volume and high-cardinality data.
  • Define strategies for metrics, events, alerts, topology, network state, and other signals required to understand network health.
  • Drive architectures that enable rapid detection, correlation, diagnosis, and isolation of network failures and performance degradation.
  • Establish standards for telemetry quality, completeness, freshness, accuracy, retention, and availability.
  • Lead the design of horizontally scalable and elastic systems supporting continued OCI infrastructure and traffic growth.
  • Identify performance, throughput, latency, storage, and scalability bottlenecks across telemetry pipelines and drive systemic improvements.
  • Architect high-throughput streaming and event-processing systems with appropriate partitioning, backpressure, buffering, aggregation, and failure-handling strategies.
  • Define resilient state-management, replication, synchronization, and recovery strategies for distributed monitoring systems.
  • Evaluate architectural trade-offs involving consistency, availability, latency, durability, and cost.
  • Establish service-level objectives and engineering standards for availability, durability, latency, data freshness, and correctness of the observability platform.
  • Architect fault-tolerant systems that continue operating through infrastructure failures, network partitions, service disruptions, and software upgrades.
  • Define key performance indicators and telemetry that measure the health of the monitoring platform itself and identify gaps or blind spots in observability.
  • Lead diagnosis and resolution of complex production issues spanning networking, distributed systems, telemetry pipelines, and infrastructure.
  • Serve as a senior technical escalation point for critical incidents and drive root-cause analysis and systemic corrective actions.
  • Design systems and operational practices that support automated deployment, upgrade, rollback, recovery, and minimal customer-visible disruption.
  • Define approaches for monitoring large-scale Layer 2 and Layer 3 network infrastructure and identifying changes in network state, topology, reachability, performance, and device health.
  • Drive correlation of telemetry across devices, network layers, and infrastructure services to improve fault localization and reduce time to detection and resolution.
  • Develop approaches for identifying abnormal network behavior, telemetry gaps, capacity risks, and emerging infrastructure failures.
  • Partner with network engineering teams to translate network behavior and operational requirements into scalable monitoring capabilities.
  • Improve signal quality and actionable alerting while reducing noise and unnecessary operational load.
  • Set technical direction and influence architecture decisions across Network Monitoring and dependent OCI organizations.
  • Lead complex and ambiguous initiatives spanning multiple systems and engineering teams.
  • Establish architectural patterns, engineering standards, and technical best practices for large-scale observability systems.
  • Mentor senior engineers and provide technical guidance on distributed systems, networking, reliability, and observability.
  • Evaluate emerging technologies and engineering approaches and drive adoption where they materially improve scale, reliability, performance, or operational efficiency.
  • Contribute to technical talent development through senior-level interviewing, candidate assessment, mentoring, and knowledge sharing.
Desired Qualifications
  • Experience building observability or telemetry platforms for large-scale cloud or network infrastructure.

Oracle provides enterprise software, cloud infrastructure, databases, and business applications for organizations. It offers cloud computing, storage, networking, and AI-enabled data management, plus Fusion Cloud Applications for ERP, HCM, supply chain, manufacturing, and customer experience, with options for hybrid and on-premises deployment. The company differentiates itself with an integrated stack that spans databases, cloud infrastructure, and enterprise apps, built on a history of acquisitions and a broad customer base. Its goal is to help organizations run operations efficiently, scale data and processes, and pursue digital transformation through an end-to-end platform.

Company Size

10,001+

Company Stage

IPO

Headquarters

Austin, Texas

Founded

1977

Get referred to Oracle

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Oracle's Q1 FY2026 revenue rose 30% to $19.3 billion, signaling strong demand.
  • OpenAI-related demand and Stargate data-center buildouts support OCI growth through 2027.
  • Oracle raised FY2027 targets to $90 billion, implying management sees sustained acceleration.

What critics are saying

  • EU regulators started reviewing Oracle licensing on September 1, 2026; SAP-like remedies follow.
  • Fiscal 2026 free cash flow was negative $23.7 billion after $55.7 billion capex.
  • Oracle's dependency on massive AI buildouts and OpenAI-style contracts creates existential financing pressure by 2027.

What makes Oracle unique

  • Oracle's 2026 RPO hit $638 billion, dwarfing annual revenue and rivals' backlogs.
  • Oracle Database and OCI combine software lock-in with cloud migration paths across enterprises.
  • Oracle's March 2026 AI Database innovations strengthen vector and agentic workloads on-premises and cloud.

Help us improve and share your feedback! Did you find this helpful?

Benefits

401(k) Savings and Investment Plan with company match

Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.

11 paid holidays

Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.

Paid parental leave

Adoption assistance

Employee Stock Purchase Plan

Financial planning and group legal

Voluntary benefits including auto, homeowner and pet insurance

Growth & Insights and Company News

Headcount

6 month growth

1%

1 year growth

1%

2 year growth

2%
Bloomberg
Sep 10th, 2026
Oracle beats earnings estimates with sales rising 30% to $19.3B

Oracle reported fiscal first-quarter earnings of $1.92 per share, surpassing analyst estimates. The enterprise software company's sales climbed 30% to $19.3 billion during the quarter. The strong results demonstrate continued growth for Oracle as businesses invest in cloud infrastructure and database services. The earnings beat suggests the company is maintaining momentum in its cloud computing business. Oracle's performance comes amid ongoing competition in the enterprise software and cloud services market. The 30% revenue increase represents significant year-over-year growth for the tech giant.

Yahoo Finance
Sep 10th, 2026
Oracle faces earnings tonight with 79% beat probability despite 52% stock decline since last year's 36% post-earnings surge

Oracle reports earnings today with shares down 18% year-to-date and trading around $159, roughly 52% below last September's post-earnings close of $328. The stock surged 36% in a single session following last year's report. Despite the decline, fundamentals remain strong. The company guided first-quarter revenue growth of 27% to 29% and cloud revenue growth of 58% to 64%. Remaining performance obligations reached $638 billion, up 363% year-over-year, with $75 billion tied to GPU arrangements. Options traders are selling calls at the $150 strike, and implied volatility sits at 73, among the highest in the S&P 500. However, 36 analysts rate the stock a buy with a $241 target, and consensus expects earnings of $1.74 per share on $19.13 billion revenue.

Yahoo Finance
Sep 10th, 2026
Big Tech burns $13.5B as AI buildout sends Treasury yields to 4.79%

US Treasury yields near 4.79% reflect strong corporate borrowing for AI investments rather than economic weakness, according to analysts Joel Litman and Rob Spivey. They argue that context matters more than absolute rate levels. Companies borrowing at 5% to fund projects returning 30-40% benefit from current rates, whilst those earning less than borrowing costs face pressure. The analysts note that AI-related corporate debt issuance reached roughly $1.5 trillion this year, driving yields higher. Alphabet posted its first negative free cash flow since 2004, burning $5.9 billion in Q2 as capital expenditure hit $44.9 billion. Amazon swung to negative $7.6 billion on a trailing basis. However, negative cash flow can signal productive investment rather than distress, the analysts suggest.

Yahoo Finance
Sep 9th, 2026
Oracle and Adobe earnings Thursday: Which AI play offers better value?

Oracle and Adobe report quarterly earnings on Thursday 10 September, offering different angles on AI investment opportunities. Oracle expects fiscal Q1 revenue of roughly $19.13 billion, up 27%-29% year-over-year, driven by cloud infrastructure demand. Total cloud revenue is forecast to surge 58%-64%, with remaining performance obligations soaring 363% to $638 billion due to large-scale AI contracts. Adobe anticipates fiscal Q3 revenue of $6.69 billion, representing approximately 12% growth. The company recently reported AI-first annual recurring revenue tripled year-over-year to over $500 million, with record Q2 revenue of $6.61 billion. Oracle benefits from AI-related cloud infrastructure demand, whilst Adobe embeds generative AI throughout its creative, document, and marketing platforms. Investors will watch whether Oracle's growth justifies its premium valuation or Adobe's discounted stock offers better value.

Yahoo Finance
Sep 9th, 2026
Options traders bet Oracle could jump 8% as AI cloud growth faces capital spending scrutiny

Options traders are positioning for significant movement in Oracle shares ahead of the company's fiscal first-quarter earnings report on Thursday. The stock closed at $162.52 on Tuesday, with call options showing particular interest in the $175 strike price — representing nearly an 8% premium. Oracle's cloud infrastructure revenue surged 93% to $5.8 billion last quarter, driven by AI-related demand. The company ended fiscal 2026 with $638 billion in remaining performance obligations, indicating substantial contracted future business. However, aggressive AI data centre investment has strained finances. Whilst Oracle generated $32 billion in operating cash flow in fiscal 2026, free cash flow turned negative at $23.7 billion due to heavy capital spending. Investors will scrutinise whether AI cloud revenue growth can justify the significant infrastructure expenditure.