If you’re passionate about building a better future for individuals, communities, and our country—and you’re committed to working hard to play your part in building that future—consider WGU as the next step in your career.
Driven by a mission to expand access to higher education through online, competency-based degree programs, WGU is also committed to being a great place to work for a diverse workforce of student-focused professionals. The university has pioneered a new way to learn in the 21st century, one that has received praise from academic, industry, government, and media leaders. Whatever your role, working for WGU gives you a part to play in helping students graduate, creating a better tomorrow for themselves and their families.
The salary range for this position takes into account the wide range of factors that are considered in making compensation decisions including but not limited to skill sets; experience and training; licensure and certifications; and other business and organizational needs.
At WGU, it is not typical for an individual to be hired at or near the top of the range for their position, and compensation decisions are dependent on the facts and circumstances of each case. A reasonable estimate of the current range is:
Grade: Technical 404Pay Range: $81,300.00 - $121,900.00
Job Description
The NOC Engineer performs tasks that ensure maximum service availability and system performance. They are responsible for monitoring systems with diversified industry tools to detect anomalies. When anomalies are detected, engineers collaborate with various product teams to resolve them. The NOC Engineer is able to diagnose issues and alert the proper team without causing extended downtime of the IT systems. They track and document all defects, system health, anomalies, and events that occur throughout our ecosystem.
The NOC Engineer safeguards service availability and system performance through continuous observability across all applications, services, and infrastructure. Skilled in enterprise monitoring and alerting tools, they detect anomalies, service degradation, and outages in real time and know how to interpret system behavior to distinguish genuine issues from noise. When impact occurs, they act quickly to open incidents and engage the engineering teams that own the affected services, gathering and analyzing supporting evidence to accelerate root cause identification and speed resolution. Their mission is to help prevent outages before they happen and to aid in restoring service as quickly as possible when they do, documenting findings throughout to strengthen the reliability and performance of the environment.
Continuously monitor system health across all applications, services, and infrastructure using enterprise observability tools, including synthetic monitoring, automatically detected issues, dashboards, and real-time alerting.
Respond to alerts and identified degradation by validating issues through active testing to distinguish genuine impact from noise, then initiate incidents and engage the Problem Management team to classify and escalate based on severity.
Investigate system-detected issues and proactively engage engineering teams to diagnose service degradation, gather supporting evidence, and push for timely resolution.
Build and maintain browser synthetic monitors to ensure complete availability insight into the end user experience.
Review and approve change requests for deployments, confirming that appropriate monitoring is in place to validate success or failure, and that deployment timing aligns with service traffic patterns.
Configure maintenance windows to suppress unnecessary alerting during expected downtime.
Create and maintain student and staff communications, including student portal notifications for planned maintenance and expected service unavailability.
Professionally notify affected service departments when an issue or service degradation has been identified.
Monitor and follow up with engineering teams to ensure timely completion of certificate renewals and other observability-related tasks.
Administer the team's observability and alerting platforms by fulfilling service requests such as building dashboards, configuring alerting integrations, managing user permissions, and provisioning accounts.
Track and document findings, incidents, and system events, and create and maintain knowledge base articles for both the NOC team and the broader university staff to keep observability and alerting knowledge accessible.
Attend cross-functional initiative meetings, vendor and partner meetings, and daily stand-ups.
Collaborate with engineering teams on a range of initiatives to deliver the monitoring coverage they need.
Perform other job-related duties as assigned.
Knowledge, Skills, and Abilities (KSAs)
Solid understanding of observability and telemetry concepts, including how metrics, logs, traces, and synthetic testing are used to assess the health and performance of applications and services.
Working knowledge of cloud infrastructure and modern application architecture, including how distributed services, dependencies, and integrations behave and fail.
Understanding of networking fundamentals such as DNS, HTTP/HTTPS, TLS certificates, load balancing, and how these components affect availability and performance.
Familiarity with incident and problem management practices, including how issues are triaged, classified by impact, escalated, and driven to resolution across multiple teams.
Ability to interpret system behavior and telemetry to distinguish genuine service impact from noise and false positives.
Knowledge of change and release management concepts, including how deployment timing and monitoring readiness affect the stability of production services.
Ability to administer and configure monitoring and alerting tooling, including designing dashboards, tuning alert thresholds, and managing user access.
Strong analytical and problem-solving skills, with the ability to remain effective under pressure and prioritize during active service events.
Exceptional verbal and written communication skills, with the ability to convey technical issues clearly and professionally to both technical and non-technical audiences.
Strong customer service orientation and the ability to build effective working relationships across engineering, service, and vendor teams.
Ability to work independently, exercise sound judgment, and make decisions with minimal management intervention.
Ability to manage multiple concurrent issues and interruptions while multitasking effectively in a fast-changing environment.
Ability to anticipate potential obstacles and develop contingency plans to address them.
Strong documentation skills, including producing clear knowledge base articles for technical and non-technical readers.
Awareness of emerging and existing technologies, architectures, and industry best practices.
Ability to work flexible hours, including participation in on-call or coverage rotations as needed.
Bachelor's degree in Computer Information Systems, Computer Science, or a related field, OR equivalent relevant experience performing the essential functions of this role. Generally, equivalent relevant experience is defined as one year of experience for one year of education, at the discretion of the hiring manager.
2 years of experience in an IT operations, monitoring, technical support, incident/problem management, or similar role.
Demonstrated experience monitoring the availability and performance of production applications, services, or infrastructure in a live operations environment.
Hands-on experience with observability or monitoring platforms and the ability to interpret dashboards, alerts, and telemetry to assess system health (e.g., Dynatrace, Datadog, New Relic).
Working experience with an IT service management or ticketing platform for incident, problem, and change workflows (e.g., ServiceNow).
Strong written and verbal communication skills, with the ability to clearly report service impact to technical and non-technical audiences.
Proven ability to remain organized and effective under pressure while managing multiple concurrent issues.
Availability to work flexible hours, including potential evening, weekend, or on-call coverage to support 24/7 operations.
Proficiency with standard business productivity tools (e.g., Microsoft Office 365).
Comfort adopting and working with modern AI-assisted tools (e.g., Amazon Kiro)
AWS Certified Practitioner Certification
ComptiaA+, ComptiaNet+ Certification
1-2 years of experience with Dynatrace
1-2 years of experience with ServiceNow
Job Description Disclaimer
This position description provides the major duties/responsibilities, requirements, and working conditions for the position. It is intended to be an accurate reflection of the current position;however, management reserves the right to revise or change as necessary to meet organizational needs. Other responsibilities may be assigned when circumstances require.
This position requires occasional travel of up to 20%, including required attendance at designated company summits (typically one to two per year). Additional travel may include conferences, visits to company locations, and other business-related events as needed. Additional travel may be assigned as needed to support business requirements.
Position & Application Details
Full-Time Regular Positions (classified as regular and working 40 standard weekly hours): This is a full-time, regular position (classified for 40 standard weekly hours) that is eligible for bonuses; medical, dental, vision, telehealth and mental healthcare; health savings account and flexible spending account; basic and voluntary life insurance; disability coverage; accident, critical illness and hospital indemnity supplemental coverages; legal and identity theft coverage; retirement savings plan; wellbeing program; discounted WGU tuition; and flexible paid time off for rest and relaxation with no need for accrual, flexible paid sick time with no need for accrual, 11 paid holidays, and other paid leaves, including up to 12 weeks of parental leave.
How to Apply: If interested, an application will need to be submitted online. Internal WGU employees will need to apply through the internal job board in Workday.
Additional Information
Disclaimer: The job posting highlights the most critical responsibilities and requirements of the job. It’s not all-inclusive.
Accommodations: Applicants with disabilities who require assistance or accommodation during the application or interview process should contact our Talent Acquisition team at [email protected].
Equal Employment Opportunity: All qualified applicants will receive consideration for employment without regard to any protected characteristic as required by law.