NOC Resiliency Specialist (Night Shift)
Role Overview
The VNOC Resiliency Specialist (Night Shift) supports day-to-day operations as the primary monitoring, routing, and triage authority alongside Incident Managers. Operating with high autonomy during off-hours, this team member serves as our first line of defense in maintaining on-premise and cloud infrastructure stability.
Our Ideal Candidate
We are seeking a highly reliable, process-driven professional who remains composed under pressure and thrives in an overnight environment. This role does not require deep system engineering; instead, it demands vigilance and execution.
Core Focus
Your daily mission focuses on dashboard surveillance, rapid alert triage, and runbook execution to support organizational resiliency.
Core Functional Responsibilities
- Dashboard Surveillance: Maintain high-vigilance monitoring of enterprise network, on-premise, and cloud infrastructure health utilizing SolarWinds, Elastic, and Application Performance Monitoring (APM) dashboards to immediately catch system degradation or outages.
- Queue Management & Triage: Triage incoming outage calls, acknowledge automated system alerts, prioritize the queue, and route events to the correct technical engineering teams based on established routing rules.
- First-Response & Escalation: Provide composed, immediate first-line response during Major Incidents by executing standard step-by-step runbooks. Coordinate closely with Incident Managers to escalate issues and engage technical Subject Matter Experts (SMEs) as required.
- Cloud Health Monitoring: Monitor high-level system alerts and service health dashboards within AWS and OCI (Oracle Cloud Infrastructure) environments to identify cloud service disruptions and initiate standard escalation workflows.
- Ticketing & SLA Compliance: Manage the administrative lifecycle of incidents within ServiceNow and Jira, ensuring precise documentation of event timelines, ticket updates, and strict adherence to established SLA response times.
- Shift Turnover & Incident Logging: Maintain precise shift turnover logs and conduct formal, detailed handovers to incoming day-shift personnel. Assist Incident Managers by compiling chronological event timeline data for post-incident reviews and After Action Reports (AARs).
- Change Window Monitoring: Support change management activities by observing dashboard health statuses via SolarWinds and APM for anomalies during scheduled maintenance and software deployment windows tracked in ServiceNow.
- Runbook & Process Adherence: Strictly follow established VNOC runbooks, Tactical Techniques and Procedures (TTPs), and SOPs. Flag outdated documentation in Jira/ServiceNow repositories to ensure instructions remain accurate.
Required Qualifications:
- 1 to 3 years of experience in an IT Operations Center (NOC/SOC) environment, Tier 1/2 IT Helpdesk, or relevant Military Operations (such as communications, cyber, or logistics).
- Fundamental Cloud Knowledge: Basic understanding of cloud infrastructure concepts and exposure to AWS or OCI environments (such as AWS Certified Cloud Practitioner or Oracle Cloud Infrastructure Foundations level).
- Familiarity with enterprise ticketing platforms (ServiceNow, Jira) and monitoring tools (such as SolarWinds, Elastic, APM) is highly preferred.
- Current DoD Secret Clearance or higher
- Ability to work night shift schedule (Pitman Schedule: 12-hour night shifts)
Ideal Candidate Profile & Background
- To succeed in this functional night-shift role, you will be a disciplined, detail-oriented operator who excels at executing established procedures.
- Process-Oriented & Composed: You respect established SOPs when critical alerts trigger. You can clearly follow step-by-step triage checklists during high-stress incidents.
- Autonomous & Reliable: You are highly self-motivated and punctual, capable of maintaining focus during quiet night-shift hours without direct leadership supervision.
- Strong Communicator: You possess excellent written communication skills, essential for drafting clear chronological turnover logs and logging precise ticket updates.