Network Operations Center Specialist at xAI

On-site - Southaven, MS; Memphis, TN

Apply
More jobs at xAI

SpaceXAI seeks a Network Operations Center Specialist to monitor campus health signals around the clock, detect and verify incidents, and coordinate incident communication. The role requires strong communication, incident management experience, and a willingness to work rotating shifts in a 24/7 operations environment.

Requirements

Skills

  • Experience in a 24/7 operations environment (NOC, SOC, dispatch, mission control, or equivalent)
  • Proven ability to acknowledge, classify, and elevate incidents under SLA in a high-signal environment
  • Experience opening and running incident bridges, including stakeholder updates on a fixed cadence and live timeline hygiene
  • Excellent written and verbal communication skills; able to write clear updates while an incident is in progress
  • Demonstrated pattern recognition across multiple domains (compute, network, storage, and/or facilities signals) and curiosity about how those systems interact
  • Experience following, maintaining, and improving operational process (runbooks, escalation matrices, handoffs, or similar)
  • Willingness and ability to work a rotating shift schedule, including nights and weekends, as part of continuous campus coverage
  • Prior NOC, data center operations, or campus reliability experience in a high-performance computing, AI/ML infrastructure, or large-scale production environment
  • Experience writing major-incident reports and driving corrective follow-ups to closed (e.g., tickets, projects, or Linear)
  • Familiarity with Linear or similar work-tracking tools for corrective action programs
  • Experience partnering with SRE, SiteOps, and Facilities on escalations and post-incident follow-through
  • Participation in game days, tabletop exercises, or runbook improvement programs
  • Prior work in a fast-paced startup or tech company like SpaceXAI

Responsibilities

  • Staff the console per shift schedule and watch the designated signal surface: cluster health, node availability, network health, facility trend panels, storage alarms, and threshold breaches
  • Acknowledge every page within SLA; classify (actionable / known / noise) and log disposition; feed noise patterns back to SRE so signal quality keeps improving
  • Detect, verify, and escalte within time budgets; operate the escalation matrix (NOC → on-call SRE → domain owners) and page correctly the first time
  • Open and run incident bridges; own stakeholder communications (first update within SLA, then fixed cadence); maintain the incident timeline in real time; call out ownership stalls
  • Produce first-pass RCA framing (what happened, when, what’s impacted, who’s engaged) and hand it to SRE / Hardware Failure Analysis for depth — the NOC does not publish root cause
  • Run structured shift handoffs and durable shift logs; maintain cross-site awareness
  • Write major-incident reports; open corrective projects in Linear and chase them to closure — the NOC is the nag of record
  • Maintain and continuously improve NOC runbooks, escalation matrices, and communications templates; participate in SRE-run game days

Technologies

Linear

See if your resume is ready for this job

See how our AI can optimize your resume and improve your chances for this role.