NOC Technician (Data Center and Site Ops) at xAI

On-site - Memphis, Tennessee

Apply
More jobs at xAI

SpaceXAI is a small, highly motivated AI systems organization focused on engineering excellence. The NOC Technician role is responsible for monitoring site health across data center campuses, escalating incidents, coordinating communications, and supporting the incident lifecycle. The position requires continuous 24/7 coverage, strong communication skills, and familiarity with monitoring dashboards and ticketing systems.

Requirements

Skills

  • High school diploma or equivalency certificate
  • 1+ year of professional experience in a Network Operations Center (NOC), Security Operations Center (SOC), mission‑control / dispatch, data center operations watch, or equivalent 24/7 monitoring and incident‑communications role
  • Demonstrated written and verbal communication skills under time pressure (stakeholder updates, handoffs, timelines)
  • Calm under pressure; excellent written and verbal communications — leadership should be able to trust your incident updates verbatim
  • Pattern recognition across domains; multi-domain curiosity (compute, network, storage, power/cooling signals)
  • Experience following and improving process: runbooks, escalation matrices, shift handoffs, post‑incident follow‑through
  • Prior NOC, SOC, or critical‑environment operations experience in a datacenter or hyperscale infrastructure environment
  • Familiarity with reading operational dashboards, acknowledging/classifying alerts, and coordinating across on‑site technicians, facilities, and engineering on‑call
  • Comfort with ticketing / project tracking systems (e.g. Linear, Jira) for opening and chasing corrective work to closure
  • Industry certifications a plus (Network+, Security+, ITIL, or similar) — not a substitute for judgment and communications quality
  • Basic familiarity with datacenter topology (racks, fabric, OOB) and how facility events affect compute availability — enough to triage and elevate correctly, not to deep‑diagnose
  • Bachelor's degree in IT, Computer Science, Cybersecurity, or STEM discipline preferred but not required
  • Must be available for on‑shift rotations supporting 24/7/365 console coverage
  • Shift structure (e.g. 12‑hour rotations) to be confirmed; nights, weekends, and holidays are part of the role
  • Must be able to work extended hours during major incidents as needed

Responsibilities

  • Continuous monitoring (the watch)
  • Staff the console per shift schedule to sustain 24/7 coverage (coverage posture: 2 on console per site)
  • Watch the designated signal surface: cluster health dashboards, node availability, network health, facility trend panels (power/cooling), storage alarms, and threshold breaches as defined by SRE monitoring standards
  • Acknowledge every page/alert within the SLA; classify it (actionable / known / noise) and log the disposition; feed noise patterns back to SRE for suppression or redesign
  • Maintain a live picture of ongoing maintenance, planned work, and degraded‑but‑accepted states so real anomalies stand out
  • Detect → verify → elevate within defined time budgets; verification is signal‑level (is it real, what's the blast radius), not deep diagnosis
  • Operate the escalation matrix: NOC → on‑call SRE → domain owners (SiteOps, Facilities, Network, Storage, HW FA, vendors); page correctly the first time
  • Recommend incident declaration and severity to the on‑call SRE; declare directly per runbook when thresholds are unambiguous
  • Open and run the bridge; get the right people on within the time‑to‑bridge SLA
  • Own stakeholder communications: first update within the SLA, then a fixed cadence until resolution
  • Maintain the incident timeline in real time — timestamps, actions, decisions, engagements
  • Track who owns what during the incident and call out stalls
  • Produce initial framing for major site outages: what happened, when it started, what's impacted (halls/racks/services), what changed recently, who is engaged
  • Hand framing to SRE / Hardware FA for depth — the NOC does not publish root cause
  • Write major‑incident reports; open corrective projects in Linear with named owners and track them to closure ("filed" is not "done")
  • Run structured shift handoffs and keep durable shift logs; maintain cross‑site awareness
  • Own and continuously improve NOC runbooks: escalation matrix, comms templates, severity ladders, per‑signal response procedures
  • Participate in game days run by SRE; every incident where the runbook was wrong or missing produces a runbook change before the incident closes

Technologies

LinearJira

See if your resume is ready for this job

See how our AI can optimize your resume and improve your chances for this role.