SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands‑on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills.
Site Reliability Engineer at xAI
On-site - Memphis, Tennessee; Southaven, Mississippi
More jobs at xAIRequirements
Skills
- Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field (or equivalent experience)
- 5+ years of experience in site reliability, systems engineering, or large-scale production operations, preferably in high-performance computing or data center environments
- Proven large-scale incident command experience and calm technical leadership on a bridge
- Demonstrated monitoring and observability design at fleet or campus scale, including alert hygiene, suppression, and signal quality
- Experience working across at least two of: compute, network, storage, power, and cooling / facilities telemetry
- Experience writing and operating playbooks or runbooks with a 24/7 operations or NOC partner
- Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Go, Rust, or similar)
- Excellent problem-solving skills with a data-driven approach to reliability engineering
- Ability to work collaboratively with cross-functional teams, including NOC, data center operations, and infrastructure engineering
- Experience in AI/ML infrastructure or supercomputing environments
- Hands-on definition and use of SLOs, SLIs, and error budgets at service or campus boundaries
- Experience running game days, dependency mapping, and closed-loop corrective action programs
- Familiarity with data center hardware and plant signals (servers, GPUs, networking, power, cooling) in addition to software telemetry
- Prior work in a fast-paced startup or tech company like SpaceXAI
Responsibilities
- Own monitoring architecture and signal quality
- Provide SEV command support
- Run blameless postmortems and drive corrective actions
- Lead cross-functional reliability projects spanning compute, network, storage, and facility signal boundaries
- Build and maintain playbooks, run game days, keep cross-discipline dependency maps current, own runbook quality
- Define error budgets and availability objectives at campus and service boundaries
- Participate in on-call rotations and incident response for SEV-class events in the Memphis / Southaven data center campus
Technologies
PythonBashCC++JavaGoRust
See if your resume is ready for this job
See how our AI can optimize your resume and improve your chances for this role.