SpaceXAI seeks a Manager, Operations to lead facilities operations and power generation for hyperscale AI compute facilities. The role owns the day‑to‑day and long‑term performance of critical data center operations, including power, cooling, mechanical, electrical, environmental systems, and fiber networking, ensuring 24/7 uptime for AI training at unprecedented scale. The position requires deep expertise in data center or hyperscale operations, leadership in fast‑paced environments, and delivering world‑class performance under aggressive growth timelines.
Manager, Operations en SpaceXAI
Presencial - Memphis, TN, USA
Más vacantes en SpaceXAIRequirements
Skills
- 5+ years of progressive experience in data center facilities operations, power generation operations, hyperscale infrastructure management, or mission-critical industrial operations
- At least 2+ years in a management or supervisor role
- Proven track record leading large-scale operations teams supporting high-density compute environments with significant on-site or dedicated power generation
- Strong experience managing fiber optic networks, dark fiber deployments, or high-bandwidth connectivity infrastructure
- Deep knowledge of power generation systems (gas turbines, reciprocating engines, cogeneration, etc.), MEP (mechanical, electrical, plumbing) systems, BMS/SCADA, liquid cooling, power redundancy topologies, and 24/7 operations best practices
- Demonstrated success delivering high reliability, rapid incident resolution, and operational excellence under aggressive scaling timelines
- Hands‑on leadership style with the ability to roll up sleeves while effectively managing teams, budgets, and cross‑functional stakeholders
- Proficiency with operations tools, CMMS (computerized maintenance management systems), monitoring platforms, and data-driven decision making
- Bachelor’s or Master’s degree in Electrical, Mechanical Engineering, Power Systems, Facilities Management, or related field (certifications such as CDCP or CDCS are a plus)
- Willingness to be primarily onsite at key facilities with on‑call responsibilities and travel to other sites as needed
- Ability to work in industrial/data center environments and lead teams during high-pressure phases
Responsibilities
- Lead and scale the facilities operations and power generation teams responsible for the reliable operation, maintenance, monitoring, and optimization of critical infrastructure including on‑site power generation assets, electrical systems, mechanical/HVAC, liquid cooling, power distribution, UPS, generators, and building management systems
- Direct the fiber teams overseeing the design, deployment, maintenance, and expansion of high-speed fiber optic networks, dark fiber, and connectivity infrastructure supporting AI compute clusters and data center interconnects
- Own key performance metrics such as uptime (targeting 99.999%+), mean time to detect/repair (MTTD/MTTR), power usage effectiveness (PUE), water usage effectiveness (WUE), power generation efficiency, and overall infrastructure availability
- Develop and enforce standard operating procedures (SOPs), preventive maintenance programs, incident response protocols, and continuous improvement processes for both facilities and power generation assets to minimize downtime and maximize efficiency
- Build, mentor, and grow multidisciplinary teams of operations technicians, power generation engineers and controls specialists while fostering a culture of ownership, safety, and excellence
- Partner closely with engineering, construction, procurement, and AI hardware teams to support new facility builds, expansions, commissioning, power integration, and smooth handovers from project to operations
- Manage operational budgets, vendor relationships (maintenance contractors, fiber providers, power generation OEMs, fuel suppliers), spare parts inventory, and risk mitigation strategies in a high‑velocity environment
- Drive innovation in operational practices, automation, predictive maintenance, power generation optimization, and sustainability initiatives to support the extreme power and cooling demands of next‑generation AI systems
- Provide regular performance reporting, root cause analyses, lessons learned, and strategic recommendations to senior leadership
Technologies
CMMSmonitoring platformsBMS/SCADAgas turbinesreciprocating enginescogenerationMEP systemsliquid coolingfiber opticdark fiberhigh‑bandwidth connectivity
Descubre si tu currículum está listo para esta vacante
Mira cómo nuestra IA puede optimizar tu currículum y aumentar tus chances en este puesto.