Technical Account Manager en xAI

Presencial - Bastrop, TX; Los Angeles, CA; Seattle, WA

Postularse
Más vacantes en xAI

SpaceXAI’s mission is to build AI systems that accurately understand the universe and aid humanity in its pursuit of knowledge. The company maintains a small, highly motivated engineering team operating in a flat organizational structure where employees contribute directly to the mission. We are looking for a Technical Account Manager / Solutions Architect to own the technical success of deployments, leading implementation, integration, and optimization across all customer accounts to ensure reliable, high‑performance compute environments from initial provisioning through steady‑state operations.

Requirements

Skills

  • Bachelor’s degree in computer science, Electrical Engineering, Computer Engineering, or a related technical discipline
  • 7+ years of hands-on professional experience in large-scale compute, data center, or hyperscale infrastructure environments
  • Experience leading technical implementations involving GPUs/CPUs and high-performance systems
  • Master’s degree or higher in a STEM discipline
  • Prior experience as a Technical Account Manager, Solutions Architect, or equivalent supporting enterprise or hyperscale compute deployments
  • Deep knowledge of InfiniBand, RoCE, NVIDIA networking technologies (NCCL, NVLink, GPU Direct RDMA), and large GPU cluster architectures
  • Experience with AI/ML training and inference infrastructure at scale
  • Strong background in performance tuning, security/compliance implementation, and hybrid/cloud integration
  • Demonstrated ability to assess risk and make decisions with incomplete data in a fast-paced environment
  • Excellent written and verbal communication skills with the ability to interface effectively with customers and internal stakeholders
  • Hands‑on experience with Linux environments, orchestration tools, and infrastructure monitoring
  • Ability to work extended hours and weekends as needed during critical implementation, testing, and incident response periods
  • Willingness to travel (up to 25%) to data center sites and customer locations
  • Must be comfortable operating with extreme ownership in a demanding, high-expectation environment

Responsibilities

  • Apply expertise to large-scale compute system development, including design validation, integration, performance tuning, and optimization
  • Act as the primary technical point of contact for compute infrastructure definition, requirements, and delivery across all customer accounts
  • Lead technical implementation, integration testing, provisioning, and go‑live activities for GPU/CPU clusters and supporting infrastructure
  • Provide deep expertise on compute-specific elements including GPUs/CPUs, high-performance networking, data center connectivity, and related systems
  • Handle technical escalations, root‑cause analysis, troubleshooting, and drive resolution to ensure SOW and SLA compliance
  • Represent compute systems in cross‑functional trades, risk discussions, and issue resolution with internal teams, management, and customers
  • Interface with Hardware, Network, Facilities, Cloud Compute, SRE, and Program Management teams to ensure successful end-to-end delivery
  • Participate in verification testing, performance characterization, reliability assessments, and continuous optimization
  • Support expansion planning, capacity scaling, and new use cases across customer accounts
  • Ensure on‑time deliverables, proactive risk mitigation, and overall mission success for all customers

Technologies

GPUCPUInfiniBandRoCENVIDIA NCCLNVIDIA NVLinkGPU Direct RDMALinuxOrchestration toolsInfrastructure monitoring

Compartir vacante

Descubre si tu currículum está listo para esta vacante

Mira cómo nuestra IA puede optimizar tu currículum y aumentar tus chances en este puesto.