SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. The team is small, highly motivated, and focused on engineering excellence. The organization is flat, with all employees expected to be hands‑on and contribute directly to the mission. The role involves building and operating large‑scale networks that underpin training and inference infrastructure, including high‑performance/supercompute fabrics that connect GPU clusters and core, edge, and datacenter networks. This is a hands‑on engineering seat, not a NOC technician or network‑software SWE role; the engineer will design and build networks, qualify platforms, ship changes safely, and keep availability and performance high as the infrastructure scales.
Network Engineer at xAI
On-site - Palo Alto, CA
More jobs at xAISalary
USD 150,000 - 250,000
Requirements
Skills
- Several years designing and/or operating production networks in a datacenter, ISP, cloud, or large enterprise environment
- Solid hands‑on experience with BGP and at least one interior routing protocol
- Working knowledge of TCP/IP, VLANs, EVPN/VXLAN or equivalent datacenter overlays, and optics / high-speed Ethernet
- Experience troubleshooting live production network incidents and participating in on‑call
- Strong written and verbal communication; clear change docs and incident notes
- Experience with modern datacenter vendors (e.g. Arista, Cisco, Juniper, Nvidia/Mellanox)
- Familiarity with high-performance or supercompute networking (RoCEv2, congestion control, GPU cluster fabrics)
- Network automation (Python, Ansible, Terraform, or similar) used in production
- Experience with EVPN, leaf‑spine, and large‑scale Ethernet fabrics
- Prior work supporting rapid datacenter or cluster capacity build‑outs
- Willing to work onsite in Palo Alto
Responsibilities
- Design, deploy, and operate production datacenter and campus/core networks at scale
- Own routing and switching configuration standards (BGP and at least one IGP such as OSPF or IS‑IS), including change design, peer reviews, and execution
- Qualify new network platforms, optics, and topologies; contribute to architecture and capacity planning
- Build and improve monitoring, alerting, and operational documentation so issues are caught and fixed quickly
- Troubleshoot Layer 2/Layer 3 incidents end to end — from link flaps and optics through routing and traffic engineering — and drive root cause and lasting fixes
- Automate repetitive network tasks with Python, Ansible, or similar tooling where it reduces toil
- Partner with compute, facilities, and software teams during cluster build‑outs and maintenance windows
- Support high‑performance / supercompute network environments (Ethernet AI/HPC fabrics, RoCE/RDMA‑capable designs) as part of the broader network estate — deep specialist RoCE/NCCL ownership is a plus, not the bar for this seat
Technologies
PythonAnsibleTerraformBGPOSPFIS‑ISEVPNVXLANRoCERoCEv2NCCLAristaCiscoJuniperNvidia/MellanoxEthernet AI/HPC fabricsRoCE/RDMA‑capable designs
See if your resume is ready for this job
See how our AI can optimize your resume and improve your chances for this role.