SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our small, highly motivated engineering team builds and operates large‑scale networks that underpin training and inference infrastructure, including high‑performance super‑compute fabrics and core, edge, and datacenter networks. This hands‑on engineering role focuses on designing, deploying, and operating production datacenter and campus/core networks, owning routing and switching standards, qualifying new platforms, and automating network tasks to maintain high availability and performance.
Network Engineer at xAI
On-site - Dublin, Ireland
More jobs at xAISalary
USD 100,000 - 150,000
Requirements
Skills
- Several years designing and/or operating production networks in a datacenter, ISP, cloud, or large enterprise environment
- Solid hands‑on experience with BGP and at least one interior routing protocol
- Working knowledge of TCP/IP, VLANs, EVPN/VXLAN or equivalent datacenter overlays, and optics / high‑speed Ethernet
- Experience troubleshooting live production network incidents and participating in on‑call
- Strong written and verbal communication; clear change docs and incident notes
- Experience with modern datacenter vendors (e.g. Arista, Cisco, Juniper, Nvidia/Mellanox)
- Familiarity with high‑performance or supercompute networking (RoCEv2, congestion control, GPU cluster fabrics)
- Network automation (Python, Ansible, Terraform, or similar) used in production
- Experience with EVPN, leaf‑spine, and large‑scale Ethernet fabrics
- Prior work supporting rapid datacenter or cluster capacity build‑outs
- Willing to work onsite in Dublin
Responsibilities
- Design, deploy, and operate production datacenter and campus/core networks at scale
- Own routing and switching configuration standards (BGP and at least one IGP such as OSPF or IS-IS), including change design, peer reviews, and execution
- Qualify new network platforms, optics, and topologies; contribute to architecture and capacity planning
- Build and improve monitoring, alerting, and operational documentation so issues are caught and fixed quickly
- Troubleshoot Layer 2/Layer 3 incidents end to end — from link flaps and optics through routing and traffic engineering — and drive root cause and lasting fixes
- Automate repetitive network tasks with Python, Ansible, or similar tooling where it reduces toil
- Partner with compute, facilities, and software teams during cluster build‑outs and maintenance windows
- Support high‑performance / supercompute network environments (Ethernet AI/HPC fabrics, RoCE/RDMA‑capable designs) as part of the broader network estate
Technologies
BGPOSPFIS-ISTCP/IPVLANsEVPNVXLANPythonAnsibleTerraformRoCERDMA
See if your resume is ready for this job
See how our AI can optimize your resume and improve your chances for this role.