Member of Technical Staff - Post-Training and RL at xAI

On-site - Palo Alto, CA

Apply
More jobs at xAI

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our small, highly motivated, engineering-focused team values hands‑on contributions, strong prioritization, communication skills, and leadership that emerges from initiative. The organization operates with a flat structure, encouraging all employees to contribute directly to the mission.

Salary

USD 180,000 - 600,000

Requirements

Skills

  • Experience with post-training, RLHF, or large-scale trained models (not required)
  • Experience with reinforcement learning and alignment methods
  • Power user of AI models
  • Belief in truth-seeking AI as a core challenge
  • Thrives in meritocratic environments and takes pride in work

Responsibilities

  • Work on critical post-training and reinforcement learning challenges including reward modeling, preference optimization (RLHF/DPO), and RL for improving reasoning, truthfulness, and real-world capabilities
  • Obtain clarity on first project before offer

Technologies

Post-trainingReinforcement learningReward modelingPreference optimizationRLHFDPOAlignment methods

See if your resume is ready for this job

See how our AI can optimize your resume and improve your chances for this role.