SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. The team is small, highly motivated, and focused on engineering excellence. As a Data Engineer in the X product engineering team, you will play a key role in providing comprehensive data solutions that serve stakeholders, maximize the potential of data, and enable product decisions. The role involves building and operating a distributed data platform, designing pipelines, creating datasets, automating workflows, ensuring data correctness, collaborating across query engines and frameworks, partnering with product and business teams, and iterating quickly on feedback.
Software Engineer - X Data at SpaceXAI
On-site - Palo Alto, CA
More jobs at SpaceXAISalary
USD 125,000 - 400,000
Requirements
Skills
- 3+ years of professional software engineering experience
- Experience in data engineering or distributed systems
- Hands‑on expertise in Python, Rust, Scala, Go or Java
- Knowledge of data pipeline tooling and distributed systems
- Experience with realtime and batch data processing tools such as Spark, Kafka, Flink, SQL
- Familiarity with storage systems in RMDBs and NoSQL
- Ability to solve large‑scale problems and build new systems for future improvements
- Proven record of translating product requirements into engineering implementation plans
- Strong communication skills across AI, product, marketing/sales, and engineering groups
Responsibilities
- Design, build, and operate production‑grade realtime and batch pipelines that ingest, process, validate, and deliver data powering user‑behavior insights and product decisions
- Create shared datasets, fact tables, and internal data products for other teams to analyze, debug, and improve product performance
- Prototype and build tooling that automates and accelerates internal data workflows (backfills, dashboards, report generation, self‑serve access to data)
- Own data correctness end to end: validate with output invariants, denominator reconciliation, and independent recomputation; lead root‑cause investigations when key metrics move unexpectedly
- Move fluidly across query engines and frameworks (e.g., BigQuery, Trino, Clickhouse for analytics; Flink, Kafka, Spark/Scalding for streaming and batch), choosing the right tool and adapting quickly to new infrastructure
- Partner across product and business teams to surface data gaps and prioritize the highest‑impact opportunities for new data acquisition and improvement
- Iterate quickly on feedback, shipping the smallest useful increment with a strong bias toward efficient, accurate, and reliable solutions
Technologies
PythonRustScalaGoJavaSparkKafkaFlinkSQLBigQueryTrinoClickhouseSpark/Scalding
See if your resume is ready for this job
See how our AI can optimize your resume and improve your chances for this role.