PB✓
PBridge

Full-time jobsthe United States

Sr. High Performance Computing (HPC) Systems Engineer

spacex · Hawthorne, CA · Full-time

About this role

SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.

SR. HIGH PERFORMANCE COMPUTING (HPC) SYSTEMS ENGINEER 

SpaceX is looking for an HPC Systems Engineer with strong knowledge and experience in a world class engineering organization. This employee will be a member of the HPC team and will support SpaceX personnel and proprietary systems. The ideal candidate will be flexible and flourish in a fast paced and challenging environment. They should be a self-starter, self-motivator and possess ingenuity to excel at this position.

RESPONSIBILITIES:

• Administer and manage HPC clusters, storage systems, and high-speed networks

• Provide application support to SpaceX employees across engineering disciplines

• Install and integrate Linux-based compute clusters

• Write instructional documentation and convey highly technical ideas in non-technical terms

BASIC QUALIFICATIONS:

• Bachelor's degree in computer science, engineering, math, or scientific discipline and 5+ years of systems engineering experience; OR 7+ years of professional experience building software in lieu of a degree

• 5+ years of hands-on experience with client and server hardware/software, management tools, enterprise networking, virtualization, and security technologies

• Experience with Kubernetes

PREFERRED SKILLS AND EXPERIENCE:

• 5+ years of professional experience building, deploying and troubleshooting Linux systems

• Experience with a scripting language (Bash, Python) to automate and solve reoccurring tasks

• Experience building, deploying and troubleshooting HPC clusters

• Familiarity with cluster resource managers (Slurm, PBS, LSF)

• Experience with monitoring and alerting technologies (Prometheus, Grafana, Nagios)

• Familiarity with scientific and engineering computing (CFD, FEA)

• Familiarity with large scale AI training

• Familiarity with GPU usage in a compute cluster and Cuda

• Experience with containers (Docker, Podman, Singularity)

• Experience deploying and maintaining automated configuration management software (Puppet, Ansible)

• Comfortable working with mission critical and sensitive systems, with a sense of urgency appropriate to the responsibilities

• Eligibility for access to classified material up to TS/SCI with Polygraph

ADDITIONAL REQUIREMENTS:

• Must be willing to work extended hours and weekends as needed

COMPENSATION AND BENEFITS:    

Pay Range:          Sr. HPC Systems Engineer: $165,000.00-$230,000.00/per year

Your actual level and base salary will be determined on a case-by-case basis and may vary based on the following considerations: job-related knowledge and skills, education, and experience.

Base salary is just one part of your to

Tired of applying one by one?

Our Career Success Team finds roles in the United States that fit you, tailors your CV to each, and submits the applications — tracked end to end. You just show up to interviews.

We apply, you interview →