AI/HPC System Performance Engineer jobs in United States
cer-icon
Apply on Employer Site
company-logo

Meta · 2 days ago

AI/HPC System Performance Engineer

Meta is a technology company that builds platforms to help people connect and find communities. They are seeking an AI/HPC System Performance Engineer to lead teams in developing solutions for large scale training systems, ensuring performance and availability of the communication system, and defining technical vision for network architecture.

Computer Software
check
Comp. & Benefits

Responsibilities

Lead multi-disciplinary teams to develop solutions for large scale training systems. Assess trade-offs of various solutions and make pragmatic decisions
Ensure timely milestone delivery with teamwork and close collaboration
Responsible for the overall performance of the communication system, including performance benchmarking, monitoring and troubleshooting production issues
Defining technical vision and driving a multi-year roadmap to make progress towards the related objectives
Work with cross functional teams and provide guidance on the AI network architecture including topologies, transport, congestion control techniques

Qualification

Host networking protocolsNetwork designDeploymentPerformance benchmarkingCommunication librariesAI training workloadsRDMA congestion controlMachine learning frameworksSystems software developmentTechnical visionTeam collaboration

Required

Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
Experience with developing, evaluating and debugging host networking protocols such as RDMA
10+ years of experience in designing, deploying and operating networks
Experience with triaging performance issues in complex scale-out distributed applications

Preferred

Experience with developing communication libraries, such as Message Passing Interface, NCCL, and UCX
Understanding of AI training workloads and demands they exert on networks
Understanding of RDMA congestion control mechanisms on InfiniBand and RoCE Networks
Understanding of the latest artificial intelligence (AI) technologies
Experience with machine learning frameworks such as PyTorch and TensorFlow
Experience in developing systems software in languages like C++

Benefits

Bonus
Equity
Benefits

Company

Meta's mission is to build the future of human connection and the technology that makes it possible.

Funding

Current Stage
Late Stage

Leadership Team

leader-logo
Kathryn Glickman
Director, CEO Communications
linkedin
leader-logo
Christine Lu
CTO Business Engineering NA
linkedin
Company data provided by crunchbase