Meta · 2 days ago
AI/HPC System Performance Engineer
Meta is a technology company that builds platforms to help people connect and find communities. They are seeking an AI/HPC System Performance Engineer to lead teams in developing solutions for large scale training systems, ensuring performance and availability of the communication system, and defining technical vision for network architecture.
Computer Software
Responsibilities
Lead multi-disciplinary teams to develop solutions for large scale training systems. Assess trade-offs of various solutions and make pragmatic decisions
Ensure timely milestone delivery with teamwork and close collaboration
Responsible for the overall performance of the communication system, including performance benchmarking, monitoring and troubleshooting production issues
Defining technical vision and driving a multi-year roadmap to make progress towards the related objectives
Work with cross functional teams and provide guidance on the AI network architecture including topologies, transport, congestion control techniques
Qualification
Required
Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
Experience with developing, evaluating and debugging host networking protocols such as RDMA
10+ years of experience in designing, deploying and operating networks
Experience with triaging performance issues in complex scale-out distributed applications
Preferred
Experience with developing communication libraries, such as Message Passing Interface, NCCL, and UCX
Understanding of AI training workloads and demands they exert on networks
Understanding of RDMA congestion control mechanisms on InfiniBand and RoCE Networks
Understanding of the latest artificial intelligence (AI) technologies
Experience with machine learning frameworks such as PyTorch and TensorFlow
Experience in developing systems software in languages like C++
Benefits
Bonus
Equity
Benefits
Company
Meta
Meta's mission is to build the future of human connection and the technology that makes it possible.
Funding
Current Stage
Late StageRecent News
Crunchbase News
2025-11-17
2025-11-16
Company data provided by crunchbase