Member of Technical Staff - Distributed Training Engineer jobs in United States
cer-icon
Apply on Employer Site
company-logo

Liquid AI · 1 week ago

Member of Technical Staff - Distributed Training Engineer

Liquid AI, spun out of MIT CSAIL, builds general-purpose AI systems for various industries. They are seeking a Distributed Training Engineer to design and optimize infrastructure for large-scale training, focusing on performance and reliability.

Artificial Intelligence (AI)Foundational AIGenerative AIInformation TechnologyMachine Learning
check
H1B Sponsor Likelynote

Responsibilities

Design and build core systems that make large training runs fast and reliable
Build scalable distributed training infrastructure for GPU clusters
Implement and tune parallelism/sharding strategies for evolving architectures
Optimize distributed efficiency (topology-aware collectives, comm/compute overlap, straggler mitigation)
Build data loading systems that eliminate I/O bottlenecks for multimodal datasets
Develop checkpointing mechanisms balancing memory constraints with recovery needs
Create monitoring, profiling, and debugging tools for training stability and performance

Qualification

Distributed training infrastructurePerformance optimizationGPU clustersData pipeline optimizationHardware acceleratorsMonitoring toolsDebugging toolsSoft skills

Required

Hands-on experience building distributed training infrastructure (PyTorch Distributed DDP/FSDP, DeepSpeed ZeRO, Megatron-LM TP/PP)
Experience diagnosing performance bottlenecks and failure modes (profiling, NCCL/collectives issues, hangs, OOMs, stragglers)
Understanding of hardware accelerators and networking topologies
Experience optimizing data pipelines for ML workloads

Preferred

MoE (Mixture of Experts) training experience
Large-scale distributed training (100+ GPUs)
Open-source contributions to training infrastructure projects

Benefits

We pay 100% of medical, dental, and vision premiums for employees and dependents
401(k) matching up to 4% of base pay
Unlimited PTO plus company-wide Refill Days throughout the year

Company

Liquid AI

twittertwittertwitter
company-logo
Build efficient general-purpose AI at every scale.

H1B Sponsorship

Liquid AI has a track record of offering H1B sponsorships. Please note that this does not guarantee sponsorship for this specific role. Below presents additional info for your reference. (Data Powered by US Department of Labor)
Distribution of Different Job Fields Receiving Sponsorship
Represents job field similar to this job
Trends of Total Sponsorships
2025 (2)

Funding

Current Stage
Growth Stage
Total Funding
$293.1M
Key Investors
AMD VenturesOSS Capital L.P.
2024-12-13Series A· $250M
2023-12-01Seed· $37.5M
2023-05-05Seed· $5.6M

Leadership Team

leader-logo
Ramin Hasani
Co-founder and CEO
linkedin
leader-logo
Mathias Lechner
Co-founder and CTO
Company data provided by crunchbase