SIGN IN
Research Engineer – Training Infra jobs in United States
info-icon
This job has closed.
company-logo

Snorkel AI · 1 month ago

Research Engineer – Training Infra

Snorkel AI is on a mission to help enterprises transform expert knowledge into specialized AI at scale. As an Applied Research Engineer, you will own the infrastructure that powers model training and evaluation, building and operating GPU cluster infrastructure and ensuring that research and engineering teams can run experiments reliably and at scale.
Artificial Intelligence (AI)Big DataEnterprise SoftwareAI InfrastructureData Collection and LabelingMachine Learning
check
Growth Opportunities
check
H1B Sponsor Likelynote

Responsibilities

Set up and manage GPU cluster infrastructure on major cloud providers (e.g., AWS HyperPod) for distributed model training, including networking, provisioning, and cost tracking
Build and operate job orchestration and scheduling systems (e.g., Kubernetes, Slurm, or cloud-native equivalents) to reliably launch and manage training, rollout, and evaluation jobs across multi-node clusters
Integrate and maintain ML training frameworks and post-training pipelines, ensuring they run stably and reproducibly at scale
Set up and maintain experiment tracking, dataset versioning, and model artifact management to support fast iteration
Monitor and optimize cluster health, inter-node communication, and resource utilization; implement fault tolerance and auto-recovery so long-running jobs survive node failures
Work closely with research scientists and ML engineers to understand requirements, unblock experiments, and evolve infrastructure as our training workloads needs change

Qualification

GPU cluster infrastructure managementAWS HyperPodKubernetesSlurmDistributed trainingML experiment trackingDataset versioningModel artifact managementPythonSoftware engineering fundamentalsSupervised fine-tuningReinforcement learning

Required

Set up and manage GPU cluster infrastructure on major cloud providers (e.g., AWS HyperPod) for distributed model training, including networking, provisioning, and cost tracking
Build and operate job orchestration and scheduling systems (e.g., Kubernetes, Slurm, or cloud-native equivalents) to reliably launch and manage training, rollout, and evaluation jobs across multi-node clusters
Integrate and maintain ML training frameworks and post-training pipelines, ensuring they run stably and reproducibly at scale
Set up and maintain experiment tracking, dataset versioning, and model artifact management to support fast iteration
Monitor and optimize cluster health, inter-node communication, and resource utilization; implement fault tolerance and auto-recovery so long-running jobs survive node failures
Work closely with research scientists and ML engineers to understand requirements, unblock experiments, and evolve infrastructure as our training workloads needs change

Preferred

Hands-on experience managing GPU clusters on major cloud providers, including provisioning, network configuration, and cost management
Experience with distributed compute orchestration tools such as Kubernetes, Slurm, or equivalent cluster management systems
Working knowledge of distributed training concepts: parallelism strategies, memory optimization techniques, and inter-node communication
Experience with setting up, managing, and integrating ML experiment tracking and data/model versioning tools
Strong Python proficiency and solid software engineering fundamentals such as version control, modular design, and automation
Ability to work in a fast-moving, iterative environment and take end-to-end ownership of ambiguous infrastructure problems
Hands-on experience with post-training workflows such as supervised fine-tuning (SFT) or reinforcement learning (RLHF, GRPO, or similar) is a strong plus, but not required

Company

Snorkel AI

twitterlinkedincrunchbase
company-logo
Snorkel AI is an AI platform that accelerates data labeling by using machine learning for faster model training.

H1B Sponsorship

Snorkel AI has a track record of offering H1B sponsorships. Please note that this does not guarantee sponsorship for this specific role. Below presents additional info for your reference. (Data Powered by US Department of Labor)
Distribution of Different Job Fields Receiving Sponsorship
90%
Represents job field similar to this job
Engineering and Development
Management and Executive
Product Management
Sales
Trends of Total Sponsorships
*2026 (19)
2025 (15)
2024 (5)
2023 (4)
2022 (4)
2021 (10)
2020 (1)

Funding

Current Stage
Late Stage
Total Funding
$235.25M
Key Investors
Accenture VenturesAdditionQBE Ventures
2025-08-06Series Unknown
2025-05-29Series D· $100M
2024-01-23Series Unknown

Leadership Team

leader-logo
Alexander Ratner
Co-Founder and CEO
linkedin
leader-logo
Henry Ehrenberg
Co-Founder
linkedin
Company data provided by crunchbase