This job has closed.

NVIDIA · 1 month ago

Engineering Manager - AI DevOps

United States

Full-time

Remote

Senior Level, Lead/Staff

$224K/yr - $426K/yr

8+ years exp

NVIDIA is looking for an outstanding AI DevOps Engineering Manager to lead and expand their next-gen inference operations infrastructure. This role is essential for transforming AI inference delivery and ensuring that NVIDIA's AI products achieve outstanding performance and reliability worldwide.

Artificial Intelligence (AI)SemiconductorConsumer GoodsHardwareSoftwareAppsAI InfrastructureConsumer ElectronicsFoundational AIGPUVirtual Reality

Growth Opportunities

H1B Sponsor Likely

Hiring Manager

Bushra S.

Responsibilities

Supervise a team of DevOps engineers with expertise in AI inference infrastructure, test automation (SDET), and Infrastructure as Code (IaC)

Architect and implement scalable test automation strategies for AI inference workloads, including performance benchmarking and automated quality gates

Lead the maintenance of our GitHub First public CI infrastructure, focusing on single/multi-GPU testing, Kubernetes multi-node GPU testing, and CSP validation

Drive Infrastructure as Code efforts by employing Terraform, Ansible, and Kubernetes to support scaling across multiple clouds and lead GPU clusters effectively

Attain operational proficiency encompassing 24x7 on-call rotations, SRE methodologies, automated monitoring, and self-repairing systems to guarantee uptime exceeding 99.9%

Lead release coordination, cost optimization, and management of multi-cloud deployments

Qualification

AI inference infrastructureInfrastructure as CodeCI/CD pipeline developmentTest automation frameworksKubernetesTerraformAnsiblePythonMonitoring toolsEffective interpersonal skillsTeam leadership

Required

Bachelor's/Master's degree in Computer Science, Engineering, or equivalent experience

4+ years leading DevOps/SRE organizations with direct SDET leadership experience

8+ years hands-on experience in software development, test automation, or infrastructure engineering with AI/ML or GPU-intensive workloads

Proficiency in Infrastructure as Code (IaC) platforms: Terraform, Ansible, or CloudFormation with exposure to multiple cloud environments (AWS, GCP, Azure, OCI)

Strong technical leadership in test automation frameworks, CI/CD pipeline development, and quality engineering practices

Familiarity with containerization and orchestration tools such as Docker and Kubernetes for leading AI/ML workloads and GPU resources

Proven success building and scaling teams in fast-paced, high-growth environments

Effective interpersonal skills to collaborate with remote teams and build agreement

Proficiency in Python, Rust, or related programming languages along with the capability to engage in architecture conversations

Demonstrated history of operational proficiency encompassing 24x7 on-call oversight, SRE methodologies, and robust high-availability infrastructures

Preferred

Experience with CI/CD (specifically GitHub Actions), releasing Open-source AI software

Proficient in Deep AI/ML infrastructure with expertise in NVIDIA technologies such as CUDA, TensorRT, Dynamo and Triton Inference Server, including coordinating GPU cluster operations and GPU workload performance benchmarking

Background in DevOps, system software testing, and previous experience leading teams on inference engines, model serving platforms, or AI acceleration frameworks

Track record with monitoring tools (Prometheus, Grafana), security scanning, static/dynamic analysis tools, and license compliance automation for critical AI inferencing frameworks

Benefits

Equity

Benefits

Company

NVIDIA

Glassdoor4.6

NVIDIA is a computing platform company operating at the intersection of graphics, HPC, and AI.

Founded in 1993

Santa Clara, California, USA

10001+ employees

https://www.nvidia.com

H1B Sponsorship

NVIDIA has a track record of offering H1B sponsorships. Please note that this does not guarantee sponsorship for this specific role. Below presents additional info for your reference. (Data Powered by US Department of Labor)

Distribution of Different Job Fields Receiving Sponsorship

Represents job field similar to this job

Trends of Total Sponsorships

2025 (1877)

2024 (1355)

2023 (976)

2022 (835)

2021 (601)

2020 (529)

Funding

Current Stage

Public Company

Total Funding

$4.09B

Key Investors

ARPA-EARK Investment ManagementSoftBank Vision Fund

2023-05-09Grant· $5M

2022-08-09Post Ipo Equity· $65M

2021-02-18Post Ipo Equity

Leadership Team

Jensen Huang

Founder and CEO

Michael Kagan

Chief Technology Officer

Recent News

legacy.thefly.com

Mixed options sentiment in NVIDIA with shares up 1.79%

2026-02-12

BizWatchNigeria.Ng

Nvidia Dethrones Tech Kings: Chipmaker Becomes World’s Most Valuable Company

2026-02-12

Livemint.com

Indian stock market: 10 key things that changed for market overnight - Gift Nifty, US nonfarm payrolls to gold prices

2026-02-12

Company data provided by crunchbase