Apply on Employer Site

NVIDIA · 22 hours ago

Senior Site Reliability Engineer - Storage

Santa Clara, CA

Full-time

Onsite

Senior Level

$168K/yr - $270K/yr

5+ years exp

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. They are seeking a Senior Site Reliability Engineer focused on HPC storage to design, implement, and optimize on-prem storage solutions while collaborating with engineering teams.

AI InfrastructureArtificial Intelligence (AI)Consumer ElectronicsFoundational AIGPUHardwareSoftwareVirtual Reality

Growth Opportunities

H1B Sponsor Likely

Responsibilities

Design, implement an on-prem HPC infrastructure supplemented with cloud computing to support the growing IT needs of NVIDIA

Design and implement advanced storage solutions, such as high-performance NFS, S3-compatible object storage, and distributed storage systems

Develop tooling to automate deployment and management of large-scale infrastructure environments, to automate operational monitoring and alerting, and to enable self-service consumption of resources

Document the general procedures and practices, perform technology evaluations, related to distributed file systems

Collaborate across teams to better understand developers' workflows and gather their infrastructure requirements

Influence and guide methodologies for building, testing, and deploying applications to ensure optimal performance and resource utilization

Qualification

HPC storage solutionsStorage protocolsContainerization technologiesProgramming languagesCloud infrastructureMonitoring toolsHPCAI technologiesRDMA fabricsHPC cluster managementCommunication skillsCollaboration skills

Required

BS in Computer Science (or equivalent experience) with 8+ years of relevant experience, MS with 5+ years of experience or Ph.D. with 3 years of experience

Deep experience with storage protocols such as nfs, NVMe/TCP, S3 and Lustre (LNet)

Experience with containerization technologies like Kubernetes and their integration with storage solutions

Proficiency in one or more programming languages (Python, GO) is a must

Experience working with monitoring and configuration management tools such as Chef, Ansible, Puppet, Saltstack, etc

Background with cloud infrastructure - AWS, Azure or Google Cloud

Experience with multiple monitoring stacks such as Prometheus+Grafana, Elasticsearch+Kibana

Excellent communication and collaboration skills

Preferred

Knowledge of HPC and AI solution technologies from CPU's and GPU's to high speed interconnects and supporting software

Experience with RDMA (InfiniBand or RoCE) fabrics

Background with HPC cluster management tools such as Slurm, PBS, LSF, etc

Passionate and experienced in AI methodologies

Benefits

Equity

Benefits

Company

NVIDIA

Glassdoor4.6

NVIDIA is a computing platform company operating at the intersection of graphics, HPC, and AI.

Founded in 1993

Santa Clara, California, USA

10001+ employees

https://www.nvidia.com

H1B Sponsorship

NVIDIA has a track record of offering H1B sponsorships. Please note that this does not guarantee sponsorship for this specific role. Below presents additional info for your reference. (Data Powered by US Department of Labor)

Distribution of Different Job Fields Receiving Sponsorship

Represents job field similar to this job

Trends of Total Sponsorships

2025 (1877)

2024 (1355)

2023 (976)

2022 (835)

2021 (601)

2020 (529)

Funding

Current Stage

Public Company

Total Funding

$4.09B

Key Investors

ARPA-EARK Investment ManagementSoftBank Vision Fund

2023-05-09Grant· $5M

2022-08-09Post Ipo Equity· $65M

2021-02-18Post Ipo Equity

Leadership Team

Jensen Huang

Founder and CEO

Michael Kagan

Chief Technology Officer

Recent News

The Motley Fool

Prediction: Nvidia Stock Is Going to Soar After Feb. 25

2026-02-04

Business Insider

Tech stocks are leading a fresh market sell-off as oil prices spike

2026-02-04

Cointelegraph

Bitcoin loses $73K as US stocks sell off: Analyst says BTC price action not ‘abnormal’

2026-02-04

Company data provided by crunchbase