Senior Datacenter System Software Architect - DGX Cloud jobs in United States
cer-icon
Apply on Employer Site
company-logo

NVIDIA · 4 months ago

Senior Datacenter System Software Architect - DGX Cloud

NVIDIA is hiring engineers to scale up its AI Infrastructure, seeking a highly motivated engineer with strong experience in system software to join the DGX Cloud Software Team. The role involves leading the architecture, design, and implementation of next-generation DGX cloud clusters, focusing on hybrid deployments and ensuring seamless integration across various engineering teams.

AI InfrastructureArtificial Intelligence (AI)Consumer ElectronicsFoundational AIGPUHardwareSoftwareVirtual Reality
check
Growth Opportunities
check
H1B Sponsor Likelynote

Responsibilities

Lead technical activities for data centers with focus on hybrid deployments between cloud and on-prem
Providing expertise in infrastructure workflows, including hardware, software release, workload orchestration and application tuning
Provide fast and creative solutions for complex problems and write effective, clear and reliable architecture specification
Translate requirements to vision, architecture and roadmap
Work with engineering teams across NVIDIA to ensure your software integrates seamlessly from the hardware all the way up to the AI training applications

Qualification

System software architectureDistributed systemsProgramming in PythonHigh-level programming languagesData SciencesDeep LearningMachine LearningGPU deep learningDocker containersKubernetesAnalytical skillsDebugging skillsProblem-solving skillsCommunication skillsContinuous learning

Required

Masters or PhD in Computer Science, Computer Engineering, Physics or equivalent experience
9+ years of experience in this field
Data Sciences, Deep Learning, or Machine Learning coursework
Ability to seamlessly shift between Linux system environments to Python programming
Programming skills in 1 or more high-level languages (C, C++, Go, Rust, etc)
System-level experience with both hardware and software
Motivated self-starter with an equal balance of strong problem-solving skills and customer-facing communication skills
Strong design, coding, analytical, debugging and problem-solving skills
Passion for continuous learning and knowledge transfer. Ability to work concurrently with multiple groups locally and abroad in the organization

Preferred

Experience with GPU deep learning and data sciences
Experience using TensorFlow, PyTorch or other DL framework
Experience working with Docker containers, Slurm, Terraform and Kubernetes
CUDA programming and NCCL experience
HPC programming experience including MPI, OpenACC, or other parallel programming tools
Hands-on experience with DGX Cloud, NVIDIA AI Enterprise AI Software, Base Command Manager, NEMO and NVIDIA Inference Microservices
Interest in crafting, analyzing and fixing large-scale distributed systems
Systematic problem-solving approach, coupled with strong communication skills and a sense of ownership and drive

Benefits

Equity
Benefits

Company

NVIDIA is a computing platform company operating at the intersection of graphics, HPC, and AI.

H1B Sponsorship

NVIDIA has a track record of offering H1B sponsorships. Please note that this does not guarantee sponsorship for this specific role. Below presents additional info for your reference. (Data Powered by US Department of Labor)
Distribution of Different Job Fields Receiving Sponsorship
Represents job field similar to this job
Trends of Total Sponsorships
2025 (1877)
2024 (1355)
2023 (976)
2022 (835)
2021 (601)
2020 (529)

Funding

Current Stage
Public Company
Total Funding
$4.09B
Key Investors
ARPA-EARK Investment ManagementSoftBank Vision Fund
2023-05-09Grant· $5M
2022-08-09Post Ipo Equity· $65M
2021-02-18Post Ipo Equity

Leadership Team

leader-logo
Jensen Huang
Founder and CEO
linkedin
leader-logo
Michael Kagan
Chief Technology Officer
linkedin
Company data provided by crunchbase