NVIDIA · 1 week ago
Principal Software Engineer - Rack Scale Systems Infrastructure
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. As a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software systems that support NVIDIA's upcoming rack-scale infrastructure products and services, translating complex hardware into manageable infrastructure.
Artificial Intelligence (AI)TransportationCloud ComputingEmbedded SoftwareGamingHardwareQuantum ComputingSemiconductorAI InfrastructureAutonomous VehiclesEmbedded SystemsFoundational AI
Responsibilities
Define the complete software architecture for rack-scale infrastructure products and services, covering control plane services, infrastructure management, firmware, operating systems, kernel drivers, networking fabrics, accelerator software, and user-mode manageability software
Use Kubernetes and cloud-native primitives as an infrastructure fabric when appropriate. This includes controllers, operators, reconciliation loops, and open source components. These components can operate safely at rack and fleet scale. Build open source infrastructure software that can be embraced in different forms, including libraries, services, controllers, operators, and integration APIs for internal deployments and CSP environments
Bridge hardware and software teams across firmware, BMC, BIOS, boot flows, OS images, drivers, networking, NVLink domains, InfiniBand, GPUs, DPUs, CPUs, and system management interfaces. Translate forward-looking infrastructure roadmaps into formal software requirements, architecture specifications, and execution plans that align teams across the organization
Partner directly with hyperscalers, CSPs, enterprise customers, internal component leads, vendors, and business partners to align infrastructure capabilities with real-world deployment and integration needs. Establish reliability, security, validation, and left-shift strategies that reduce risk before hardware reaches production environments
Mentor senior engineers and technical leads, raising the engineering bar for large-scale networked systems, foundational software, and rack-scale control plane development
Make high-quality technical decisions in ambiguous environments, balancing customer needs, schedule, hardware realities, software maintainability, open source adoption, and long-term infrastructure evolution
Qualification
Required
BS or MS in Computer Engineering, Computer Science, Electrical Engineering, or a related field, or equivalent experience
Proven experience (15+ years) in systems architecture, system software, distributed systems, infrastructure control planes, or infrastructure engineering
Solid architectural knowledge of coordination frameworks, state machines, declarative APIs, reconciliation loops, lifecycle orchestration, failure handling, upgrade and rollback workflows, and distributed systems tradeoffs
Practical coding skills in Go, C++, or Rust, encompassing the capability to write, review, and direct production-quality infrastructure software
Experience with Kubernetes or similar orchestration systems, especially as a fabric for managing infrastructure, hardware resources, or large-scale infrastructure services
Experience with Linux-based infrastructure software, OS rollout and image management, kernel or driver interactions, firmware lifecycle, and hardware bring-up workflows
Strong understanding of data center networking technologies and protocols, such as Ethernet, InfiniBand, RDMA, and fabric-level manageability
Experience with complex accelerator-based systems, including GPUs, DPUs, FPGAs, custom silicon, or other high-performance computing systems
Expertise in in-band and out-of-band management architectures, including BMCs, Redfish, IPMI, and related system management protocols
Ability to work with security experts to define practical tradeoffs across secure boot, attestation, access control, update safety, serviceability, and ease of operation
Experience crafting software intended for open source release, including API stability, modularity, documentation, community usability, and clean separation between shared software and deployment-specific integrations
Experience using AI-assisted development tools responsibly as an engineering multiplier for coding, test generation, debugging, build iteration, and documentation
Established skill in specifying requirements, guiding architecture, and managing delivery across various engineering teams and organizations
Strong written and verbal communication skills, enabling clear explanation of complex hardware/software tradeoffs to engineering leaders, customers, partners, and executives
Preferred
Built software supporting multiple adoption models — internal services, CSP-integrated offerings, reusable libraries, and customer-extensible APIs
Strong Rust skills in systems, infrastructure, or hardware-adjacent software
Multiplied team impact through reference implementations, design reviews, shared libraries, architecture docs, dev workflows, and AI-assisted engineering
Hands-on with fleet-scale provisioning, updates, rollback, observability, health, and remediation
Led across the full data center product lifecycle: inception, pre- and post-silicon, manufacturing, deployment, and operations
Familiar with open source ecosystems, contribution models, and balancing community collaboration with product needs
Deep experience with rack- or cluster-scale systems spanning compute, networking, storage, accelerators, firmware, and infra management as one operational domain
Skilled at finding simple, durable abstractions in complex systems to align teams, customers, and long-term direction
Benefits
You will also be eligible for equity and [benefits](https://www.nvidia.com/en-us/benefits/).
Company
NVIDIA
NVIDIA is a computing platform company operating at the intersection of graphics, HPC, and AI.
H1B Sponsorship
NVIDIA has a track record of offering H1B sponsorships. Please note that this does not
guarantee sponsorship for this specific role. Below presents additional info for your
reference. (Data Powered by US Department of Labor)
Distribution of Different Job Fields Receiving Sponsorship
89%
Represents job field similar to this job
Engineering and Development
Management and Executive
Product Management
Marketing
Sales
Accounting and Finance
Customer Service and Support
Creatives and Design
Human Resources
Arts and Entertainment
Legal and Compliance
Trends of Total Sponsorships
*2026 (1247)
2025 (1868)
2024 (1353)
2023 (976)
2022 (835)
2021 (601)
2020 (529)
Funding
Current Stage
Public CompanyTotal Funding
$29.09BKey Investors
ARPA-EARK Investment ManagementSoftBank Vision Fund
2026-06-15Post Ipo Debt· $25B
2023-05-09Grant· $5M
2022-08-09Post Ipo Equity· $65M
Recent News
Brazil Journal
2026-07-11
2026-07-11
Company data provided by crunchbase