SIGN IN
Senior Deep Learning Frameworks CUDA Software Engineer jobs in United States
info-icon
This job has closed.
company-logo

NVIDIA · 2 weeks ago

Senior Deep Learning Frameworks CUDA Software Engineer

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. They are seeking a motivated Deep Learning engineer to integrate advanced CUDA features and Distributed Runtime technologies into AI stacks, while collaborating with teams to enhance performance and programmability of AI applications.
Artificial Intelligence (AI)TransportationCloud ComputingEmbedded SoftwareGamingHardwareQuantum ComputingSemiconductorAI InfrastructureAutonomous VehiclesEmbedded SystemsFoundational AI
check
Growth Opportunities
check
H1B Sponsor Likelynote

Responsibilities

Integrate new CUDA features and Runtime abstractions in AI frameworks: from PoC to performance analysis to production
Perform deep analysis of AI workloads and frameworks to identify requirements and opportunities to innovate in the lower layers of the stack. Collaborate hands-on with teams working on the latest AI models
Own and drive improvements in the AI Compiler-Runtime interface to build speed-of-light multi-GPU multi-node solutions
Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads
Influence the roadmap of core CUDA to facilitate building next-gen DL frameworks
Collaborate with a very dynamic team across multiple time zones
Collaborate closely with AI researchers, HW and SW architects, kernel and compiler authors and CUDA driver experts to co-design systems and frameworks that enhance performance and programmability
Develop exploratory tools and runtime systems to profile and accelerate new paradigms in deep learning
Write clean, effective, and maintainable code, ensuring exploratory prototypes can smoothly transition into open-source releases, upstream framework integrations, internal tools, or closed-source commercial products

Qualification

CUDADeep Learning Frameworks - PyTorchDeep Learning Frameworks - JAXInference Engines - TRT-LLMInference Engines - vLLMInference Engines - SGLangPythonC++AI ModelsParallelismCompiler Technologies - torch.compilePerformance BenchmarkingPyTorch ProfilerNVIDIA Nsight SystemsHPC/AI CommunicationComputer System ArchitectureOperating Systems PrinciplesNCCLMPIUCXDistributed Machine Learning TechniquesPipeline ParallelismTensor ParallelismKernel Authoring - TritonKernel Authoring - cuTeDeep Learning Compilers - XLADeep Learning Compilers - torch compileDistributed Runtime Programming

Required

BS, MS, or PhD degree in Computer Science, Computer Engineering, Electrical Engineering, or related field (or equivalent experience)
8+ years of relevant industry experience or equivalent academic experience after completed degree
Development experience with Deep Learning Frameworks such PyTorch, JAX, and Inference Engines such as TRT-LLM, vLLM, SGLang
Rapid prototyping and development with Python, C++, CUDA or related DSLs
Solid grasp of AI models, parallelisms, and/or compiler technologies (e.g. torch.compile)
Experience conducting performance benchmarking on AI clusters. Familiarity with at least one performance profiler toolchain (PyTorch profiler, NVIDIA Nsight Systems)
Understanding of HPC/AI communication concepts
Good understanding of computer system architecture, HW-SW interactions and operating systems principles (aka systems software fundamentals)
Adaptability and passion to learn new frameworks and tools
Flexibility to work and communicate effectively across different teams and timezones

Preferred

Deep expertise in the performance internals and execution graphs of major deep learning autograd, training and inference frameworks (e.g., PyTorch, JAX, TensorRT, vLLM, sgLang, Nemo, Megatron, MaxText, etc.)
Hands-on experience with CUDA, specific communication libraries (e.g., NCCL, MPI, UCX) and distributed machine learning techniques (e.g., pipeline parallelism, tensor parallelism)
Expertise in one or more of these areas: Training, Distributed inference, MoE, Reinforcement Learning, kernel authoring (on CUDA, Triton, cuTe, etc)
Background in deep learning compilers, both graph-level and codegen (e.g., Triton, XLA, torch compile)
Experience with programming for compute & communication overlap in distributed runtime

Benefits

You will also be eligible for equity and [benefits](https://www.nvidia.com/en-us/benefits/).

Company

NVIDIA is a computing platform company operating at the intersection of graphics, HPC, and AI.

H1B Sponsorship

NVIDIA has a track record of offering H1B sponsorships. Please note that this does not guarantee sponsorship for this specific role. Below presents additional info for your reference. (Data Powered by US Department of Labor)
Distribution of Different Job Fields Receiving Sponsorship
89%
Represents job field similar to this job
Engineering and Development
Management and Executive
Product Management
Marketing
Sales
Accounting and Finance
Customer Service and Support
Creatives and Design
Human Resources
Arts and Entertainment
Legal and Compliance
Trends of Total Sponsorships
*2026 (1247)
2025 (1868)
2024 (1353)
2023 (976)
2022 (835)
2021 (601)
2020 (529)

Funding

Current Stage
Public Company
Total Funding
$29.09B
Key Investors
ARPA-EARK Investment ManagementSoftBank Vision Fund
2026-06-15Post Ipo Debt· $25B
2023-05-09Grant· $5M
2022-08-09Post Ipo Equity· $65M

Leadership Team

leader-logo
Jensen Huang
Founder and CEO
linkedin
leader-logo
Michael Kagan
Chief Technology Officer
linkedin
Company data provided by crunchbase