Bright Vision Technologies · 1 month ago
GPU Software Engineer (CUDA)
United States
Full-time
Remote
Senior Level
$100K/yr - $150K/yr
6+ years exp
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. They are seeking a GPU Software Engineer (CUDA) to design and optimize workloads on GPU platforms, focusing on performance improvements and collaborating with cross-functional teams.
Artificial Intelligence (AI)Cyber SecurityInformation TechnologySoftware
No H1B
Responsibilities
Design and optimize compute-intensive workloads on modern accelerator hardware
Extract maximum performance from GPU platforms for AI training, inference, scientific computing, and high-throughput data processing workloads
Translate ambiguous requirements into well-engineered solutions
Raise the bar through code review, design review, and mentorship of more junior engineers
Qualification
CUDA C/C++GPU programmingGPU architecturesGPU memory hierarchiesGPU execution modelsGPU workload profilingGPU workload optimizationNCCLMPIHigh-performance interconnect technologiesC++Systems programmingLinear algebraNumerical methodsML framework kernel integration
Required
Bachelor's or Master's degree in Computer Science, Computer Engineering, or a related field
Six or more years of experience in GPU programming and performance engineering
Deep expertise in CUDA C/C++ and GPU programming models
Strong understanding of modern GPU architectures, memory hierarchies, and execution models
Hands-on experience profiling and optimizing GPU workloads in production
Familiarity with NCCL, MPI, and high-performance interconnect technologies
Experience integrating custom kernels into ML frameworks
Strong C++ skills and familiarity with modern systems programming practices
Solid grounding in linear algebra and numerical methods
Strong communication and collaboration skills with research and engineering teams
Preferred
Experience with Triton, CUTLASS, or other GPU kernel authoring frameworks
Familiarity with TensorRT, FasterTransformer, or vLLM internals
Exposure to compiler infrastructure such as LLVM or MLIR
Open-source contributions to GPU or ML performance libraries
Experience with large-scale distributed training infrastructure
Company
Funding
Current Stage
Growth StageCompany data provided by crunchbase