SIGN IN
AI Inference Engineer jobs in United States
cer-icon
Apply on Employer Site
company-logo

Fuse Energy · 2 weeks ago

AI Inference Engineer

Fuse Energy is a forward-thinking renewable energy startup on a mission to deliver a terawatt of renewable energy - fast. They are seeking a Founding AI Inference Engineer to define and build how Fuse serves AI inference workloads at scale, focusing on architecture and performance optimization.
IndustrialEnergyRenewable EnergyElectrical DistributionEnergy Management
check
H1B Sponsor Likelynote

Responsibilities

Define Fuse's inference serving strategy and architecture from first principles
Design and build the serving stack: request routing, batching, scheduling, and autoscaling for high-throughput, latency-sensitive inference workloads
Own model-level optimisation strategy for serving - deciding where and how to apply quantisation, distillation, speculative decoding, and similar techniques to improve throughput and cost per token, partnering with the CUDA/GPU engineers
Make the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or equivalents)
Translate throughput, latency, and uptime commitments into concrete technical specifications and serving capacity plans
Act as a direct technical owner of inference performance and reliability
Work closely with the CUDA and GPU engineering teams to ensure custom kernels and hardware performance work are integrated cleanly into the serving layer
Set the standards, tooling, and benchmarks this function will run on as it grows

Qualification

Inference serving frameworksBatchingKV-cache managementQuantisationSpeculative decodingGPU/CUDA integrationServing system architectureTriton Inference ServerAutoscalingCapacity planningMulti-tenant servingSLA-driven infrastructureKubernetesSlurm

Required

4+ years of experience building or operating large-scale inference serving systems, or equivalent strong project/industry experience
Deep, hands-on experience with inference serving frameworks and the techniques used to optimise them (batching, KV-cache management, quantisation, speculative decoding)
Strong systems thinking - able to reason about the full path from incoming request to served response across a large cluster
Comfortable working directly with GPU/CUDA engineers to integrate low-level performance work into a serving system
A track record of making high-stakes architecture calls and owning the outcome
Comfort operating without a playbook - this is a founding role shaping a new function around architecture that's still early-stage, not joining an established one

Preferred

Experience with Triton or custom ML inference/training frameworks
Experience with autoscaling or capacity planning for large-scale inference workloads
Exposure to multi-tenant serving or SLA-driven infrastructure
Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system
Familiarity with Kubernetes/Slurm for cluster orchestration
Interest or experience in energy markets, grid systems, or sustainability-focused compute

Benefits

Competitive salary and an equity sign-on bonus
Biannual bonus scheme
Fully expensed tech to match your needs
Breakfast and dinner allowance for office based employees

Company

Fuse Energy

twitterlinkedincrunchbase
company-logo
Fuse Energy is an innovative energy company that simplifies gas and electricity management by providing accurate bill forecasts.

H1B Sponsorship

Fuse Energy has a track record of offering H1B sponsorships. Please note that this does not guarantee sponsorship for this specific role. Below presents additional info for your reference. (Data Powered by US Department of Labor)
Distribution of Different Job Fields Receiving Sponsorship
Represents job field similar to this job
Product Management
Trends of Total Sponsorships
*2024 (1)

Funding

Current Stage
Late Stage
Total Funding
$199.86M
Key Investors
Future FiftyBalderton Capital,Lowercarbon CapitalMulticoin Capital
2026-06-04Series B· $30M
2026-03-19Non Equity Assistance
2025-12-18Series B· $69.18M

Leadership Team

leader-logo
Alan Chang
Founder & CEO
linkedin
leader-logo
Charles Orr
Founder & COO
linkedin
Company data provided by crunchbase