Fuse Energy · 2 weeks ago
AI Inference Engineer
Fuse Energy is a forward-thinking renewable energy startup on a mission to deliver a terawatt of renewable energy - fast. They are seeking a Founding AI Inference Engineer to define and build how Fuse serves AI inference workloads at scale, focusing on architecture and performance optimization.
IndustrialEnergyRenewable EnergyElectrical DistributionEnergy Management
Responsibilities
Define Fuse's inference serving strategy and architecture from first principles
Design and build the serving stack: request routing, batching, scheduling, and autoscaling for high-throughput, latency-sensitive inference workloads
Own model-level optimisation strategy for serving - deciding where and how to apply quantisation, distillation, speculative decoding, and similar techniques to improve throughput and cost per token, partnering with the CUDA/GPU engineers
Make the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or equivalents)
Translate throughput, latency, and uptime commitments into concrete technical specifications and serving capacity plans
Act as a direct technical owner of inference performance and reliability
Work closely with the CUDA and GPU engineering teams to ensure custom kernels and hardware performance work are integrated cleanly into the serving layer
Set the standards, tooling, and benchmarks this function will run on as it grows
Qualification
Required
4+ years of experience building or operating large-scale inference serving systems, or equivalent strong project/industry experience
Deep, hands-on experience with inference serving frameworks and the techniques used to optimise them (batching, KV-cache management, quantisation, speculative decoding)
Strong systems thinking - able to reason about the full path from incoming request to served response across a large cluster
Comfortable working directly with GPU/CUDA engineers to integrate low-level performance work into a serving system
A track record of making high-stakes architecture calls and owning the outcome
Comfort operating without a playbook - this is a founding role shaping a new function around architecture that's still early-stage, not joining an established one
Preferred
Experience with Triton or custom ML inference/training frameworks
Experience with autoscaling or capacity planning for large-scale inference workloads
Exposure to multi-tenant serving or SLA-driven infrastructure
Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system
Familiarity with Kubernetes/Slurm for cluster orchestration
Interest or experience in energy markets, grid systems, or sustainability-focused compute
Benefits
Competitive salary and an equity sign-on bonus
Biannual bonus scheme
Fully expensed tech to match your needs
Breakfast and dinner allowance for office based employees
Company
Fuse Energy
Fuse Energy is an innovative energy company that simplifies gas and electricity management by providing accurate bill forecasts.
H1B Sponsorship
Fuse Energy has a track record of offering H1B sponsorships. Please note that this does not
guarantee sponsorship for this specific role. Below presents additional info for your
reference. (Data Powered by US Department of Labor)
Distribution of Different Job Fields Receiving Sponsorship
Represents job field similar to this job
Product Management
Trends of Total Sponsorships
*2024 (1)
Funding
Current Stage
Late StageTotal Funding
$199.86MKey Investors
Future FiftyBalderton Capital,Lowercarbon CapitalMulticoin Capital
2026-06-04Series B· $30M
2026-03-19Non Equity Assistance
2025-12-18Series B· $69.18M
Recent News
2026-06-16
2026-06-12
Business Live
2026-05-15
Company data provided by crunchbase