Bitdeer (NASDAQ: BTDR) · 1 week ago
AI Cloud Senior DevOps Engineer
United States
Full-time
Remote
Senior Level
5+ years exp
Bitdeer is a technology company focused on AI and Bitcoin mining infrastructure, offering cloud capabilities for artificial intelligence workloads. The Cloud Senior DevOps Engineer will support the AI Cloud team by automating deployment and infrastructure operations, managing cloud-native and AI infrastructure, and establishing reliable MLOps and DevOps practices. The role also leads high-availability architecture, observability, security, compliance, and incident resolution.
Artificial Intelligence (AI)Crypto & Web3Cloud ComputingInformation TechnologyAI InfrastructureBitcoinBlockchainCryptocurrency
Responsibilities
Design, implement, and maintain end-to-end CI/CD pipelines for both software applications and machine learning models. Automate build, test, deployment, and rollback processes to ensure seamless transitions from innovation to production
Build, optimize, and scale cloud-native infrastructure using Kubernetes (K8s) and Docker. Manage and provision specialized computing resources (e.g., GPU clusters) to support high-performance AI workloads and model inferencing
Take ownership of high-availability design in production environments. Implement disaster recovery (DR) strategies, self-healing mechanisms, capacity planning, and performance tuning to meet stringent business SLAs
Champion IaC practices utilizing tools such as Terraform, Ansible, and Helm to achieve fully automated, reproducible, and auditable infrastructure provisioning across multiple cloud environments
Architect and refine comprehensive monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, ELK/EFK stack) to provide deep visibility into system health, application performance, and AI model metrics
Work closely with R&D, Data Science, Security, and Business teams to streamline workflows, eliminate bottlenecks, and continuously elevate engineering efficiency through Internal Developer Platforms (IDP) and Platform Engineering initiatives
Establish and enforce robust system stability and security standards. Manage release workflows, implement Zero Trust access controls, oversee secrets management, and ensure compliance with industry frameworks (e.g., SOC2, ISO27001)
Act as the technical lead during complex system anomalies and major incidents. Spearhead rapid troubleshooting, conduct thorough root cause analysis (RCA), and implement preventative remediation plans
Qualification
LinuxTCP/IPDNSHTTPLoad BalancingVirtual Private Cloud (VPC)DockerKubernetesAWSGoogle Cloud Platform (GCP)Microsoft AzureAlibaba CloudGoPythonShell ScriptingCI/CDInfrastructure as Code (IaC)Site Reliability Engineering (SRE)MLOpsVLLMText Generation Inference (TGI)Triton Inference ServerGPU Cluster ManagementDistributed SystemsInternal Developer Platforms (IDP)Zero Trust ArchitectureDevSecOpsSOC 2ISO 27001Cross-team Communication
Required
Experience & Education: Bachelor's degree or above in Computer Science, Engineering, or a related technical field, with 5+ years of hands-on experience in DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure roles
Networking & OS: Expert-level knowledge of Linux operating systems and core networking principles (TCP/IP, DNS, HTTP, Load Balancing, VPCs)
Containerization & Orchestration: Deep mastery of Docker and Kubernetes orchestration, including a thorough understanding of underlying principles, cluster management, and production-level best practices
Cloud Platforms: Proven proficiency in designing and managing infrastructure on major Public or Hybrid Cloud platforms (e.g., AWS, GCP, Azure, Alibaba Cloud), including multi-cloud and hybrid-cloud strategies
Programming Skills: Strong coding and scripting capabilities in at least one major language (Go, Python, Shell, etc.) with a solid engineering-oriented mindset focused on automation and tooling development
Domain Knowledge: Systematic and practical understanding of CI/CD methodologies, Infrastructure as Code (IaC), Observability paradigms, and Site Reliability Engineering (SRE) principles
Soft Skills: Exceptional problem-solving abilities, sharp technical judgment, and excellent cross-team communication skills to effectively collaborate in a fast-paced, dynamic environment
Preferred
AI/ML Infrastructure Experience: Familiarity with MLOps practices, model serving/inferencing frameworks (e.g., vLLM, TGI, Triton Inference Server), and experience managing GPU clusters for AI/ML workloads
Large-Scale Systems: Proven track record working with large-scale distributed systems or high-concurrency environments (e.g., Fintech, Trading, Real-time processing, or AI platforms)
Platform Engineering: Hands-on experience in designing and building Internal Developer Platforms (IDP) to enhance developer autonomy and productivity
Advanced Security: Deep familiarity with Zero Trust architecture, automated security testing (DevSecOps), and implementing strict compliance frameworks (e.g., SOC2, ISO27001)
Leadership: Prior experience acting as a Technical Lead, mentoring junior engineers, or managing DevOps teams
Benefits
Full-Time employment
Remote (within locations)
Company
Funding
Current Stage
Public CompanyTotal Funding
$2.07BKey Investors
BITTetherGenimous
2026-02-20Post Ipo Debt· $325M
2025-11-13Post Ipo Debt· $400M
2025-06-17Post Ipo Debt· $330M
Recent News
2026-09-03
Morningstar.com
2026-08-25
2026-08-21
Company data provided by crunchbase