Health Business Solutions · 1 month ago
LLM Operations Engineer
Health Business Solutions is seeking an LLM Operations Engineer to build, automate, and scale machine learning delivery pipelines on the Lakehouse. The role involves managing the entire model lifecycle, collaborating closely with leadership and data engineers to ensure reliable, auditable, and secure ML systems using Databricks and modern DevOps practices.
Responsibilities
Design and maintain Databricks workspaces, clusters, SQL Warehouses, cluster policies, and workspace governance (RBAC, SCIM, SSO, secret scopes)
Implement robust data pipelines with Delta Lake (ACID tables, Z‑ordering, OPTIMIZE/VACUUM), Delta Live Tables (DAGs, expectations), and Workflows (jobs, task orchestration)
Set up Unity Catalog for cross-workspace governance: data & model lineage, permissions, catalogs/schemas, data tags, and auditability
Operationalize ML models using MLflow (tracking, artifacts, metrics, model registry, approvals, stages: Staging/Production)
Build/maintain Feature Store entities and feature pipelines; enforce reproducibility and feature governance
Establish model deployment patterns (batch scoring, streaming, microservices) using Model Serving
Create scalable CI/CD for notebooks, repos, and jobs using Azure DevOps, including unit/integration tests, data/feature validation, and registry promotions
Implement data quality and ML quality controls (e.g., Great Expectations/Delta expectations, statistical tests, drift detection, canary releases)
Build robust monitoring & alerting for data freshness, pipeline SLAs, model performance, drift, and operational metrics
Optimize performance and cost (autoscaling, spot instances, DBR runtimes, caching, storage tiers)
Enforce compliance and security best practices (PII handling, encryption at rest/in transit, network controls, secret management)
Partner with data engineers and subject matter experts to standardize templates for experiments, pipelines, model packaging, and deployment
Document patterns and build internal tooling (CLI utilities, Python packages) to streamline model release and observability
Contribute to incident response, post‑mortems, and continuous improvements
Design, deploy, and operate LLMOps pipelines for Retrieval‑Augmented Generation (RAG), including document ingestion, embedding generation, vector storage, retrieval strategies, prompt/version management, and evaluation, using Databricks (Delta Lake, MLflow, Model Serving) to ensure secure, auditable, and production‑grade GenAI systems
Qualification
Required
BS/MS in Computer Science, Engineering, Data Science, or equivalent practical experience
3+ years of MLOps/ML Engineering/Platform Engineering experience in Databricks
Hands‑on expertise with Databricks: Delta Lake, Unity Catalog, MLflow (Tracking/Registry), Feature Store, Workflows/Jobs, Repos, and Model Serving
Strong Python engineering skills (packaging, testing, virtual environments); familiarity with Spark (PySpark) and SQL
Experience with CI/CD (GitHub Actions/Azure DevOps/GitLab), artifact registries, and environment management
Solid understanding of data/machine learning pipeline design (batch/streaming), data quality checks, and ML evaluation/monitoring
Excellent communication and organizational abilities
Ability to work independently and as a part of cross-functional teams
Comfortable operating in a fast-paced, changing environment
Strong analytical and problem-solving skills, with the ability to interpret data and drive recommendations