Deep Tech | MLOps | Model Deployment | Model Serving | ML Infrastructure
Our client, a growing Deep Tech / AI company based in Munich, is looking for an ML Deployment Engineer to build the deployment layer that takes machine-learning models from experimentation into scalable, reliable production services.
You'll work at the intersection of ML Engineering, MLOps, and Platform Engineering , creating the tooling and infrastructure that makes model deployment repeatable, observable, and production-ready.
What You'll Work On
Build production deployment pipelines for machine-learning models
Deploy and operate model-serving workloads on Kubernetes
Build scalable inference services using KServe
Containerise ML workloads using Docker
Develop deployment tooling and automation in Python
Manage model versions, artefacts, and deployment workflows with MLflow
Build CI/CD pipelines for testing and releasing ML services
Deploy workloads across AWS and/or GCP environments
Implement rollout, rollback, and model versioning strategies
Improve deployment reliability, scalability, and observability
Automate the path from approved model to production endpoint
Collaborate with ML Engineers to productionise new models without requiring them to manage the underlying infrastructure
Core Skills
3+ years in MLOps, ML Engineering, ML Infrastructure, Platform Engineering, or similar roles
KServe or comparable model-serving technology
AWS and/or GCP
Strong understanding of production ML systems
Nice to Have
NVIDIA Triton Inference Server
Ray Serve
PyTorch / TensorFlow
Terraform
Canary or blue-green deployments
GPU-enabled inference workloads
Model monitoring and drift detection
Experience operating real-time inference APIs
#J-18808-Ljbffr