We are seeking a highly skilled MLOps Engineer to join our dynamic team of 13 experts for a fully remote mission. In this role, you will play a pivotal part in streamlining our AI infrastructure for advanced scientific research. Your primary focus will be to standardize, scale, and optimize the deployment of machine learning models from our newly consolidated model registry to various internal production environments.
Here's what we offer
- Attractive salary and long-term job security through affiliation with a large corporation.
- Up to 30 days of vacation per year
- Company pension scheme contribution after the end of the probationary period
- Extensive social benefits, including Christmas and holiday pay
- Reimbursement of travel expenses
- Usually an open-ended employment contract
- Good opportunities for acquisition with our business partners
- Tailored professional development opportunities and free language courses
- A wide range of employee benefits
Your tasks
Primary Focus : Standardize and enable the deployment of machine learning models from the newly consolidated model registry to a variety of internal deployment targets and frameworks in a secure and scalable way
Deployment Pipeline Development: Design, implement, and document robust, repeatable pipelines for deploying models stored in the consolidated model registry.
Deployment Target Integration: Establish and ensure connectivity and deployment capabilities to various target environments, including:
- Orbital Pipelines
- Kubeflow/KServe
Framework Adaptation: Develop and maintain necessary wrappers, SDK extensions, and configuration templates to ensure models (potentially in various formats/frameworks like TensorFlow, PyTorch, Scikit-learn) can be correctly packaged and served by the target deployment frameworks/environments.
Standardization: Collaborate with the core AIE MLOps team and other stakeholders in the organization to define best practices, tooling, and standards for model serving configurations, monitoring hooks, and endpoint health checks.
Documentation & Knowledge Transfer: Create comprehensive technical documentation, runbooks, and provide training/knowledge transfer sessions to the internal MLOps team for long-term ownership and maintenance of the deployed solutions.
Troubleshooting: Address and resolve technical challenges related to model packaging, deployment failures, latency issues, and framework compatibility across diverse environments.
Your profile
Must-haves:
- Degree: Bachelor or Master
- Location: can be completely remote
- Languages: only English; no German required
- At least 8 years of relevant experience
- Hands-on experience working with model registry (preferably the AI developer platform “Weights and Biases”)
- Advanced proficiency in Python. Familiarity with shell scripting
- Expert proficiency in Docker and container orchestration, preferably Kubernetes
- Proven ability to implement and optimize high-availability model serving solutions (REST APIs, gRPC).
Nice to Haves:
- Deep hands-on experience with: Kubeflow, KServe, Dagster.
- Strong experience with CI/CD tools (eg, Jenkins, GitLab CI, ArgoCD) and pipeline automation.
- Experience working within cloud environments (preferably AWS) and infrastructure-as-code principles (eg, Terraform) is a plus.