LiveFull-timeApply by 1 Nov 2026
Infrastructure Engineer for ML Models
Pune, Maharashtra, India
- Experience
- 4–7 years
- Employment
- Full-time
- Work mode
- Onsite
- Salary
- ₹12 L–20 L / year
- Deadline
- Apply by 1 Nov 2026
- Posted
- 2025-02-11
Required skills
| Skill | Experience | Level |
|---|---|---|
| Machine Learning | 4+ years | Not specified |
| ML Model | 3+ years | Intermediate |
| CI/CD | 2+ years | Not specified |
| Kubernetes | 1+ years | Not specified |
About the role
We are looking for a highly skilled Infrastructure Engineer to design and implement scalable and efficient infrastructure for machine learning models and AI applications.
Key Responsibilities
- Design and implement highly scalable infrastructure for machine learning models and AI applications.
- Develop and maintain infrastructure-as-code (IaC) using Terraform and configuration management tools such as Ansible and Puppet for automated provisioning and orchestration.
- Develop and maintain CI/CD pipelines to ensure seamless deployment and monitoring of machine learning models.
- Collaborate with data scientists and software engineers to integrate machine learning solutions into production environments.
- Automate routine tasks and workflows to improve efficiency and reduce manual intervention.
- Monitor and troubleshoot production systems to ensure high availability and performance.
- Implement security best practices to protect sensitive data and ensure compliance with industry standards.
- Identify and resolve infrastructure gaps to ensure reliable, efficient, and scalable solutions.
- Develop advanced AI/ML infrastructure solutions to enhance the efficiency of skilled ML teams.
- Deploy and manage open-source GenAI components, such as vector databases and various AI/ML models, ensuring seamless integration and optimal performance within the Kubernetes environment.
Requirements
- 4 to 7 years of hands-on experience as Infrastructure Engineer with a strong emphasis on Kubernetes knowledge.
- Experience in designing and implementing the scalable and efficient infrastructure with Kubernetes based automation.
- Experience in developing and maintaining CI/CD pipelines to ensure seamless deployment and monitoring of machine learning models.
- Experience to automate routine tasks and workflows to improve efficiency and reduce manual intervention.
- Experience to monitor and troubleshoot production systems to ensure high availability and performance.
- Experience to implement security best practices to protect sensitive data and ensure compliance with industry standards.
- Design and implement solutions for critical areas, including Distributed storage systems, Scheduling systems, High availability capabilities, Core reliability issues within large-scale GPU clusters.
- Knowledge of infrastructure as code (IaC) tools (e.g., Terraform, Ansible).
- Experience to implement and manage GPU infrastructure within Kubernetes clusters to support high-performance computing and AI/ML tasks, ensuring scalability and efficiency.