Software DevOps Engineer Gen Software and ML software
Job Summary
The ePlane Company is at the forefront of Indias urban air mobility revolution. Incubated at IIT Madras we are a deep-tech startup dedicated to designing and building the worlds most compact electric flying taxi. Our mission is to make door-to-door flying a reality drastically reducing commute times and decongesting our cities for a cleaner greener future. Were a passionate team of engineers designers and visionaries working on cutting-edge technology and were looking for brilliant minds to help us take flight.
We are looking for a person with a deep understanding of the tooling and lifecycle of ML based systems algorithms including taking them from prototype to production. The selected candidate will own the integrated DevOps and MLOps infrastructure. You will bridge the gap between software engineering IT infrastructure and AI-driven model deployment. This also involves work that makes our entire CI/CD pipeline from general software services to safety-critical AI/ML models robust scalable compliant and qualified for regulated engineering processes. This includes instrumenting and hardening live capabilities and the strategic work of designing a path toward on-premise deployment and formal AI qualification.
Build and maintain data pipelines for model training validation and continuous retraining
Instrument monitor and manage the operational health of the AI/ML capabilities in production including performance drift latency and data quality.
Build CI/CD pipelines model versioning rollback procedures and A/B testing infrastructure
Manage the container orchestration layer and server-side configurations to ensure high availability for internal tools and product operations suites
Lead the feasibility assessment and the implementation of on-premise AI deployment
Own the qualification and governance documentation process across data provenance model architecture explainability and human oversight procedures robustness testing with applicable regulatory frameworks
Establish and operate a recurring governance review process across all deployed capabilities
4 years MLOps or ML infrastructure engineering with production systems experience
CI/CD design and implementation for ML systems (MLflow DVC Weights & Biases or equivalent)
Model monitoring: drift detection performance degradation alerting data quality checks
Deep expertise in Kubernetes cluster management service mesh and cloud/on-prem hybrid infrastructure
Proven experience in setting up CI/CD pipelines for both non-ML software and ML models
Containerised ML deployment (Docker Kubernetes)
Experience deploying ML systems in Safety-critical or regulated domain background where AI output quality must be explainable
Familiarity with change management processes in regulatory environments
RAG system infrastructure experience at scale
Cloud compute job queue management for computationally intensive workloads
On-premise ML/AI deployment experience: quantisation inference optimisation GPU cluster management
Required Skills:
Required Qualifications 4 years MLOps or ML infrastructure engineering with production systems experience CI/CD design and implementation for ML systems (MLflow DVC Weights & Biases or equivalent) Model monitoring: drift detection performance degradation alerting data quality checks Deep expertise in Kubernetes cluster management service mesh and cloud/on-prem hybrid infrastructure Proven experience in setting up CI/CD pipelines for both non-ML software and ML models Containerised ML deployment (Docker Kubernetes) Preferred Qualifications Experience deploying ML systems in Safety-critical or regulated domain background where AI output quality must be explainable Familiarity with change management processes in regulatory environments RAG system infrastructure experience at scale Cloud compute job queue management for computationally intensive workloads On-premise ML/AI deployment experience: quantisation inference optimisation GPU cluster management