Principal Engineer
Job Summary
About the AI Technology Innovation Centre
We are establishing a world class AI research center called AI Technology Innovation Center (TIC) in Bengaluru. The TIC will drive research in frontier AI technology with top AI talent and state-of-the-art infrastructure. The research breakthrough will lead to commercialization of next generation AI solutions. The AI TIC is being strategically designed to establish a globally recognized leader in AI research and development driving innovation and setting industry benchmarks. The TIC aims to attract nurture and retain top-tier AI talent in India by fostering a dynamic inclusive and high-impact work environment that encourages continuous learning and breakthrough thinking.
By seamlessly integrating commercial agility with a pioneering scientific mindset we will deliver scalable AI solutions that address real-world challenges while pushing the frontiers of technology. This ensures that our AI-TIC excels in cutting-edge research and translates discoveries into tangible business value reinforcing our position as a hub for AI excellence.
The Role Mandate
As the Principal Engineer of AI TIC youll own the end-to-end system architecture and design and delivery of production-grade agentic and Generative AI systems (on-premise/cloud). This is a highly hands-on role requiring deep architectural insight coding proficiency and an obsession with performance scalability and reliability. Youll architect secure cost-efficient AI platforms guide developers through complex debugging and optimization and ensure all systems are observable governed and production ready. You will create a reference platform and AI pipeline on which TIC research will get translated to real product.
Why This Role - What You Gain
- Founding-stage platform ownership
Define architecture standards and SDKs that every Centre team will build on.
- Direct link to global products
Carry frontier research into dependable observable and secure deployed systems.
- The hardest form of the problem
Engineer memory emotion control and self-learning under real compute and reliability constraints.
- Technical leadership without leaving code
Mentor review debug and ship as a true principal engineer.
The Larger Impact
- Determine whether empathic agents become dependable products rather than demonstrations.
- Create a secure modular and cost-efficient platform for multimodal memory-enabled tool-using agents.
- Lead model adaptation compression and optimisation for self-managed GPU environments.
- Create reusable SDKs connectors CI/CD templates evaluation harnesses and observability.
- Set the engineering standards that become the Centres technical reputation.
Key Responsibilities:
- Architect Production AI Systems: Design robust overall architectures for agentic systems (planning reasoning tool-calling) GenAI/RAG pipelines and evaluation workflows. Create detailed design documents including flow/UML/sequence diagrams and AWS deployment topologies. Additionally ensure architectures support advanced LLM training and inference workflows incorporating distributed strategies for scalability.
- Optimize for Cost & Performance: Model throughput latency concurrency autoscaling CPU/GPU sizing and vector index performance to ensure scalable efficient deployments. Include optimization for multi-node GPU clusters and distributed training efficiency to reduce compute overhead.
- Lead Debugging & Stability Efforts: Conduct deep-dive debugging fix critical defects and resolve production incidents; pair-program with developers to improve code quality and performance. Apply MLOps-driven stability practices leveraging configuration management and automated recovery for high availability.
- Standardize Agentic Frameworks: Build reference implementations using Semantic Kernel (preferred) LangGraph AutoGen or CrewAI with strong schema validation grounding and memory management.
- Implement Observability & Monitoring: Set up distributed tracing metrics and logging via OpenTelemetry and Datadog. Standardize dashboards alerts and incident response workflows.
- Govern Evaluation & Rollouts: Build test and evaluation frameworksgolden sets A/B experiments regression suites and controlled rolloutsto ensure consistent quality across releases.
- Establish Engineering Standards: Create reusable SDKs connectors CI/CD templates and architecture review checklists to promote consistency across teams.
- Cross-Functional Leadership: Collaborate with product data and SRE teams for capacity planning DR strategies and post-incident RCA reviews. Mentor engineers to strengthen design and reliability practices
How the Centre Works
The TIC follows a research-to-product operating model. Scientists and engineers work as one team across hypothesis formation algorithm design model development evaluation platform engineering and technology transfer. Every role combines technical depth with ownership reproducibility responsible AI and clear communication of outcomes.
Qualifications and Experience
Experience and leadership
- Education: Bachelors/Masters from a top-tier institute (IIT/Tier-1) in Computer Science AI or related field.
- 710 years in software/AI engineering including 4 years in GenAI application development and 2 years architecting agentic AI systems.
- Expert in Python 3.11 (asyncio typing packaging profiling pytest).
- Hands-on experience with Semantic Kernel LangGraph AutoGen or CrewAI or equivalent.
- Proven delivery of GenAI/RAG systems on AWS Bedrock or equivalent vector-based platforms (OpenSearch Serverless Pinecone Redis).
- Deep understanding of AWS ecosystem: EKS Bedrock S3 SQS/SNS RDS ElastiCache Secrets Manager IAM/Okta Kong API Gateway management and automated recovery for high availability.
What Success Looks Like (18-24 Months)
- A hardened observable agentic system and platform having custom models on on-premise infrastructure.
- Reference implementations SDKs and CI/CD templates become the Centre default.
- Quality safety latency and cost are measurable and continuously improved.
- Research lands in production safely through a regular governed release cadence.
- Production problems are detected early and converted into systemic improvements.
Required Experience:
Staff IC
About Company
As a global leader, Wipro blends consulting and AI expertise across design, engineering and operations to accelerate business transformation and deliver future-ready technology.