AI Engineer
Job Summary
Design develop and maintain GenAI-powered applications such as chatbots customer support assistants recommendation assistants and document/query intelligence systems.
Build and optimize RAG pipelines including document ingestion chunking embeddings vector search retrieval logic reranking and response generation.
Work on LLM orchestration using frameworks such as LangChain LangGraph LlamaIndex or similar tools.
Develop backend AI services using Python FastAPI REST APIs and microservice-based architecture.
Implement intent classification entity extraction query parsing and structured JSON output generation for business workflows.
Work with open-source and cloud-based LLMs including local model setup model serving and inference optimization.
Implement guardrails to prevent hallucination unsafe responses competitor comparisons irrelevant answers prompt injection and data leakage.
Optimize AI systems for latency accuracy reliability token usage and scalability.
Integrate AI workflows with backend systems databases APIs CRM/ERP systems payment flows and business applications.
Create evaluation datasets regression test cases accuracy reports and failure analysis for AI responses.
Collaborate with product managers backend developers DevOps teams and business stakeholders to convert business requirements into scalable AI solutions.
Strong hands-on experience with Python and backend development.
Practical experience in building LLM-based applications using OpenAI Claude Gemini Llama Mistral or other open-source models.
Good understanding of RAG architecture embeddings vector databases semantic search hybrid search and retrieval optimization.
Experience with frameworks such as LangChain LangGraph LlamaIndex or similar orchestration tools.
Knowledge of vector databases such as FAISS ChromaDB Pinecone Weaviate Qdrant Milvus or pgvector.
Experience with prompt engineering system prompts structured output generation function calling/tool calling and JSON schema-based responses.
Understanding of LLM guardrails safety filters fallback handling confidence scoring and hallucination control.
Experience with FastAPI / Flask / Django REST APIs and backend integration.
Understanding of Docker Git Linux basics and deployment workflows.
Ability to debug production AI issues related to latency incorrect responses token limits retrieval failure context mismatch and model output inconsistency.
Experience with vLLM Ollama Hugging Face Transformers TensorRT-LLM or other model serving frameworks.
Knowledge of model quantization inference optimization batching GPU utilization and token streaming.
Experience with AWS Azure or GCP for AI/ML deployment.
Exposure to telecom fintech customer support billing recharge payments or high-scale consumer applications.
Experience in multilingual AI systems especially Hinglish or Indian language handling.
Understanding of ASR/transcription-based input normalization will be a plus.
Experience with monitoring tools logging prompt/version management and AI evaluation frameworks.
Knowledge of MCP agentic workflows tool-based reasoning or multi-step AI orchestration will be an advantage.
Experience: 2 to 5 years
Education: BE/BTech/MTech/MCA in Computer Science AI/ML Data Science IT or equivalent practical experience.
The candidate should have built at least one real-world AI/GenAI application involving LLMs backend integration RAG or production deployment.
The candidate should be able to explain implementation details clearly not just theoretical AI concepts.