Launched in 2007 BookMyShow owned and operated by Big Tree Entertainment Pvt. Ltd. (founded in 1999) is Indias leading entertainment destination with global operations and the one-stop shop for every entertainment need. The firm is present in over 650 towns and cities in India and works with partners across the industry to provide unmatched entertainment experiences to millions of customers.
Over the years the company has evolved from a purely online ticketing platform for movies across 6000 plus screens to end-to-end management of live entertainment events including music concerts live performances theatricals sports and more. Some of the key properties that BookMyShow has brought to its markets include U2s The Joshua Tree Tour NBAs debut games in India Disneys Aladdin Cirque du Soleil BAZZAR as well as international artists such as Coldplay Ed Sheeran and Justin Bieber. BookMyShow is invested in providing the best user experience whether on-ground or online.
The company has developed BookMyShow Stream Indias largest home-grown transactional video-on-demand (TVOD) platform.
Role Overview
Were looking for a Data Enthusiast who combines strong data engineering fundamentals with hands-on experience applying machine learning in production. Youll build and scale data pipelines on Databricks and partner closely with Data Science/Business Intelligence teams to operationalize models - from feature engineering through deployment and monitoring as well as generating insights that create business impact
Your Profile
Design build and maintain scalable ETL/ELT pipelines using Databricks (Spark Delta Lake Lakeflow/DLT Unity Catalog)
Develop and productionize feature pipelines for ML use cases ensuring reliability freshness and reproducibility
Collaborate with Data Scientists/ML Engineers to deploy models (batch and/or real-time) including via Databricks Model Serving or ML flow
Optimize Spark jobs SQL warehouses and cluster configurations for cost and performance
Build and maintain system-level observability for pipelines and ML jobs (usage cost quality drift)
Implement data quality checks testing and monitoring across the medallion architecture (bronze/silver/gold)
Own Unity Catalog governance for datasets and features access controls lineage and PII masking
Partner with platform/infra teams on job orchestration CI/CD for data & ML pipelines and cost optimization
Contribute to architecture decisions around lakehouse design streaming vs. batch tradeoffs and tool selection (build vs. buy)
Perform analysis on top of the data you build answer ad-hoc business questions validate metrics and spot data quality issues before they reach stakeholders
Evaluate and apply LLMs/agentic frameworks responsibly within the team balancing accuracy cost and governance (e.g. row/column-level access control on what an AI agent can query)
Your Checklist
1-3 years of data engineering experience working on Databricks in production
Strong proficiency in PySpark/Spark SQL and Python
Solid understanding of Delta Lake Unity Catalog Lakeflow/DLT and Databricks system tables (billing compute query history) Experience building and maintaining feature pipelines or ML data infrastructure (feature stores training/serving data parity)
Familiarity with MLflow (experiment tracking model registry) and/or Databricks Model Serving
Strong SQL skills and experience with warehouse performance tuning (query optimization materialization strategies cluster sizing)
Understanding of ML fundamentals enough to have real conversations with Data Scientists about features drift and model lifecycle (you dont need to be building models yourself but you should understand what good looks like)
Preferred Skills
Experience with real-time/streaming architectures (Structured Streaming Kafka Lakebase or similar)
Exposure to LLM/agentic tooling (Databricks Genie RAG pipelines vector search)
Experience with cost governance/FinOps for Databricks workloads
Background with experimentation platforms or A/B testing infrastructure
Familiarity with orchestration tools (Databricks Jobs Airflow) and CI/CD for data pipelines
Required Experience:
IC
The World of BookMyShowLaunched in 2007 BookMyShow owned and operated by Big Tree Entertainment Pvt. Ltd. (founded in 1999) is Indias leading entertainment destination with global operations and the one-stop shop for every entertainment need. The firm is present in over 650 towns and cities in India...
The World of BookMyShow
Launched in 2007 BookMyShow owned and operated by Big Tree Entertainment Pvt. Ltd. (founded in 1999) is Indias leading entertainment destination with global operations and the one-stop shop for every entertainment need. The firm is present in over 650 towns and cities in India and works with partners across the industry to provide unmatched entertainment experiences to millions of customers.
Over the years the company has evolved from a purely online ticketing platform for movies across 6000 plus screens to end-to-end management of live entertainment events including music concerts live performances theatricals sports and more. Some of the key properties that BookMyShow has brought to its markets include U2s The Joshua Tree Tour NBAs debut games in India Disneys Aladdin Cirque du Soleil BAZZAR as well as international artists such as Coldplay Ed Sheeran and Justin Bieber. BookMyShow is invested in providing the best user experience whether on-ground or online.
The company has developed BookMyShow Stream Indias largest home-grown transactional video-on-demand (TVOD) platform.
Role Overview
Were looking for a Data Enthusiast who combines strong data engineering fundamentals with hands-on experience applying machine learning in production. Youll build and scale data pipelines on Databricks and partner closely with Data Science/Business Intelligence teams to operationalize models - from feature engineering through deployment and monitoring as well as generating insights that create business impact
Your Profile
Design build and maintain scalable ETL/ELT pipelines using Databricks (Spark Delta Lake Lakeflow/DLT Unity Catalog)
Develop and productionize feature pipelines for ML use cases ensuring reliability freshness and reproducibility
Collaborate with Data Scientists/ML Engineers to deploy models (batch and/or real-time) including via Databricks Model Serving or ML flow
Optimize Spark jobs SQL warehouses and cluster configurations for cost and performance
Build and maintain system-level observability for pipelines and ML jobs (usage cost quality drift)
Implement data quality checks testing and monitoring across the medallion architecture (bronze/silver/gold)
Own Unity Catalog governance for datasets and features access controls lineage and PII masking
Partner with platform/infra teams on job orchestration CI/CD for data & ML pipelines and cost optimization
Contribute to architecture decisions around lakehouse design streaming vs. batch tradeoffs and tool selection (build vs. buy)
Perform analysis on top of the data you build answer ad-hoc business questions validate metrics and spot data quality issues before they reach stakeholders
Evaluate and apply LLMs/agentic frameworks responsibly within the team balancing accuracy cost and governance (e.g. row/column-level access control on what an AI agent can query)
Your Checklist
1-3 years of data engineering experience working on Databricks in production
Strong proficiency in PySpark/Spark SQL and Python
Solid understanding of Delta Lake Unity Catalog Lakeflow/DLT and Databricks system tables (billing compute query history) Experience building and maintaining feature pipelines or ML data infrastructure (feature stores training/serving data parity)
Familiarity with MLflow (experiment tracking model registry) and/or Databricks Model Serving
Strong SQL skills and experience with warehouse performance tuning (query optimization materialization strategies cluster sizing)
Understanding of ML fundamentals enough to have real conversations with Data Scientists about features drift and model lifecycle (you dont need to be building models yourself but you should understand what good looks like)
Preferred Skills
Experience with real-time/streaming architectures (Structured Streaming Kafka Lakebase or similar)
Exposure to LLM/agentic tooling (Databricks Genie RAG pipelines vector search)
Experience with cost governance/FinOps for Databricks workloads
Background with experimentation platforms or A/B testing infrastructure
Familiarity with orchestration tools (Databricks Jobs Airflow) and CI/CD for data pipelines