Experience

Professional engineering track record, infrastructure design, and measured impact.

Data Scientist @ Bajaj Finserv Health

10/2024 — 06/2026
Pune, India•Full-time

Spearheading real-time streaming infrastructure, distributed data marts on Azure Databricks, graph-based recommendation engines, and self-serve AI analytics platforms.

  • ›Engineered a Kafka-based streaming pipeline, enabling real-time data flow to Power BI and powering live dashboards with <5s insight latency.
  • ›Streamlined Airflow orchestration for data mart creation by enabling concurrent execution of multiple Azure Databricks jobs with dependency-aware scheduling, reducing refresh time by 67% (3 hrs → 1 hr) and ensuring 9 AM SLA compliance for MIS and business reports.
  • ›Implemented ETL pipeline with Azure Databricks & SQL, reducing revenue processing time by 85%. Enhanced Power BI reports with trend analysis. Improved data accuracy by 70% through stakeholder collaboration.
  • ›Constructed a cross-sell recommendation engine using Neo4j, leveraging historical booking transactions and user activity data, which deployed targeted offers and intelligent bundling strategies to drive a 27% lift in GMV.
  • ›Delivered an AI-powered analytics platform for non-tech users with LLM insights, persona analysis, data summarization and agent-driven exploration, improving self-serve analytics adoption by 3x.
Key Results & Technical Milestones
  • <5s insight latency on real-time Power BI streaming
  • 67% faster refresh (3 hrs → 1 hr) meeting 9 AM SLA compliance
  • 85% reduction in revenue processing time via Databricks
  • 27% lift in GMV via Neo4j cross-sell recommendation engine
  • 3x improvement in self-serve analytics adoption with LLM platform
PythonApache KafkaAzure DatabricksApache AirflowNeo4jPower BISQLLLMsGenerative AI

Associate Data Scientist @ Bajaj Finserv Health

07/2023 — 10/2024
Pune, India•Full-time

Designed high-throughput Python ELT pipelines, AI assistant backends for healthcare triage, and automated lab scoring engines.

  • ›Built and tested a scalable Azure ETL pipeline to store 50+ KPIs in MySQL. Delivered insights via Power BI dashboard for visualization. Mechanized reporting workflows, boosting data accuracy and availability by 70%.
  • ›Architected a high-throughput Python ELT pipeline to extract over 1 million rows from Google BigQuery into Azure Data Lake, leveraging multiprocessing and batch processing to achieve an average ingestion time of under 10 minutes.
  • ›Created a flexible backend for AI bots integration, deployed OpenAI Assistant for lab bookings and triage, and established monitoring databases, cutting booking times by 30% and boosting triage accuracy by 20%.
  • ›Devised scalable backend for LLM models with conversation history and session tracking, added file upload with automated insights, and enabled agent-based creation, increasing self-serve insight generation by 60%.
  • ›Formulated a Lab Service Excellence Score to evaluate labs on price, volume, coverage, and catalog diversity, integrating it into booking to boost high-performing labs and improving user experience and efficiency by 30%.
Key Results & Technical Milestones
  • 1M+ rows extracted from BigQuery to Azure Data Lake in <10 mins
  • 30% cut in booking times & 20% boost in triage accuracy via OpenAI Assistant
  • 60% increase in self-serve insight generation with LLM sessions
  • 70% increase in data accuracy & availability across 50+ KPIs
PythonOpenAI AssistantGoogle BigQueryAzure Data LakeMySQLPower BIMultiprocessingLangChain

Data Science Intern @ Bajaj Finserv Health

07/2022 — 07/2023
Pune, India•Internship

Developed configuration-driven ETL workflows, cloud storage lifecycle archival pipelines, and cost-optimized data marts.

  • ›Developed a configuration-driven ETL pipeline with Azure Data Factory. Consolidated datasets into data marts, reducing query time by 65% and lowering costs.
  • ›Orchestrated an archival pipeline to compress and migrate aged transactional datasets to Azure Blob Cold Storage, slashing data storage costs by 40% while maintaining full accessibility for audits.
  • ›Enhanced query costs by 20% and execution time by 15% through data models and data marts development.
Key Results & Technical Milestones
  • 65% query time reduction via configuration-driven data marts
  • 40% data storage cost reduction through Cold Storage archival
  • 20% query cost reduction and 15% faster execution
Azure Data FactoryAzure Blob StorageSQLData ModelingPython