Experience
Professional engineering track record, infrastructure design, and measured impact.
Data Scientist @ Bajaj Finserv Health
10/2024 — 06/2026
Pune, India•Full-time
Spearheading real-time streaming infrastructure, distributed data marts on Azure Databricks, graph-based recommendation engines, and self-serve AI analytics platforms.
- ›Engineered a Kafka-based streaming pipeline, enabling real-time data flow to Power BI and powering live dashboards with <5s insight latency.
- ›Streamlined Airflow orchestration for data mart creation by enabling concurrent execution of multiple Azure Databricks jobs with dependency-aware scheduling, reducing refresh time by 67% (3 hrs → 1 hr) and ensuring 9 AM SLA compliance for MIS and business reports.
- ›Implemented ETL pipeline with Azure Databricks & SQL, reducing revenue processing time by 85%. Enhanced Power BI reports with trend analysis. Improved data accuracy by 70% through stakeholder collaboration.
- ›Constructed a cross-sell recommendation engine using Neo4j, leveraging historical booking transactions and user activity data, which deployed targeted offers and intelligent bundling strategies to drive a 27% lift in GMV.
- ›Delivered an AI-powered analytics platform for non-tech users with LLM insights, persona analysis, data summarization and agent-driven exploration, improving self-serve analytics adoption by 3x.
Key Results & Technical Milestones
- <5s insight latency on real-time Power BI streaming
- 67% faster refresh (3 hrs → 1 hr) meeting 9 AM SLA compliance
- 85% reduction in revenue processing time via Databricks
- 27% lift in GMV via Neo4j cross-sell recommendation engine
- 3x improvement in self-serve analytics adoption with LLM platform
PythonApache KafkaAzure DatabricksApache AirflowNeo4jPower BISQLLLMsGenerative AI
Associate Data Scientist @ Bajaj Finserv Health
07/2023 — 10/2024
Pune, India•Full-time
Designed high-throughput Python ELT pipelines, AI assistant backends for healthcare triage, and automated lab scoring engines.
- ›Built and tested a scalable Azure ETL pipeline to store 50+ KPIs in MySQL. Delivered insights via Power BI dashboard for visualization. Mechanized reporting workflows, boosting data accuracy and availability by 70%.
- ›Architected a high-throughput Python ELT pipeline to extract over 1 million rows from Google BigQuery into Azure Data Lake, leveraging multiprocessing and batch processing to achieve an average ingestion time of under 10 minutes.
- ›Created a flexible backend for AI bots integration, deployed OpenAI Assistant for lab bookings and triage, and established monitoring databases, cutting booking times by 30% and boosting triage accuracy by 20%.
- ›Devised scalable backend for LLM models with conversation history and session tracking, added file upload with automated insights, and enabled agent-based creation, increasing self-serve insight generation by 60%.
- ›Formulated a Lab Service Excellence Score to evaluate labs on price, volume, coverage, and catalog diversity, integrating it into booking to boost high-performing labs and improving user experience and efficiency by 30%.
Key Results & Technical Milestones
- 1M+ rows extracted from BigQuery to Azure Data Lake in <10 mins
- 30% cut in booking times & 20% boost in triage accuracy via OpenAI Assistant
- 60% increase in self-serve insight generation with LLM sessions
- 70% increase in data accuracy & availability across 50+ KPIs
PythonOpenAI AssistantGoogle BigQueryAzure Data LakeMySQLPower BIMultiprocessingLangChain
Data Science Intern @ Bajaj Finserv Health
07/2022 — 07/2023
Pune, India•Internship
Developed configuration-driven ETL workflows, cloud storage lifecycle archival pipelines, and cost-optimized data marts.
- ›Developed a configuration-driven ETL pipeline with Azure Data Factory. Consolidated datasets into data marts, reducing query time by 65% and lowering costs.
- ›Orchestrated an archival pipeline to compress and migrate aged transactional datasets to Azure Blob Cold Storage, slashing data storage costs by 40% while maintaining full accessibility for audits.
- ›Enhanced query costs by 20% and execution time by 15% through data models and data marts development.
Key Results & Technical Milestones
- 65% query time reduction via configuration-driven data marts
- 40% data storage cost reduction through Cold Storage archival
- 20% query cost reduction and 15% faster execution
Azure Data FactoryAzure Blob StorageSQLData ModelingPython