Sharad Babar.

Data Scientist.

01

// about.me

I Build Models
That Move Decisions.

I'm a Data Scientist with 4+ years of experience across retail and financial services. At Walmart, I build demand forecasts, pricing models, and customer segments with Python, SQL, PySpark, and Databricks, and design A/B tests that guide merchandising and search decisions. Previously at Accenture, I built fraud and risk models, near real-time transaction pipelines, and Power BI reporting for 100+ enterprise banking clients. I own each question end to end: pulling the data, validating the model or experiment, and explaining the result to people who don't read code.

0+
Years in Data Science
~18%
Forecast Error Reduction
~12M
Customers Segmented

> specializes_in:

Demand ForecastingCustomer SegmentationExperimentationFraud & Risk Modeling

> focus_points[]:

->Use statistical validation, explainability, and clear reporting to support business decisions.
->Deploy models on AWS SageMaker with MLflow tracking and drift monitoring; evaluate RAG summaries for factual accuracy.
->M.S. in Computer Science, State University of New York, May 2025.
Let's Talk Data
~/portfolio/sharad - bash
> whoami
data_scientist

> cat stack.config
{
languages: ["Python", "SQL", "PySpark"],
ml: ["scikit-learn", "XGBoost", "K-Means", "SHAP"],
analytics: ["A/B Testing", "Power BI", "Tableau"],
platforms: ["Databricks", "Spark", "AWS SageMaker", "MLflow"],
genai: ["Hugging Face Transformers", "RAG", "LLM Evaluation"],
}

> impact.metrics
forecast_error_reduction: ~18%, customers_segmented: ~12M, daily_transactions_processed: ~2M

> approach
data → test → decision|
02

Experience.

Forecasting, experimentation, and fraud analytics across retail and financial services.

Walmart

Walmart

Apr 2025 - Present

Data Scientist

USA

  • Built weekly demand forecasts across ~40 store-item categories in PySpark on Databricks, comparing gradient boosting with seasonal ARIMA and choosing boosting for promotional spikes; forecast error fell ~18%.
  • Segmented roughly 12M customers using K-Means, RFM, and basket-mix features, helping improve campaign response by about 15%.
  • Modeled price elasticity across seasonal promotions using regression on ~3 years of transaction history, distinguishing markdowns that moved units from those that eroded margin; category margin improved ~6%.
  • Designed and analyzed A/B tests for search-ranking changes, using power calculations and guardrail metrics to support evidence-based decisions and improve add-to-cart rate by about 4%.
  • Diagnosed repeat stockouts across regional distribution centers in SQL and Power BI, identifying replenishment lead-time as the dominant driver; store-level availability improved ~7%.
  • Deployed forecasting and segmentation models on AWS SageMaker with MLflow tracking and drift alerts, replacing manual refreshes and reducing retraining cycle time to under 24 hours.
  • Automated weekly merchandising summaries with SQL result sets and a retrieval-augmented LLM workflow, evaluated in MLflow for factual accuracy; saved analysts ~10 hours of write-up per week.
Accenture

Accenture

Jun 2020 - Jul 2023

Data Scientist

India

  • Developed a transaction fraud classifier in Python and PySpark, comparing XGBoost with logistic regression and choosing XGBoost for rare-event recall; AUC rose from 0.82 to ~0.90.
  • Engineered Apache Spark, Kafka, and SQL ingestion and ETL pipelines over ~2M daily transactions, restructuring joins for near real-time scoring and reducing decision latency by ~35%.
  • Created behavioral, temporal, and merchant-level features in Pandas and Spark SQL, raising precision and recall on high-risk fraud segments by ~20%.
  • Validated rule and model changes with A/B tests and hypothesis testing, tuning alert thresholds to reduce false positives without suppressing genuine fraud; review volume dropped ~18%.
  • Surfaced emerging fraud patterns with K-Means and Isolation Forest on unlabeled transaction streams, catching schemes the rules engine missed; undetected incidents fell ~27%.
  • Published self-serve Power BI dashboards on fraud loss, alert volume, and risk scores for 100+ enterprise banking clients, replacing weekly manual decks.
  • Partnered with risk and compliance leads on audit-defensible thresholds and wrote SQL views for monthly reviews; scoring covered ~14M flagged transactions yearly.
03

Technical Skills.

Tools and methods I use for forecasting, experimentation, customer analytics, and production ML.

Programming & Databases

SQLPythonPySparkPandasNumPySpark SQLQuery Optimization

Statistics & Experimentation

A/B TestingHypothesis TestingExperimental DesignRegression AnalysisCausal InferenceTime-Series ForecastingFeature EngineeringStatistical Validation

Analytics & BI

Power BITableauKPI DefinitionDashboard DesignCohort & Funnel AnalysisSQL AnalyticsData Storytelling

Machine Learning

scikit-learnXGBoostRandom ForestLogistic RegressionClustering & SegmentationK-MeansIsolation ForestAnomaly DetectionModel ValidationSHAP ExplainabilityPyTorchTensorFlow

Data Platforms & Cloud

SnowflakeDatabricksAWS S3AWS EC2AWS SageMakerDelta LakeApache SparkKafkaAirflowETL PipelinesData Quality Validation

Deployment & Applied GenAI

MLflowDockerGitCI/CDModel MonitoringDrift DetectionHugging Face TransformersRetrieval-Augmented GenerationLLM Evaluation
04

Selected Works.

Churn modeling, retention experiments, and e-commerce analytics that turn data into business decisions.

Subscription Churn Prediction & Retention Test Design

Predicted subscriber churn from 18 months of order and support history, choosing logistic regression over XGBoost for interpretable coefficients and reaching ~0.78 recall. Validated on a held-out quarter with stratified time splits after finding random-split AUC was inflated by roughly 6 points. A two-arm holdout test on the top churn decile compared discounts with concierge outreach; a chi-square readout showed outreach retained ~11% more accounts.

PythonPandasscikit-learnLogistic RegressionXGBoostFeature EngineeringTime-Split ValidationA/B & Holdout TestingChi-Square TestChurn Modeling
~0.78 recall · ~11% more accounts retained with outreach

E-Commerce Funnel & Cohort Analytics Dashboard

Modeled clickstream and order data into a session-grain SQL fact table covering ~2M sessions, defining activation, repeat purchase, and 90-day retention as shared KPIs. Used window functions to trace monthly signup cohorts through a five-step purchase funnel, isolating a ~34% drop between cart and payment. Delivered a Tableau dashboard with a Power BI companion for finance and recommended a guest-checkout pilot that recovered ~8% of abandoned carts.

SQLData ModelingKPI DefinitionCohort AnalysisFunnel AnalysisRetention AnalysisTableauPower BIDashboard Design
~2M sessions analyzed · ~8% of abandoned carts recovered
05

Let's Build Smarter Systems.

Open to data science opportunities in forecasting, experimentation, customer analytics, and machine learning.

Fill in your details to open a draft in your email app, then send it from there. You can also use the email link above to get in touch directly.

(c) 2026 Sharad BabarBuilt with Next.js & Framer Motion