Building Your First ML Model for a Business Problem: What the Bootcamp Doesn't Tell You
Building Your First ML Model for a Business Problem: What the Bootcamp Doesn't Tell You
Most ML courses teach you how to train models. Very few teach you how to decide whether you should.
After years of building models for real business problems — churn prediction at Amazon, propensity models at Cartesian Consulting, recommender systems for hospitality clients — I've come to believe that the biggest failure mode for early-career data scientists isn't technical. It's problem framing.
Here's what I wish someone had told me before I built my first production model.
The problem you think you're solving is usually not the problem
In my first year doing analytics professionally, a stakeholder asked me to build a model to predict which customers were about to churn. I spent weeks cleaning data, engineering features, comparing algorithms. I delivered a model with solid precision and recall.
It sat unused for three months.
The problem wasn't the model. The problem was that I had answered the question I was given, not the question that actually mattered. The business didn't need to *predict* who was churning — they needed to know which customers could be *saved* and at what intervention cost. Those are fundamentally different problems. One requires a churn model. The other requires a churn model plus an uplift model plus an economics layer.
Before you write a single line of code, answer these questions:
Feature engineering is where the real work lives
Bootcamps and Kaggle competitions tend to reward clever algorithm choices. Production ML tends to reward better features.
The dirty secret of applied data science is that a logistic regression with thoughtful, domain-informed features will usually outperform a gradient boosting model trained on raw data. More importantly, it will be far easier to debug, explain to stakeholders, and maintain over time.
Domain knowledge is a feature engineering superpower. When I was building the winback model at Cartesian Consulting, the most predictive signal wasn't anything sophisticated — it was *how long it had been since the customer's last purchase relative to their historical purchase cadence*. That feature required understanding the customer, not knowing the algorithm.
Spend twice as long on your features as you think you need to. The model will thank you.
Offline metrics are a starting point, not a finish line
A model that scores well on your holdout set is not a model that works in production. This is the hardest lesson to learn because it requires you to think about how your model will interact with the real world over time.
A few things that will break your offline metrics in production:
The conversation you have to be willing to have
At some point in almost every ML project, the right answer is: "a simpler rule-based system will work better for this problem."
This is a hard thing to say when you've been hired as a data scientist. But the credibility you build by recommending the right solution — even when it's not the flashy one — is worth far more than delivering a neural network nobody trusts.
The goal is better decisions. Not better models.
Previous
From BI Engineer to Analytics Leader: How the Role Evolves and Where Most People Get Stuck
Next
Why Marketplace Analytics Is a Different Beast Than E-Commerce

Saurabh skipped presentations and built real AI products.
Saurabh Deshpande was part of the January 2026 cohort at Curious PM, alongside 13 other talented participants.
