Putting it to work — a forecast someone can act on
The whole path, turned into something a marketing lead can approve: which channel to cut, what next quarter will sell, how wrong that number will be, and the one experiment that would settle what no model here can.
Seven lessons ago there was a marketing lead with a budget, three channels and one question. No new technique in this lesson — just the answer, and the care it takes to give one.
The recommendation
Cut the newspaper line. It takes 15% of the media budget — $30.6k in the average market. And contributes -0.03k units. Move it to radio, in a test rather than in one move. Next quarter, a market spending $150k on TV and $30k on radio should sell about 15.4k units. Anything between 12k and 18.8k would not surprise us.
Four sentences, and every one of them is defensible. The rest of this lesson is why. And, more usefully, what a careful person would push back on.
One model explains, another forecasts, and they are not the same model
There is a temptation to find the model and use it for everything. The path has actually produced two, and they are for different questions.
| What it is for | Why that one | |
|---|---|---|
| All three channels | Explaining. What is each channel worth, holding the others constant? | You cannot say newspaper contributes nothing unless newspaper is in the model. Its coefficient of -0.001 is the finding. |
| TV + radio | Forecasting. What will a market sell on this plan? | Chosen in Which features earn their place on adjusted R² and AIC, then scored honestly in The honest number. |
All three channels
- What it is for
- Explaining. What is each channel worth, holding the others constant?
- Why that one
- You cannot say newspaper contributes nothing unless newspaper is in the model. Its coefficient of -0.001 is the finding.
TV + radio
- What it is for
- Forecasting. What will a market sell on this plan?
- Why that one
- Chosen in Which features earn their place on adjusted R² and AIC, then scored honestly in The honest number.
Keep a column so you can report that it is worthless. Drop it from the model you forecast with. Both are correct at once. The model you use to explain is not automatically the model you use to predict.
Quote the forecast with the range that comes with it
A point estimate on its own is not a forecast. It is a guess wearing a suit. What makes it usable is the second number.
| Value | Where it came from | |
|---|---|---|
| Base sales | 2.92k units | the intercept — what a market sells with no media at all |
| Per $1,000 of TV | 0.0458k units | the TV coefficient |
| Per $1,000 of radio | 0.188k units | the radio coefficient |
| Typical miss | 1.69k units | cross-validated in The honest number — measured on markets no coefficient ever saw |
Base sales
- Value
- 2.92k units
- Where it came from
- the intercept — what a market sells with no media at all
Per $1,000 of TV
- Value
- 0.0458k units
- Where it came from
- the TV coefficient
Per $1,000 of radio
- Value
- 0.188k units
- Where it came from
- the radio coefficient
Typical miss
- Value
- 1.69k units
- Where it came from
- cross-validated in The honest number — measured on markets no coefficient ever saw
So the plan of $150k TV and $30k radio forecasts 15.4k units, quoted as 12k to 18.8k. The point estimate plus and minus twice the cross-validated error. Note which error: the one measured on markets the model had never seen. Using the training error here would narrow the range and flatter the forecast. That is the commonest way a good model becomes a broken promise.
And the range is not the same width everywhere. Is it any good? measured the error spread at 1.86k units among low-budget markets and 4.42k among high-budget ones. The interval quoted above is an average over both. On your largest market it is optimistic — which is exactly where somebody is most likely to plan against it.
What this evidence does not support
This is the part that earns the trust. Say it before you are asked, and say it in the room.
- Not a causal claim. Nobody randomised these budgets. Everything here is associated with, not caused by. A market that spends more on radio may differ from one that does not in a dozen ways nobody recorded.
- Not a promised return on the money you move. Shifting newspaper's budget into radio would take the average market to $53.8k of radio, and the largest radio budget in the data is $49.6k. Past that, the model is extrapolating. And Look before you fit already showed the TV cloud bending, which is what saturation looks like before anyone models it.
- Not a base-sales number to quote elsewhere. Base is 2.92k units in this model. The single-channel model in The line put it at 7.03k, because with a channel missing the intercept absorbs its work. Base sales is a property of the model rather than of the business. So the honest answer to "what is our base?" is another question. Base, after accounting for what?
- Not a permanent verdict on newspaper. It is a verdict on newspaper in these markets, at these budgets, at this time. A channel can stop working and start again, and nothing in this path would notice.
Each of those is a sentence somebody will try to put in the deck.
One experiment would settle it
Everything above is inference from budgets somebody else set. The way out of that is not a better model. It is a better dataset, and it is cheap now that you know where to point it.
- Pick markets that are alike on what you can measure, and split them at random.
- Hold newspaper flat in one group. Cut it to zero in the other.
- Change nothing else, and wait a full purchase cycle.
- Compare sales between the groups. That difference is causal, because you assigned it.
That is a geo holdout test, and it is the reason this path's recommendation ends in "test it" rather than "do it". The model's job was never to make the decision. It was to find the one line in the budget worth spending an experiment on. And out of three channels and 200 markets, it did.
Where to find this in the product
You do not have to build this to meet it. Open CRM → Deals → Pipeline Analytics and look at the card titled 3-Month Revenue Forecast. It plots two lines: total pipeline value, and a weighted value. The pipeline discounted by the rate at which deals have actually been won recently, spread across the next three months.
Read it with the eyes this path just gave you, and the whole anatomy is visible:
| What it is doing | In this path's language |
|---|---|
| Predicting revenue, a quantity | a regression problem, not a classification one — When the answer is a number |
| Multiplying pipeline value by one historical win rate | a model with exactly one coefficient, fitted on every deal at once |
| Using the same rate for every deal | no per-deal features at all. Nothing playing the part radio played in More than one thing matters |
| Quoting a single line, with no band around it | a point estimate with no error attached. The gap The honest number is about |
Predicting revenue, a quantity
- In this path's language
- a regression problem, not a classification one — When the answer is a number
Multiplying pipeline value by one historical win rate
- In this path's language
- a model with exactly one coefficient, fitted on every deal at once
Using the same rate for every deal
- In this path's language
- no per-deal features at all. Nothing playing the part radio played in More than one thing matters
Quoting a single line, with no band around it
- In this path's language
- a point estimate with no error attached. The gap The honest number is about
That is not a criticism of it. A weighted-pipeline forecast is a genuinely reasonable first model. It is transparent, it needs no training, and everyone in the room understands it. It is the flat line from The line, promoted: better than nothing, and the thing anything cleverer has to beat.
The useful question to ask of it, and of any forecast handed to you. What is it multiplying? What would change its answer? And. The one nobody asks — how wrong has it been? A forecast whose past errors nobody has measured is a habit rather than a forecast.
What you can now do
The capability this path was built to hand over, stated as things rather than topics:
- Turn a business question into a prediction. Name the decision, the number that would change it, and what is known before that moment.
- Look at data before modelling it, and read a scatter for the two things that will bite later. A bend and a fan.
- Fit a regression and say what it claims, in base sales and incremental units per $1,000, not in coefficients.
- Score it three ways and distrust all three, then plot the residuals — which tell you where the model is wrong, and no score can.
- Read a coefficient as conditional on what else is in the model, which is what saved you from 15% of the budget.
- Choose between models without lying to yourself. Never on R². On adjusted R² or AIC over the same rows, and finally on data the model has never seen.
- Quote a number with a range, and say out loud what it does not support.
Eight lessons, one dataset, and a recommendation somebody can act on.
What this path deliberately left out
Three things were visible in the data and refused all the way through. Each is a different capability rather than a harder version of this one.
- The bend. Sales climb fast at low TV budgets and then flatten. Spotted in Look before you fit, measured as curvature in Is it any good?, and never fixed. Modelling saturation, and the fact that advertising keeps working after you stop paying for it, is the whole of media mix modelling.
- Columns that travel together. Newspaper and radio correlate 0.35, which was enough to make a dead channel look alive. With five channels across regions and weeks, it stops being an anecdote and starts making coefficients unstable.
- Choosing among many columns at once. Which features earn their place scored all 8 models because three channels allow it. Forty columns do not, and the answer is not a cleverer search. It is shrinking every coefficient toward zero and letting the data decide which survive.
Next in this track: Regression II. Saturation and adstock, correlated channels, and the regularisers — ridge and lasso — that make a model with forty columns behave. It uses a five-channel digital media file. It starts where this lesson stops: a recommendation you believe, and a budget too complicated to defend by hand.
Related lessons
Base rates — what a piece of evidence is actually worth
A face-recognition system that is 99.9% accurate and almost entirely wrong, and a number that sent an innocent woman to prison. Both are the same arithmetic, and it is the arithmetic that decides what any piece of evidence is worth.
ReadConfirmation and survivorship — what you never looked for
Two questions about evidence you did not go looking for. One is a rule you have to discover, and one is a pattern in five famous people — and in both, the thing that would have told you the truth is the thing nobody checks.
ReadLoss aversion, sunk cost and regression — what it costs you
Four questions you answer about yourself rather than about a scenario, and your own answers are the finding. Then the pattern that makes praise look useless and criticism look like it works, whatever you actually do.
Read
