Skip to main content
Data Statistics Math
View as Markdown Suggest changes

Condition on What the Client Already Knows

· Reading time: 20 min
Condition on What the Client Already Knows

TL;DR

In a factory’s sensor data, the strongest predictor of scrap was the ball type: already known, and not something the client can change. So instead of asking which signals predict scrap, we asked which are linked to scrap once the ball type is known. Five methods answered, and after correcting for testing twelve signals, none found evidence of such a link. The methods miss different kinds of link, so together they cover more kinds than any one of them. The one method that tests each signal on its own also bounds what it didn’t find: for eleven of the twelve signals, the data rule out an average within-type slope steep enough to matter.

The factory

On a current data science project, a client wants to know where its scrap comes from. To protect the client, everything in this post is fictional: the factory, the balls, the signals, and every number and figure. What carries over from the real project is the structure of the problem, the methods, and the shape of the result.

A factory makes balls: footballs, volleyballs, basketballs. Every order runs through the same four machines, M1 to M4, one after the other, and each machine records three signals, such as a temperature, a spin speed or a pressure. Some balls come off the line below standard and are thrown away. That is the scrap.

Ball type ZM1X1–X3M2X4–X6M3X7–X9M4X10–X12GoodScrap

Orders of three ball types ZZ pass four machines. Each machine records three signals, twelve in all, X1X_1 to X12X_{12}. At the end, some balls are scrap.

We look at orders, not single balls. Each order has a ball type ZZ, a scrap rate YY, the share of its balls thrown away, and a value for each signal XiX_i, its average over the time the order was on the line. There are n=400n = 400 orders: 180 of footballs, 140 of volleyballs and 80 of basketballs. Within a ball type, the scrap rate varies from order to order with a standard deviation of about 0.6 percentage points.

The obvious answer

One thing predicted scrap well: the ball type. Basketball orders lose 4.9% of their balls on average, volleyballs 3.2%, footballs 1.8%.

The client knew that already, and can’t act on it: an order for basketballs has to be filled with basketballs.

The ball type also gets in the way. The machines run differently for each ball type, so nearly every signal tracks the ball type, and through it, scrap. Ask which signals predict scrap, and the data keeps answering with the ones that give away the ball type.

0%2%4%6%150160170180190Temperature at M2 (°C)Scrap rateFootballVolleyballWithin typeBasketballAll orders

Across all orders, a hotter M2 goes with more scrap. Within each ball type, scrap shows no trend with temperature: the trend across all orders is the ball type.

The question we asked instead

The question is not whether XiX_i predicts YY, but whether it says anything about YY that ZZ doesn’t already: is the signal independent of scrap given the ball type,

Xi⊥ ⁣ ⁣ ⁣⊥Y∣Z  ?X_i \perp\!\!\!\perp Y \mid Z \;?

Because ZZ takes only three values, this reads concretely: among basketball orders, XiX_i says nothing about scrap, and the same among footballs and among volleyballs.

?Ball type ZSignal XiScrap rate Y

The ball type drives both the signals and the scrap rate. The question is whether a signal is linked to scrap beyond that.

A signal for which the data reject this independence is a candidate source of scrap, not yet a cause. A change of shift could move both the signal and scrap, or the signal could sit downstream of the defect: a ball damaged at M2 can still look different to a sensor at M4.

What independence implies

Conditional independence is a statement about whole distributions. For every ball type zz, signal value xx and scrap rate yy,

P(Y≤y∣Xi=x, Z=z)=P(Y≤y∣Z=z).P(Y \le y \mid X_i = x,\, Z = z) = P(Y \le y \mid Z = z).

That is rarely what gets tested directly. Most methods test something it implies, usually one of these two:

  1. The best prediction doesn’t change. Once the ball type is known, knowing the signal doesn’t improve the best prediction of scrap. Which prediction is best depends on the error you measure: the conditional mean minimises the expected squared error, the conditional median med⁡(Y∣Xi,Z)\operatorname{med}(Y \mid X_i, Z) the expected absolute error. For the mean:

    E[Y∣Xi,Z]=E[Y∣Z].\mathbb{E}[Y \mid X_i, Z] = \mathbb{E}[Y \mid Z].
  2. The conditional covariance is zero. Within each ball type, signal and scrap are uncorrelated:

    Cov⁡(Xi,Y∣Z)=0.\operatorname{Cov}(X_i, Y \mid Z) = 0.

Independence gives the first, for the mean and the median alike, since both are read off the conditional distribution above. The second follows from the mean version of the first: write μi(Z)=E[Xi∣Z]\mu_i(Z) = \mathbb{E}[X_i \mid Z], condition on XiX_i inside, and use the first:

Cov⁡(Xi,Y∣Z)=E[(Xi−μi(Z)) Y∣Z]=E[(Xi−μi(Z)) E[Y∣Xi,Z]∣Z]=E[(Xi−μi(Z)) E[Y∣Z]∣Z]=E[Y∣Z]  E[Xi−μi(Z)∣Z]=0.\begin{aligned} \operatorname{Cov}(X_i, Y \mid Z) &= \mathbb{E}\big[(X_i - \mu_i(Z))\, Y \mid Z\big] \\ &= \mathbb{E}\big[(X_i - \mu_i(Z))\, \mathbb{E}[Y \mid X_i, Z] \mid Z\big] \\ &= \mathbb{E}\big[(X_i - \mu_i(Z))\, \mathbb{E}[Y \mid Z] \mid Z\big] \\ &= \mathbb{E}[Y \mid Z]\; \mathbb{E}\big[X_i - \mu_i(Z) \mid Z\big] = 0. \end{aligned}

Neither step reverses, so either implication can hold while the signal still matters:

  • Zero covariance, but the mean moves. Scrap often rises when a temperature drifts from its set point in either direction. Within a ball type that is a U, and if both the U and the spread of the temperature around the set point are symmetric, the covariance is zero. A correlation sees nothing; the conditional mean does.
  • The same mean, but not the same distribution. A signal can make scrap more erratic from order to order without moving its average. A comparison of mean predictions sees nothing. For the client this case matters less: they pay for the expected scrap, which a change in spread leaves alone.
0%2%4%6%8%−10−50+5+10Temperature − set point (°C)Scrap rateConditional meanBest straight line

A hypothetical signal, for basketball orders: scrap rises as the temperature leaves its set point either way. The best straight line is flat, so the covariance is about zero; the conditional mean E[Y∣Xi, Z=basketball]\mathbb{E}[Y \mid X_i,\, Z = \text{basketball}] is not flat.

Five methods and what each tests

We answered the question five ways. Write X=(X1,…,X12)X = (X_1, \dots, X_{12}) for all signals together and X−iX_{-i} for all of them but XiX_i. For order k=1,…,nk = 1, \dots, n, write yky_k for its scrap rate, zkz_k for its ball type, and yˉz\bar y_z for the average scrap rate of the orders of ball type zz.

Every method that needed a regression model ran with three learners: XGBoost, a random forest, and TabICL, a pretrained transformer for tables that predicts from the training rows in context instead of being fitted to them.

MethodTests
Feature importance: SHAP and permutation importance of a model of YY on (X,Z)(X, Z)E[Y∣X,Z]=E[Y∣X−i,Z]{\mathbb{E}[Y \mid X, Z]} = {\mathbb{E}[Y \mid X_{-i}, Z]}, a neighbouring question
Model comparison: models of YY on ZZ and on (X,Z)(X, Z) that predict the median, scored on absolute errormed⁡(Y∣X,Z)=med⁡(Y∣Z){\operatorname{med}(Y \mid X, Z)} = {\operatorname{med}(Y \mid Z)}
Residual model: a model of Y−E[Y∣Z]{Y - \mathbb{E}[Y \mid Z]} on (X,Z)(X, Z)E[Y−E[Y∣Z]∣X,Z]=0{\mathbb{E}\big[Y - \mathbb{E}[Y \mid Z] \mid X, Z\big]} = 0
Models per ball type: a model of YY on XX for each ball type zzE[Y∣X,Z=z]=E[Y∣Z=z]{\mathbb{E}[Y \mid X, Z = z]} = {\mathbb{E}[Y \mid Z = z]} for each zz
GCM: the generalised covariance measure, for each signalE[Cov⁡(Xi,Y∣Z)]=0{\mathbb{E}[\operatorname{Cov}(X_i, Y \mid Z)]} = 0

The residual model and the models per ball type test the same equation, E[Y∣X,Z]=E[Y∣Z]\mathbb{E}[Y \mid X, Z] = \mathbb{E}[Y \mid Z], written two ways; the model comparison tests its median version. All three take the signals together. Each of their equations holds if X⊥ ⁣ ⁣ ⁣⊥Y∣ZX \perp\!\!\!\perp Y \mid Z, which implies Xi⊥ ⁣ ⁣ ⁣⊥Y∣ZX_i \perp\!\!\!\perp Y \mid Z for every ii, but not conversely: two signals can be linked to scrap together while each alone is independent of it. So a rejection by these three would say that the signals together are linked, not which one, and perhaps no single one. They differ in how they estimate the equation, and so in how much noise they carry. Only the GCM takes one signal at a time, and feature importance answers a neighbouring question.

Feature importance asks whether a model of YY on (X,Z)(X, Z) uses XiX_i. For a perfect model, that is whether E[Y∣X,Z]\mathbb{E}[Y \mid X, Z] changes with XiX_i while the other signals stay fixed: whether XiX_i adds anything to the other signals, not to the ball type alone. That is a neighbouring question, not an implication of ours, and the two differ in both directions. Two thermometers on one machine make each other redundant, so each can score low while both are linked to scrap. And a signal independent of scrap given the ball type can score high. Say a thermometer reads too high on hot days, while the machine’s true temperature, which drives scrap, doesn’t depend on the weather. If the outside temperature is recorded, a model uses it to correct the thermometer, although it is independent of scrap given the ball type.

Importance also comes with no threshold. And permutation importance, which usually shuffles XiX_i across all orders, can make a stand-in for the ball type look important: the shuffle breaks the signal’s link to the ball type too. Shuffling within each ball type avoids that, and on held-out orders it gives an exact test, but of Xi⊥ ⁣ ⁣ ⁣⊥(Y,X−i)∣ZX_i \perp\!\!\!\perp (Y, X_{-i}) \mid Z, which is stronger than ours: a signal correlated with another within a ball type can be rejected although Xi⊥ ⁣ ⁣ ⁣⊥Y∣ZX_i \perp\!\!\!\perp Y \mid Z.

The model comparison scores the model on (X,Z)(X, Z) against the model on ZZ alone, by absolute error on held-out orders. Each model predicts a conditional median: XGBoost and the random forest are fitted to minimise absolute error, and TabICL is asked for the 0.5 quantile of its predictive distribution. So the comparison tests the median version of the first implication. Absolute error keeps a few badly scrapped orders from dominating the scores. The orders must be held out: on the orders they were fitted to, the model with more inputs looks better whether or not the inputs matter. A third model, on XX alone, shows how much of the ball type the signals carry.

The residual model removes the ball type first. With three ball types, the best prediction from the ball type alone is the average per type, so the residual yk−yˉzky_k - \bar y_{z_k} estimates Y−E[Y∣Z]Y - \mathbb{E}[Y \mid Z]. The first implication says nothing should predict it. A second model, on (X,Z)(X, Z), tries anyway. Unlike the model comparison, it doesn’t have to relearn the ball type, and it isn’t judged by the difference between two noisy scores: plotted against the actual residuals of held-out orders, its predictions show the answer at a glance.

The models per ball type condition on ZZ by construction. Each sees orders of one ball type only and is scored, by squared error on held-out orders, against that type’s average scrap rate, so it can only find what the signals add within the type. Comparing them shows whether the signals matter in every ball type or in one. The price is fewer orders per model: the basketball model learns from 80.

A test per signal

Conditional independence is hard to test in general. Shah and Peters (2020) show that when (Xi,Y,Z)(X_i, Y, Z) has a joint density, which requires ZZ to be continuous, a test that rejects at most 5% of the time under every such distribution with Xi⊥ ⁣ ⁣ ⁣⊥Y∣ZX_i \perp\!\!\!\perp Y \mid Z also rejects at most 5% of the time under every such distribution without it, at any sample size. Every practical test buys its power with assumptions about how XiX_i and YY depend on ZZ.

With three ball types, that problem doesn’t arise: the question splits into one ordinary independence question per ball type, Xi⊥ ⁣ ⁣ ⁣⊥Y∣Z=zX_i \perp\!\!\!\perp Y \mid Z = z. With a continuous ZZ, such as the date, it does: the averages per ball type below become regressions of XiX_i and YY on ZZ, and the test is only as good as those regressions.

The same paper proposes the test we used, the generalised covariance measure (GCM). With a categorical ZZ it is short. Fix a signal XiX_i and write xkx_k for its value in order kk and xˉz\bar x_z for its average over the orders of ball type zz. The sample quantities below, from RkR_k to β^\hat\beta, belong to this one signal, so they carry no index ii. Subtract each order’s ball-type averages, multiply, and check whether the products average to zero:

Rk=(xk−xˉzk)(yk−yˉzk),T=n Rˉσ^R,R_k = \big(x_k - \bar x_{z_k}\big)\big(y_k - \bar y_{z_k}\big), \qquad T = \frac{\sqrt{n}\,\bar R}{\hat\sigma_R},

where Rˉ\bar R and σ^R\hat\sigma_R are the mean and standard deviation of R1,…,RnR_1, \dots, R_n. If Xi⊥ ⁣ ⁣ ⁣⊥Y∣ZX_i \perp\!\!\!\perp Y \mid Z, then for large nn, TT is approximately standard normal: its p-value is 2 (1−Φ(∣T∣))2\,(1 - \Phi(|T|)), with Φ\Phi the standard normal distribution function, and ∣T∣>1.96|T| > 1.96 rejects at the 5% level.

Rˉ\bar R estimates the covariance within ball types, averaged over them with weights pz=P(Z=z)p_z = P(Z = z):

E[Cov⁡(Xi,Y∣Z)]=∑zpzCov⁡(Xi,Y∣Z=z).\mathbb{E}[\operatorname{Cov}(X_i, Y \mid Z)] = \sum_{z} p_z \operatorname{Cov}(X_i, Y \mid Z = z).

Rˉ\bar R is in units of the signal times percentage points. Divided by the signal’s standard deviation within ball types, s^\hat s, with s^2=1n∑k(xk−xˉzk)2\hat s^2 = \frac{1}{n} \sum_{k} (x_k - \bar x_{z_k})^2, it becomes comparable across signals measured in degrees, revolutions per minute and bar:

β^=Rˉs^,β^±1.96 σ^Rn s^.\hat\beta = \frac{\bar R}{\hat s}, \qquad \hat\beta \pm 1.96\, \frac{\hat\sigma_R}{\sqrt{n}\, \hat s}.

β^\hat\beta is the least-squares slope of yk−yˉzky_k - \bar y_{z_k} on xk−xˉzkx_k - \bar x_{z_k}, times s^\hat s. On that line, an order whose signal lies one within-type standard deviation above its type’s average has a scrap rate β^\hat\beta percentage points above its type’s average. If the three ball types share one slope, β^\hat\beta estimates it, in these units; if not, the average of the three slopes weighted by pzVar⁡(Xi∣Z=z)p_z \operatorname{Var}(X_i \mid Z = z). The interval is an approximate 95% confidence interval for β=E[Cov⁡(Xi,Y∣Z)]/E[Var⁡(Xi∣Z)]\beta = \mathbb{E}[\operatorname{Cov}(X_i, Y \mid Z)] / \sqrt{\mathbb{E}[\operatorname{Var}(X_i \mid Z)]}, and it excludes zero exactly when the GCM rejects.

The GCM tests the covariance implication, averaged over ball types, so it misses any link whose average covariance is zero. The U is one: its covariance is zero within each ball type. The average adds another: covariances of opposite sign in different ball types can cancel.

−10+1−8−40+4+8Temperature − type average (°C)Scrap rate − type average (percentage points)FootballVolleyballCommon slope

A hypothetical signal, with each order centred on its ball type’s averages. Scrap rises with the temperature in footballs and falls in volleyballs. The covariances cancel, so the common slope, and with it the GCM, sees nothing; a model per ball type sees both slopes.

A GCM per ball type would close the second gap: compute TzT_z from the orders of ball type zz alone, and compare ∑zTz2\sum_z T_z^2 with a χ2\chi^2 distribution with three degrees of freedom, which it approximately follows under the independence. We relied on the models per ball type instead.

Where the GCM and the models differ

As nn grows, Rˉ\bar R settles at E[Cov⁡(Xi,Y∣Z)]\mathbb{E}[\operatorname{Cov}(X_i, Y \mid Z)] and σ^R\hat\sigma_R at some σR>0\sigma_R > 0, so ∣T∣|T| grows like n ∣E[Cov⁡(Xi,Y∣Z)]∣/σR\sqrt{n}\,\big|\mathbb{E}[\operatorname{Cov}(X_i, Y \mid Z)]\big| / \sigma_R: any average covariance that isn’t exactly zero is eventually found. With our 400 orders, the GCM finds a link with ∣β∣=0.09|\beta| = 0.09, in percentage points of scrap per within-type standard deviation, four times out of five.

A model has a harder job with the same orders: it has to find that link among twelve signals without knowing its shape, while the GCM estimates one number per signal. And a ranking of importances puts a small link at the bottom, next to the signals with no link at all, with nothing to tell them apart. The GCM reports β^\hat\beta with an interval; whether a slope of that size is worth acting on is the client’s call.

Correlated signals

Feature importance divides credit among the inputs of one model, and how correlated signals share it depends on the learner: a random forest tends to spread credit over all of them, while XGBoost tends to pick one and leave the rest near zero. SHAP explains whichever model it is given, so it inherits the split.

The GCM doesn’t have this problem. Each test conditions only on the ball type, so correlated signals don’t compete. The flip side is that both of two correlated signals are flagged, like the two thermometers above. Asking which of them matters means conditioning on the other signals too, Xi⊥ ⁣ ⁣ ⁣⊥Y∣(Z,X−i)X_i \perp\!\!\!\perp Y \mid (Z, X_{-i}), and since they are continuous, that brings back the hardness result above.

Why five methods

Each method misses some kinds of link, and a null result from one method leaves its blind spot open. Here a small link is too weak for the models: it lowers the model comparison’s held-out mean absolute error by less than the standard deviation of the per-fold difference between two models, about 0.015 percentage points. For a straight-line link, that means ∣β∣|\beta| below about 0.15. The GCM finds a link four times out of five from about ∣β∣=0.09|\beta| = 0.09.

Can it find…ImportanceComparisonResidualPer typeGCM
a small link with ∣β∣≥0.09\lvert\beta\rvert \ge 0.09NoNoNoNoYes
a large link with zero covariance within each type, like the UYesYesYesYesNo
a large link whose covariances cancel across typesYesYesYesYesNo
a small link with β=0\beta = 0NoNoNoNoNo
a link only in the spread, with mean and median unchangedNoNoNoNoNo
which signal is linkedNoNoNoNoYes

A Yes for a model-based method assumes that at least one of the three learners can fit the link’s shape. Feature importance does name signals, but the ones a model uses given the other signals, which can differ from the linked ones in both directions. The comparison, the residual model and the models per ball type see the same kinds of link in this grid, the comparison through the median and the other two through the mean; they differ in how much noise they carry, so each checks the others. Two kinds of link have no Yes: one only in the spread matters less to the client, as above, and the small one with β=0\beta = 0 is listed below.

What we found

After correcting for testing twelve signals, none of the five methods found evidence that a signal is linked to scrap beyond the ball type. The three learners agreed for every method that used them.

  • Feature importance: in each of the three models, the ball type took about 90% of the total absolute SHAP value, and no signal more than 2%. Shuffling any one signal within its ball type raised the held-out error by less than 0.005 percentage points.
  • The model comparison, with five-fold cross-validation, gave a held-out mean absolute error of 1.08 percentage points with no inputs, 0.48 with the ball type alone, 0.49 with the ball type and the signals, and 0.51 with the signals alone. Per fold, adding the signals to the ball type raised the error by 0.01 on average, with a standard deviation of 0.015. On their own, the signals came close to the ball type because they carry it.
  • The residual model’s predictions for held-out orders had a standard deviation of 0.08 percentage points, against 0.6 for the residuals they tried to predict, and were uncorrelated with them: their squared error was about 2% higher than that of predicting zero.
  • The models per ball type had a held-out squared error 0% to 3% above that of their type’s average scrap rate, in all three ball types, so there was nothing to compare across types.
  • The GCM rejected one of the twelve signals at the 5% level: X7X_7, a spin speed at M3, with p=0.02p = 0.02. With twelve tests, 0.6 such rejections are expected by chance if no signal is linked. After a Bonferroni correction for twelve tests, which holds however the tests depend on each other, its p-value is 12×0.02=0.2412 \times 0.02 = 0.24; Benjamini–Hochberg gives the same.
−2−10+1+2−2−10+1+2Actual residual (percentage points)Predicted residual (percentage points)Predicted = actual

Held-out orders. A link would tilt them towards the dashed line.

A flag from one method alone deserves a look before it is dismissed, if it falls in a blind spot of the others. X7X_7’s does: a small link with nonzero covariance is what the models miss. Its β^\hat\beta is 0.075, with a 95% interval from 0.01 to 0.14.

What no evidence means

The GCM also bounds the links it didn’t find. The client set the bar for a signal worth acting on at ∣β∣=0.15|\beta| = 0.15: 0.15 percentage points of scrap per within-type standard deviation of the signal, a quarter of the 0.6-point standard deviation of scrap between orders of one ball type.

Too small to act on (±0.15)−0.2−0.10+0.1+0.2Percentage points of scrapper within-type standard deviation of the signalM1M2M3M4X1X2X3X4X5X6p = 0.02X7X8X9X10Wide intervalX11X12

β^\hat\beta and its 95% interval for each signal, grouped by machine. The shaded band holds slopes too small to act on. X7X_7’s is the only interval that excludes zero, the one rejection before correction; X11X_{11}‘s reaches past the band.

Eleven of the twelve intervals lie inside the band, X7X_7‘s included: for each of those signals on its own, the data rule out ∣β∣≥0.15|\beta| \ge 0.15 at 95% confidence. If X7X_7‘s link is real, it is too small to act on. X11X_{11}, a pressure at M4, has a few orders with extreme values that make its interval wide, from −0.20 to 0.14, so the data can’t rule out a link that matters there.

What the null leaves open:

  • X11X_{11}. More orders, or a look at its extreme orders, would narrow its interval.
  • Slopes of opposite sign in different ball types. β^\hat\beta averages them, so the bound doesn’t cover them; the models per ball type would have found large ones.
  • Small links with zero average covariance. The bound covers β\beta only, and the models miss small links of any shape. More orders would help the models; the GCM misses such a link at any nn.
  • Causes no sensor measures. The raw material, or handling between machines. If scrap comes from there, no analysis of the recorded signals will find it, unless it leaves a trace in them.

On held-out orders, no model predicted any of the variance of scrap within a ball type. That variance comes from something the signals don’t measure, from links too small or oddly shaped for these methods, or from chance; a straight-line link just below the client’s bar would account for about 6% of it. So the next place to look is outside the twelve signals: what isn’t measured yet, and how much of the spread is chance.

The general move

The strongest predictor is often something the client already knows and can’t change: the product, the site, the season. Reporting it again is easy and worthless. Condition on it, and ask what is left. Sometimes the answer is nothing the data can see, and that is worth knowing too.

AI Chat

Messages you send are processed by the Google Gemini API to generate responses. Do not share sensitive personal data. See the privacy policy for details.