Influence Analysis Through the Value of Information Framework
- 32 mins All opinions are my own.Influence analysis asks: how much does a specific piece of information affect our model or decision? The “piece of information” can be a training data point, an input feature, a model parameter, or new data. Different fields have developed different approaches to answer this question, each from a distinct intellectual tradition:
- Gradient-based (Influence Functions): from robust statistics and optimization. Uses gradients and Hessians of the loss to measure how the model changes under small data perturbations. Originally developed for classical statistics (Cook’s distance, leverage), extended to deep learning by Koh and Liang (2017). Also the basis of feature attribution methods like saliency maps and integrated gradients.
- Game-theoretic (Shapley Values): from cooperative game theory. Measures each player’s fair share of the total payoff by averaging marginal contributions across all possible coalitions. Applied to data points (Data Shapley) and to features (SHAP) as separate instantiations of the same Shapley value formula.
- Decision-theoretic (Value of Information): from Bayesian decision theory. Measures the expected reduction in loss from learning a piece of information. Connects to classical test statistics (F-test), observation diagnostics (Cook’s distance), and variance-based sensitivity analysis (Sobol indices) as special cases.
All three frameworks can evaluate influence from both the feature and data perspectives. This post develops the Value of Information framework in detail, then connects it to influence functions and Data Shapley, showing how classical and modern methods relate under a common lens.
The central object is the Expected Value of Information Ratio (EVOIR), which measures the realized influence of a piece of information relative to its expected influence. Under squared loss, EVOIR reduces to familiar quantities: the F-statistic for parameter groups, Cook’s distance for observations, and variance-based sensitivity indices for features.
The Value of Information Framework
Let $\theta$ denote the unknown parameters, $Y$ the observed data, and $\phi$ a piece of information we want to evaluate. Here $\phi$ can be a subset of $\theta$, a function of $\theta$, or a random variable conditional on $\theta$ (such as new data $Y_{\text{new}}$).
Given a loss function $L(a, \theta)$ and an action $a$, define:
- $a_{Y}$: the optimal action given data $Y$ alone.
- $a_{Y,\phi}$: the optimal action given both $Y$ and $\phi$.
Retrospective Value of Information
The retrospective value of information (RVOI) measures the observed reduction in expected loss from learning $\phi$, after $\phi$ is revealed:
$$ \text{RVOI} = E\{L(a_Y, \theta) - L(a_{Y,\phi}, \theta) \mid \phi, Y\}. $$
RVOI answers: “Now that we know $\phi$, how much did it actually help?” It is always non-negative because $a_{Y,\phi}$ optimizes the loss with strictly more information than $a_{Y}$.
Prospective Value of Information
The prospective value of information (PVOI) is the expected RVOI before $\phi$ is observed:
$$ \text{PVOI} = E[\text{RVOI} \mid Y] = E\{L(a_Y, \theta) - L(a_{Y,\phi}, \theta) \mid Y\}. $$
PVOI answers: “How much do we expect $\phi$ to help, on average before we observed it?” In health economics, PVOI is called the Value of Partial Perfect Information (for parameters) or the Value of Sample Information (for new data). It measures the expected influence of new information.
Expected Value of Information Ratio
The EVOIR is the ratio of realized to expected influence:
$$ \text{EVOIR}(\phi \mid Y) = \frac{\text{RVOI}}{\text{PVOI}} = \frac{E\{L(a_Y, \theta) - L(a_{Y,\phi}, \theta) \mid \phi, Y\}}{E\{L(a_Y, \theta) - L(a_{Y,\phi}, \theta) \mid Y\}}. $$
The expected value of EVOIR is 1 (by the law of total expectation). Values much larger than 1 indicate that $\phi$ was surprisingly influential; values near 0 indicate it had little effect. This ratio normalizes influence to a common scale, making it comparable across different types of information.
EVOIR Under Squared Loss
Assume
$$ Y = f(X) + \epsilon, \quad E[\epsilon] = 0, $$
where $f(X)$ is parameterized by $\theta$. Under squared loss
$$ L(\hat{f}, f) = (\hat{f}(X) - f(X))^\top (\hat{f}(X) - f(X)), $$
the Bayes estimator is the posterior mean:
$$ \hat{f}_Y(X) = E[f(X) \mid Y], \quad \hat{f}_{Y,\phi}(X) = E[f(X) \mid Y, \phi]. $$
By the law of total expectation, $\hat{f}_{Y}(X) = E[\hat{f}_{Y,\phi}(X) \mid Y]$.
Under squared loss, RVOI, PVOI, and EVOIR simplify to clean geometric quantities:
$$ \text{RVOI} = (\hat{f}_Y(X) - \hat{f}_{Y,\phi}(X))^\top (\hat{f}_Y(X) - \hat{f}_{Y,\phi}(X)), $$
$$ \text{PVOI} = \text{tr}(\text{Var}(\hat{f}_{Y,\phi}(X) \mid Y)), $$
$$ \text{EVOIR} = \frac{(\hat{f}_Y(X) - \hat{f}_{Y,\phi}(X))^\top (\hat{f}_Y(X) - \hat{f}_{Y,\phi}(X))}{\text{tr}(\text{Var}(\hat{f}_{Y,\phi}(X) \mid Y))}. $$
RVOI is the squared distance between the original and updated predictions. PVOI is the total posterior variance of the updated predictions. EVOIR is the ratio of realized squared change to expected squared change.
Derivation of RVOI, PVOI, and EVOIR under squared loss
RVOI derivation. Starting from the definition:
$$\text{RVOI} = E\{L(\hat{f}_Y, f) - L(\hat{f}_{Y,\phi}, f) \mid Y, \phi\}.$$Expanding the squared loss terms:
$$= E\{(\hat{f}_Y - f)^\top(\hat{f}_Y - f) - (\hat{f}_{Y,\phi} - f)^\top(\hat{f}_{Y,\phi} - f) \mid Y, \phi\}.$$Expanding and using the fact that $E[f(X) \mid Y, \phi] = \hat{f}_{Y,\phi}(X)$:
$$= \hat{f}_Y^\top \hat{f}_Y - 2\hat{f}_Y^\top \hat{f}_{Y,\phi} - \hat{f}_{Y,\phi}^\top \hat{f}_{Y,\phi} + 2\hat{f}_{Y,\phi}^\top \hat{f}_{Y,\phi}$$ $$= \hat{f}_Y^\top \hat{f}_Y - 2\hat{f}_Y^\top \hat{f}_{Y,\phi} + \hat{f}_{Y,\phi}^\top \hat{f}_{Y,\phi}$$ $$= (\hat{f}_Y - \hat{f}_{Y,\phi})^\top (\hat{f}_Y - \hat{f}_{Y,\phi}).$$RVOI is the squared Euclidean distance between the original and updated Bayes estimates.
PVOI derivation. Taking the expectation of RVOI over $\phi \mid Y$:
$$\text{PVOI} = E\{(\hat{f}_Y - \hat{f}_{Y,\phi})^\top (\hat{f}_Y - \hat{f}_{Y,\phi}) \mid Y\}.$$Since $\hat{f}_{Y} = E[\hat{f}_{Y,\phi} \mid Y]$, this is the expected squared deviation of $\hat{f}_{Y,\phi}$ from its mean, which is the trace of its covariance:
$$= \text{tr}(E\{(\hat{f}_{Y,\phi} - \hat{f}_Y)(\hat{f}_{Y,\phi} - \hat{f}_Y)^\top \mid Y\}) = \text{tr}(\text{Var}(\hat{f}_{Y,\phi}(X) \mid Y)).$$End of expanded note.
Linear Regression Example
In linear regression, $f(X) = X\beta$ with the Bayesian model:
$$ Y \mid \beta, \sigma^2 \sim N(X\beta, \sigma^2 I_n), \quad \pi(\beta, \sigma^2) \propto \sigma^{-2}. $$
The posterior is:
$$ \beta \mid \sigma^2, Y \sim N(\hat{\beta}_Y, \sigma^2 (X^\top X)^{-1}), \quad \sigma^2 \mid Y \sim \chi^{-2}(n-p, S_n^2), $$
where $\hat{\beta}_{Y} = (X^\top X)^{-1} X^\top Y$ and $S_{n}^2 = (Y - X\hat{\beta}_{Y})^\top (Y - X\hat{\beta}_{Y}) / (n-p)$.
Case A: Learning a Parameter Group
Let $\phi = \beta_{g}$, a subvector of $\beta$ with $p_{2}$ components. Partition:
$$ \beta \mid \sigma^2, Y \sim N\left(\begin{pmatrix} \hat{\beta}_{Y,-g} \ \hat{\beta}_{Y,g} \end{pmatrix}, \sigma^2 \begin{pmatrix} A & B \ B^\top & C \end{pmatrix}\right). $$
The EVOIR for learning $\beta_{g}$ is:
$$ \text{EVOIR} = \frac{n-p-2}{n-p} \cdot F, $$
where $F$ is the classical F-statistic for testing $\beta_{g}$:
$$ F = \frac{(\hat{\beta}_{Y,g} - \beta_g)^\top C^{-1} (\hat{\beta}_{Y,g} - \beta_g) / p_2}{S_n^2} \sim F_{p_2, n-p}. $$
The F-statistic, one of the most widely used quantities in regression, is (up to a known scale factor) the EVOIR for a parameter group under squared loss. Large F means $\beta_{g}$ was more influential than expected; small F means learning $\beta_{g}$ changed the predictions less than expected.
Derivation: EVOIR equals scaled F-statistic
The conditional distribution of $\beta_{-g}$ given $\beta_{g}$ is:
$$\beta_{-g} \mid Y, \beta_g, \sigma^2 \sim N(\hat{\beta}_{Y,-g} + BC^{-1}(\beta_g - \hat{\beta}_{Y,g}),\; \sigma^2(A - BC^{-1}B^\top)).$$The Bayes estimates are $\hat{\beta}_{Y}$ (given $Y$ alone) and
$$\hat{\beta}_{Y,\beta_g} = \begin{pmatrix} \hat{\beta}_{Y,-g} + BC^{-1}(\beta_g - \hat{\beta}_{Y,g}) \ \beta_g \end{pmatrix}$$(given $Y$ and $\beta_{g}$). Computing RVOI:
$$\text{RVOI} = (X\hat{\beta}_Y - X\hat{\beta}_{Y,\beta_g})^\top (X\hat{\beta}_Y - X\hat{\beta}_{Y,\beta_g}).$$After simplification using the partitioned structure of $X^\top X$:
$$\text{RVOI} = (\hat{\beta}_{Y,g} - \beta_g)^\top C^{-1} (\hat{\beta}_{Y,g} - \beta_g).$$For PVOI, we take the expectation over $\beta_{g} \mid Y$:
$$\text{PVOI} = E\{(\hat{\beta}_{Y,g} - \beta_g)^\top C^{-1} (\hat{\beta}_{Y,g} - \beta_g) \mid Y\}.$$Since $\beta_{g} \mid \sigma^2, Y \sim N(\hat{\beta}_{Y,g}, \sigma^2 C)$, the quadratic form $(\hat{\beta}_{Y,g} - \beta_{g})^\top C^{-1} (\hat{\beta}_{Y,g} - \beta_{g}) / \sigma^2 \sim \chi^2_{p_{2}}$ conditional on $\sigma^2$. Taking the double expectation:
$$\text{PVOI} = E\{E\{(\hat{\beta}_{Y,g} - \beta_g)^\top C^{-1} (\hat{\beta}_{Y,g} - \beta_g) \mid Y, \sigma^2\} \mid Y\} = E\{p_2 \sigma^2 \mid Y\} = p_2 \frac{n-p}{n-p-2} S_n^2.$$Therefore:
$$\text{EVOIR} = \frac{(\hat{\beta}_{Y,g} - \beta_g)^\top C^{-1} (\hat{\beta}_{Y,g} - \beta_g)}{p_2 \frac{n-p}{n-p-2} S_n^2} = \frac{n-p-2}{n-p} \cdot \frac{(\hat{\beta}_{Y,g} - \beta_g)^\top C^{-1} (\hat{\beta}_{Y,g} - \beta_g) / p_2}{S_n^2} = \frac{n-p-2}{n-p} \cdot F.$$End of expanded note.
Case B: Learning New Data
Let $\phi = Y_{\text{new}}$, a vector of $m$ new observations. Let $X$ denote the combined design matrix for both the original $n$ and new $m$ samples (an $(n+m) \times p$ matrix).
$$ \text{RVOI} = \sum_{i=1}^{n+m} (x_i^\top \hat{\beta}_Y - x_i^\top \hat{\beta}_{Y, Y_{\text{new}}})^2, $$
$$ \text{PVOI} = \text{tr}(\text{Var}(X\hat{\beta}_{Y,Y_{\text{new}}} \mid Y)), $$
$$ \text{EVOIR} = \frac{\sum_{i=1}^{n+m} (x_i^\top \hat{\beta}_Y - x_i^\top \hat{\beta}_{Y,Y_{\text{new}}})^2}{\text{tr}(\text{Var}(X\hat{\beta}_{Y,Y_{\text{new}}} \mid Y))}. $$
This measures how much new data changes the predictions across all observation points, relative to the expected change. It connects to the Value of Sample Information in health economics.
Case C: Influence of a Single Observation (Cook’s Distance)
Let $\phi = (x_{k}, y_{k})$, the $k$-th observation. The RVOI for removing observation $k$ measures how much the fitted values change:
$$ \text{RVOI}_k = (X\hat{\beta}_Y - X\hat{\beta}_{Y}^{(-k)})^\top (X\hat{\beta}_Y - X\hat{\beta}_{Y}^{(-k)}), $$
where $\hat{\beta}_{Y}^{(-k)}$ is the OLS estimate with observation $k$ deleted. Cook’s distance is this quantity scaled by the variance:
$$ D_k = \frac{(\hat{\beta}_Y - \hat{\beta}_Y^{(-k)})^\top (X^\top X) (\hat{\beta}_Y - \hat{\beta}_Y^{(-k)})}{p \, S_n^2} = \frac{(X\hat{\beta}_Y - X\hat{\beta}_Y^{(-k)})^\top (X\hat{\beta}_Y - X\hat{\beta}_Y^{(-k)})}{p \, S_n^2}. $$
So Cook’s distance is a scaled RVOI:
$$ D_k = \frac{\text{RVOI}_k}{p \, S_n^2}. $$
The numerator is the realized squared change in predictions (RVOI). The denominator $p \, S_{n}^2$ is proportional to the expected prediction variance, which plays the same normalizing role as PVOI. Large $D_{k}$ means observation $k$ had an outsized influence on the fitted values.
Cook's distance closed form and connection to leverage and residuals
Using the Sherman-Morrison-Woodbury formula for rank-1 updates, the leave-one-out estimate has a closed form:
$$\hat{\beta}_Y^{(-k)} = \hat{\beta}_Y - \frac{(X^\top X)^{-1} x_k e_k}{1 - h_{kk}},$$where $e_{k} = y_{k} - x_{k}^\top \hat{\beta}_{Y}$ is the ordinary residual and $h_{kk} = x_{k}^\top (X^\top X)^{-1} x_{k}$ is the leverage (the $k$-th diagonal element of the hat matrix $H = X(X^\top X)^{-1}X^\top$). Substituting into Cook's distance:
$$D_k = \frac{e_k^2}{p \, S_n^2} \cdot \frac{h_{kk}}{(1 - h_{kk})^2}.$$This reveals that Cook's distance is the product of two factors:
- Residual magnitude: $e_{k}^2 / (p \, S_{n}^2)$ measures how poorly the model fits observation $k$. Large residuals indicate the point is an outlier in the $y$-direction.
- Leverage: $h_{kk} / (1 - h_{kk})^2$ measures how extreme observation $k$'s position is in the feature space. Points far from the center of the data have high leverage and exert more pull on the fitted line.
A point needs both a large residual and high leverage to be truly influential. An outlier in $y$ with low leverage barely moves the fit. A high-leverage point that falls on the regression line has a small residual and low Cook's distance despite its extreme position.
Rule of thumb. $D_{k} > 4/n$ or $D_{k} > 1$ are common thresholds for flagging influential observations, though these are guidelines rather than formal tests.
Connection to EVOIR. Cook's distance normalizes the RVOI by $p \, S_{n}^2$ (a variance estimate). Compare with the F-statistic case where EVOIR = $\frac{n-p-2}{n-p} \cdot F$: both are ratios of a realized squared change to a variance scale. Cook's distance is the observation-level analogue of the F-statistic: F measures parameter-group influence, $D_{k}$ measures observation influence, and both are scaled RVOIs under the VoI framework.
End of expanded note.
How to choose the Cook's distance threshold in practice
Two common rules of thumb exist: $D_{k} > 1$ (conservative, flags only extreme cases) and $D_{k} > 4/n$ (liberal, flags roughly the most influential observations in moderate-sized samples). Neither is a formal significance test. Which to use depends on your goal.
Start with $4/n$, then inspect. The $4/n$ rule corresponds roughly to the 50th percentile of the $F_{p, n-p}$ distribution, so it flags points that shift the fit by more than the "median expected shift." Use it as a screening step, not a deletion criterion. Plot $D_{k}$ against observation index and look for points that separate from the bulk, rather than relying on a fixed cutoff.
Leverage deserves its own check. Cook's distance conflates leverage ($h_{kk}$) and residual size. A point can have moderate $D_{k}$ because of high leverage with a small residual (it is pulling the line toward itself). Check leverage separately: the standard threshold is $h_{kk} > 2p/n$. In R, hatvalues(model) returns leverages; in Python, OLSResults.get_influence().hat_matrix_diag does the same. Half-normal plots of leverages (R: faraway::halfnorm(hatvalues(model))) are more informative than a fixed cutoff because they reveal the distribution's shape and make outliers visually obvious.
Software. In R: cooks.distance(model) or plot(model, which=4) for the index plot. In Python: OLSResults.get_influence().cooks_distance returns both $D_{k}$ values and p-values based on the $F_{p, n-p}$ distribution. Always examine flagged points substantively (measurement error, different subpopulation, data entry mistake) before deciding whether to remove, downweight, or keep them.
End of expanded note.
Properties of PVOI
The prospective value of information satisfies four fundamental properties under two assumptions: (A1) Bayes estimators are used as decision rules, and (A2) the Bayes estimator is unique under the given loss function.
Non-negativity. For any information $I$, $\text{PVOI}(I \mid Y) \geq 0$.
Uniqueness. Under A2, $\text{PVOI}(I \mid Y) = 0$ if and only if $I$ is empty (contains no information).
Monotonicity. If $I_{2} \subset I_{1}$, then $\text{PVOI}(I_{2} \mid Y) \leq \text{PVOI}(I_{1} \mid Y)$. More information is always at least as valuable.
Additivity. If $I_{2} \subset I_{1}$ and $I_{1-2} = I_{1} \setminus I_{2}$, then
$$ \text{PVOI}(I_1 \mid Y) = \text{PVOI}(I_2 \mid Y) + E_{I_2 \mid Y}\, \text{PVOI}(I_{1-2} \mid I_2, Y). $$
The value of the full information set equals the value of a subset plus the expected conditional value of the remainder. This decomposition is symmetric: you can also write it as
$$ \text{PVOI}(I_1 \mid Y) = \text{PVOI}(I_{1-2} \mid Y) + E_{I_{1-2} \mid Y}\, \text{PVOI}(I_2 \mid I_{1-2}, Y). $$
Proofs of PVOI properties
Non-negativity. Since $d_{Y,I}$ is a Bayes estimator, by definition $E_{\theta\lvert Y,I} L(d_{Y,I}, \theta) \leq E_{\theta\rvertY,I} L(d_{Y}, \theta)$ for all $I$. Hence
$$\text{PVOI}(I \mid Y) = E_{I|Y}\, E_{\theta|Y,I}\{L(d_Y, \theta) - L(d_{Y,I}, \theta)\} \geq 0.$$Uniqueness. If $\text{PVOI}(I \mid Y) = 0$, then by non-negativity, $E_{\theta\lvert Y,I}\{L(d_{Y}, \theta) - L(d_{Y,I}, \theta)\} = 0$ for almost every $I \mid Y$. By uniqueness of the Bayes estimator (A2), $d_{Y,I} = d_{Y}$ for almost every $I \mid Y$, which means $I$ provides no information.
Monotonicity. For $I_{2} \subset I_{1}$ with $I_{1-2} = I_{1} \setminus I_{2}$:
$$\text{PVOI}(I_1|Y) - \text{PVOI}(I_2|Y) = E_{\theta,I_1|Y}\{L(d_{Y,I_2}, \theta) - L(d_{Y,I_1}, \theta)\}.$$Since $I_{2} \subset I_{1}$, we can write $d_{Y,I_{1}} = d_{Y,I_{2},I_{1-2}}$ and condition on $I_{2}$:
$$= E_{I_2|Y}\, E_{\theta,I_{1-2}|Y,I_2}\{L(d_{Y,I_2}, \theta) - L(d_{Y,I_2,I_{1-2}}, \theta)\} = E_{I_2|Y}\, \text{PVOI}(I_{1-2} \mid Y, I_2) \geq 0.$$Additivity. The monotonicity proof directly gives the additivity result:
$$\text{PVOI}(I_1|Y) = \text{PVOI}(I_2|Y) + E_{I_2|Y}\, \text{PVOI}(I_{1-2} \mid I_2, Y).$$By symmetry (swapping the roles of $I_{2}$ and $I_{1-2}$), the same argument gives:
$$\text{PVOI}(I_1|Y) = \text{PVOI}(I_{1-2}|Y) + E_{I_{1-2}|Y}\, \text{PVOI}(I_2 \mid I_{1-2}, Y).$$End of expanded note.
RVOI Additivity
There is a corresponding additivity result for retrospective value. If $I_{2} \subset I_{1}$ and $I_{1-2} = I_{1} \setminus I_{2}$:
$$ E_{I_{1-2}|Y,I_2}\, \text{RVOI}(I_1 \mid Y, I_1) = \text{RVOI}(I_2 \mid Y, I_2) + \text{PVOI}(I_{1-2} \mid Y, I_2). $$
The expected retrospective value of the full set, given a revealed subset, equals the retrospective value of that subset plus the prospective value of the remainder. This bridges the retrospective and prospective perspectives.
Proof of RVOI additivity
Starting from
$$E_{I_{1-2}|Y,I_2}\, \text{RVOI}(I_1|Y,I_1) - \text{PVOI}(I_{1-2}|Y,I_2),$$expanding both terms:
$$= E_{\theta,I_{1-2}|Y,I_2}\{L(d_Y,\theta) - L(d_{Y,I_1},\theta)\} - E_{\theta,I_{1-2}|Y,I_2}\{L(d_{Y,I_2},\theta) - L(d_{Y,I_2,I_{1-2}},\theta)\}.$$Since $d_{Y,I_{1}} = d_{Y,I_{2},I_{1-2}}$, the terms involving $d_{Y,I_{1}}$ cancel:
$$= E_{\theta,I_{1-2}|Y,I_2}\{L(d_Y,\theta) - L(d_{Y,I_2},\theta)\} = E_{\theta|Y,I_2}\{L(d_Y,\theta) - L(d_{Y,I_2},\theta)\} = \text{RVOI}(I_2|Y,I_2).$$End of expanded note.
Connection 1: Variance-Based Sensitivity Analysis
Variance-based sensitivity analysis (Sobol indices) asks: how much of the output variance is attributable to each input variable? The VoI framework provides a decision-theoretic interpretation of these indices.
Sobol Indices
The first-order variance (main effect) of input $X_{i}$ is
$$ V_i = \text{Var}_{X_i}[E_{X_{\sim i}}(y \mid X_i)], $$
and the total-effect variance is
$$ V_{T_i} = E_{X_{\sim i}}\, \text{Var}_{X_i}[y \mid X_{\sim i}]. $$
Their normalized versions are the first-order and total-effect sensitivity indices:
$$ S_i = \frac{V_i}{\text{Var}[y]}, \quad S_{T_i} = \frac{V_{T_i}}{\text{Var}[y]}. $$
$S_{i}$ measures the main effect of $X_{i}$ alone. $S_{T_{i}}$ measures the total contribution of $X_{i}$ including all interactions with other variables.
Foundations: law of total variance and ANOVA-like decomposition
Variance-based sensitivity analysis rests on two results. The first is the law of total variance:
$$\text{Var}[y] = \text{Var}[E[y \mid X_i]] + E[\text{Var}[y \mid X_i]].$$The first term is $V_{i}$ (variance explained by $X_{i}$); the second is the residual variance.
The second is the high-dimensional model representation (HDMR): any function $y = \eta(x_{1}, \ldots, x_{n})$ can be decomposed as
$$y = E[y] + \sum_i z_i(x_i) + \sum_{i < j} z_{ij}(x_i, x_j) + \cdots,$$where
$$z_i(x_i) = E[y \mid x_i] - E[y], \quad z_{ij}(x_i, x_j) = E[y \mid x_i, x_j] - E[y \mid x_i] - E[y \mid x_j] + E[y].$$This gives $2^n - 1$ terms, which is why the total-effect index $S_{T_{i}}$ is useful: it summarizes all interactions involving $X_{i}$ without enumerating them.
Special case. When $y = a + bx_{1} + cx_{2}$ and $x_{1}, x_{2}$ are independent and centered:
$$S_i = \frac{b^2 \text{Var}[x_1]}{\text{Var}[y]} = \rho_{y,x_1}^2 = S_{T_i}.$$The first-order and total-effect indices coincide because there are no interactions in a linear model with independent inputs.
End of expanded note.
VoI Interpretation of Sensitivity Indices
Under squared loss with decision $\hat{d}$, the Sobol indices have direct VoI interpretations:
First-order index as PVOI:
$$ \text{PVOI}(X_i \mid \text{Prior}) = E_{X_i}(d_{\text{prior}} - d_{X_i})^2 = \text{Var}_{X_i}[E_{X_{\sim i}}[d_X \mid X_i]] \approx V_i. $$
Total-effect index as conditional PVOI:
$$ E_{X_{\sim i}}\, \text{PVOI}(X_i \mid X_{\sim i}) = E_{X_{\sim i}}[\text{Var}_{X_i}[d_X \mid X_{\sim i}]] \approx V_{T_i}. $$
The first-order Sobol index measures how much learning $X_{i}$ alone reduces expected loss. The total-effect index measures the average reduction from learning $X_{i}$ after all other variables are already known. The VoI framework generalizes these beyond squared loss to any loss function, which is useful when the parameter of interest is highly skewed or multimodal.
Connection 2: Estimation Stability With Cross-Validation (ESCV)
The ESCV criterion for choosing a regularization parameter $\lambda$ is
$$ \text{ESCV}(\lambda) = \frac{\hat{\text{var}}(\hat{Y}_{[\lambda]})}{||\bar{\hat{Y}}_{[\lambda]}||_2^2} = \frac{\frac{1}{V}\sum_{k=1}^{V} ||\hat{Y}_{[-k,\lambda]} - \bar{\hat{Y}}_{[\lambda]}||_2^2}{||\bar{\hat{Y}}_{[\lambda]}||_2^2}, $$
where $\hat{Y}_{[-k,\lambda]}$ is the prediction from the model trained without fold $k$, and $\bar{\hat{Y}}_{[\lambda]} = \frac{1}{V}\sum_{k=1}^{V} \hat{Y}_{[-k,\lambda]}$.
In the VoI framework, each fold removal is an information perturbation. The numerator is an averaged RVOI:
$$ ||\hat{Y}_{[-k,\lambda]} - \bar{\hat{Y}}_{[\lambda]}||_2^2 \approx \text{RVOI}((\mathbf{X}, \mathbf{y})_{[k,\lambda]} \mid (\mathbf{X}, \mathbf{y})_{[\lambda]}), $$
and the variance estimate connects RVOI to PVOI:
$$ \hat{\text{var}}(\hat{Y}_{[\lambda]}) \approx E_{(\mathbf{X},\mathbf{y})_{[-k,\lambda]}}\, \text{PVOI}((\mathbf{X},\mathbf{y})_{[k,\lambda]} \mid (\mathbf{X},\mathbf{y})_{[-k,\lambda]}). $$
Derivation: ESCV variance as expected PVOI
Starting from the variance estimate:
$$\hat{\text{var}}(\hat{Y}_{[\lambda]}) = \frac{1}{V}\sum_{k=1}^{V} ||\hat{Y}_{[-k,\lambda]} - \bar{\hat{Y}}_{[\lambda]}||_2^2 \approx E_{(\mathbf{X},\mathbf{y})_{[\lambda]}}\, \text{RVOI}((\mathbf{X},\mathbf{y})_{[k,\lambda]} \mid (\mathbf{X},\mathbf{y})_{[\lambda]}).$$By the law of total expectation, conditioning on the leave-one-fold-out data:
$$E_{(\mathbf{X},\mathbf{y})_{[\lambda]}}\, \text{RVOI} = E_{(\mathbf{X},\mathbf{y})_{[-k,\lambda]}}\, E_{(\mathbf{X},\mathbf{y})_{[k,\lambda]} | (\mathbf{X},\mathbf{y})_{[-k,\lambda]}}\, \text{RVOI} = E_{(\mathbf{X},\mathbf{y})_{[-k,\lambda]}}\, \text{PVOI}((\mathbf{X},\mathbf{y})_{[k,\lambda]} \mid (\mathbf{X},\mathbf{y})_{[-k,\lambda]}).$$The ESCV variance is the expected prospective value of each fold's data, averaged across leave-one-out configurations. A large ESCV indicates that the model's predictions are sensitive to which data points are included, and the regularization parameter $\lambda$ should be chosen to reduce this sensitivity.
End of expanded note.
ESCV selects $\lambda$ by minimizing prediction instability. Viewed through VoI, it chooses the regularization that minimizes the expected influence of any single fold on the predictions. Stable models have low PVOI for each fold; unstable models have high PVOI.
How to choose the regularization parameter with ESCV in practice
When to prefer ESCV over standard CV. Standard cross-validation (minimizing prediction error) can select models that are accurate on average but unstable: small changes in the training set produce large swings in predictions. ESCV penalizes this instability. Use ESCV when prediction stability matters more than raw predictive accuracy, for example in scientific inference (where you want robust coefficient estimates) or in high-dimensional settings where the CV error curve is flat and multiple $\lambda$ values give similar accuracy.
Number of folds. ESCV uses the same fold structure as standard $V$-fold CV. $V = 5$ or $V = 10$ are typical defaults. Larger $V$ gives a less biased variance estimate but increases computation. Leave-one-out ($V = n$) is the most stable estimator of ESCV but costs $n$ model fits.
Key difference from standard CV. In standard CV, each model predicts only its held-out fold, producing a single $n$-vector of out-of-sample predictions. In ESCV, each model predicts all $n$ points. With $V = 5$ folds, you get 5 full $n$-dimensional prediction vectors $\hat{Y}_{[-1,\lambda]}, \dots, \hat{Y}_{[-5,\lambda]}$, and ESCV measures pointwise variance across these 5 vectors. Some predictions are in-sample (the point was in training) and some are out-of-sample (the point was held out), but that is fine because ESCV measures stability (do different training sets give similar predictions?), not prediction accuracy.
Implementation. ESCV is not built into standard packages, but it is straightforward to compute from any CV loop. Fit the model on each of the $V$ training folds, predict all $n$ points from each model, compute their pointwise variance, and normalize by the squared norm of their mean. In R, cv.glmnet(..., keep=TRUE)$fit.preval stores the full prediction matrix (all $n$ points for each fold's model). In Python, fit each fold's model manually and call predict on the full dataset; sklearn.model_selection.cross_val_predict only returns out-of-fold predictions, so it does not directly give the ESCV prediction matrix. Select the $\lambda$ that minimizes ESCV, or use the "one-standard-error" heuristic applied to the ESCV curve instead of the CV error curve.</p>
Combining stability with accuracy. Pure ESCV can over-regularize (very stable but poor predictions). Two practical approaches:
- ESCV as tiebreaker (recommended). Find all $\lambda$ within one standard error of the minimum CV error, then among those pick the $\lambda$ with minimum ESCV. This is the stability-aware version of the classic 1-SE rule, requires no extra tuning, and works best when the CV error curve is flat (common in high-dimensional settings) where multiple $\lambda$ values give similar accuracy and stability should decide.
- Combined criterion. Minimize $\text{CV}_{\text{error}}(\lambda) + \alpha \cdot \text{ESCV}(\lambda)$, where $\alpha$ controls the accuracy-stability tradeoff. Set $\alpha$ by normalizing both terms to the same scale (e.g., divide each by its range across the $\lambda$ grid) and starting with $\alpha = 1$ (equal weight). Increase $\alpha$ when stability matters more (scientific inference, regulatory settings); decrease when prediction accuracy dominates (forecasting, competitions).
ESCV is model-agnostic. The formula uses only prediction vectors, not model internals. It works for any model with a tuning parameter: LASSO, ridge, random forest, gradient boosting, neural networks. It also doubles as an outlier detector: the per-point contribution $\text{ESCV}_{i} = \frac{1}{V}\sum_{k=1}^V (\hat{f}_{-k}(x_{i}) - \bar{f}(x_{i}))^2$ flags points that drive instability. Points with high $\text{ESCV}_{i}$ are the ones where removing a fold changes the prediction most, which is exactly the PVOI interpretation. This avoids masking because each fold's model excludes that fold's points.
End of expanded note.
</details> ## Connection 3: High-Dimensional Influence Measure In high-dimensional linear models, the classical influence measure for observation $k$ is $$ D_k = \frac{1}{p} \sum_{j=1}^{p} (\hat{\rho}_j - \hat{\rho}_j^{(k)})^2, $$ where $\hat{\rho}_{j}$ is the marginal correlation between feature $j$ and $y$, and $\hat{\rho}_{j}^{(k)}$ is the same quantity with observation $k$ removed. In the VoI framework, this is exactly an RVOI when the decision is the vector of marginal correlations under squared loss: $$ \text{RVOI}((\mathbf{X}, \mathbf{y})_{[k]} \mid (\mathbf{X}, \mathbf{y})) = \frac{1}{p} \sum_{j=1}^{p} (\hat{\rho}_j - \hat{\rho}_j^{(k)})^2 = D_k. $$ Under regularity conditions, when there is no influential point and $\min\{n, p\} \to \infty$: $$ n^2 D_k \to \chi^2(1). $$ This provides a formal test: compute $n^2 D_{k}$ for each observation and obtain a p-value for the hypothesis that observation $k$ is not influential.How to choose the threshold for the high-dimensional influence test in practice
Use the $\chi^2(1)$ p-value with a multiple testing correction. Compute $n^2 D_{k}$ for each observation and compare to $\chi^2_{1}$ quantiles. With $n$ simultaneous tests, apply a Bonferroni correction (reject if p-value $< \alpha / n$) or, less conservatively, control the false discovery rate with the Benjamini-Hochberg procedure. At $\alpha = 0.05$ with Bonferroni, the effective threshold on $n^2 D_{k}$ is $\chi^2_{1}(1 - 0.05/n)$, which is approximately $(\log n)^2$ for large $n$.
Practical considerations. The asymptotic $\chi^2(1)$ result requires $\min\{n, p\} \to \infty$ and assumes no truly influential point exists under the null. In finite samples, the approximation can be liberal (too many false positives) when $n/p$ is close to 1. Simulate from the fitted model to calibrate the null distribution if you are in this regime. Compute $D_{k}$ for all observations and plot the sorted $n^2 D_{k}$ values against $\chi^2(1)$ quantiles (a Q-Q plot). Points that deviate sharply from the diagonal are candidate influential observations.
Software. No standard R or Python package implements this specific test. The computation itself is simple: for each observation $k$, recompute all $p$ marginal correlations $\hat{\rho}_{j}^{(k)}$ by the leave-one-out update formula $\hat{\rho}_{j}^{(k)} = (n \hat{\rho}_{j} - x_{kj} y_{k} / s_{xj} s_{y}) / (n-1)$ (with appropriate rescaling), then sum the squared differences. Vectorized over $k$, this costs $\mathcal{O}(np)$ total, which is feasible even when $p$ is large.
End of expanded note.
Derivation, computational tricks, and limitations of influence functions
Where the formula comes from. Consider upweighting training point $z$ by a small amount $\epsilon$. The perturbed parameters are:
$$\hat{\theta}_{\epsilon,z} = \arg\min_\theta \frac{1}{n} \sum_{i=1}^{n} L(z_i, \theta) + \epsilon \, L(z, \theta).$$By a first-order Taylor expansion around $\epsilon = 0$:
$$\hat{\theta}_{\epsilon,z} \approx \hat{\theta} - \epsilon \, H_{\hat{\theta}}^{-1} \nabla_\theta L(z, \hat{\theta}).$$The change in test loss is:
$$L(z_{\text{test}}, \hat{\theta}_{\epsilon,z}) - L(z_{\text{test}}, \hat{\theta}) \approx -\epsilon \, \nabla_\theta L(z_{\text{test}}, \hat{\theta})^\top H_{\hat{\theta}}^{-1} \nabla_\theta L(z, \hat{\theta}).$$Dividing by $\epsilon$ gives the influence function.
Computational challenge. The Hessian $H_{\hat{\theta}}$ is a $p \times p$ matrix. For models with millions or billions of parameters, explicitly forming and inverting it is impossible ($\mathcal{O}(p^3)$ cost). Instead, the key quantity $s_{\text{test}} = H_{\hat{\theta}}^{-1} \nabla_\theta L(z_{\text{test}}, \hat{\theta})$ is computed by solving the linear system $H_{\hat{\theta}} \mathbf{x} = \nabla_\theta L(z_{\text{test}}, \hat{\theta})$ using conjugate gradient (CG) iterations. Each CG step requires one Hessian-vector product $H_{\hat{\theta}} \mathbf{v}$, computable via automatic differentiation in $\mathcal{O}(p)$ time without forming $H$ explicitly.
Assumptions and limitations. The influence function assumes: (1) local convexity around $\hat{\theta}$ (the Hessian is positive definite), (2) the perturbation $\epsilon$ is small enough for the Taylor approximation to hold, and (3) the model is trained to convergence. Deep learning models are non-convex globally, but the local curvature around a converged $\hat{\theta}$ is often well-behaved enough for the approximation to be useful.
Scalability concerns. Even with HVP tricks, influence functions remain expensive for very large models (GPT-scale). Follow-up methods include SGD-based influence tracing (retrace the training trajectory) and representer points (decompose predictions into linear combinations of training point activations).
End of expanded note.
Computation, approximations, and comparison with influence functions
Game-theoretic intuition. Think of each player joining a coalition one at a time, in a random order. The Shapley value is the average marginal contribution across all $n!$ orderings. This is equivalent to the subset formula but sometimes easier to reason about.
Computational challenge. The exact formula requires evaluating $U$ on $2^n$ subsets (Data Shapley) or $2^d$ subsets (SHAP), which is infeasible for large $n$ or $d$. Approximation strategies include:
- Monte Carlo permutation sampling. Sample random orderings and average marginal contributions. Works for both data and feature Shapley.
- KernelSHAP. Approximates feature Shapley values using a weighted linear regression on binary coalition vectors. Efficient for moderate $d$.
- TreeSHAP. Exact Shapley values for tree-based models in $\mathcal{O}(TLD^2)$ time, where $T$ is the number of trees, $L$ is the maximum number of leaves, and $D$ is the maximum depth. The key insight is that a decision tree already partitions the feature space into regions, so the combinatorial sum over feature subsets can be computed by a single recursive pass through the tree. For a given input $x$, TreeSHAP pushes $x$ down the tree while tracking, at each internal node, two quantities: (1) the proportion of feature subsets $S$ that include the split feature (in which case the split is applied using $x$'s value), and (2) the proportion that exclude it (in which case both branches are followed, weighted by the training data proportion going left vs right). At each leaf, the algorithm has accumulated the exact weight that each feature subset assigns to that leaf's prediction. These weights are aggregated into Shapley values by a polynomial-time recursion that avoids enumerating all $2^d$ subsets explicitly. For an ensemble of $T$ trees (e.g., XGBoost, LightGBM, random forest), the Shapley values from each tree are simply averaged (by the linearity axiom). This makes TreeSHAP fast enough for production use on large ensembles with hundreds of trees.
- GradientSHAP. Uses model gradients to approximate feature Shapley values, bridging gradient-based and game-theoretic methods.
- KNN-Shapley. Exact data Shapley values for KNN models in $\mathcal{O}(n \log n)$ time.
Data Shapley vs influence functions. Both measure data-point influence, but from different perspectives:
- Influence functions measure local, infinitesimal perturbation: "if I slightly upweight $z$, how does the test loss change?" This is a first-order approximation around the current model.
- Data Shapley measures global contribution: "across all possible training subsets, what is $z$'s average marginal contribution?" This considers the effect of $z$ at all possible model states.
Influence functions are cheaper (one Hessian-vector solve) but rely on local approximations. Data Shapley is more principled (axiomatic guarantees) but requires many model retrains. For large-scale deep learning, neither is cheap, but influence functions are more practical.
End of expanded note.
How the three frameworks relate to each other
Gradient-based and decision-theoretic. Cook's distance is fundamentally a decision-theoretic quantity: it is the exact RVOI (squared prediction change from removing an observation) under squared loss, computed via the closed-form Sherman-Morrison update of OLS. No gradient approximation is needed. The influence function is a gradient-based method that approximates this same quantity for general models via a first-order Taylor expansion. For linear regression, the approximation happens to be exact because the loss is quadratic (the Taylor expansion has no higher-order terms). So Cook's distance belongs natively to the VoI framework, and the influence function recovers it as a special case only in the linear setting. For non-linear models, RVOI is the exact quantity and the influence function is the practical approximation.
Game-theoretic and decision-theoretic. The PVOI additivity property ($\text{PVOI}(I_{1}\lvert Y) = \text{PVOI}(I_{2}\rvertY) + E_{I_{2}\lvert Y}\text{PVOI}(I_{1-2}\rvertI_{2},Y)$) mirrors the Shapley value's marginal contribution structure. The Shapley value can be viewed as a specific way of averaging conditional PVOIs across all possible conditioning sets, with combinatorial weights ensuring fairness. Both ask "how much does this information contribute?" but VoI uses a Bayesian criterion while Shapley uses axiomatic game theory.
Gradient-based and game-theoretic. GradientSHAP (a variant of SHAP) uses gradients to approximate Shapley values for features, blending the two traditions. For data influence, influence functions give a local approximation of what Data Shapley measures globally: the effect of a single data point on the model.
Sobol indices bridge all three. First-order Sobol indices equal PVOI for features under squared loss (decision-theoretic), can be estimated via gradient-based methods, and relate to the ANOVA decomposition that underlies balanced Shapley values. They sit at the intersection of all three frameworks.
The unified template. All influence methods fit a common structure:
- Define what information $\phi$ you care about (parameter, feature, data point, new sample).
- Define a measure of model quality (loss, utility, variance).
- Compare the model with and without $\phi$.
- Aggregate across uncertainty in $\phi$ (expectation, averaging over subsets, or conditioning).
The three frameworks differ in step 4: gradient-based methods use local Taylor expansion, Shapley values average over all coalitions, and VoI takes a Bayesian expectation. The choice depends on whether you need computational efficiency (gradients), axiomatic fairness (Shapley), or decision-theoretic coherence (VoI).
End of expanded note.