outpro.Rdoutpro computes an out-of-distribution (OOD) distance for new
inputs using the predictor representation and variable priorities from a
fitted varpro object, or the predictor representation from an
rfsrc grow object. The score is subspace aware: departures are
measured in selected coordinates, optionally weighted by model-derived
variable priority, rather than by an unweighted global distance over all
features.
outpro(object,
newdata,
neighbor = NULL,
distancef = "knn",
reduce = TRUE,
cutoff = NULL,
max.rules.tree = 150,
max.tree = 150,
knn.chunk.size = 100L,
newdata.xscale = FALSE)
outpro.null(object,
nulldata = NULL,
neighbor = NULL,
distancef = "knn",
reduce = TRUE,
cutoff = NULL,
max.rules.tree = 150,
max.tree = 150,
knn.chunk.size = 100L,
nulldata.xscale = FALSE)A fitted varpro object or an rfsrc object
with classes c("rfsrc", "grow").
New data to score. If omitted, the training design
matrix is scored. For varpro objects and
newdata.xscale = FALSE, newdata should be supplied on
the original data scale and is aligned to the fitted training design by
the package's hot-encoding helper. If newdata.xscale = TRUE,
newdata is assumed to already be on the fitted varpro
x-scale.
Number of training neighbors per case. If NULL,
the default is min(n/10, 5000), where n is the number of
training rows. For distancef = "knn", this is the number of
ordinary nearest neighbors used in the selected standardized subspace.
For the other distance functions, this is the number of model-derived
neighbors requested from varpro.strength.
Distance function. One of "knn",
"prod", "euclidean", "mahalanobis",
"manhattan", "minkowski", or "kernel". The
default is "knn". The "knn" option uses an ordinary
weighted Manhattan nearest-neighbor distance in the selected subspace
and does not call varpro.strength. The remaining options use
model-derived neighbor frames and then aggregate coordinate-wise
deviations across those neighbors.
Controls variable selection and weighting. If
TRUE with a varpro object, variables are selected using
model-based variable priority and the threshold cutoff. If
TRUE with an rfsrc object, all predictors are used with
unit weights. A character vector selects variables by name. A named
numeric vector selects variables by name and supplies their weights.
Any other value uses all predictors with unit weights.
Threshold used with varpro variable-priority
z values when reduce = TRUE. If NULL, a default
based on the number of fitted predictor columns is used: .79
for moderate dimension and 0 for high dimension. This is not an
OOD acceptance threshold; it controls only the reduced scoring
subspace.
Maximum number of rules per tree used by
varpro.strength for model-derived neighbor extraction. Ignored
when distancef = "knn".
Maximum number of trees used by varpro.strength
for model-derived neighbor extraction. Ignored when
distancef = "knn".
Positive integer giving the number of scored rows
processed at one time by the "knn" distance calculation. Smaller
values use less memory; larger values can be faster. Ignored by the
non-KNN distance functions.
Logical value used only for varpro
objects. If FALSE, the default, newdata is treated as raw
user data and hot-encoded/aligned to the fitted training design. If
TRUE, newdata is assumed to already contain the fitted
x-scale columns, such as a matrix created internally from
object$x. This option is mainly intended for internal package
use, for example when scoring virtual data already constructed on the
fitted varpro x-scale.
For outpro.null, optional in-distribution
reference data used to form a null/reference distance distribution. If
omitted, the training design matrix is used.
Logical value used only for varpro
objects when nulldata is supplied. It has the same meaning for
nulldata as newdata.xscale has for newdata.
For a varpro object, the training design object$x is the
canonical fitted x-scale. When the original data include factors,
object$x stores the hot-encoded representation used by the fitted
forest. Ordinary user calls should normally leave
newdata.xscale = FALSE, in which case new data are hot-encoded and
aligned to the fitted training columns. Internal code can set
newdata.xscale = TRUE when the data to score have already been
constructed on the same fitted x-scale.
The default distance function is "knn". It standardizes the
selected variables using training means and standard deviations, computes
weighted Manhattan distances from each scored input to the training
inputs in the selected subspace, and returns the average distance to the
neighbor closest training points. When newdata is omitted
and the training data are scored against themselves, the exact self match
is excluded from the nearest-neighbor set.
The non-KNN distance functions use model-derived reference neighborhoods.
For these metrics, outpro calls varpro.strength on the
fitted forest, extracts local neighbor frames, computes standardized
coordinate-wise absolute deviations from each scored case to its
neighbors, and aggregates those deviations by the requested metric.
Available metrics are multiplicative product, weighted Euclidean,
Mahalanobis, Manhattan, Minkowski, and kernel distance.
For varpro objects with reduce = TRUE, the selected subspace
is determined from variable-priority z values using cutoff.
If one or fewer variables pass the cutoff, all variables from the
priority table are used. A character reduce value bypasses this
step and directly specifies the scoring variables. A named numeric
reduce value directly specifies both the scoring variables and
their relative weights.
All distances are computed after standardizing selected variables with
training means and scales. Variables with zero training standard
deviation are dropped automatically. Larger distance values mean
that the scored case is farther from the in-distribution reference set in
the selected subspace.
outpro.null calls outpro on a reference sample and augments
the result with the empirical CDF of the reference distances. This is
useful for calibrating a distance to a reference quantile or for
constructing a support score such as 1 - F_0(distance).
outpro returns a list with components:
distance: numeric vector with one OOD distance per scored case.
distance.object: ingredients used for distance computation, including
score: neighbor frames returned by varpro.strength; NULL for distancef = "knn".
neighbor: neighbor count used.
oob.bits: indicator of whether scoring was done on training rows or new data.
xvar.names: selected variable names after zero standard-deviation removal.
xvar.wt: variable weights after internal normalization and squaring.
dist.xvar: list of absolute coordinate-difference matrices in standardized units; NULL for distancef = "knn".
xorg.scale, xnew.scale: standardized training and scored matrices for the selected variables.
means, sds: training means and scales for the selected variables.
dropped.zero.sd.variables: variables removed because their training standard deviation was zero.
distance.args: list of metric arguments actually used. For distancef = "knn", this includes the effective neighbor count, whether self matches were excluded, and knn.chunk.size.
score: the neighbor information returned by varpro.strength; NULL for distancef = "knn".
neighbor: neighbor setting used.
cutoff: cutoff used for variable-priority reduction.
oob.bits: indicator of whether scoring was done on training rows or new data.
selected.variables: variables used in scoring after all filters.
selected.weights: final selected-variable weights.
dropped.zero.sd.variables: variables removed due to zero standard deviation.
means, sds: training means and scales for convenience.
newdata.xscale: logical value indicating whether supplied newdata was treated as already on the fitted x-scale.
call: the matched call.
outpro.null returns the same list with two additional components:
cdf: the empirical distribution function of distance.
quantile: the empirical cumulative probability for each reference case.
The method follows a model-centered view of OOD detection. The goal is not to estimate a full feature density, but to measure whether an input is well supported in the feature directions used by the fitted prediction rule. The default KNN score provides a fast selected-subspace support distance, while the non-KNN metrics use the forest structure to define model-derived local neighborhoods before aggregating coordinate deviations.
# \donttest{
## ------------------------------------------------
## fit a varPro model
data(BostonHousing, package = "mlbench")
smp <- sample(1:nrow(BostonHousing), size = nrow(BostonHousing) * .75)
train.data <- BostonHousing[smp,]
test.data <- BostonHousing[-smp,]
vp <- varpro(medv ~ ., data = train.data)
## Score new data with the default KNN metric
op <- outpro(vp, newdata = test.data)
head(op$distance)
## Forest-neighborhood multiplicative distance
op.prod <- outpro(vp, newdata = test.data, distancef = "prod")
head(op.prod$distance)
## Calibrate a null/reference distribution using training data
op.null <- outpro.null(vp)
head(op.null$quantile)
# }