outpro computes an out-of-distribution (OOD) distance for new inputs using the predictor representation and variable priorities from a fitted varpro object, or the predictor representation from an rfsrc grow object. The score is subspace aware: departures are measured in selected coordinates, optionally weighted by model-derived variable priority, rather than by an unweighted global distance over all features.

outpro(object,
       newdata,
       neighbor = NULL,
       distancef = "knn",
       reduce = TRUE,
       cutoff = NULL,
       max.rules.tree = 150,
       max.tree = 150,
       knn.chunk.size = 100L,
       newdata.xscale = FALSE)

outpro.null(object,
            nulldata = NULL,
            neighbor = NULL,
            distancef = "knn",
            reduce = TRUE,
            cutoff = NULL,
            max.rules.tree = 150,
            max.tree = 150,
            knn.chunk.size = 100L,
            nulldata.xscale = FALSE)

Arguments

object

A fitted varpro object or an rfsrc object with classes c("rfsrc", "grow").

newdata

New data to score. If omitted, the training design matrix is scored. For varpro objects and newdata.xscale = FALSE, newdata should be supplied on the original data scale and is aligned to the fitted training design by the package's hot-encoding helper. If newdata.xscale = TRUE, newdata is assumed to already be on the fitted varpro x-scale.

neighbor

Number of training neighbors per case. If NULL, the default is min(n/10, 5000), where n is the number of training rows. For distancef = "knn", this is the number of ordinary nearest neighbors used in the selected standardized subspace. For the other distance functions, this is the number of model-derived neighbors requested from varpro.strength.

distancef

Distance function. One of "knn", "prod", "euclidean", "mahalanobis", "manhattan", "minkowski", or "kernel". The default is "knn". The "knn" option uses an ordinary weighted Manhattan nearest-neighbor distance in the selected subspace and does not call varpro.strength. The remaining options use model-derived neighbor frames and then aggregate coordinate-wise deviations across those neighbors.

reduce

Controls variable selection and weighting. If TRUE with a varpro object, variables are selected using model-based variable priority and the threshold cutoff. If TRUE with an rfsrc object, all predictors are used with unit weights. A character vector selects variables by name. A named numeric vector selects variables by name and supplies their weights. Any other value uses all predictors with unit weights.

cutoff

Threshold used with varpro variable-priority z values when reduce = TRUE. If NULL, a default based on the number of fitted predictor columns is used: .79 for moderate dimension and 0 for high dimension. This is not an OOD acceptance threshold; it controls only the reduced scoring subspace.

max.rules.tree

Maximum number of rules per tree used by varpro.strength for model-derived neighbor extraction. Ignored when distancef = "knn".

max.tree

Maximum number of trees used by varpro.strength for model-derived neighbor extraction. Ignored when distancef = "knn".

knn.chunk.size

Positive integer giving the number of scored rows processed at one time by the "knn" distance calculation. Smaller values use less memory; larger values can be faster. Ignored by the non-KNN distance functions.

newdata.xscale

Logical value used only for varpro objects. If FALSE, the default, newdata is treated as raw user data and hot-encoded/aligned to the fitted training design. If TRUE, newdata is assumed to already contain the fitted x-scale columns, such as a matrix created internally from object$x. This option is mainly intended for internal package use, for example when scoring virtual data already constructed on the fitted varpro x-scale.

nulldata

For outpro.null, optional in-distribution reference data used to form a null/reference distance distribution. If omitted, the training design matrix is used.

nulldata.xscale

Logical value used only for varpro objects when nulldata is supplied. It has the same meaning for nulldata as newdata.xscale has for newdata.

Details

For a varpro object, the training design object$x is the canonical fitted x-scale. When the original data include factors, object$x stores the hot-encoded representation used by the fitted forest. Ordinary user calls should normally leave newdata.xscale = FALSE, in which case new data are hot-encoded and aligned to the fitted training columns. Internal code can set newdata.xscale = TRUE when the data to score have already been constructed on the same fitted x-scale.

The default distance function is "knn". It standardizes the selected variables using training means and standard deviations, computes weighted Manhattan distances from each scored input to the training inputs in the selected subspace, and returns the average distance to the neighbor closest training points. When newdata is omitted and the training data are scored against themselves, the exact self match is excluded from the nearest-neighbor set.

The non-KNN distance functions use model-derived reference neighborhoods. For these metrics, outpro calls varpro.strength on the fitted forest, extracts local neighbor frames, computes standardized coordinate-wise absolute deviations from each scored case to its neighbors, and aggregates those deviations by the requested metric. Available metrics are multiplicative product, weighted Euclidean, Mahalanobis, Manhattan, Minkowski, and kernel distance.

For varpro objects with reduce = TRUE, the selected subspace is determined from variable-priority z values using cutoff. If one or fewer variables pass the cutoff, all variables from the priority table are used. A character reduce value bypasses this step and directly specifies the scoring variables. A named numeric reduce value directly specifies both the scoring variables and their relative weights.

All distances are computed after standardizing selected variables with training means and scales. Variables with zero training standard deviation are dropped automatically. Larger distance values mean that the scored case is farther from the in-distribution reference set in the selected subspace.

outpro.null calls outpro on a reference sample and augments the result with the empirical CDF of the reference distances. This is useful for calibrating a distance to a reference quantile or for constructing a support score such as 1 - F_0(distance).

Value

outpro returns a list with components:

  • distance: numeric vector with one OOD distance per scored case.

  • distance.object: ingredients used for distance computation, including

    • score: neighbor frames returned by varpro.strength; NULL for distancef = "knn".

    • neighbor: neighbor count used.

    • oob.bits: indicator of whether scoring was done on training rows or new data.

    • xvar.names: selected variable names after zero standard-deviation removal.

    • xvar.wt: variable weights after internal normalization and squaring.

    • dist.xvar: list of absolute coordinate-difference matrices in standardized units; NULL for distancef = "knn".

    • xorg.scale, xnew.scale: standardized training and scored matrices for the selected variables.

    • means, sds: training means and scales for the selected variables.

    • dropped.zero.sd.variables: variables removed because their training standard deviation was zero.

  • distance.args: list of metric arguments actually used. For distancef = "knn", this includes the effective neighbor count, whether self matches were excluded, and knn.chunk.size.

  • score: the neighbor information returned by varpro.strength; NULL for distancef = "knn".

  • neighbor: neighbor setting used.

  • cutoff: cutoff used for variable-priority reduction.

  • oob.bits: indicator of whether scoring was done on training rows or new data.

  • selected.variables: variables used in scoring after all filters.

  • selected.weights: final selected-variable weights.

  • dropped.zero.sd.variables: variables removed due to zero standard deviation.

  • means, sds: training means and scales for convenience.

  • newdata.xscale: logical value indicating whether supplied newdata was treated as already on the fitted x-scale.

  • call: the matched call.

outpro.null returns the same list with two additional components:

  • cdf: the empirical distribution function of distance.

  • quantile: the empirical cumulative probability for each reference case.

Background

The method follows a model-centered view of OOD detection. The goal is not to estimate a full feature density, but to measure whether an input is well supported in the feature directions used by the fitted prediction rule. The default KNN score provides a fast selected-subspace support distance, while the non-KNN metrics use the forest structure to define model-derived local neighborhoods before aggregating coordinate deviations.

See also

Examples


# \donttest{

## ------------------------------------------------

## fit a varPro model
data(BostonHousing, package = "mlbench")
smp <- sample(1:nrow(BostonHousing), size = nrow(BostonHousing) * .75)
train.data <- BostonHousing[smp,]
test.data <- BostonHousing[-smp,]
vp <- varpro(medv ~ ., data = train.data)

## Score new data with the default KNN metric
op <- outpro(vp, newdata = test.data)
head(op$distance)

## Forest-neighborhood multiplicative distance
op.prod <- outpro(vp, newdata = test.data, distancef = "prod")
head(op.prod$distance)

## Calibrate a null/reference distribution using training data
op.null <- outpro.null(vp)
head(op.null$quantile)

# }