library("xgboost")
library("ParBayesianOptimization")
data(agaricus.train, package = "xgboost")
Folds <- list(
Fold1 = as.integer(seq(1,nrow(agaricus.train$data),by = 3))
, Fold2 = as.integer(seq(2,nrow(agaricus.train$data),by = 3))
, Fold3 = as.integer(seq(3,nrow(agaricus.train$data),by = 3))
)Package Process
Machine learning projects will commonly require a user to “tune” a model’s hyperparameters to find a good balance between bias and variance. Several tools are available in a data scientist’s toolbox to handle this task, the most blunt of which is a grid search. A grid search gauges the model performance over a pre-defined set of hyperparameters without regard for past performance. As models increase in complexity and training time, grid searches become unwieldly.
Idealy, we would use the information from prior model evaluations to guide us in our future parameter searches. This is precisely the idea behind Bayesian Optimization, in which our prior response distribution is iteratively updated based on our best guess of where the best parameters are. The ParBayesianOptimization package does exactly this in the following process:
- Initial parameter-score pairs are found
- Gaussian Process is fit/updated
- Numerical methods are used to estimate the best parameter set
- New parameter-score pairs are found
- Repeat steps 2-4 until some stopping criteria is met
Practical Example
In this example, we will be using the agaricus.train dataset provided in the XGBoost package. Here, we load the packages, data, and create a folds object to be used in the scoring function.
Now we need to define the scoring function. This function should, at a minimum, return a list with a Score element, which is the model evaluation metric we want to maximize. We can also retain other pieces of information created by the scoring function by including them as named elements of the returned list. In this case, we want to retain the optimal number of rounds determined by the xgb.cv:
scoringFunction <- function(max_depth, min_child_weight, subsample) {
# Coerce types explicitly
max_depth <- as.integer(max_depth)[1]
min_child_weight <- as.numeric(min_child_weight)[1]
subsample <- as.numeric(subsample)[1]
dtrain <- xgboost::xgb.DMatrix(
data = agaricus.train$data,
label = agaricus.train$label
)
Pars <- list(
booster = "gbtree",
eta = 0.01,
max_depth = max_depth,
min_child_weight = min_child_weight,
subsample = subsample,
objective = "binary:logistic",
eval_metric = "auc"
)
xgbcv <- xgboost::xgb.cv(
params = Pars,
data = dtrain,
nrounds = 100, # <- canonical argument name
folds = Folds,
prediction = FALSE, # set TRUE only if you actually need CV preds
showsd = TRUE,
early_stopping_rounds = 5,
maximize = TRUE,
verbose = 0
)
# Compute Score robustly (scalar)
score_vec <- xgbcv$evaluation_log$test_auc_mean
Score <- as.numeric(max(score_vec, na.rm = TRUE))[1]
# Derive a scalar "best nrounds" robustly
bi <- xgbcv$best_iteration
if (is.null(bi) || length(bi) != 1 || is.na(bi)) {
# fallback: the iteration where test AUC is maximized
bi <- which.max(score_vec)
if (length(bi) != 1 || is.na(bi)) bi <- 1L
}
BestNrounds <- as.integer(bi)[1]
list(
Score = Score,
BestNrounds = BestNrounds
)
}Some other objects we need to define are the bounds, GP kernel and acquisition function. In this example, the kernel and acquisition function are left as the default.
- The
boundswill tell our process its search space. - The kernel is passed to the
GauProfunctionGauPro_kernel_modeland defines the covariance function. - The acquisition function defines the utility we get from using a certain parameter set.
We are now ready to put this all into the bayesOpt function.
The console informs us that the process initialized by running scoringFunction 4 times. It then fit a Gaussian process to the parameter-score pairs, found the global optimum of the acquisition function, and ran scoringFunction again. This process continued until we had 7 parameter-score pairs. You can interrogate the bayesOpt object to see the results:
optObj$scoreSummary
#> Epoch Iteration max_depth min_child_weight subsample gpUtility acqOptimum inBounds Elapsed Score BestNrounds errorMessage
#> <num> <int> <num> <num> <num> <num> <lgcl> <lgcl> <num> <num> <int> <lgcl>
#> 1: 0 1 9 5.863591 0.2585819 NA FALSE TRUE 0.063 0.9984933 5 NA
#> 2: 0 2 4 10.154185 0.5230172 NA FALSE TRUE 0.068 0.9981565 14 NA
#> 3: 0 3 6 24.487949 0.8622225 NA FALSE TRUE 0.048 0.9957170 5 NA
#> 4: 0 4 2 17.988070 0.6821260 NA FALSE TRUE 0.066 0.9880582 14 NA
#> 5: 1 5 2 2.905824 1.0000000 0.7331915 TRUE TRUE 0.048 0.9871588 8 NA
#> 6: 2 6 9 11.730571 0.2774807 0.6660313 TRUE TRUE 0.065 0.9969579 14 NA
#> 7: 3 7 9 1.000000 0.2500000 0.7402508 TRUE TRUE 0.073 0.9999173 17 NA
getBestPars(optObj)
#> $max_depth
#> [1] 9
#>
#> $min_child_weight
#> [1] 1
#>
#> $subsample
#> [1] 0.25