# Multiple Outputs

**Contents**

> * [Training with One-Model-Per-Target](#training-with-one-model-per-target)
> * [Training with Vector Leaf](#training-with-vector-leaf)
> * [Using Reduced Gradient (Sketch Boost)](#using-reduced-gradient-sketch-boost)
> * [Brief History](#brief-history)
> * [References](#references)

#### Versionadded
Added in version 1.6.

Starting from version 1.6, XGBoost has experimental support for multi-output regression
and multi-label classification with Python package.  Multi-label classification usually
refers to targets that have multiple non-exclusive class labels.  For instance, a movie
can be simultaneously classified as both sci-fi and comedy.  For detailed explanation of
terminologies related to different multi-output models please refer to the
[scikit-learn user guide](https://scikit-learn.org/stable/modules/multiclass.html).

#### NOTE
As of XGBoost 3.4.0, the feature is experimental.

The `hist` tree method is feature complete. In addition, the Python interface along with the Dask interface are supported. See the end of the doc for a brief history.

## Training with One-Model-Per-Target

By default, XGBoost builds one model for each target similar to sklearn meta estimators,
with the added benefit of reusing data and other integrated features like SHAP.  For a
worked example of regression, see
[A demo for multi-output regression](../python/examples/multioutput_regression.html.md#sphx-glr-python-examples-multioutput-regression-py). For multi-label classification,
the binary relevance strategy is used.  Input `y` should be of shape `(n_samples,
n_classes)` with each column having a value of 0 or 1 to specify whether the sample is
labeled as positive for respective class. Given a sample with 3 output classes and 2
labels, the corresponding y should be encoded as `[1, 0, 1]` with the second class
labeled as negative and the rest labeled as positive. At the moment XGBoost supports only
dense matrix for labels.

```python
from sklearn.datasets import make_multilabel_classification
import numpy as np

X, y = make_multilabel_classification(
    n_samples=32, n_classes=5, n_labels=3, random_state=0
)
clf = xgb.XGBClassifier(tree_method="hist")
clf.fit(X, y)
np.testing.assert_allclose(clf.predict(X), y)
```

The feature is still under development with limited support from objectives and metrics.

## Training with Vector Leaf

#### Versionadded
Added in version 2.0.0.

XGBoost can optionally build multi-output trees with the size of leaf equals to the number
of targets when the tree method hist is used. The behavior can be controlled by the
`multi_strategy` training parameter, which can take the value one_output_per_tree (the
default) for building one model per-target or multi_output_tree for building
multi-output trees.

```python
clf = xgb.XGBClassifier(tree_method="hist", multi_strategy="multi_output_tree")
```

See [A demo for multi-output regression](../python/examples/multioutput_regression.html.md#sphx-glr-python-examples-multioutput-regression-py) for a worked example with
regression.

## Using Reduced Gradient (Sketch Boost)

#### Versionadded
Added in version 3.2.0.

#### NOTE
This is experimental. It is documented here for early testers to provide feedback. Related
interface might change without notice.

When the number of targets is large, training a gradient boosting tree model using the
full gradient matrix becomes challenging. The training procedure may run out of memory for
storing the histogram, or run extremely slowly due to the amount of computation needed. As
an optimization, XGBoost implements an interface for using two types of gradients based on
the concepts from Sketch Boost [[1]](#references).

The key insight is that we can use different gradients for two distinct purposes:

- **Split gradient**: A reduced-dimension gradient used to determine the tree structure.
- **Value gradient**: The full gradient used to calculate the final leaf values for
  accurate predictions.

This separation allows the expensive histogram building and split finding to operate on a
smaller gradient matrix, while still producing valid predictions using the full loss
function for leaf values. The Sketch Boost paper proposes using dimensionality reduction
on the gradient matrix. In practice, one can also define a different but related loss with
a small gradient matrix for finding the tree structure.

To access this feature, create a custom objective that inherits from `TreeObjective` and
implement the `split_grad` method.

```python
from xgboost.objective import TreeObjective
from cuml.decomposition import TruncatedSVD

import cupy as cp

class LsObj(TreeObjective):
    def __call__(self, iteration: int, y_pred, dtrain):
        """Least squared error."""
        y_true = dtrain.get_label()
        grad = y_pred - y_true
        hess = cp.ones(grad.shape)
        return cp.array(grad), cp.array(hess)

    def split_grad(self, iteration: int, grad, hess):
        svd_params = {"algorithm": "jacobi", "n_components": 2, "n_iter": 8}
        svd = TruncatedSVD(output_type="cupy", **svd_params)
        svd.fit(grad)
        grad = svd.transform(grad)
        hess = svd.transform(hess)
        hess = cp.clip(hess, 0.01, None)

        return grad, hess
```

See [A demo for multi-output regression using reduced gradient](../python/examples/multioutput_reduced_gradient.html.md#sphx-glr-python-examples-multioutput-reduced-gradient-py) for a complete worked
example. The feature supports only the `multi_strategy=multi_output_tree`.

## Brief History

Some milestones of multi-output support are recorded below:

- XGBoost introduced basic concepts for multi-output in v1.6.0.
- v2.0.0 introduced a prototype for training CPU-based vector-leaf models, which established the vector-leaf model format.
- v3.2.0 saw a major upgrade to the prototype for CUDA implementation, added reduced gradient, vector intercept, along with external memory.
- v3.4.0 completes the `hist` tree method implementation, for both CPU, GPU, distributed training (Dask), and categorical features, along with various metrics, objectives, and model visualization.

## References

[1] Leonid Iosipoi, Anton Vakhrushev. “[Fast Gradient Boosted Decision Tree for Multioutput Problems](https://proceedings.neurips.cc/paper_files/paper/2022/file/a36c3dbe676fa8445715a31a90c66ab3-Paper-Conference.pdf)”. NeurIPS 2022, pp 25422 - 25435.
