Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
82 changes: 66 additions & 16 deletions docs/user_guide/outliers/OutlierTrimmer.rst
Original file line number Diff line number Diff line change
Expand Up @@ -125,6 +125,13 @@ and percentile methods stay closer to where the observations actually lie:
they are true outliers or faithful data points. That requires further examination
and domain knowledge.

.. note::

If all or most of the values of a variable are the same, the method may return a
spread of 0 (for example, an IQR of 0 when over half of the values are 0). The
variable then has no outliers, so :class:`OutlierTrimmer()` sets its limits to infinity
in `right_tail_caps_` and `left_tail_caps_`, and leaves it untouched.

Let’s move on to removing outliers in Python.

Removing outliers in Python
Expand Down Expand Up @@ -293,7 +300,7 @@ In the following output, we see the maximum of the variables after removing the
.. code:: python

fare 65.0
age 53.0
age 74.0
dtype: float64

Finally, we can check the boxplot of the transformed variables to corroborate the effect on their distribution.
Expand Down Expand Up @@ -521,7 +528,7 @@ We see the adjusted data size compared to the original size here:

.. code:: python

((916, 8), (736, 76))
((916, 8), (828, 142))

Feature-engine's pipeline can also adjust the target:

Expand All @@ -535,7 +542,7 @@ We see the adjusted data size compared to the original size here:

.. code:: python

((916,), (736,))
((916,), (828,))

To wrap up, let's add a machine learning algorithm to the pipeline. We'll use logistic regression to predict survival:

Expand Down Expand Up @@ -565,7 +572,7 @@ We see the following output:

.. code:: python

array([1, 1, 1, 0, 1, 0, 1, 1, 0, 1], dtype=int64)
array([1, 1, 0, 1, 0, 1, 0, 0, 1, 0])

We can obtain the probability of survival:

Expand All @@ -580,16 +587,16 @@ We see the following output:

.. code:: python

array([[0.13027536, 0.86972464],
[0.14982143, 0.85017857],
[0.2783799 , 0.7216201 ],
[0.86907159, 0.13092841],
[0.31794531, 0.68205469],
[0.86905145, 0.13094855],
[0.1396715 , 0.8603285 ],
[0.48403632, 0.51596368],
[0.6299007 , 0.3700993 ],
[0.49712853, 0.50287147]])
array([[0.23320943, 0.76679057],
[0.22089305, 0.77910695],
[0.85469885, 0.14530115],
[0.28510312, 0.71489688],
[0.85468117, 0.14531883],
[0.0494853 , 0.9505147 ],
[0.58079146, 0.41920854],
[0.536129 , 0.463871 ],
[0.36885157, 0.63114843],
[0.81102131, 0.18897869]])

We can obtain the accuracy of the predictions over the test set:

Expand All @@ -601,7 +608,7 @@ That returns the following accuracy:

.. code:: python

0.7823343848580442
0.804093567251462

We can obtain the names of the features after the transformation:

Expand Down Expand Up @@ -635,7 +642,7 @@ We see the resulting sizes here:

.. code:: python

((393, 8), (317, 76))
((393, 8), (342, 142))


Setting up the stringency (param `fold`)
Expand All @@ -656,6 +663,49 @@ The default values for fold are as follows:
You can manually adjust the fold value to make the outlier detection process more or less
conservative, thus customising the extent of outlier trimming.

With polars
-----------

:class:`OutlierTrimmer()` works in the same way with a polars dataframe:

.. code:: python

import polars as pl
from feature_engine.outliers import OutlierTrimmer

df = pl.DataFrame({
"Age": [20, 21, 19, 18, 95],
"Marks": [0.9, 0.8, 0.7, 0.6, 0.1],
})

transformer = OutlierTrimmer(
capping_method="quantiles",
tail="both",
fold=0.2,
)

print(transformer.fit_transform(df))

Only the rows where both `Age` and `Marks` fall within the 20th-80th
percentile range survive; the other three rows breach the bound on at
least one of the two variables:

.. code:: text

shape: (2, 2)
┌─────┬───────┐
│ Age ┆ Marks │
│ --- ┆ --- │
│ i64 ┆ f64 │
╞═════╪═══════╡
│ 21 ┆ 0.8 │
│ 19 ┆ 0.7 │
└─────┴───────┘

`transform_x_y()` and `get_feature_names_out()` work identically to the
pandas examples above.


Additional resources
--------------------

Expand Down
64 changes: 50 additions & 14 deletions feature_engine/outliers/trimmer.py
Original file line number Diff line number Diff line change
@@ -1,9 +1,10 @@
# Authors: Soledad Galli <solegalli@protonmail.com>
# License: BSD 3 clause

import pandas as pd
import narwhals as nw
import narwhals.dependencies as nwd
from narwhals.typing import IntoDataFrame, IntoSeries

from feature_engine._base_transformers.mixins import TransformXyMixin
from feature_engine._docstrings.fit_attributes import (
_feature_names_in_docstring,
_left_tail_caps_docstring,
Expand All @@ -23,6 +24,7 @@
)
from feature_engine._docstrings.methods import _fit_transform_docstring
from feature_engine._docstrings.substitute import Substitution
from feature_engine.dataframe_checks import check_X_y
from feature_engine.outliers.base_outlier import WinsorizerBase


Expand All @@ -41,7 +43,7 @@
n_features_in_=_n_features_in_docstring,
fit_transform=_fit_transform_docstring,
)
class OutlierTrimmer(WinsorizerBase, TransformXyMixin):
class OutlierTrimmer(WinsorizerBase):
"""The OutlierTrimmer() removes observations with outliers from the dataset.

The OutlierTrimmer() first calculates the maximum and/or minimum values
Expand Down Expand Up @@ -174,29 +176,63 @@ class OutlierTrimmer(WinsorizerBase, TransformXyMixin):
9 0.54256
"""

def transform(self, X: pd.DataFrame) -> pd.DataFrame:
def transform(self, X: IntoDataFrame) -> IntoDataFrame:
"""
Remove observations with outliers from the dataframe.

Parameters
----------
X : pandas dataframe of shape = [n_samples, n_features]
X : dataframe of shape = [n_samples, n_features]
The data to be transformed.

Returns
-------
X_new: pandas dataframe of shape = [n_samples, n_features]
X_new: dataframe of shape = [n_samples, n_features]
The dataframe without outlier observations.
"""
nw_X = self._check_transform_input_and_state(X)
return self._remove_outliers(nw_X).to_native()

X = self._check_transform_input_and_state(X)
def transform_x_y(self, X: IntoDataFrame, y: IntoSeries):
"""
Remove observations with outliers from the dataframe and the target.

Parameters
----------
X: dataframe of shape = [n_samples, n_features]
The dataframe to transform.

y: Series or Dataframe of length = n_samples
The target variable to transform. Can be multi-output.

Returns
-------
X_new: dataframe
The dataframe without outlier observations. It may contain less rows
than the original dataset.

y_new: Series or DataFrame
The target variable, with as many rows as those left in X_new.
"""
_, y = check_X_y(X, y)

row_index = "__row_index__"
nw_X = self._check_transform_input_and_state(X).with_row_index(row_index)
nw_X = self._remove_outliers(nw_X)
rows = nw_X.get_column(row_index).to_list()

if nwd.is_into_series(y):
y = nw.from_native(y, series_only=True)[rows].to_native()
else:
y = nw.from_native(y, eager_only=True)[rows].to_native()

return nw_X.drop(row_index).to_native(), y

for feature in self.right_tail_caps_.keys():
inliers = X[feature].le(self.right_tail_caps_[feature])
X = X.loc[inliers]
def _remove_outliers(self, nw_X: nw.DataFrame) -> nw.DataFrame:
conditions = [nw.col(f) <= c for f, c in self.right_tail_caps_.items()]
conditions += [nw.col(f) >= c for f, c in self.left_tail_caps_.items()]

for feature in self.left_tail_caps_.keys():
inliers = X[feature].ge(self.left_tail_caps_[feature])
X = X.loc[inliers]
if len(conditions) > 0:
nw_X = nw_X.filter(nw.all_horizontal(*conditions, ignore_nulls=False))

return X
return nw_X
Loading