October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PCOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk3 min

How to Use `scipy.stats.gaussian_kde` in Python

Learn to fit and evaluate scipy.stats.gaussian_kde, format one- and multidimensional data, compare bandwidth choices, and use density, resampling, and integration methods.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To use Python’s scipy.stats.gaussian_kde, pass it observed sample values, then evaluate the fitted density at the points you want to inspect. For one variable, provide a one-dimensional array; for multiple variables, arrange the data as dimensions by observations. The main modeling choice is bandwidth: SciPy defaults to Scott’s rule, but that is a starting point rather than a universally best setting.

Fit a KDE and evaluate it on a grid

A kernel density estimate (KDE) uses smooth kernels centered on observed samples to estimate a probability density. This minimal example fits a univariate estimate and evaluates it across a grid:

As an Amazon Associate I earn from qualifying purchases.

import numpy as np
from scipy.stats import gaussian_kde

samples = np.array([1.2, 1.5, 1.7, 2.0, 2.4, 2.8])
kde = gaussian_kde(samples)  # Scott's rule is the default

grid = np.linspace(samples.min() - 1, samples.max() + 1, 200)
density = kde(grid)

density contains estimated density values corresponding to the locations in grid. You can also call kde.evaluate(grid); calling the fitted object is a shorthand for evaluating it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Arrange the data correctly

One variable

For univariate observations, pass a one-dimensional array, with one sample per entry, as in the example above.

Multiple variables

For multivariate data, SciPy expects an array shaped (number of dimensions, number of samples). Each row is a variable, and each column is one observation. For two variables measured over N observations, the shape is (2, N), not (N, 2). See the SciPy gaussian_kde API reference for the documented input shape and parameters.

Choose and compare the bandwidth

The bandwidth controls how much the estimate smooths the data. Too much smoothing can hide meaningful modes or local structure; too little can leave a noisy curve. SciPy cautions that its estimator works best for unimodal distributions and that multimodal distributions tend to be oversmoothed. Compare plausible choices against the same data and grid, considering how many features remain visible and how smooth or noisy the result looks.

Built-in rules and custom factors

With bw_method=None, SciPy uses Scott’s rule. The documented choices also include 'scott', 'silverman', a scalar factor, or a callable. Scott’s factor is n**(-1. / (d + 4)), where n is the sample count and d is the number of dimensions. SciPy’s multivariate Silverman factor is (n * (d + 2) / 4.)**(-1. / (d + 4)).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A scalar passed as bw_method is a factor, not a bandwidth in the measurement units of your data. SciPy forms the kernel covariance by multiplying the data covariance by factor**2. Consequently, a scalar changes the smoothing relative to the data covariance rather than specifying a fixed width such as “0.5 units.” The formulas and parameter behavior are described in the API reference.

Compare Scott, Silverman, and a scalar on the same grid

kde = gaussian_kde(samples)  # Scott's rule
scott_density = kde(grid)

kde.set_bandwidth(bw_method="silverman")
silverman_density = kde(grid)

kde.set_bandwidth(bw_method=0.5)  # scalar factor, not data units
custom_density = kde(grid)

Here, 0.5 is an illustrative scalar factor, not a recommended value for every dataset. For a fair visual comparison, plot each returned array against the same grid. SciPy’s set_bandwidth reference documents changing the bandwidth and demonstrates comparing built-in rules with a scalar. The best choice depends on the data and the goal; SciPy notes cross-validation and plug-in approaches as other possible selection methods, without identifying one method as universally best.

Use weights when observations should not count equally

If some observations should contribute more than others, provide sample weights matching the dataset shape. Without weights, samples are equally weighted. SciPy’s documented Scott and Silverman factors use the effective sample count, neff, for unequal weights rather than the raw count n. Check the API reference for the current signature and weight details.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Use the fitted KDE for other tasks

After fitting, choose a method that matches the calculation you need:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • kde(points) or kde.evaluate(points) returns estimated density values at points.
  • kde.logpdf(points) returns log-density values.
  • kde.resample(...) draws samples from the estimated density.
  • kde.integrate_box_1d(low, high) integrates a univariate estimate over an interval.
  • kde.integrate_box(low_bounds, high_bounds) integrates over a rectangular region.
  • kde.integrate_gaussian(mean, cov) integrates the KDE against a multivariate Gaussian; the mean and covariance dimensions must match the KDE.
  • kde.integrate_kde(other) integrates the product of two KDEs. SciPy documents a ValueError if the estimates have different dimensionality.

For details on the product integral’s behavior, see the SciPy integrate_kde reference.

Check the SciPy version behind the API details

The cited class, bandwidth, and product-integral references are for SciPy 1.16.0, 1.18.0, and 1.17.0 respectively. These documentation versions do not establish which version is installed in a particular environment. If an argument or method behaves differently than expected, check your installed SciPy version and consult its matching documentation.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.