Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Sweetviz is an open-source Python library that turns a pandas DataFrame into a visual exploratory data analysis (EDA) report. With a small amount of code, it summarizes types, distributions, missing values, duplicates, descriptive statistics, associations, target relationships, and differences between datasets or groups. That speed makes it an excellent first-pass audit—not a replacement for validation, domain analysis, leakage checks, or production monitoring.

What Sweetviz does

Sweetviz is designed for analysts, students, and machine-learning practitioners who already have tabular data in pandas. Its main functions are analyze() for one dataset, compare() for two compatible datasets, and compare_intra() for two subgroups within one dataset. Reports can be saved as self-contained HTML or embedded in a notebook.

The project is open source under the MIT license. The package page is available at PyPI.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What “EDA in seconds” means

The slogan describes how quickly you can generate a report, not how quickly meaningful analysis is completed. Sweetviz can expose suspicious columns and relationships almost immediately, but you still need to investigate data quality, business meaning, leakage, causality, and model suitability.

Version and compatibility notes

A version-specific PyPI page exists for Sweetviz 2.3.3, while project metadata also contains older compatibility text and an April 2026 note referring to 2.3.2. Do not assume a page’s number is the latest release or copy the older Python support statement blindly. Check the package index and the version installed in your environment:

python -m pip index versions sweetviz
python -m pip show sweetviz

Current PyPI classifiers begin at Python 3.7, but compatibility with your exact Python, pandas, and NumPy versions should be tested. See the Sweetviz 2.3.3 page for the documented API.

Install Sweetviz in an isolated environment

  1. Create an environment:
    python -m venv .venv
  2. Activate it on macOS or Linux:
    source .venv/bin/activate

    On Windows PowerShell:

    .venvScriptsActivate.ps1
  3. Install the package and pandas:
    python -m pip install -U pip
    python -m pip install sweetviz pandas
  4. Confirm which version and module are being used:
    python -c "import sweetviz as sv; print(sv.__version__)"
    python -c "import sweetviz; print(sweetviz.__file__)"

In Jupyter, install into the active kernel with %pip install sweetviz, then restart the kernel if the import still fails.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Generate your first HTML report

The expanded two-step form keeps the report object available for later display options:

import pandas as pd
import sweetviz as sv

df = pd.read_csv("data.csv")

report = sv.analyze(df)
report.show_html("sweetviz_report.html")

Sweetviz writes sweetviz_report.html. Depending on the environment and display settings, it may open a browser. The HTML is intended to be portable, so you can inspect it separately from the Python process. A compact equivalent is sv.analyze(df).show_html("sweetviz_report.html").

Analyze a target column

For supervised-learning data, pass the real target column name with target_feat:

import pandas as pd
import sweetviz as sv

df = pd.read_csv("titanic.csv")
report = sv.analyze(df, target_feat="Survived")
report.show_html("titanic_target_report.html")

The report organizes feature summaries around how variables differ with the selected target. This is descriptive: it does not prove predictive performance, causation, statistical significance, or the absence of target leakage.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Compare training and test data

train_df = pd.read_csv("train.csv")
test_df = pd.read_csv("test.csv")

report = sv.compare(
    [train_df, "Training Data"],
    [test_df, "Test Data"],
    target_feat="target"
)
report.show_html("train_test_comparison.html")

The comparison can reveal differences in distributions, missingness, unique-value counts, summary statistics, associations, and target behavior where the target is present.

Check the schema first

print(train_df.shape, test_df.shape)
print(train_df.columns.tolist())
print(test_df.columns.tolist())
print(train_df.dtypes)
print(test_df.dtypes)

Resolve missing or extra columns, renamed fields, incompatible dtypes, inconsistent missing-value markers, and targets that appear in only one input. A visual comparison does not prove that a split is valid: it cannot detect every temporal leak, duplicate entity, or contaminated label. Differences may also be intentional after stratified sampling.

Compare two groups in one DataFrame

compare_intra() splits one frame with a Boolean mask. The names identify the true and false sides of that mask:

report = sv.compare_intra(
    df,
    df["gender"] == "male",
    ["Male", "Female"],
    target_feat="target"
)
report.show_html("group_comparison.html")

This pattern works for treatment versus control, churned versus retained customers, converted versus non-converted users, or one region versus the remainder. The result is observational; group differences do not establish that membership caused them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Render reports in browsers, scripts, and notebooks

HTML display controls

report.show_html(
    filepath="report.html",
    open_browser=False,
    layout="vertical",
    scale=0.8
)
  • filepath sets the output path.
  • open_browser=False is appropriate for CI, containers, remote servers, and headless jobs.
  • layout accepts widescreen or vertical.
  • scale changes visual sizing.

Notebook display

report = sv.analyze(df)
report.show_notebook(
    w="100%",
    h=700,
    scale=0.8,
    layout="widescreen"
)

Large reports can overwhelm a notebook cell. Reduce the scale, use a vertical layout, or save HTML instead:

report.show_notebook(
    w="100%",
    h=700,
    scale=0.7,
    layout="vertical"
)
report.show_html("report.html", open_browser=False, layout="vertical")

What appears in a Sweetviz report

Area Examples How to use it
Schema and cardinality Data type, unique values, frequent values Find identifiers, unexpected categories, and type errors.
Quality indicators Missing values and duplicate rows Prioritize cleaning and investigate why values are absent.
Numerical summaries Minimum, maximum, range, quartiles, mean, median, mode, standard deviation, sum, median absolute deviation, coefficient of variation, skewness, and kurtosis Spot scale, spread, skew, and possible outliers.
Relationships Pearson correlation, uncertainty coefficient, and correlation ratio for mixed data types Choose follow-up plots and questions; do not treat scores as proof.
Target and comparisons Feature-versus-target views and side-by-side dataset or subgroup summaries Look for distribution changes and candidate features for investigation.

Pearson correlation can miss nonlinear dependence, and every association measure has assumptions. Use these signals to decide what to examine next, not as a substitute for formal tests or domain knowledge.

Prepare the data before profiling

  • Parse dates and extract meaningful date features rather than profiling raw timestamp strings.
  • Normalize markers such as "N/A" before counting missing values.
  • Convert low-cardinality numeric codes to categorical types when they represent labels, not measurements.
  • Check that Boolean fields encoded as 0/1 have the intended meaning.
  • Review numeric columns accidentally imported as strings.
  • Exclude or separately handle IDs, UUIDs, hashes, raw URLs, log messages, full addresses, and near-unique text fields.
  • Confirm the target’s dtype and values.

Automatic type inference is helpful, but an incorrect schema can make a chart or statistic misleading.

Large datasets, privacy, and interpretation limits

Memory and runtime

Sweetviz profiles pandas objects, so the data normally must already fit in memory. Runtime depends on rows, columns, dtypes, hardware, and report complexity; there is no universal “seconds” guarantee. For very large data, start with a representative sample, remove unnecessary columns, improve inefficient object dtypes where appropriate, and run profiling separately from production pipelines.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the report cannot prove

  • Causality or direction of effect.
  • That an outlier is invalid.
  • That a feature is safe or appropriate for production.
  • That leakage, fairness problems, or label contamination are absent.
  • That a train/test split will remain stable in production.
  • That a model will perform well.

Sharing risk

A self-contained HTML file is convenient, but it can contain personal information, rare categories, free text, internal fields, subgroup differences, and target labels. Inspect the report before emailing, uploading, or publishing it.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot common failures

ModuleNotFoundError: No module named 'sweetviz'

Install with the same interpreter that runs your code:

python -m pip install sweetviz
python -c "import sweetviz; print(sweetviz.__file__)"

In a notebook, use %pip install sweetviz in the active kernel and restart it if necessary.

AttributeError: module 'sweetviz' has no attribute 'analyze'

Check that your script is not named sweetviz.py, which shadows the installed package. Rename it and remove stale .pyc files or the related __pycache__ directory.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Notebook or browser output is unusable

Set open_browser=False, reduce scale, choose vertical, or save the report as HTML. In Docker, SSH, CI, and remote notebook environments, retrieve the generated file through the environment’s normal artifact mechanism.

Non-Latin characters show missing glyphs

Missing-glyph warnings generally indicate a font or rendering limitation rather than corrupted data. Use an environment with fonts containing the required characters.

Sweetviz versus other tools

Tool Best fit Trade-off
Sweetviz Fast local visual EDA on pandas data, target analysis, and train/test or subgroup comparisons Not a governance, causal-analysis, or monitoring system; large or high-cardinality data can be unwieldy.
YData Profiling Broader automated profiling, data-quality diagnostics, and documented pandas and Spark workflows Prefer it when exhaustive profiling matters more than Sweetviz’s compact comparison style.
pandas plus Matplotlib, Seaborn, or Plotly Exact plot control, custom transformations, aggregations, and statistical tests Requires more code and analysis decisions.
Deepchecks Systematic dataset/model validation and production-oriented monitoring A different category from a lightweight local EDA report.

Sweetviz documentation also describes optional Comet.ml integration for logging reports with an API key. Comet is not required for local use; see Comet if centralized experiment history is already part of your workflow.

When Sweetviz is the right choice

  • Your data is already in a pandas DataFrame.
  • You need a quick, shareable HTML overview.
  • Target, train/test, or subgroup comparisons are central to the first investigation.
  • You want a lightweight MIT-licensed library without a paid account.

Choose another approach when you need custom scientific analysis, very large-scale processing, governed data-quality checks, fairness assessment, or continuous production drift monitoring.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The Bottom Line

Bottom line: Sweetviz is a strong first-pass EDA tool for pandas users. Generate the report quickly, verify its inferred schema, investigate its signals with targeted analysis, and treat the HTML artifact as potentially sensitive. It accelerates discovery; it does not replace sound data work.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.