LightAutoML is an open-source Python library for automated machine learning on tabular and text data. It supports binary and multiclass classification, as well as regression, through ready-made tabular, text, and WhiteBox presets or modular custom pipelines. Pipeline creation can automate data processing and typing, feature selection, hyperparameter tuning, time utilization, and report creation. Documented model classes include linear models, LightGBM and CatBoost boosted trees, neural networks, and a WhiteBox scorecard. Its basic tutorial describes automatic handling of missing values and outliers. Optional installation extras cover NLP, computer vision, and reports, and a tutorial demonstrates a SQL data source. Installation is available from PyPI with `pip install lightautoml`; the project is licensed under Apache License, Version 2.0. The maker describes it as a way for data scientists and analysts to reduce routine data preparation and model selection work. The current package handles datasets with independent samples in each row; multitable datasets and sequences remain a work in progress.
Who it is for
LightAutoML is intended for data scientists and analysts working with tabular or text data. It suits users comfortable installing and working with a Python library.
What is good
- Open-source under Apache License 2.0.
- Supports classification and regression.
- Offers ready-made presets and custom pipelines.
- Can handle missing values and outliers automatically.
- Optional extras cover NLP and computer vision.
What to know first
- Multitable datasets remain a work in progress.
- Sequences remain a work in progress.
- Current package expects independent samples in each row.
Freedom251 review
LightAutoML: the full review
LightAutoML provides automated pipelines and several model types for tabular and text machine-learning tasks. Its current dataset constraint matters if your work depends on multitable data or sequences.
LightAutoML is a free, open-source Python library for automating machine learning on tabular and text data. It is best suited to data scientists and analysts who want to reduce routine pipeline work while keeping options to customize how models are built. It is a capable choice for independent-row datasets, but not for multitable data or sequences.
Overview
The library automates data processing and typing, feature selection, hyperparameter tuning, time utilization, and report creation. That breadth can help teams move through routine preparation and model selection with less manual work. Its focus is narrower than a general-purpose machine-learning platform: the current package expects independent samples in rows, leaving multitable datasets and sequences out of scope.
LightAutoML is installed from PyPI with pip install lightautoml and is licensed under Apache License 2.0. Users can seek advice through Slack or Telegram, and report bugs or request features through GitHub issues. The project makes no security or compliance claims, so buyers with those requirements will need to assess them separately.
Key features
Presets and custom pipelines
Ready-made tabular, text, and WhiteBox presets offer practical starting points; modular custom pipeline creation gives Python users room to adapt workflows when presets are not a fit. This combination is useful for practitioners who want automation without surrendering pipeline choices, but it is not a no-code modeling service.
Model coverage and preparation
Documented model classes span linear models, LightGBM and CatBoost boosted trees, neural networks, and a WhiteBox scorecard model. The project supports binary and multiclass classification as well as regression. Automatic handling of missing values and outliers helps with common cleanup, while feature engineering, automated model selection, and model explainability round out its workflow capabilities.
Extensions and data access
Optional installation extras cover NLP, computer vision, and reports, and a tutorial demonstrates a SQL data source. These extend the library's uses, though they do not change its stated dataset constraint: independent samples in rows remain the supported structure.
Pricing
LightAutoML costs 0.00 USD per free. The plan is the open-source Python library, installable from PyPI and released under Apache License 2.0. There is no paid tier or trial to weigh against it; the tradeoff is that users work with a library rather than buying a managed service.
Platforms
LightAutoML supports Linux, macOS, Windows, self-hosted use, and web workflows. Its PyPI installation and Python-library format make it a code-oriented option rather than a standalone desktop app.
Who it's for
LightAutoML fits data scientists and analysts working on classification or regression with independent-row tabular or text data. It is especially relevant when they want automated preparation and model selection alongside preset and custom pipeline options. Teams handling relational multitable data or sequences should look elsewhere.
Pros and cons
- Broad automation: preparation, feature selection, tuning, time utilization, and reporting can reduce repeated pipeline work.
- Choice of approach: presets and modular custom pipelines serve both quick starts and tailored Python workflows.
- Useful model range: linear, boosted-tree, neural-network, and scorecard classes cover several modeling approaches.
- Dataset constraint: multitable data and sequences are not currently supported, limiting its fit for those projects.
- Code-first format: installation through PyPI suits Python practitioners, not users seeking a turnkey no-code interface.
Alternatives
FEDOT is another free, open-source AutoML framework; consider it if you want an alternative framework with API and self-hosted platform options.
AutoGluon is a free, open-source Python library; it may suit readers comparing Apache-licensed libraries for Linux, macOS, self-hosted, or Windows use.
AutoKeras is a free Python package installed with pip, another option for readers seeking a package-based workflow.
Amazon SageMaker Autopilot uses pay-as-you-go pricing with no minimum fees or upfront commitments; choose it if you prefer an API or web offering over a free local library.
Auto-PyTorch is a free project developed by the AutoML Groups of the University of Freiburg and Hannover, and supports Linux and self-hosted use.
BigML has a free plan with a 16 MB maximum dataset per task, two parallel tasks, and one user; it suits readers who want a web or API service and can work within those caps.
JADBio offers a free Basic plan limited to one seat, three projects, 50 MB uploads, 500 MB storage, and one model export; consider it if those caps suit your needs and you prefer a web or API service.
EvalML is another free option, with API and self-hosted platform support.
Browse more options in AutoML Software.
Verdict
Choose LightAutoML if you are a Python-capable data scientist or analyst seeking a free, customizable way to automate common tabular or text classification and regression pipelines. Its combination of presets, model variety, and workflow automation is the main draw. Look elsewhere if your work depends on multitable datasets or sequences, or if you need a managed, no-code product.
LightAutoML plans and pricing
All plansCompared on AutoML software
- Feature engineering
- Yes
- Automated model selection
- Yes
- Model explainability
- Yes
- Workflow interface
- both
- Hosting model
- both

