Splink

Windows · Mac · Linux · Self-hosted

Freedom report

Three barsScore 6.5

  • Free tierA free tier is on its own pricing page
  • Open codeNo open-source code on record
  • Runs widely3 of 6 device platforms
  • DocumentedPlans, terms and facts published

Splink is a free Python package for linking and deduplicating records when unique identifiers are unavailable. Its probabilistic matching method follows the Fellegi-Sunter model and can be trained without labeled data. Users can apply term-frequency adjustments and define fuzzy matching logic. The package generates SQL for a selected backend, with documented support for DuckDB, Spark, SQLite and PostgreSQL. The maker recommends DuckDB for most users, while Spark is suggested for very large linkages or when a Spark cluster is easier to access. The maker says Splink can handle linkages of 100+ million records on DuckDB or big-data backends such as Spark. It works best with standardized data across multiple columns that are not highly correlated, rather than a single bag-of-words column. Interactive visualizations help users examine and diagnose models, predictions and clusters. Install it with pip or conda; optional backend-specific installation guidance is documented for Spark and PostgreSQL. Users with remaining questions can turn to the GitHub discussion forum.

Who it is for

Splink suits teams and analysts who need to link or deduplicate records without unique identifiers, and can work with Python and a supported SQL backend. It is a poor fit for data represented by only one bag-of-words column.

What is good

  • Can train without labeled data.
  • Supports term-frequency adjustments and user-defined fuzzy logic.
  • Offers DuckDB, Spark, SQLite and PostgreSQL backends.
  • Interactive visualizations help diagnose models and clusters.
  • Free to install with pip or conda.

What to know first

  • Works best with standardized, multiple-column data.
  • Not designed for a single bag-of-words column.
  • The team may struggle with Databricks-specific issues.

Freedom251 review

Splink: the full review

Splink offers a free, configurable approach to probabilistic record linkage, with several SQL backend options and diagnostic visualizations. Data shape and backend setup matter, and Databricks-specific help may be limited.

Splink is a free Python package for linking and deduplicating records without unique identifiers. It is best suited to teams with well-structured data and the capacity to run a SQL backend. Its strongest case is flexible probabilistic matching that can scale, not a managed service or a way to match a single field of unstructured text.

Overview

Splink uses the Fellegi-Sunter model and can learn matching rules without labeled data. That makes it useful when verified examples are scarce. It generates SQL for a backend chosen by the user—DuckDB, Spark, SQLite or PostgreSQL—so it can fit different data environments, but the team remains responsible for choosing and operating that backend.

Data shape matters: Splink works best with standardized records that have several useful, not highly correlated columns. It is not designed to resolve records from one bag-of-words column, so teams without structured attributes should look elsewhere.

Key features

Flexible matching and scale

Term-frequency adjustments account for how common values are, and user-defined fuzzy matching logic lets teams tailor comparisons. The maker says Splink can handle linkages of 100+ million records on DuckDB or big-data backends such as Spark. DuckDB is recommended for most users; Spark is the better fit for very large linkages or when a Spark cluster is easier to access.

Diagnostics and setup

Interactive visualizations, including dashboards for predictions and clusters, help users inspect and diagnose linkage models rather than treating the output as a black box. Installation is through pip or conda, with optional backend-specific installs documented for Spark and PostgreSQL. Questions that remain after reading the documentation can go to the GitHub discussion forum. Databricks users should factor in a support limitation: the development team says it lacks access to a Databricks environment and may struggle to resolve platform-specific issues.

Pricing

Splink is free: its open-source Python package costs 0.00 USD per free and can be installed with pip or conda. There is no free trial to weigh against a paid plan. The package is the offer; teams still need to select and operate a backend suited to their data and infrastructure.

Platforms

Splink supports Linux, macOS and Windows, and can be self-hosted. It is a Python package that generates SQL for the selected backend, rather than a hosted service.

Who it's for

Splink is a strong fit for analysts and engineering teams in government, academia or other sectors that need probabilistic linkage across structured records and can manage their own data environment. Its maker names the Office for National Statistics, NHS England and the Australian Bureau of Statistics among its users. It is a weaker choice for teams whose data is mainly unstructured text, or who need a managed service or dependable Databricks-specific help.

Pros and cons

  • Pros: Unsupervised training can reduce reliance on labeled matches; configurable fuzzy matching and term-frequency adjustments suit varied record sets; multiple SQL backends and diagnostics support both deployment flexibility and model inspection; the package is free.
  • Cons: Results depend on standardized, multi-column data; users must choose and operate a backend; Databricks-specific support may be difficult to obtain.

Alternatives

Browse the Identity Resolution Software category to compare options. InterSystems IRIS is a paid product with a free plan and trial, a possible alternative for teams considering a broader platform. Hightouch has a free plan with two active syncs per month and hourly syncs; consider it when those sync limits fit the job. Zingg offers a free open-source community edition for batch probabilistic matching across millions of records, making it another option for that workload.

Open Client Registry (OpenCR) is an alternative developed by IntraHealth International with USAID support through the MEASURE Evaluation project. Senzing Entity Resolution is a paid alternative with a free plan and trial; its listed 10M DSR subscription costs 58560.00 USD per year. Treasure AI Composable CDP is a paid alternative with warehouse credits applying to query execution. NVECTA is another paid option. Precisely Data Matching and Entity Resolution is a paid alternative with custom pricing.

Verdict

Choose Splink if you need free, configurable probabilistic linkage for structured records and can take responsibility for backend setup and operation. Its combination of unsupervised learning, adjustable matching and diagnostics is compelling; look elsewhere if your data is mostly unstructured or you need a managed service with stronger Databricks-specific support.

Splink plans and pricing

All plans
Splink Free Open-source Python package · install via pip or conda github.com · 29 Sept 2026

Compared on identity resolution software

Free plan
Yes
Matching approach
hybrid
Real-time API
Yes
Batch file import
Yes
Organization matching
Yes

Best Splink alternatives

See all 20