The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →“From Data Mining to Knowledge Discovery in Databases” explains that data mining is one important step within the broader knowledge discovery in databases (KDD) process. Mining algorithms identify patterns; KDD also covers choosing, preparing, and understanding the data, then evaluating and interpreting the patterns so they can become useful knowledge.
What does “From Data Mining to Knowledge Discovery” refer to?
It refers to the 1996 article From Data Mining to Knowledge Discovery in Databases by Usama M. Fayyad, Gregory Piatetsky-Shapiro, and Padhraic Smyth. Published in AI Magazine, volume 17, issue 3, pages 37–54, the article clarified the relationship between data mining, KDD, and fields including machine learning, statistics, and databases. View the article’s DOI record.
The authors’ central distinction is practical: finding a pattern is not the same as discovering knowledge. A pattern may be statistically detectable yet irrelevant, misleading, or unusable unless the data and the result are considered in context.
How are data mining and KDD different?
KDD is the whole discovery process; data mining is the central step that applies methods to extract patterns. KDD includes work before mining, such as selecting and preparing data, and work after it, such as assessing and interpreting the results. Calling the entire workflow “data mining” can obscure how much useful discovery depends on these surrounding tasks.
#1 Best Overall
| Aspect | Data mining | KDD |
|---|---|---|
| Scope | Applying methods to discover or extract patterns from data. | The broader process of turning data into useful knowledge, including mining. |
| Typical work | Classification, prediction, clustering, association, or other pattern-finding tasks. | Defining a goal, selecting and preparing data, mining, evaluating and interpreting patterns, and using the result. |
| Role of context | Methods operate on selected data to produce patterns. | Prior and domain knowledge help guide discovery and judge whether patterns matter. |
| Result | Candidate patterns or models. | Patterns assessed and interpreted in relation to a discovery objective. |
In short, data mining answers “What patterns can these methods find?” KDD asks the larger question: “How can this data support a meaningful discovery?”
What are the steps in the KDD process?
The article presents KDD as an iterative workflow, not a one-click sequence. The stages can be revisited when results reveal a problem with the data, assumptions, or objective.
- Define the discovery objective. Specify the question or decision the analysis is meant to support. A clear objective helps determine which data and patterns are relevant.
- Select and understand data. Identify the sources and records related to the objective, and consider what the fields represent and how they were collected.
- Clean and preprocess. Address missing, inconsistent, noisy, or otherwise unsuitable data so the analysis is not built on avoidable data problems.
- Transform or reduce the data. Prepare the data in a form suitable for analysis, including transformations or reductions where appropriate.
- Apply data-mining methods. Choose methods suited to the intended pattern, such as classification, prediction, clustering, or association.
- Evaluate and interpret patterns. Assess whether results are valid, interesting, understandable, and relevant to the objective. Use prior and domain knowledge to help interpret them.
- Use the resulting knowledge. Relate accepted findings to the application or decision that motivated the discovery.
Fayyad, Piatetsky-Shapiro, and Smyth stress why the surrounding stages matter: “The additional steps in the KDD process, such as data preparation, data selection, data cleaning, incorporation of appropriate prior knowledge, and proper interpretation of the results of mining, are essential to ensure that useful knowledge is derived from the data.”
Why does the distinction matter in practice?
An algorithm can return a pattern without establishing that the pattern is dependable or useful. KDD makes the analyst account for the data’s quality, the purpose of the analysis, and the meaning of the result. That matters in application areas named in the article’s opening, including science, marketing, finance, health care, and retail, where large collections of data may contain patterns relevant to real problems.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteRank #3
When assessing a KDD method or tool, consider more than whether it can run an algorithm. Ask what preparation it requires, what kinds of patterns it can find, how domain knowledge shapes the analysis, whether its output can be interpreted, how it evaluates interestingness, whether it can handle the data volume, and how directly its results can support a decision.
How does KDD relate to machine learning, statistics, and databases?
KDD is an interdisciplinary field rather than a synonym for any one of these areas. Machine learning and statistics contribute techniques for modeling and pattern discovery; database systems support the storage, access, and management of data. KDD brings such capabilities into a larger, application-oriented process that includes data preparation and the evaluation and interpretation of findings.
The article’s point is not that one discipline replaces the others. It is that successful discovery draws on methods and infrastructure while keeping the objective and meaning of the result in view.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Who wrote the article, and where can you read more?
Usama M. Fayyad, Gregory Piatetsky-Shapiro, and Padhraic Smyth wrote the article, published on September 1, 1996, in AI Magazine. Its DOI is 10.1609/aimag.v17i3.1230.
Free tools Windows power users keep installed
One-click scans. No signup required.
Best Value
A related reference volume is Advances in Knowledge Discovery and Data Mining, published by AAAI Press in 1996. The 611-page book includes the related overview chapter, listed on pages 1–34. Its ISBN is 0-262-56097-6. Check the book’s Google Books record; availability and format can vary by library and bookseller.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




