Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsClassification is a machine-learning task that predicts a category, such as whether an email is spam or not spam. It differs from regression, which predicts a numerical value. To understand whether a classifier works well, look beyond its predictions: the class structure, the types of mistakes it makes, and the cost of those mistakes all matter.
What classification means
A classification model assigns an example to one or more categories. For an email filter, the categories might be “spam” and “not spam.” Once the actual label is known, you can compare it with the model’s prediction and count correct decisions and errors. Regression instead predicts a number, such as a house price. Google’s machine-learning glossary distinguishes these two prediction tasks.
Binary, multiclass and multilabel classification
The names describe how many labels are possible and whether an example can receive more than one.
- Binary classification: There are two classes, such as spam and not spam.
- Multiclass classification: There are more than two possible classes, and the model selects one. Identifying a handwritten digit from 0 through 9 is an example when each image represents exactly one digit.
- Multilabel classification: A single example can receive multiple labels. An image, for instance, might be tagged with several subjects at once.
Multiclass and multilabel are not interchangeable: multiclass selects one class from a set, while multilabel can assign several nonexclusive labels. scikit-learn’s classification guide also describes related multioutput settings, where a model predicts multiple targets.
#1 Best Overall
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
How to read a binary confusion matrix
A confusion matrix compares predicted labels with observed labels. For a spam filter, treat “spam” as the positive class and “not spam” as the negative class:
| Actual label | Predicted spam (positive) | Predicted not spam (negative) |
|---|---|---|
| Spam (positive) | True positive (TP): spam correctly flagged | False negative (FN): spam missed |
| Not spam (negative) | False positive (FP): legitimate email incorrectly flagged | True negative (TN): legitimate email correctly passed |
The matrix separates a model’s decision from the ground truth. As Google’s thresholding guide puts it, “The probability score is not reality, or ground truth.” The observed label is what lets you tell whether a prediction was right.
Rank #2
What accuracy, precision, recall and F1 tell you
Each metric answers a different question. For a binary classifier, let TP, FP, FN and TN mean the counts in the confusion matrix above.
| Metric | Formula | Question it answers |
|---|---|---|
| Accuracy | (TP + TN) / (TP + TN + FP + FN) | What share of all predictions were correct? |
| Precision | TP / (TP + FP) | Among examples predicted positive, how many were actually positive? |
| Recall | TP / (TP + FN) | Among actual positive examples, how many did the model find? |
| F1 | 2 × (precision × recall) / (precision + recall) | What is the equal-weight harmonic mean of precision and recall? |
Accuracy gives the proportion of all classifications that were correct. Precision focuses on the reliability of positive predictions; recall focuses on finding positive cases. F1 combines precision and recall, giving them equal weight. More generally, scikit-learn’s F-beta metric is a weighted harmonic mean, so the beta value changes the balance between the two. scikit-learn’s metric reference explains these measures and their averaging options.
Why accuracy can mislead on imbalanced data
When one class has far more examples than another, accuracy can conceal poor performance on the less common class. A model that always predicts the majority class may appear accurate while never finding a rare positive case. In that situation, compare class-wise precision and recall rather than relying on a single accuracy figure. Google’s classification metrics guide warns that majority-class predictions can produce high accuracy while failing on the rare class.
Choose metrics according to the consequences of errors. In disease screening, a false negative—a missed positive—may be more serious than a false positive that leads to follow-up. In spam filtering, a false positive can be especially disruptive if a legitimate message is hidden. State which error matters more for the task instead of treating one metric as universally best.
Rank #4
How the classification threshold changes results
Many classifiers produce a score and use a threshold to turn it into a class prediction. A higher threshold generally makes positive predictions harder to trigger: false positives tend to fall, while false negatives tend to rise. Lowering the threshold generally moves the balance in the opposite direction. The right operating point depends on the application’s relative cost of false alarms and missed positives; report the threshold when comparing models or results. Google illustrates the trade-off in its thresholding guide.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to compare classification results fairly
A headline metric is not enough to explain how a classifier behaves. When comparing models or operating points, report the choices that shape the result:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Best Value
- Task and labels: Say whether the problem is binary, multiclass or multilabel, and explain how many labels an example can receive.
- Class balance: Show or describe whether classes are similarly represented, especially when rare cases matter.
- Error priority: Identify whether false positives or false negatives are more costly for the intended use.
- Threshold and score policy: State the threshold used to turn scores into predictions. If scores are calibrated or otherwise transformed before applying it, describe that policy.
- Averaging method: For multiclass or multilabel metrics, name whether results are micro-, macro- or weighted-averaged. These summarize class-level performance differently: macro averaging gives each class equal influence, weighted averaging accounts for class support, and micro averaging aggregates decisions across classes.
- Operational impact: Explain what each kind of error means in practice, not just how many occurred.
scikit-learn supports per-label metrics and different averaging strategies for multiclass and multilabel evaluation; the selected average affects how much each class influences the summary. Its evaluation documentation describes these options.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




