The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →A complete machine-learning project is more than model.fit() and an accuracy score. This walkthrough builds a reproducible binary-classification project from a tabular dataset through data contracts, leakage-safe preprocessing, model selection, evaluation, persistence, batch inference, and an optional API.
The example uses Titanic-style passenger data because it contains both numerical and categorical fields. The same structure applies to churn, fraud, risk, and many other problems—provided you define what is known at prediction time and choose metrics that match the decision.
What the finished project should contain
At the end, you should have code that can be run outside a notebook, a documented data assumption, a fitted preprocessing-and-model pipeline, an evaluation report, a saved artifact, and a prediction interface. A high score on a classroom dataset is not proof of production readiness: real systems also require input validation, monitoring, security, governance, and retraining.
1. Define the prediction contract first
Write down the decision before choosing an algorithm. For the example project:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
- One row: one passenger.
- Target:
survived, where 1 means survived and 0 means did not. - Prediction-time inputs: fields available before the outcome, such as sex, age, passenger class, fare, family counts, and embarkation port.
- Action: an educational prediction; in a real project, specify the operational action that follows.
- Primary metric: selected after considering the cost of false positives and false negatives, not chosen automatically as accuracy.
For churn, the equivalent contract might be “predict whether a customer cancels within 30 days using only data available on the scoring date.” A false positive consumes retention capacity; a false negative misses a customer who might have been saved.
Classify the task
State whether the problem is classification, regression, ranking, forecasting, clustering, or anomaly detection. The split strategy, metrics, and validation method depend on that choice.
Define the data boundary
List every field that exists at prediction time. Post-outcome fields, future aggregates, duplicated records, and identifiers that encode collection order can create leakage even when the code runs without an error.
2. Create a small, reproducible repository
ml-project/
├── data/
│ ├── raw/
│ └── processed/
├── models/
├── reports/
├── src/
│ ├── load_data.py
│ ├── train.py
│ ├── evaluate.py
│ └── predict.py
├── tests/
├── notebooks/
├── requirements.txt
├── README.md
└── .gitignore
Use notebooks for exploration, but keep the final training and inference paths in scripts. This prevents hidden notebook state from becoming an undocumented dependency.
Free tools Windows power users keep installed
One-click scans. No signup required.
Create an isolated environment
python -m venv .venv
Activate it on macOS or Linux:
source .venv/bin/activate
On Windows PowerShell:
.venvScriptsActivate.ps1
Install the basic local stack:
python -m pip install --upgrade pip
pip install pandas numpy scikit-learn matplotlib seaborn joblib
Python documents venv as a lightweight way to isolate project dependencies: https://docs.python.org/3/library/venv.html. Pin the versions actually tested in your companion repository rather than copying unverified version numbers. The Python, scikit-learn, and pandas documentation pages listed Python 3.14.7, scikit-learn 1.9.0, and pandas 3.0.5 on August 18, 2026; these are not permanent compatibility guarantees.
numpy==<tested-version>
pandas==<tested-version>
scikit-learn==<tested-version>
joblib==<tested-version>
matplotlib==<tested-version>
seaborn==<tested-version>
See the official package documentation for current releases: scikit-learn and pandas.
3. Load and audit the data
import pandas as pd
df = pd.read_csv("data/raw/train.csv")
print(df.head())
print(df.shape)
print(df.info())
print(df.describe(include="all").T)
print(df.isna().mean().sort_values(ascending=False))
Pandas’ introductory tutorials cover loading tabular data, inspecting a DataFrame, selecting subsets, plotting, combining tables, and working with time series: https://pandas.pydata.org/docs/getting_started/intro_tutorials/index.html.
Questions your audit must answer
- How many rows and columns are present?
- Which fields are numeric, categorical, dates, identifiers, or free text?
- Where are values missing, and is missingness concentrated in a subgroup?
- Is the target imbalanced?
- Are there duplicate rows or repeated entities?
- Are any values impossible, such as negative ages or invalid dates?
- Does an ID encode time, geography, customer, or collection order?
- Could any column have been created after the outcome?
Document every removal
Do not silently drop columns. A Titanic-style dataset may contain fields such as name, ticket, cabin, boat, or body. Remove a field only with a stated reason: it is unavailable at prediction time, a unique identifier, high-cardinality text outside the tutorial scope, too incomplete, or a leakage risk. If a field is retained, document how it will be represented.
4. Explore without contaminating the experiment
import matplotlib.pyplot as plt
import seaborn as sns
sns.countplot(data=df, x="survived")
plt.show()
sns.histplot(data=df, x="age", hue="survived", kde=True)
plt.show()
print(df.groupby("sex")["survived"].mean())
Use a few purposeful plots to identify imbalance, missingness patterns, outliers, possible class separation, and fields that would not exist when a prediction is made. Group averages describe associations in this dataset; they do not prove that changing a feature would cause an outcome.
5. Separate features and target
target = "survived"
X = df.drop(columns=[target])
y = df[target]
drop_columns = ["name", "ticket", "cabin", "boat", "body"]
X = X.drop(columns=[c for c in drop_columns if c in X.columns])
Keep this decision in your README. A reproducible project records not only what was used, but why other fields were excluded.
6. Split before learning preprocessing statistics
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.20,
stratify=y,
random_state=42,
)
stratify=y preserves class proportions for an ordinary classification split. It is not a universal rule:
- Use a time-based split for forecasting or temporally ordered records.
- Use a group split when several rows belong to the same person, household, patient, account, or device.
- Use a spatial split when nearby observations could otherwise appear in both partitions.
A random split can be optimistically biased when related records cross the boundary.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallRank #3
7. Build leakage-safe mixed-type preprocessing
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler
numeric_features = ["age", "fare", "sibsp", "parch"]
categorical_features = ["sex", "class", "embarked"]
numeric_pipeline = Pipeline(
steps=[
("imputer", SimpleImputer(strategy="median")),
("scaler", StandardScaler()),
]
)
categorical_pipeline = Pipeline(
steps=[
("imputer", SimpleImputer(strategy="most_frequent")),
("onehot", OneHotEncoder(handle_unknown="ignore")),
]
)
preprocessor = ColumnTransformer(
transformers=[
("numeric", numeric_pipeline, numeric_features),
("categorical", categorical_pipeline, categorical_features),
],
remainder="drop",
)
SimpleImputerlearns replacement values from training data.StandardScalerstandardizes numeric values for estimators that benefit from comparable scales.OneHotEncoderturns categories into numeric columns.handle_unknown="ignore"prevents a new category from crashing inference.ColumnTransformerapplies different transformations to different column groups.
Put the preprocessor and estimator in one Pipeline. During cross-validation, each fold then learns imputation, scaling, and encoding only from its own training portion. Scikit-learn describes this composition at https://scikit-learn.org/stable/getting_started.html and https://scikit-learn.org/stable/modules/compose.html.
8. Establish a baseline before tuning
from sklearn.dummy import DummyClassifier
baseline = DummyClassifier(strategy="prior")
baseline.fit(X_train, y_train)
print(baseline.score(X_test, y_test))
This baseline predicts according to the observed class prior. It answers whether a model learns anything useful beyond the majority distribution.
Now fit a simple, interpretable candidate:
from sklearn.linear_model import LogisticRegression
logistic_pipeline = Pipeline(
steps=[
("preprocessor", preprocessor),
("model", LogisticRegression(max_iter=1000)),
]
)
logistic_pipeline.fit(X_train, y_train)
Exceeding 50% accuracy on a binary problem is not, by itself, evidence of success. Compare with the baseline and report metrics tied to the decision.
9. Compare candidate estimators
from sklearn.ensemble import RandomForestClassifier
models = {
"logistic_regression": LogisticRegression(max_iter=1000),
"random_forest": RandomForestClassifier(
n_estimators=300,
random_state=42,
n_jobs=-1,
),
}
pipelines = {
name: Pipeline(
steps=[
("preprocessor", preprocessor),
("model", model),
]
)
for name, model in models.items()
}
| Estimator | Strengths | Trade-offs |
|---|---|---|
| Logistic regression | Fast, interpretable baseline; often useful for probability estimates | Needs feature engineering for complex nonlinear interactions |
| Random forest | Captures nonlinearities and interactions; scaling is usually unnecessary | Less transparent, larger artifacts, and probabilities may need calibration |
| Gradient boosting | Often strong on tabular data | More tuning-sensitive and easier to overfit |
No algorithm is universally best. Dataset size, feature types, missingness, class balance, temporal structure, and error costs determine the appropriate comparison.
10. Evaluate with metrics that match the decision
from sklearn.metrics import (
accuracy_score, classification_report, confusion_matrix,
f1_score, precision_score, recall_score, roc_auc_score,
)
predictions = logistic_pipeline.predict(X_test)
probabilities = logistic_pipeline.predict_proba(X_test)[:, 1]
print("Accuracy:", accuracy_score(y_test, predictions))
print("Precision:", precision_score(y_test, predictions))
print("Recall:", recall_score(y_test, predictions))
print("F1:", f1_score(y_test, predictions))
print("ROC AUC:", roc_auc_score(y_test, probabilities))
print(confusion_matrix(y_test, predictions))
print(classification_report(y_test, predictions))
- Accuracy is the fraction of all predictions that are correct.
- Precision is the fraction of predicted positives that are truly positive.
- Recall is the fraction of actual positives found.
- F1 is the harmonic mean of precision and recall.
- ROC AUC measures ranking across thresholds, not the quality of one chosen threshold.
- PR AUC is often more informative when the positive class is rare.
- Confusion matrix shows true positives, false positives, true negatives, and false negatives.
- Calibration asks whether predicted probabilities match observed frequencies.
Scikit-learn’s scoring reference is at https://scikit-learn.org/stable/modules/model_evaluation.html. MLflow’s evaluation documentation lists common classification outputs, including accuracy, precision, recall, F1, ROC AUC, PR AUC, log loss, Brier score, confusion matrices, and classification reports: https://mlflow.org/docs/latest/ml/evaluation.
Regression uses different metrics
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score
predictions = model.predict(X_test)
mae = mean_absolute_error(y_test, predictions)
rmse = mean_squared_error(y_test, predictions) ** 0.5
r2 = r2_score(y_test, predictions)
print({"mae": mae, "rmse": rmse, "r2": r2})
MAE is expressed in target units. RMSE penalizes large errors more heavily. R² is not percentage accuracy and can be negative on unseen data.
Rank #4
11. Cross-validate on the training partition
from sklearn.model_selection import StratifiedKFold, cross_validate
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
scores = cross_validate(
logistic_pipeline,
X_train,
y_train,
cv=cv,
scoring=["accuracy", "precision", "recall", "f1", "roc_auc"],
n_jobs=-1,
)
for metric in ["test_accuracy", "test_precision", "test_recall", "test_f1", "test_roc_auc"]:
print(metric, scores[metric].mean(), scores[metric].std())
Report the mean and standard deviation, not just the best fold. Keep the final test set untouched until model selection is complete. If rows are grouped or time-ordered, replace ordinary stratified folds with an appropriate grouped or temporal strategy. Scikit-learn documents these strategies at https://scikit-learn.org/stable/modules/cross_validation.html.
12. Tune hyperparameters without touching the test set
from sklearn.model_selection import RandomizedSearchCV
search_pipeline = Pipeline(
steps=[
("preprocessor", preprocessor),
("model", RandomForestClassifier(random_state=42, n_jobs=-1)),
]
)
param_distributions = {
"model__n_estimators": [100, 300, 500],
"model__max_depth": [None, 5, 10, 20],
"model__min_samples_leaf": [1, 2, 5, 10],
"model__max_features": ["sqrt", "log2", None],
}
search = RandomizedSearchCV(
search_pipeline,
param_distributions=param_distributions,
n_iter=20,
scoring="roc_auc",
cv=cv,
random_state=42,
n_jobs=-1,
refit=True,
)
search.fit(X_train, y_train)
print(search.best_params_)
print(search.best_score_)
The model__parameter syntax addresses a parameter inside the named pipeline step. Use GridSearchCV for a small deliberate grid and RandomizedSearchCV for a larger search space. Search the complete pipeline, not an estimator detached from preprocessing.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →13. Make one final test evaluation
best_model = search.best_estimator_
test_predictions = best_model.predict(X_test)
test_probabilities = best_model.predict_proba(X_test)[:, 1]
final_metrics = {
"accuracy": accuracy_score(y_test, test_predictions),
"precision": precision_score(y_test, test_predictions),
"recall": recall_score(y_test, test_predictions),
"f1": f1_score(y_test, test_predictions),
"roc_auc": roc_auc_score(y_test, test_probabilities),
}
print(final_metrics)
Your report should state the dataset snapshot, split strategy, random seed, cross-validation design, tuning metric, test-set size, final metrics, and any uncertainty estimate that is practical. Results depend on the dataset version, retained rows, feature choices, seed, split, library versions, and missing-value policy. Do not repeatedly inspect the test result and modify the model; that turns the test set into validation data.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.14. Inspect errors, thresholds, and subgroups
The default 0.5 threshold is a convention, not a law.
import numpy as np
thresholds = np.arange(0.10, 0.91, 0.05)
for threshold in thresholds:
adjusted = (test_probabilities >= threshold).astype(int)
print(
threshold,
precision_score(y_test, adjusted, zero_division=0),
recall_score(y_test, adjusted, zero_division=0),
)
Lowering a threshold generally raises recall and can reduce precision; raising it generally does the opposite. Select a threshold with validation data or a separate calibration set, not by repeatedly optimizing the final test set.
errors = X_test.copy()
errors["actual"] = y_test
errors["predicted"] = test_predictions
errors["probability"] = test_probabilities
print(errors[errors["actual"] != errors["predicted"]].head())
Review false positives and false negatives, and calculate performance for relevant subgroups. A model can have acceptable overall metrics while failing materially for one population. Feature importance indicates model association, not causation; correlated variables can divide importance among themselves.
Best Value
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
15. Persist the complete pipeline
import joblib
joblib.dump(best_model, "models/classifier_pipeline.joblib")
loaded_model = joblib.load("models/classifier_pipeline.joblib")
new_predictions = loaded_model.predict(new_data)
new_probabilities = loaded_model.predict_proba(new_data)[:, 1]
Save the entire pipeline so inference applies the same imputation, encoding, scaling, and estimator steps used during training. Record Python, scikit-learn, pandas, NumPy, and other dependency versions beside the artifact. Cross-version loading is not automatically safe or guaranteed. Never load a pickle or joblib artifact from an untrusted source: deserialization can execute code. Scikit-learn’s persistence guidance is at https://scikit-learn.org/stable/modules/model_persistence.html.
16. Add a batch prediction script
# src/predict.py
import sys
import joblib
import pandas as pd
model = joblib.load("models/classifier_pipeline.joblib")
input_path = sys.argv[1]
data = pd.read_csv(input_path)
predictions = model.predict(data)
output = data.copy()
output["prediction"] = predictions
if hasattr(model, "predict_proba"):
output["prediction_probability"] = model.predict_proba(data)[:, 1]
output.to_csv("reports/predictions.csv", index=False)
python src/predict.py data/raw/new_samples.csv
Validate inputs deliberately
Test the script with missing columns, extra columns, unknown categories, incorrect numeric types, empty files, null values, and an artifact produced by a different dependency version. handle_unknown="ignore" prevents one-hot encoding from failing on a new category, but your application should still decide whether to log, reject, or review that input.
17. Optional prediction API
from typing import Literal
import joblib
import pandas as pd
from fastapi import FastAPI
from pydantic import BaseModel
app = FastAPI()
model = joblib.load("models/classifier_pipeline.joblib")
class Passenger(BaseModel):
age: float | None = None
fare: float | None = None
sibsp: int = 0
parch: int = 0
sex: Literal["female", "male"]
passenger_class: str
embarked: str | None = None
@app.post("/predict")
def predict(passenger: Passenger):
row = pd.DataFrame([passenger.model_dump()])
prediction = int(model.predict(row)[0])
response = {"prediction": prediction}
if hasattr(model, "predict_proba"):
response["probability"] = float(model.predict_proba(row)[0, 1])
return response
uvicorn app:app --reload
FastAPI’s documentation is at https://fastapi.tiangolo.com/. A production API also needs authentication, rate limiting, request IDs, structured logs, health and readiness endpoints, model-version logging, input-size limits, safe error handling, and monitoring for missingness, category drift, latency, and prediction distributions. An API is an interface, not a complete deployment plan.
18. Optional containerization
FROM python:3.14-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY app.py .
COPY models ./models
EXPOSE 8000
CMD ["uvicorn", "app:app", "--host", "0.0.0.0", "--port", "8000"]
docker build -t ml-api .
docker run --rm -p 8000:8000 ml-api
Docker’s getting-started guide is at https://docs.docker.com/get-started/. Add a container after the local training and inference path works; it does not replace testing, security, or monitoring.
19. Optional experiment tracking
MLflow can record parameters, metrics, plots, and model artifacts, but it is not required for a first local run. Useful references are MLflow tracking, its scikit-learn integration, and its evaluation workflow. Introduce tracking after the core project is understandable rather than adding infrastructure as decoration.
20. Reproducibility and production checklist
- Dataset URL, version or snapshot date, and provenance are recorded.
- Python and dependency versions are pinned or locked.
- Feature names, types, and removal reasons are documented.
- Prediction-time availability and leakage checks are explicit.
- Split type, seed, fold count, and tuning search are recorded.
- Baseline and final metrics include variability and test-set size.
- Threshold choice and probability calibration are documented when decisions use scores.
- Subgroup performance and important error cases have been reviewed.
- The complete pipeline, not only the estimator, is persisted.
- Batch or API inference validates schema and malformed input.
- Artifact versions and checksums are traceable.
- Monitoring covers data freshness, missingness, drift, latency, errors, and prediction distributions.
- Retraining triggers, rollback procedures, privacy controls, and access policies are defined.
What to build next
Once this workflow runs, extend it with a time-based churn project, a regression model with residual analysis, text classification using TF-IDF, probability calibration, subgroup and fairness evaluation, scheduled batch scoring, or monitoring and retraining. Each extension should preserve the same discipline: define the data boundary, keep learned transformations inside validation, protect the final test set, and document what the metrics do not prove.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




