Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

To integrate machine learning into a Flask app, save a fitted preprocessing-and-model pipeline, load it when the application starts, validate incoming data against the training feature contract, and return a clear prediction response. Flask handles HTTP requests; a library such as scikit-learn performs inference. For production, run Flask behind a production WSGI server—not with Flask’s development server.

This guide builds a small JSON prediction API, explains how to add an HTML form, and covers testing, deployment, security, and the point at which a separate inference service is a better fit.

What Flask does in a machine-learning app

A Flask integration usually takes one of three forms: an HTML form that returns a rendered result, a JSON API consumed by another application, or a hybrid that provides both. In each case, Flask is the web layer around a model built with scikit-learn, PyTorch, TensorFlow, XGBoost, or another ML framework.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Client → Flask route → input validation → preprocessing pipeline → model → response

Flask does not train, version, monitor, or automatically scale a model. Those responsibilities belong to your training and deployment workflow or additional services.

1. Train and save the full pipeline

Serving bugs often come from reproducing training-time transformations separately in the web app. If training scales numeric columns or encodes categories, inference must apply the same fitted transformations in the same order. A scikit-learn Pipeline containing both preprocessing and estimator is a strong default.

from pathlib import Path

import joblib
import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.ensemble import RandomForestClassifier
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler

DATA_PATH = Path("data/training.csv")
MODEL_PATH = Path("artifacts/model.joblib")

df = pd.read_csv(DATA_PATH)
X = df[["age", "income", "city"]]
y = df["approved"]

preprocessor = ColumnTransformer([
    ("numeric", StandardScaler(), ["age", "income"]),
    ("categorical", OneHotEncoder(handle_unknown="ignore"), ["city"]),
])

pipeline = Pipeline([
    ("preprocessor", preprocessor),
    ("model", RandomForestClassifier(n_estimators=200, random_state=42)),
])

pipeline.fit(X, y)
MODEL_PATH.parent.mkdir(parents=True, exist_ok=True)
joblib.dump(pipeline, MODEL_PATH)

The feature names and their meanings are part of the model contract. A DataFrame with named columns makes that contract more visible than an unlabelled list of numbers. handle_unknown="ignore" prevents an unseen category from breaking one-hot encoding, but it does not guarantee a useful prediction for a category the model never learned.

Evaluate the model before saving it, using an evaluation design appropriate to the task. Keep the artifact, training code, feature schema, dependency versions, data identifier, and evaluation results together. A prediction endpoint returning a value does not by itself mean the model is accurate or production-ready.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

2. Create the Flask prediction API

Install the basic dependencies in a virtual environment. Use a lockfile or pinned dependency set for reproducible deployments rather than relying indefinitely on unconstrained package names.

python -m venv .venv

Activate it with source .venv/bin/activate on macOS/Linux or .venvScriptsActivate.ps1 in Windows PowerShell, then install packages:

python -m pip install Flask pandas scikit-learn joblib

Save this as app.py. It loads the artifact once at startup, checks the request shape and required fields, converts values, and returns JSON.

Rank #2
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
from pathlib import Path

import joblib
import pandas as pd
from flask import Flask, jsonify, request

app = Flask(__name__)
MODEL_PATH = Path(__file__).parent / "artifacts" / "model.joblib"
model = joblib.load(MODEL_PATH)

@app.get("/health")
def health():
    return jsonify({"status": "ok"})

@app.post("/predict")
def predict():
    payload = request.get_json(silent=True)
    if not isinstance(payload, dict):
        return jsonify({"error": "Request body must be a JSON object"}), 400

    required = ["age", "income", "city"]
    missing = [name for name in required if name not in payload]
    if missing:
        return jsonify({"error": "Missing required fields", "fields": missing}), 400

    try:
        age = float(payload["age"])
        income = float(payload["income"])
        city = str(payload["city"])
    except (TypeError, ValueError):
        return jsonify({"error": "Invalid input types"}), 400

    if not 0 <= age <= 120:
        return jsonify({"error": "age must be between 0 and 120"}), 400
    if income < 0:
        return jsonify({"error": "income must not be negative"}), 400
    if not city.strip():
        return jsonify({"error": "city must not be empty"}), 400

    row = pd.DataFrame([{"age": age, "income": income, "city": city}])
    prediction = model.predict(row)[0]
    if hasattr(prediction, "item"):
        prediction = prediction.item()

    response = {"prediction": prediction}
    if hasattr(model, "predict_proba"):
        probabilities = model.predict_proba(row)[0]
        response["probabilities"] = [float(value) for value in probabilities]
    return jsonify(response)

In Python source, write the range operators as ordinary <= and < characters; the escaped forms shown above are for HTML display. For example, the actual condition is if not 0 <= age <= 120: after HTML decoding.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Replace the demonstration feature names and range rules with the model’s actual schema and domain rules. Check nulls, empty strings, units, finite numeric values, categorical constraints, and payload size. Reject unexpected fields too if the endpoint needs a strict schema. A validation library such as Pydantic or Marshmallow can help centralize complex contracts, but is not required for a small route.

Loading the model at startup avoids repeated disk reads on requests. The path is based on the file location rather than the shell’s current directory, which makes it more reliable in containers. In a multi-worker WSGI deployment, each worker may hold its own model copy; large artifacts can multiply memory use. For a larger application, use an application factory or explicit model-loading layer so tests can inject a fake model and readiness can reflect whether the model loaded successfully.

3. Define a useful API contract

Start the local development server with:

flask --app app run --debug

Send a request from another terminal:

curl -X POST http://127.0.0.1:5000/predict 
  -H "Content-Type: application/json" 
  -d '{"age":35,"income":75000,"city":"Boston"}'

If the artifact and feature schema match, the response contains a prediction field and, when supported, a list of class probabilities. Not every estimator implements predict_proba(). A probability is not automatically a calibrated likelihood or a guarantee of correctness; expose it only when it is meaningful for the task and label its classes clearly.

NumPy scalar outputs are often not JSON serializable without conversion. Convert scalar values to Python types and arrays to lists. For a public API, make the response schema stable—for example, include a model identifier and a named probability mapping rather than relying on an undocumented array ordering. Do not return internal model objects, filesystem paths, tracebacks, or raw exception text.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use distinct client and server errors. Malformed JSON, missing fields, or invalid values are commonly reported as 400 Bad Request; some APIs use 422 Unprocessable Entity for syntactically valid but semantically invalid input. Oversized request handling may use 413. Unexpected failures should be logged on the server and returned as a generic 500; a service that is running but not ready to predict may return 503. Configure request-size limits and rate limits at the application, proxy, or platform layer.

4. Add an HTML form when users need a page

A browser form is a separate interface to the same model, not a replacement for server-side validation. Form field names must match the keys read by Flask.

from flask import render_template

@app.get("/")
def index():
    return render_template("index.html")

@app.post("/predict-form")
def predict_form():
    try:
        row = pd.DataFrame([{
            "age": float(request.form["age"]),
            "income": float(request.form["income"]),
            "city": request.form["city"],
        }])
        prediction = model.predict(row)[0]
        error = None
    except (KeyError, TypeError, ValueError):
        prediction = None
        error = "Please provide valid values."

    return render_template("index.html", prediction=prediction, error=error)

HTML attributes such as required and type="number" improve the user experience but can be bypassed, so validate on the server too. Render user-controlled values through Jinja’s normal escaping rather than marking them safe. If the form is part of an authenticated browser session and changes state, add CSRF protection.

5. Test success cases and failure cases

Use Flask’s test client to check the route contract without starting a network server:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
def test_predict(client):
    response = client.post(
        "/predict",
        json={"age": 35, "income": 75000, "city": "Boston"},
    )
    assert response.status_code == 200
    assert "prediction" in response.get_json()

Also test malformed JSON, invalid types, boundary and out-of-range values, empty categories, unknown categories, response serialization, artifact loading, and model-unavailable behavior. Keep at least one regression test with known inputs and expected output when its fixture and artifact are deterministic and version-controlled. Otherwise, assert the response structure and status rather than a prediction that may legitimately change when the model is retrained.

A health endpoint alone may only prove that the process responds. Distinguish liveness (the process is running) from readiness (the model and required dependencies are loaded and can serve requests). Avoid reporting the service ready before artifact loading has succeeded.

6. Protect the model and the endpoint

  • Trust the artifact: scikit-learn warns that pickle-based formats such as joblib can execute arbitrary code when loaded. Never load a file supplied by an untrusted user; control and verify artifact provenance. See the scikit-learn model persistence guidance.
  • Keep environments compatible: scikit-learn does not support relying on model loading across arbitrary library versions. Record dependency versions and test the saved artifact in the deployment image.
  • Do not expose debug mode: Flask’s debugger and development server are not production serving tools. Use a production WSGI server or hosting platform, as the Flask deployment documentation explains.
  • Limit access: An internal endpoint may need an API key, OAuth, mutual TLS, or network restrictions. An obscure URL is not authentication. Restrict CORS to origins that actually need browser access; CORS is not an access-control system for non-browser clients.
  • Limit input abuse: Bound body size and batch size, reject NaN and infinity explicitly, and consider quotas or throttling for expensive inference.
  • Protect sensitive data: Prefer logging request IDs, model versions, validation outcomes, and timing—not raw personal, financial, or health features.
  • Use HTTPS: Terminate TLS through a managed platform or reverse proxy. If the app sits behind a proxy, configure trusted forwarded headers carefully rather than trusting arbitrary client-provided values. See Gunicorn’s proxy-related settings.

Keep secret keys, API credentials, model paths, and environment-specific settings out of source control. Flask’s deployment tutorial shows generating a random secret key with python -c 'import secrets; print(secrets.token_hex())'. Use environment configuration or a secrets manager for production values.

7. Run it with a production WSGI server

Flask is a WSGI application. The built-in command flask run is for development, not public production traffic. Flask recommends a production WSGI server or managed hosting; see its deployment options.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A simple Gunicorn command for an app.py module containing a Flask object named app is:

python -m pip install gunicorn
gunicorn --bind 0.0.0.0:8000 app:app

In app:app, the first name identifies the Python module and the second the application object. Configure worker count and timeouts for the environment; do not assume adding workers always helps. Each worker may load another model copy, while the limiting factor could instead be CPU, GPU, memory, or the model itself. Check the installed server’s documentation and measure the deployed workload.

For an application factory, define create_app() and use Gunicorn’s factory support if available in the installed version, for example gunicorn --bind 0.0.0.0:8000 --factory app:create_app. Verify the exact option against that version’s documentation before adopting the command. Flask’s production tutorial also demonstrates Waitress, which supports Windows and Linux: Flask tutorial: deploy to production.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

8. Containerize and deploy

A minimal container can install locked dependencies, include the artifact, run as a non-root user, and start Gunicorn:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
FROM python:3.12-slim
WORKDIR /app
ENV PYTHONDONTWRITEBYTECODE=1
ENV PYTHONUNBUFFERED=1
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY app.py .
COPY artifacts ./artifacts
RUN useradd --create-home appuser
USER appuser
CMD ["gunicorn", "--bind", "0.0.0.0:8080", "app:app"]

Pin or lock compatible package versions in requirements.txt rather than copying version ranges from an unrelated tutorial. Build and test the image with the exact artifact it will serve. Keep large artifacts in an appropriate controlled storage or image strategy and account for their impact on image size and startup time.

For Google Cloud Run, Google’s Flask quickstart documents source deployment with gcloud run deploy --source . and uses Gunicorn to handle HTTP in the container: Cloud Run Flask quickstart. During deployment, review region and access settings rather than accepting public access by default. A managed deployment is not automatically private, free, or cost-effective: request volume, CPU and memory allocation, cold starts, logs, storage, networking, and current pricing all matter. Consult the provider’s current pricing page or calculator before estimating cost.

Other hosting choices include a conventional VM with Gunicorn, Google App Engine, AWS Elastic Beanstalk, Azure, and PythonAnywhere. Their operational model, controls, scaling behavior, and prices differ; choose based on model resources, data residency, access requirements, and team capabilities rather than a generic claim that one is best.

9. Version, monitor, and maintain

Track code version, model artifact version, data snapshot or identifier, feature schema, dependency lockfile, evaluation metrics, and deployment image. They are related but distinct. You can return a model identifier to help clients and operators understand which artifact produced a prediction.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

After deployment, monitor request count, latency, error rate, resource use, model version, input distribution, and prediction distribution. Add drift indicators where they are meaningful, but do not log sensitive raw feature values by default. Keep an explicit rollback path to a known artifact and deployment revision. Test dependency upgrades against the saved artifact before rollout; version mismatches can cause loading errors or changed behavior.

Measure model loading, preprocessing, inference, serialization, and network time rather than guessing where latency comes from. For expensive predictions that may outlast a normal request timeout, use an asynchronous design: accept a job, return an identifier, then let the client check status and retrieve the result. Flask can expose those endpoints, but it is not itself a task queue.

When Flask is enough—and when to separate serving

Flask is a practical fit for a small or medium model, synchronous inference that completes within the service’s request budget, modest traffic, a simple form or API, and business logic that belongs beside the web layer. It is often a sensible first production service when the team wants control without adding a model-serving platform.

Consider a separate inference service, queue-based workers, or managed ML platform when models need GPUs, take a long time to load or run, require independent scaling, or need specialized registry and rollout workflows. A separate system adds operations and cost, so it is not automatically an upgrade for a small tabular model. FastAPI can be attractive when an API-first application benefits from typed request schemas and generated OpenAPI documentation; it is not automatically faster for every inference workload. Model compute, serialization, process layout, and infrastructure often matter more than the web framework.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Pickle-compatible formats are convenient for Python-native scikit-learn deployments but require trusted files and compatible environments. scikit-learn also discusses alternatives such as skops.io and ONNX, each with portability and security trade-offs; ONNX conversion is not available for every pipeline and does not eliminate input or resource-abuse risks. See the official persistence guide.

Before launch

  • The saved artifact contains fitted preprocessing as well as the estimator.
  • The inference feature names, types, units, and constraints match the training contract.
  • Artifacts come from a trusted, controlled source; runtime dependencies are pinned and tested.
  • Valid, missing, malformed, out-of-range, and unknown-category inputs have defined outcomes.
  • JSON outputs are serializable and follow a documented stable schema.
  • Model readiness is distinct from process liveness.
  • Debug mode is off; production serving uses a WSGI server or managed equivalent.
  • Endpoint access, request sizes, secrets, HTTPS, logging, and rate limits are addressed.
  • Worker memory, startup time, inference latency, monitoring, and rollback have been considered.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.