Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

You can deploy a trained machine-learning model on Heroku by wrapping inference in a Python web API, packaging the model and its dependencies, and running the API as a web process. For a small or moderate CPU-based model, a Python buildpack and Git deployment are usually the simplest route. This guide builds a FastAPI example, deploys it with the Heroku CLI, and explains when Docker, background workers, external storage, or a different inference platform is a better fit.

What this guide builds

The result is a web application with a health endpoint and a prediction endpoint. A client sends JSON to POST /predict; the app validates the input, applies the same preprocessing used during training, runs inference, and returns a JSON response.

Deployment is not model training. Heroku hosts the application and manages its runtime processes; you remain responsible for the trained model, API behavior, compatible dependencies, validation, security, and operational monitoring. Training is the fitting of a model, inference is using a fitted model to produce predictions, and model serving is making inference available through an application interface. MLOps is the broader work of versioning, testing, monitoring, retraining, and governing that system.

A basic architecture is:

Client → POST /predict → Heroku web dyno → validate and preprocess → model inference → JSON response

For work that cannot complete within a web request, use a job flow instead: the web process validates and queues the request, a worker performs inference, and the result is saved to durable storage for the client to retrieve.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is Heroku suitable for your model?

Heroku is a practical choice for conventional HTTP prediction services, including prototypes, demos, internal tools, and modest production APIs. Heroku’s Python materials describe data-science and machine-learning applications as a use case, while positioning ordinary dynos for smaller models and Managed Inference and Agents for more demanding AI workloads. That positioning is not a guarantee that a particular model will fit: measure its memory, startup time, inference time, dependencies, and expected concurrency. See Heroku’s Python platform information.

  • Usually a good candidate: a small scikit-learn model, tabular regression or classification, a modest CPU-based NLP or vision model, low-to-moderate traffic, and stateless predictions.
  • Needs careful testing: a larger dependency stack, a model that takes a long time to load, memory-heavy inference, or traffic with sharp concurrency peaks.
  • Usually a poor fit for a basic web dyno: GPU-dependent inference, very large transformer or diffusion models, strict low-latency requirements, or synchronous predictions that can exceed Heroku’s router response window.

Heroku’s router expects response data within an initial 30-second window; increasing a server timeout does not extend that limit. For longer inference, use a background-job architecture or a service designed for that workload. See Heroku’s request timeout documentation and guidance for preventing H12 timeouts.

Prepare the model artifact

Package the estimator and the transformations it needs together. If training applied scaling, encoding, feature selection, or other preprocessing, save the fitted preprocessing pipeline with the estimator. Otherwise, the deployed API can receive valid-looking input but produce incorrect predictions because training-time and inference-time transformations differ.

For a scikit-learn workflow, one option is to save a single fitted pipeline:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import joblib

joblib.dump(pipeline, "model.joblib")

Load it once when the application starts, rather than once per request:

from pathlib import Path
import joblib

MODEL_PATH = Path(__file__).with_name("model.joblib")
model = joblib.load(MODEL_PATH)

Record the library versions used to create the artifact, verify the feature order and data types, and test the exact artifact with the production dependency set. Serialized artifacts should be loaded only from trusted sources; treat an untrusted model file as unsafe. A tested environment can be captured with pip freeze > requirements.txt, then reviewed and pinned rather than relying on unconstrained upgrades. Heroku’s Python workflow supports dependency files including requirements.txt, Pipfile.lock, poetry.lock, and uv.lock, and a .python-version file can select the runtime version. Check Heroku’s current Python documentation when selecting a version because supported runtimes change over time.

Create a FastAPI prediction service

FastAPI is one suitable choice, not a Heroku requirement; Flask and other supported Python frameworks can also serve an API. The example below expects the model artifact to contain a fitted pipeline with a predict method. Its fields and values must match the model actually trained.

from pathlib import Path

import joblib
import numpy as np
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel

MODEL_PATH = Path(__file__).with_name("model.joblib")
model = joblib.load(MODEL_PATH)
app = FastAPI(title="ML Prediction API")

class PredictionRequest(BaseModel):
    # Replace these with the exact features expected by your model.
    age: float
    income: float
    account_balance: float

@app.get("/health")
def health():
    return {"status": "ok"}

@app.post("/predict")
def predict(request: PredictionRequest):
    values = np.array([[
        request.age,
        request.income,
        request.account_balance,
    ]], dtype=float)
    try:
        prediction = model.predict(values)
        return {"prediction": prediction.tolist()}
    except Exception:
        # Log diagnostic details server-side; do not expose secrets or internals.
        raise HTTPException(status_code=400, detail="Prediction failed")

Explicit named fields help prevent clients from silently sending features in the wrong order. Add domain-specific bounds and validation where appropriate, reject non-finite values, and return only outputs that your model supports. For example, return classification probabilities only if the estimator provides them and the API contract defines what each probability means. Keep detailed exceptions in server logs rather than returning paths, credentials, or sensitive implementation details to clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Test locally before deploying

Create a virtual environment, install the pinned dependencies, and start the API locally:

python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
uvicorn app:app --reload --host 127.0.0.1 --port 8000

In Windows PowerShell, activate the environment with .venvScriptsActivate.ps1. Check the health route and submit a sample whose fields and values match your trained model:

curl http://127.0.0.1:8000/health

curl -X POST http://127.0.0.1:8000/predict 
  -H "Content-Type: application/json" 
  -d '{"age":35,"income":65000,"account_balance":1200}'

FastAPI’s interactive documentation is available locally at http://127.0.0.1:8000/docs. Before deployment, test missing fields, wrong types, out-of-range and non-finite values, model-loading failure, prediction latency, concurrent requests, and the production dependency lockfile. The sample numbers above are illustrative only; substitute values suitable for the model.

Set up the Heroku Python app

A minimal project can look like this:

ml-heroku-app/
├── app.py
├── model.joblib
├── requirements.txt
├── Procfile
├── .python-version
└── .gitignore

Include the tested versions of your API server, framework, model libraries, and numerical dependencies in the dependency file. Keep credentials, private certificates, and user data out of Git. Heroku config vars are intended for runtime configuration and secrets; Heroku’s platform runtime documentation describes the configuration model.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Create a file named exactly Procfile with no extension:

web: gunicorn -k uvicorn.workers.UvicornWorker app:app --bind 0.0.0.0:$PORT
  • web declares the process type that receives HTTP traffic.
  • gunicorn is the production process manager and uvicorn.workers.UvicornWorker runs the ASGI app.
  • app:app points to the app object in app.py.
  • $PORT is supplied by Heroku. The process must bind to it instead of hard-coding local port 8000.

Heroku’s Python getting-started guide explains the Procfile and Git deployment flow. The same guide applies whether the service wraps a model or another Python application.

Deploy with Git and the Heroku CLI

Install and authenticate with the Heroku CLI, then create an app. Replace the example name with an available app name:

heroku login
heroku create my-ml-api

If the project is not already a Git repository, initialize it and commit the files, then deploy the branch:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
git init
git add .
git commit -m "Deploy machine learning API"
git push heroku main

If your local branch is called master, use git push heroku master. After deployment, check the process and logs:

heroku ps -a my-ml-api
heroku logs --tail -a my-ml-api
heroku open -a my-ml-api

A successful deployment includes a completed build and release, a running web process, and an app that binds to its assigned port. Visit /health and send a real test request to /predict; a successful Git push alone does not prove that model inference works.

Configure runtime secrets and settings

Set per-app configuration through config vars instead of embedding secrets in source code or the image:

heroku config:set MODEL_VERSION=2026-08-01 -a my-ml-api
heroku config:set STORAGE_BUCKET=my-model-bucket -a my-ml-api
heroku config:set API_KEY=replace-me -a my-ml-api

Read a variable in Python with os.environ.get("MODEL_VERSION", "development"). Use heroku config to inspect configuration, but do not print secret values in application logs or exception messages. The example API does not require an API key by itself: for a non-public service, implement authentication and authorization appropriate to its clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use Docker when the buildpack is not enough

For an ordinary Python API, start with Heroku’s Python buildpack. Docker is useful when you need operating-system packages, a custom base image, native libraries, or tighter control over the runtime. Heroku describes its container workflow as an advanced option and recommends buildpacks for typical apps; see Container Registry and Runtime.

Example Dockerfile:

FROM python:3.12-slim

WORKDIR /app
ENV PYTHONDONTWRITEBYTECODE=1
ENV PYTHONUNBUFFERED=1

COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt

COPY app.py .
COPY model.joblib .

CMD ["sh", "-c", "gunicorn -k uvicorn.workers.UvicornWorker app:app --bind 0.0.0.0:${PORT}"]

Select a Python base image compatible with the artifact and dependencies; do not assume any one Python version will remain supported indefinitely. Test the image locally, then push and release it:

docker build -t ml-heroku-api .
docker run --rm -p 8000:8000 -e PORT=8000 ml-heroku-api

heroku container:login
heroku create my-ml-api --stack container
heroku container:push web -a my-ml-api
heroku container:release web -a my-ml-api
heroku open -a my-ml-api

For Heroku containers, the app still needs to bind to $PORT; EXPOSE does not choose that port. Do not rely on a Docker VOLUME for durable data, and do not treat a Docker health check as a replacement for Heroku runtime behavior. Registry-deployed images need to be rebuilt to receive operating-system updates; they are not automatically rebased.

Handle memory, startup time, and request duration

Memory and worker count

The Python runtime, imported libraries, model, and web workers all use dyno memory. A common surprise is that each worker process may load its own model copy. Start with a conservative worker count, measure resident memory and request performance, then adjust. If memory remains too high, reduce model size, remove unnecessary dependencies, or choose a suitable dyno capacity. Heroku’s pricing page lists dyno families and memory information; the appropriate capacity depends on the runtime and plan.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Memory pressure may appear as an R14 - Memory quota exceeded event, process crashes, or slow requests. A larger dyno can help, but it will not fix memory leaks, duplicate model copies, or an unnecessarily large model.

Startup and cold starts

Load the model once during process startup, outside the request handler. Heroku’s current limits documentation says a web process must bind to its assigned port within 60 seconds. Large downloads or lengthy initialization during startup can therefore prevent the process becoming available. Keep stable artifacts in the build or image when appropriate, and consider a dedicated inference service if model initialization is too heavy. See Heroku’s limits documentation.

Request timeouts

Measure typical and worst-case prediction latency. Heroku’s initial router response window is 30 seconds; an application-server timeout cannot raise it. A Gunicorn timeout can instead be set lower to fail a request sooner and free capacity, for example:

web: gunicorn -k uvicorn.workers.UvicornWorker app:app --bind 0.0.0.0:$PORT --timeout 20

Choose a value based on measured behavior. If a legitimate prediction may exceed the router window, return a job identifier after enqueueing work and let the client retrieve the result later.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Use durable storage for files and state

Dynos are isolated and their filesystems are ephemeral: files written at runtime are not durable, are not shared between dynos, and can disappear when a dyno restarts or is replaced. Do not keep uploads, generated files, prediction history, mutable model versions, or application state only on the dyno. Store durable data in a database or object-storage service, and keep any shared queue or cache in an appropriate external service. See How Heroku Works and Dyno Isolation.

Diagnose common deployment failures

Symptom Likely cause First response
Build cannot install a dependency Incompatible Python or package version, or a native build requirement Pin a tested dependency set and runtime; use Docker if system libraries are required.
App crashes immediately Import error, missing model artifact, or invalid startup command Inspect heroku logs --tail -a my-ml-api and verify the committed files and Procfile.
App does not become available Process is not listening on Heroku’s assigned port Bind to $PORT, not a fixed local port.
H12 timeout Slow inference, a blocked request, or queueing under load Profile inference, reduce blocking work, or move long jobs to a worker.
Memory quota event or process crash Model, dependencies, or worker count exceed available memory Reduce worker count, measure memory, and consider a smaller model or appropriate dyno.
Predictions differ from local results Preprocessing or dependency versions differ Package the fitted preprocessing pipeline and test the deployed artifact and pinned environment.
Uploaded files disappear Files were written to an ephemeral dyno filesystem Use a durable external storage service.
First request is slow Process wake-up or model loading is on the request path Load the model at startup and assess whether an always-on runtime or a different architecture is needed.

Useful operational commands include:

heroku logs --tail -a my-ml-api
heroku logs -p web --tail -a my-ml-api
heroku ps -a my-ml-api
heroku releases -a my-ml-api
heroku releases:info -a my-ml-api
heroku restart -a my-ml-api

Heroku combines application, system, and platform-related logs, but retained log history is limited. A production service may need an external log drain or observability system. See Heroku logging documentation and platform limits.

Scale the API and move long jobs to workers

Adding web dynos can increase concurrent request capacity, but it does not make an individual prediction faster. Each dyno can load its own model copy, so scaling out also increases the aggregate memory used by model processes. Heroku supports horizontal and vertical scaling; for example, scale the web process to two dynos with:

heroku ps:scale web=2 -a my-ml-api

Use a worker architecture for batch inference, document parsing, image processing, or other work that cannot reliably finish in a web request. A robust flow validates and enqueues a job in the web process, runs inference in a worker, writes the result to durable storage, and lets the client check job status or retrieve the result. The queue, result store, retries, and idempotency behavior still need to be designed; adding a worker alone does not solve those concerns.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Version and operate model releases

Give each model artifact an identifiable version and record its training code and data lineage where appropriate. A checksum can help confirm which artifact is running. Keep the API schema compatible with the model, expose non-sensitive version metadata through an endpoint such as /model-info, and test a rollback path before relying on it. Deploy model changes as application releases rather than manually overwriting files on a running dyno. Heroku’s runtime platform includes releases and configuration management.

A successful deployment is not by itself production readiness. Protect the API with suitable authentication and authorization, validate inputs, minimize sensitive data, monitor errors and latency, and define how model quality or drift will be reviewed. Hosting does not automatically provide these application-level controls.

Consider cost and alternatives

Heroku is a paid platform; do not plan around an assumption of free, always-on model hosting. Its pricing page, checked August 18, 2026, listed Eco at $5 per month with 0.5 GB RAM and sleeping after 30 minutes of inactivity, and Basic at $7 per month. These are plan details from that date, not a guarantee of current availability or suitability; check Heroku pricing before choosing a plan. Sleeping can make an inactive app’s first subsequent request slower, and a continuously available API may need a different plan.

Heroku’s main advantage is a short path from Python code to a managed web process, with Git deployment, config vars, logs, and optional container deployment. Trade-offs include ephemeral local storage, the router time limit, memory duplicated across workers, and the operational responsibility for model-specific behavior. Compare platforms according to the constraint you actually have:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Need Options to investigate
Simple app deployment with an alternative web-service workflow Render or Railway
More container placement and region control Fly.io
Request-driven container inference integrated with cloud services Google Cloud Run
Managed enterprise machine-learning workflows AWS SageMaker, Azure Machine Learning, or Google Vertex AI
GPU-oriented or model-serving-focused workflows Modal or Replicate
More control and potentially lower nominal infrastructure cost, with more operations work A self-managed VPS

Alternative pricing, sleep behavior, hardware, and service limits are not compared here; verify those details directly before selecting a platform. Choose Heroku when the model fits its runtime constraints and simplicity is valuable. Choose a specialized inference platform when hardware, model size, latency, or managed ML lifecycle features are the deciding requirement.

Production readiness checklist

  • Model and fitted preprocessing are packaged together, and the artifact comes from a trusted source.
  • Dependency and Python versions are tested and pinned.
  • The API validates the real feature schema, data types, and value ranges.
  • The process binds to Heroku’s $PORT and the model loads once at startup.
  • Startup time, memory, concurrency, and worst-case inference duration have been measured.
  • Long-running work uses a queue and worker instead of blocking a web request.
  • Secrets are config vars; durable files and shared state live outside the dyno filesystem.
  • Logs, latency, errors, authentication, privacy, model versioning, and rollback have an owner and an operating plan.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.