October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
async crawling

Build a Flask Callback Server for Async Crawling with MySQL

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build the callback receiver as a short-lived Flask endpoint: validate the crawler’s request while Flask’s request context is active, write the callback and job state in one MySQL transaction, commit, acknowledge according to the crawler’s contract, and send any slow follow-up work to a durable queue. Do not try to keep an asyncio task alive after returning the response.

Decide what the callback endpoint owns

Before writing code, obtain the crawler’s actual callback contract. The title alone cannot determine the route, HTTP method, signature or authentication scheme, payload fields, stable crawl identifier, retry behavior, or required acknowledgment body. Treat those as integration inputs rather than defaults.

  • Receiver: authenticate the sender, parse and validate the payload, and reject malformed data.
  • Persistence: save the raw (or appropriately redacted) callback, result metadata, and job-state transition atomically.
  • Continuation: enqueue serialized data for work that is too slow for the request.
  • Operations: correlate logs, expose state transitions, and alert on database or queue failures.

Keep credentials in deployment configuration, never in source code or ordinary logs. Decide how long callback payloads and job records are retained, which queue and worker runtime you will operate, and what delivery guarantees your crawler requires.

Why an async Flask view is not a durable worker

Flask is a WSGI application. Its documentation explains that one worker handles one request/response cycle. An async view can perform concurrent I/O during that cycle, but it does not increase the number of requests that worker handles at once and does not turn unfinished work into a durable job.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
CanaKit Raspberry Pi 5 Starter Kit PRO - Turbine Black (128GB Edition) (8GB RAM)
  • Includes Raspberry Pi 5 with 2.4Ghz 64-bit quad-core CPU (8GB RAM)
  • Includes 128GB Micro SD Card pre-loaded with 64-bit Raspberry Pi OS, USB MicroSD Card Reader
  • CanaKit Turbine Black Case for the Raspberry Pi 5
  • CanaKit Low Noise Bearing System Fan
  • Mega Heat Sink - Black Anodized

The Flask project specifically advises: “If you wish to use background tasks it is best to use a task queue to trigger background work, rather than spawn tasks in a view function.” A task started with asyncio.create_task() in a normal view can be cancelled when the view’s event loop ends or disappear when the process restarts.

Choose the execution pattern

Persist before acknowledging

For bounded validation and database work, perform both in the request, commit, and then return the crawler’s documented acknowledgment. The caller’s success response now means the data was durably written, but the request remains open for the database operation.

Queue continued work

If parsing, enrichment, exports, or other processing can be slow, commit the callback and an initial state such as queued, then submit explicit serialized task data to a durable queue. A separate worker moves the record through running, succeeded, or failed. The queue and database failure behavior must be designed together: a database commit can succeed while queue submission fails, so record that condition and provide a reconciliation or retry path rather than silently losing the job.

Pattern Acknowledgment latency Restart durability Operational cost
Validate and write in Flask Includes database time Persisted after commit Small; request must handle database failures
Commit then queue a worker task Shorter after the commit Depends on durable queue and state records Higher; operate queue, worker, retries, and reconciliation

MySQL schema for callbacks and jobs

Use a stable identifier supplied by the crawler (for example, a callback or crawl ID) and enforce uniqueness. Confirm that the identifier is truly stable across retries before relying on it for deduplication.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
CanaKit Raspberry Pi 4 4GB Starter PRO Kit - 4GB RAM
  • Includes Raspberry Pi 4 4GB Model B with 1.5GHz 64-bit quad-core CPU (4GB RAM)
  • Includes Pre-Loaded 32GB EVO+ Micro SD Card (Class 10), USB MicroSD Card Reader
  • CanaKit Premium High-Gloss Raspberry Pi 4 Case with Integrated Fan Mount, CanaKit Low Noise Bearing System Fan
  • CanaKit 3.5A USB-C Raspberry Pi 4 Power Supply (US Plug) with Noise Filter, Set of Heat Sinks, Display Cable - 6 foot (Supports up to 4K60p)
  • CanaKit USB-C PiSwitch (On/Off Power Switch for Raspberry Pi 4)
CREATE TABLE crawl_jobs (
  id BIGINT UNSIGNED NOT NULL AUTO_INCREMENT,
  crawler_job_id VARCHAR(191) NOT NULL,
  status ENUM('received','queued','running','succeeded','failed') NOT NULL,
  result_json JSON NULL,
  last_error TEXT NULL,
  created_at TIMESTAMP(6) NOT NULL DEFAULT CURRENT_TIMESTAMP(6),
  updated_at TIMESTAMP(6) NOT NULL DEFAULT CURRENT_TIMESTAMP(6)
    ON UPDATE CURRENT_TIMESTAMP(6),
  PRIMARY KEY (id),
  UNIQUE KEY uq_crawler_job (crawler_job_id)
) ENGINE=InnoDB;

CREATE TABLE crawl_callbacks (
  id BIGINT UNSIGNED NOT NULL AUTO_INCREMENT,
  crawler_job_id VARCHAR(191) NOT NULL,
  callback_id VARCHAR(191) NOT NULL,
  payload_json JSON NOT NULL,
  received_at TIMESTAMP(6) NOT NULL DEFAULT CURRENT_TIMESTAMP(6),
  PRIMARY KEY (id),
  UNIQUE KEY uq_callback (callback_id),
  KEY ix_callback_job (crawler_job_id),
  CONSTRAINT fk_callback_job FOREIGN KEY (crawler_job_id)
    REFERENCES crawl_jobs (crawler_job_id)
) ENGINE=InnoDB;

If the crawler has no callback ID, use another documented, stable event key; do not manufacture an identifier from mutable payload text without considering collisions. Decide whether duplicate callbacks should return the same accepted outcome, update the existing result, or be recorded as a separate event.

Install dependencies and configure the pool

python -m venv .venv
. .venv/bin/activate
pip install Flask mysql-connector-python
export MYSQL_HOST=127.0.0.1
export MYSQL_PORT=3306
export MYSQL_DATABASE=crawling
export MYSQL_USER=crawler_app
export MYSQL_PASSWORD='use-a-secret-manager'
export CALLBACK_TOKEN='replace-with-your-crawler-secret'

Connector/Python has autocommit disabled by default, so successful related writes require an explicit commit(). Roll back on exceptions. Its pooling module creates a fixed-size pool; requesting a connection when every slot is in use raises PoolError. Size the pool against your deployment’s concurrency and MySQL connection limits, and handle exhaustion deliberately.

Complete Flask receiver

import json
import os
from flask import Flask, jsonify, request
from mysql.connector import pooling, Error

app = Flask(__name__)

pool = pooling.MySQLConnectionPool(
    pool_name="crawler_pool",
    pool_size=int(os.getenv("MYSQL_POOL_SIZE", "5")),
    pool_reset_session=True,
    host=os.environ["MYSQL_HOST"],
    port=int(os.getenv("MYSQL_PORT", "3306")),
    database=os.environ["MYSQL_DATABASE"],
    user=os.environ["MYSQL_USER"],
    password=os.environ["MYSQL_PASSWORD"],
)

def valid_token(req):
    # Replace this with the crawler's documented signature or auth protocol.
    expected = os.environ["CALLBACK_TOKEN"]
    return req.headers.get("Authorization") == f"Bearer {expected}"

@app.post("/callbacks/crawl")
def crawl_callback():
    if not valid_token(request):
        return jsonify(error="unauthorized"), 401

    data = request.get_json(silent=True)
    if not isinstance(data, dict):
        return jsonify(error="JSON object required"), 400

    crawler_job_id = data.get("job_id")
    callback_id = data.get("callback_id")
    if not isinstance(crawler_job_id, str) or not crawler_job_id:
        return jsonify(error="job_id is required"), 400
    if not isinstance(callback_id, str) or not callback_id:
        return jsonify(error="callback_id is required"), 400

    conn = None
    try:
        conn = pool.get_connection()
        cur = conn.cursor()
        cur.execute(
            "INSERT INTO crawl_jobs (crawler_job_id, status) "
            "VALUES (%s, 'received') "
            "ON DUPLICATE KEY UPDATE updated_at = CURRENT_TIMESTAMP(6)",
            (crawler_job_id,),
        )
        cur.execute(
            "INSERT INTO crawl_callbacks "
            "(crawler_job_id, callback_id, payload_json) VALUES (%s, %s, %s)",
            (crawler_job_id, callback_id, json.dumps(data)),
        )
        cur.execute(
            "UPDATE crawl_jobs SET result_json=%s, status='queued' "
            "WHERE crawler_job_id=%s",
            (json.dumps(data.get("result")), crawler_job_id),
        )
        conn.commit()
    except Error as exc:
        if conn is not None:
            conn.rollback()
        # Log a correlation ID and the database error without secrets or payloads.
        return jsonify(error="persistence_failed"), 503
    finally:
        if conn is not None:
            conn.close()  # returns a pooled connection

    # Enqueue explicit data here, after commit, if continued work is required.
    # queue.publish({"crawler_job_id": crawler_job_id, "callback_id": callback_id})
    return jsonify(accepted=True, callback_id=callback_id), 202

if __name__ == "__main__":
    app.run()

The field names and 202 response above are examples only. Match them to the crawler’s documented protocol. If the sender expects a different status or body, use that contract. For a callback that is already present, catch the duplicate-key condition and return your defined idempotent outcome instead of inserting a second result. Do not acknowledge a callback before the transaction commits.

Keep deferred work outside Flask

At enqueue time, copy primitive values such as IDs, not Flask’s request proxy. Flask pushes the request context for handling and pops it after response processing; teardown callbacks run even when an exception escapes. A worker cannot safely access that context later.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
UCTRONICS 19” 1U Rack Mount for Raspberry Pi with SSD Mounting Brackets, Thumbscrews Front Removable Bracket Supports Up to 4 Raspberry Pi 5, 3B/3B+, 4B and 4 SSDs, Option SD Card Adapter
  • Design for Raspberry Pi: Supports installation of 4 Raspberry Pis and 4 ssds, compatible with any 2.5” Solid State Drive (7mm/9mm) and Rpi 4B/3B+, and other B/B+ models.
  • The SSD mounting bracket also has two holes reserved for the SD card extension adapter ASIN: B09CKRDFTH, which allows you to access the SD card from the front of the rack.
  • Easy to Setup: Just use two included thumbscrews to mount the rackmount, which adopts a screw-in design, which helps you install and replace quickly and easily, no tools needed!
  • Applications: This is a hardware solution to get ingenious use of the Raspberry Pi, with this kit and open source software OpenMediaVault, you can use the Pi as a NAS Server, Surveillance station, or even a Web server.
  • Optional accessories: Single mounting bracket: B09GFQLPTY; Micro SD card extension adapter ASIN: B09CKRDFTH. I/O Panel: B09FXRQPFM
# worker.py (queue library is deployment-specific)
def process(task):
    job_id = task["crawler_job_id"]
    # Open its own database connection, mark running, do bounded work,
    # then commit succeeded or failed and record a safe error message.
    # Make the operation idempotent: a redelivered task must not duplicate output.
    pass

Use queue acknowledgments and retry limits appropriate to the selected queue. Persist state transitions so a process restart leaves enough information to resume or reconcile. Add alerts for repeatedly failed jobs, queue backlog, database connection-pool exhaustion, and callbacks stuck in an intermediate state.

Connection management and transaction rules

One connection per operation

Opening a connection for each callback is simple and can be adequate at low volume, but connection setup adds overhead and can pressure MySQL under bursts.

Use a fixed pool

A pool reuses connections and makes capacity explicit. Always close the acquired connection in a finally block, including validation, SQL, and queue-error paths. A pool does not remove the database limit: when exhausted, fail or retry according to a bounded policy and return an acknowledgment that causes the crawler to retry only if its contract defines that behavior.

Security, validation, and observability

  • Require the crawler’s documented authentication or signature check before parsing sensitive fields.
  • Use parameterized SQL; never concatenate callback values into statements.
  • Apply payload-size limits and schema validation suitable for the crawler.
  • Redact credentials, authorization headers, and unnecessary personal data from logs.
  • Include a correlation ID in logs and state records, and record receive, commit, enqueue, start, success, and failure times.
  • Use TLS at the public edge and restrict database credentials to the required schema operations.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common failures and fixes

The crawler times out and sends duplicates

Measure the time spent validating and committing, then verify the crawler’s retry policy. Keep the transaction short and enforce the callback uniqueness key. Return the documented acknowledgment for an already-processed callback.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
Pironman 5-MAX Raspberry Pi 5 Case Dual NVMe M.2 SSD PCIe, Mini PC NAS RAID 0/1 Hailo-8L AI Accelerator PWM Tower Cooler+Dual RGB Fans, OLED Module, Safe Shutdown, Standard HDMI (RPI5 Not Included)
  • [ULTIMATE RASPBERRY PI 5 CASE & MINI PC] - Unlock the full potential of your Raspberry Pi 5 with the Pironman 5-MAX — the most advanced Raspberry Pi 5 Case for power users. This high-performance Raspberry Pi 5 Cooling Case features dual NVMe M.2 slots with RAID 0/1 support, AI accelerator compatibility ( e.g. Hailo-8l M.2 AI), a PCIe Gen2 switch, a PWM tower cooler + dual RGB fans and a smart OLED display. With its dual transparent panels and optimized cable management (including full-size HDMI), it’s the ideal Raspberry Pi 5 Enclosure for building a high-speed NAS, AI edge computing device, or Home Assistant hub. (Raspberry Pi NOT Included)
  • [DUAL NVMe M.2 SLITS & NAS RAID SUPPORT] - Supercharge your storage with the best Raspberry Pi 5 NVMe Case solution. Featuring two expandable NVMe M.2 slots (2230-2280) powered by a built-in PCIe Gen2 switch, this Raspberry Pi 5 NAS Case supports RAID 0/1 for ultra-fast data setups. Whether you're using a high-speed NVMe SSD or a Hailo-8L AI accelerator, Pironman 5-MAX delivers the ultimate performance boost for advanced Raspberry Pi 5 AI applications and edge computing
  • [ADVANCED COOLING SYSTEM] - Engineered for high-performance builds, Pironman 5-MAX features a powerful tower cooler, one PWM fan, and dual RGB fans for enhanced airflow. The dual transparent panel design improves ventilation while showcasing vibrant RGB lighting. Ideal for cooling both the Raspberry Pi 5 and dual NVMe SSDs or AI accelerators like Hailo-8L, it ensures stable operation under heavy workloads with low noise and long-term durability
  • [SMART OLED DISPLAY WITH VIBRATION WAKE-UP] - Pironman 5-MAX features a 0.96" OLED screen that delivers real-time system insights including CPU usage, memory, temperature, IP address, and disk status. With customizable display options and auto sleep mode, the screen can be instantly reactivated by a light tap thanks to the built-in vibration sensor—offering a smarter and more interactive experience
  • [ENHANCED FUNCTIONALITY] - Pironman 5-MAX empowers your Raspberry Pi 5 with advanced features like safe shutdown via a metal power button, customizable RGB lighting, dual full-size HDMI ports, vibration-triggered OLED wake-up, and an external GPIO extender. It also includes RTC battery support for timekeeping and seamless Home Assistant integration. With detailed guides, online tutorials, and full technical support from SunFounder, setup and use are effortless and worry-free

Rows are missing after a successful response

Check that conn.commit() runs after both writes. Connector/Python does not commit automatically by default. Ensure no exception path returns success before rollback or commit.

PoolError appears under load

All pool slots are busy. Close every acquired connection, inspect transaction duration, and size the fixed pool within the MySQL server’s connection budget. Do not hide exhaustion with unbounded retries.

Worker code raises “working outside of request context”

The worker retained Flask’s request proxy. Extract and validate the needed values in the route and pass a serialized task object instead.

Callbacks remain queued forever

Inspect queue publication, worker health, and the database state transition. Record queue-publication failures after the database commit and run a reconciliation process that republishes eligible records.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A browser-based crawl result is blocked or polluted

Consent dialogs, newsletter overlays, chat widgets, bot checks, and failed page loads are capture concerns rather than Flask transaction concerns. Keep the callback receiver independent of the browser automation provider so either side can be replaced.

Or skip the browser setup

If your crawler’s callback ultimately needs website screenshots, ScreenshotNeo can produce the capture with one request while your Flask service remains focused on receiving and persisting completion events. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing result.

For a direct call, see the ScreenshotNeo API documentation:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

It also provides an MCP server with take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots. Sign up at ScreenshotNeo’s free account page.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Verification checklist

  1. Confirm the crawler’s authentication, schema, stable identifier, retry rules, and acknowledgment.
  2. Send an invalid-authentication request and verify it is rejected without a database write.
  3. Send a valid callback and verify the callback row and job transition commit together.
  4. Replay the same callback and verify the defined idempotent result.
  5. Force a SQL error and confirm rollback plus a retry-compatible response.
  6. Exhaust the pool in a staging load test and verify bounded handling and alerts.
  7. Stop the worker after enqueueing and confirm durable queue recovery and state reconciliation.
  8. Inspect logs to ensure secrets and sensitive payload fields are absent.

Frequently Asked Questions

Should the callback endpoint return 200 or 202?

Use the status and response body required by the crawler’s documented callback contract; the example’s 202 is not universal.

Can I use Flask’s development server in production?

No. Deploy Flask behind a production WSGI server and an appropriate TLS-capable edge; the receiver/worker separation remains the same.

How long should callback payloads be retained?

Choose retention from your debugging, audit, privacy, and storage requirements; the crawler contract and deployment determine the correct period.

Quick Recap

Bestseller No. 1
CanaKit Raspberry Pi 5 Starter Kit PRO - Turbine Black (128GB Edition) (8GB RAM)
CanaKit Raspberry Pi 5 Starter Kit PRO - Turbine Black (128GB Edition) (8GB RAM)
Includes Raspberry Pi 5 with 2.4Ghz 64-bit quad-core CPU (8GB RAM); CanaKit Turbine Black Case for the Raspberry Pi 5
$259.95
Bestseller No. 2
CanaKit Raspberry Pi 4 4GB Starter PRO Kit - 4GB RAM
CanaKit Raspberry Pi 4 4GB Starter PRO Kit - 4GB RAM
Includes Raspberry Pi 4 4GB Model B with 1.5GHz 64-bit quad-core CPU (4GB RAM); Includes Pre-Loaded 32GB EVO+ Micro SD Card (Class 10), USB MicroSD Card Reader
$159.99

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.