October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PCOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk4 min

Same nvJPEG2000, Different Numbers: Timer Boundaries and Frames in Flight

nvJPEG2000 decode is asynchronous, timers cover different work, and concurrency changes throughput. Here is how to make timings comparable.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Two nvJPEG2000 decode timings can disagree even when both call the same library on the same GPU. The usual causes are what the clock brackets, whether asynchronous GPU work had finished when the clock stopped, and how many frames were being decoded at once. A figure is a measurement of one pipeline on one machine, not a constant of the codec.

Cause 1: the host call returns before the decode is done

NVIDIA’s documentation describes nvjpeg2kDecode() as asynchronous with respect to the host. GPU tasks are submitted to the CUDA stream you supply, so the function returning only means the work was queued. A timer that wraps just the call can report submission cost and call it decode time.

NVIDIA’s Quick Start Guide — nvJPEG2000 says: “cudaDeviceSynchronize() is required to complete the decoding process since nvjpeg2kDecode is asychronous with respect to the host.” (The spelling “asychronous” is NVIDIA’s.) The same guide says the input bitstream buffer must not be overwritten until decoding has completed. That matters in benchmark loops that recycle buffers.

  • Stop the clock only after completion: a stream synchronize, a device synchronize, or a CUDA event recorded on the stream and waited on.
  • If you use CUDA events, say so. Event timing measures stream time between two points, which differs from wall-clock time around the whole application loop.
  • Check the decoded output after completion, so a fast number is not a failed or incomplete decode.

Cause 2: the two timers cover different work

“Decode” can mean the codec alone or a whole file-to-pixels path. Before comparing numbers, list what sits inside each interval: bitstream parsing, host-to-device input transfer, CPU preparation, the GPU decode, device-to-host output transfer, any raw-pixel copy, and disk I/O.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The Fastvideo benchmark repository (2026) shows how much this matters, because it uses two modes:

Mode Boundaries Raw-pixel copy CPU work Disk
Single image Codec-side input and output boundaries Excluded Inside Outside
Multithreaded Host memory to host memory Included Inside Outside

The authors note that with concurrency you cannot isolate one frame’s stage from neighbouring work, so multithreaded numbers describe throughput of the whole pipeline, not the cost of a single stage.

Cause 3: different numbers of frames in flight

Frames in flight is how many images are being decoded concurrently. The benchmark writes “8×2” for eight CPU threads with two concurrent GPU frames per thread. Its nvJPEG2000 concurrency uses multiple decoder states, multiple streams, and asynchronous calls so that CPU work, transfers and GPU work overlap.

The authors tested 8×1, 8×2, 16×2, 8×4, 32×1 and 32×2. At a fixed thread count, raising frames in flight from one to two or four changed throughput by 1.02–1.20× for encoding and 1.12–2.06× for decoding across the included results. These are outcomes of that test, not expected gains elsewhere. The 2K lossy decode 8×1 point is left out of the decode range because it is unsettled (see below).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Report latency for one frame and throughput under concurrent load separately. Adding frames in flight can raise frames per second while each frame takes longer to finish, so they answer different questions.

Cause 4: different workloads and machines

The benchmark’s setup shows how many variables a comparable figure needs to pin down:

  • GPU and driver: NVIDIA GeForce RTX 4090 (24 GB), driver 610.88, 450 W maximum power.
  • Host: AMD Ryzen 9 7950X (16 cores, 32 logical), 128 GB RAM, Windows 11; measured CPU-to-GPU bus speed 25.2 GB/s.
  • Software: nvJPEG2000 0.11.0.51; Fastvideo SDK 0.23.1.0 with CUDA 13.3.
  • Images: 1920×1080 and 3840×2160, three channels, 8-bit.
  • Encoding settings: 32×32 code blocks, six levels, one quality layer, LRCP progression, no tiles.
  • Date: August 31, 2026. Each point had three series and a median; points whose repeats differed by more than 7% were re-measured up to two more times.

The authors do not cover other bit depths, 8K, multi-tile workloads or Jetson, and warn that numbers age with driver and library versions.

What the published decode figures look like

At the best multithreaded configuration, the benchmark reports decode throughput in frames per second as follows. Pairs are given as Fastvideo versus nvJPEG2000. Note the authors are the vendor of one of the compared SDKs, so treat these as owner-attributed results.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Task Fastvideo nvJPEG2000
2K lossy 1,024 1,033
2K lossless 436 438
4K lossy 394 428
4K lossless 145 134

In single-image mode the same benchmark reports nvJPEG2000 ahead in decode throughput on all four tasks. The ranking therefore depends on the timer mode, which is the article’s point in miniature.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

An unresolved cell: two states in one configuration

For nvJPEG2000 2K lossy decode at 8×1, the benchmark observed two clusters: 309 frames/s in nine process launches and 539 frames/s in eleven. The state persisted for an entire launch. Clock and temperature were the same, but the slower state used 45% more CPU time per frame. The authors say the cause is CPU-side and not established; the table reports the median, 310. Do not cite this point as a settled performance result, and do not assume a single run of your own sits in the faster cluster. Run several separate process launches.

A different experiment: multi-stream tile decoding

NVIDIA’s developer blog (2021) describes multi-tile decoding of Sentinel-2 imagery sized 10,980×10,980 and split into 121 tiles, with tiles decoded on separate streams. On a Quadro GV100 it reports an average decode time of 0.888854 ms with one stream and 0.227408 ms with ten, a reported 75% reduction for that dataset. This is a tiled, single-image workload on different hardware, so do not merge it with the RTX 4090 frame-throughput results.

Checklist for a reproducible nvJPEG2000 timing

  1. Define the interval: host-call, CUDA-event, or end-to-end. Never label a bare host-side duration around nvjpeg2kDecode() as completed decode.
  2. List what is inside it: parsing, input and output transfers, CPU preparation, output copy, disk.
  3. Synchronize before stopping the clock, and keep input buffers intact until then.
  4. State CPU threads, decode states, streams and frames in flight.
  5. Verify decoded output after completion.
  6. Repeat across separate process launches and publish the median and spread, not the best run.
  7. Record GPU, driver, library version, image dimensions, channels, bit depth, lossy or lossless mode and bitstream settings.
  8. Re-measure after changing any of the above.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.