Recommended Free Tools
Two nvJPEG2000 decode timings can disagree even when both call the same library on the same GPU. The usual causes are what the clock brackets, whether asynchronous GPU work had finished when the clock stopped, and how many frames were being decoded at once. A figure is a measurement of one pipeline on one machine, not a constant of the codec.
Cause 1: the host call returns before the decode is done
NVIDIA’s documentation describes nvjpeg2kDecode() as asynchronous with respect to the host. GPU tasks are submitted to the CUDA stream you supply, so the function returning only means the work was queued. A timer that wraps just the call can report submission cost and call it decode time.
NVIDIA’s Quick Start Guide — nvJPEG2000 says: “cudaDeviceSynchronize() is required to complete the decoding process since nvjpeg2kDecode is asychronous with respect to the host.” (The spelling “asychronous” is NVIDIA’s.) The same guide says the input bitstream buffer must not be overwritten until decoding has completed. That matters in benchmark loops that recycle buffers.
- Stop the clock only after completion: a stream synchronize, a device synchronize, or a CUDA event recorded on the stream and waited on.
- If you use CUDA events, say so. Event timing measures stream time between two points, which differs from wall-clock time around the whole application loop.
- Check the decoded output after completion, so a fast number is not a failed or incomplete decode.
Cause 2: the two timers cover different work
“Decode” can mean the codec alone or a whole file-to-pixels path. Before comparing numbers, list what sits inside each interval: bitstream parsing, host-to-device input transfer, CPU preparation, the GPU decode, device-to-host output transfer, any raw-pixel copy, and disk I/O.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →#1 Best Overall
The Fastvideo benchmark repository (2026) shows how much this matters, because it uses two modes:
| Mode | Boundaries | Raw-pixel copy | CPU work | Disk |
|---|---|---|---|---|
| Single image | Codec-side input and output boundaries | Excluded | Inside | Outside |
| Multithreaded | Host memory to host memory | Included | Inside | Outside |
The authors note that with concurrency you cannot isolate one frame’s stage from neighbouring work, so multithreaded numbers describe throughput of the whole pipeline, not the cost of a single stage.
Cause 3: different numbers of frames in flight
Frames in flight is how many images are being decoded concurrently. The benchmark writes “8×2” for eight CPU threads with two concurrent GPU frames per thread. Its nvJPEG2000 concurrency uses multiple decoder states, multiple streams, and asynchronous calls so that CPU work, transfers and GPU work overlap.
The authors tested 8×1, 8×2, 16×2, 8×4, 32×1 and 32×2. At a fixed thread count, raising frames in flight from one to two or four changed throughput by 1.02–1.20× for encoding and 1.12–2.06× for decoding across the included results. These are outcomes of that test, not expected gains elsewhere. The 2K lossy decode 8×1 point is left out of the decode range because it is unsettled (see below).
Rank #2
Report latency for one frame and throughput under concurrent load separately. Adding frames in flight can raise frames per second while each frame takes longer to finish, so they answer different questions.
Cause 4: different workloads and machines
The benchmark’s setup shows how many variables a comparable figure needs to pin down:
- GPU and driver: NVIDIA GeForce RTX 4090 (24 GB), driver 610.88, 450 W maximum power.
- Host: AMD Ryzen 9 7950X (16 cores, 32 logical), 128 GB RAM, Windows 11; measured CPU-to-GPU bus speed 25.2 GB/s.
- Software: nvJPEG2000 0.11.0.51; Fastvideo SDK 0.23.1.0 with CUDA 13.3.
- Images: 1920×1080 and 3840×2160, three channels, 8-bit.
- Encoding settings: 32×32 code blocks, six levels, one quality layer, LRCP progression, no tiles.
- Date: August 31, 2026. Each point had three series and a median; points whose repeats differed by more than 7% were re-measured up to two more times.
The authors do not cover other bit depths, 8K, multi-tile workloads or Jetson, and warn that numbers age with driver and library versions.
What the published decode figures look like
At the best multithreaded configuration, the benchmark reports decode throughput in frames per second as follows. Pairs are given as Fastvideo versus nvJPEG2000. Note the authors are the vendor of one of the compared SDKs, so treat these as owner-attributed results.
Rank #3
| Task | Fastvideo | nvJPEG2000 |
|---|---|---|
| 2K lossy | 1,024 | 1,033 |
| 2K lossless | 436 | 438 |
| 4K lossy | 394 | 428 |
| 4K lossless | 145 | 134 |
In single-image mode the same benchmark reports nvJPEG2000 ahead in decode throughput on all four tasks. The ranking therefore depends on the timer mode, which is the article’s point in miniature.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.An unresolved cell: two states in one configuration
For nvJPEG2000 2K lossy decode at 8×1, the benchmark observed two clusters: 309 frames/s in nine process launches and 539 frames/s in eleven. The state persisted for an entire launch. Clock and temperature were the same, but the slower state used 45% more CPU time per frame. The authors say the cause is CPU-side and not established; the table reports the median, 310. Do not cite this point as a settled performance result, and do not assume a single run of your own sits in the faster cluster. Run several separate process launches.
A different experiment: multi-stream tile decoding
NVIDIA’s developer blog (2021) describes multi-tile decoding of Sentinel-2 imagery sized 10,980×10,980 and split into 121 tiles, with tiles decoded on separate streams. On a Quadro GV100 it reports an average decode time of 0.888854 ms with one stream and 0.227408 ms with ten, a reported 75% reduction for that dataset. This is a tiled, single-image workload on different hardware, so do not merge it with the RTX 4090 frame-throughput results.
Quick Recap
Checklist for a reproducible nvJPEG2000 timing
- Define the interval: host-call, CUDA-event, or end-to-end. Never label a bare host-side duration around
nvjpeg2kDecode()as completed decode. - List what is inside it: parsing, input and output transfers, CPU preparation, output copy, disk.
- Synchronize before stopping the clock, and keep input buffers intact until then.
- State CPU threads, decode states, streams and frames in flight.
- Verify decoded output after completion.
- Repeat across separate process launches and publish the median and spread, not the best run.
- Record GPU, driver, library version, image dimensions, channels, bit depth, lossy or lossless mode and bitstream settings.
- Re-measure after changing any of the above.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problems




