Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

To pad a dataset, extend each variable-length sequence or array to a chosen target length or shape by adding a fill value. For a batch, the usual policy is to pad to that batch’s longest item, or to a fixed maximum when a predictable tensor shape is required. Keep the original lengths (or a padding mask), decide what happens to overlong inputs, and use a fill value that your model can distinguish from real data.

Padding makes samples stackable; it does not add genuine records, balance classes, or create new information.

What “padding a dataset” means

In machine-learning pipelines, padding normally changes representation shape. A short sequence such as [4, 7, 2] can become [4, 7, 2, 0, 0] so it can share a tensor with a five-item sequence. The added positions are placeholders, not observations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

This is different from synthetic oversampling or class balancing. If the problem is too few minority examples, use a data-balancing method; padding alone cannot fix it.

Choose the target length or shape

Pad to the longest item in each batch

Find the largest length among the samples that will be processed together and pad shorter items to that value. This minimizes filler and computation for that batch. The target can change from batch to batch, so downstream code must support dynamic dimensions.

Pad to a fixed maximum

Choose a documented maximum length or shape when a model, accelerator, export format, or serving interface needs predictable tensors. You must also define what happens when an input exceeds the maximum: reject it, truncate it under an explicit policy, or increase the maximum. Never silently discard values.

Leave data unpadded

If your framework and model accept ragged or packed inputs, no padding may be preferable. This avoids filler but shifts complexity to batching and downstream operations.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Strategy Shape Advantages Costs and risks
Batch-longest Changes per batch Less wasted memory and compute Requires dynamic-shape support; batches can still be dominated by an outlier
Fixed maximum Predictable Simple interfaces and static tensor shapes More filler; overlong items need truncation, rejection, or a larger limit
No padding Variable No placeholder values Requires ragged, packed, or per-item processing

Pad a NumPy array in Python

For one-dimensional numeric data, right-padding is a clear default: preserve the original values at the beginning and append the fill value. This pattern follows the documented mirdata 1.0.0 PyTorch Dataset example, which uses constant 0.0.

import numpy as np

def right_pad_1d(values, target_length, fill_value=0.0):
    """Right-pads a 1D array to target_length."""
    values = np.asarray(values)
    if values.ndim != 1:
        raise ValueError("values must be one-dimensional")
    if len(values) > target_length:
        raise ValueError("target_length is shorter than the input")
    return np.pad(
        values,
        (0, target_length - len(values)),
        mode="constant",
        constant_values=fill_value,
    )

samples = [np.array([1.2, 2.4]), np.array([3.1, 4.8, 5.0, 6.2])]
target = max(len(x) for x in samples)
padded = np.stack([right_pad_1d(x, target) for x in samples])
lengths = np.array([len(x) for x in samples])
mask = np.arange(target)[None, :] < lengths[:, None]

print(padded.shape)  # (2, 4)
print(lengths)       # [2 4]
print(mask)          # True for real values, False for padding

Compute the target from the partition you are actually preparing. For training, that may mean each batch; for a materialized dataset, it may mean the training split or a declared global maximum. Apply the same policy at validation, test, and inference time, while avoiding leakage from future or held-out data when the target is learned from data statistics.

Choose a fill value and preserve meaning

Numeric arrays

Zero is convenient and is the fill value in the mirdata example, but it is not universally correct. If zero is a valid measurement, a downstream model may confuse a real zero with a placeholder. Retain lengths or a Boolean mask, and ensure every operation that aggregates, normalizes, or scores the sequence ignores masked positions where appropriate.

Tokenized text

Use the tokenizer’s configured pad-token ID, not an assumed integer such as zero. Confirm that the model has a pad token and that the tokenizer’s padding side (left or right) matches the model and generation workflow.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Multidimensional samples

Define a target for every padded axis. For images or feature maps, padding height and width may require a tuple of per-axis widths; for time-frequency data, keep time and feature axes aligned. Pad labels or targets consistently when they are time-aligned with the inputs.

Padding tokenized sequences

Tokenizer APIs commonly expose three conceptual choices: pad to the longest sequence in the current batch, pad to a specified maximum length, or do not pad. Padding and truncation are separate settings. Configure truncation explicitly when using a maximum; otherwise an overlong sequence may fail or be shortened according to an unintended default.

  1. Tokenize the examples with the model’s tokenizer.
  2. Select batch-longest padding for efficient dynamic batches, or a fixed maximum for static shapes.
  3. Set truncation separately if inputs may exceed the maximum.
  4. Inspect the returned attention mask or equivalent mask and verify that padded positions are ignored.

Because tokenizer behavior and defaults can vary by installed library version, check the reference for the exact version in your environment.

Batch APIs and padded shapes

Dataset frameworks can perform padding while batching. MindSpore’s versioned padded_batch API uses pad_info to describe padded shapes and values; leaving shape entries unspecified allows padding to the largest sample shape in the batch. The references are version-specific (2.1 and 2.3.0), so verify argument names and defaults against the version you install.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Whatever API you use, make the resulting shape, fill value, padding side, and mask part of the pipeline contract rather than relying on an implicit default.

Lengths, masks, and aligned labels

Store the original length for every item, or construct a mask with one element per padded position. A mask is usually true for real data and false for filler, but follow your framework’s convention. Use it when computing attention, loss, pooling, metrics, normalization, or sequence reductions.

In sequence labeling and time-series work, pad features and aligned labels with compatible lengths. A feature sequence padded to 100 steps while its label sequence remains at 83 steps creates an indexing error or, worse, a silently misaligned target. Decide whether padded labels should be ignored with a special mask or ignore index supported by the loss.

Validate the transformation

  • Check the final rank and shape for a short item, a typical item, and the longest item.
  • Check dtype; a fill value should not unexpectedly convert integer data to floating point.
  • Verify left-versus-right padding and the pad-token ID.
  • Assert that no input longer than the target was silently truncated.
  • Compare a few rows before and after transformation to confirm that real values remain unchanged and in order.
  • Run a model or metric test with the mask enabled and disabled; the result should change only when padded positions would otherwise affect the calculation.

Common failure modes and fixes

The whole dataset is padded to an outlier

A single unusually long sample can fill every other sample with thousands of placeholders. Use batch-longest padding, length bucketing, or a justified fixed maximum instead.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A fixed length has no overlength policy

Choose explicitly between rejecting, truncating, or increasing the target. Log how many items take that branch so a data change cannot go unnoticed.

The fill value is a real value

Keep lengths or masks and make reductions mask-aware. For tokens, configure and verify the tokenizer’s pad token rather than treating zero as universal.

Padding direction is wrong

Right padding is common for stored numeric sequences; some language-model generation workflows require left padding. Match the model’s documented expectation and test an actual batch.

Labels are not padded with features

Pad or mask aligned targets under the same length policy. Add an assertion that every feature and label pair has the expected post-padding length.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Padding is mistaken for augmentation

Padding adds placeholders only. It does not increase the number of training examples, add class diversity, or create new measurements.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance, memory, and reproducibility

Padding increases tensor size, so memory and compute grow with the number of filler positions. Batch-longest padding reduces waste when lengths vary widely; grouping similarly sized samples can reduce it further. Fixed maxima simplify deployment but should be chosen from an explicit service requirement, not an arbitrary round number.

Record the target policy, maximum, truncation rule, fill value, padding side, tokenizer configuration, and library versions with the model or dataset metadata. Reproducibility depends on these details just as much as on the raw records.

Or skip the browser setup

If your workflow also needs repeatable screenshots of a web-based dataset view or report, ScreenshotNeo provides a one-request screenshot API. It is separate from array padding, but removes browser orchestration when a URL is the input:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Before capture it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools named take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Frequently Asked Questions

Should I pad before or after splitting a dataset?

Choose padding policy from the training or serving partition and apply it consistently after the split; do not use held-out data to set a global limit when that would leak information.

Can I use different padding lengths for training and inference?

Yes when the model and serving contract support it, but keep token IDs, masks, padding side, and truncation behavior compatible.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What is the difference between padding and truncation?

Padding adds positions to shorter items. Truncation removes positions from longer items; it must be configured as a separate, explicit decision.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.