Manipulate scraped data as a pipeline of record objects: use map() to normalize every item, filter() to keep only valid rows, reduce() to calculate totals or build indexes, and deliberate copy or mutation methods when editing positions. This approach makes filtering, de-duplication, pagination and export predictable instead of turning a scraper into a tangle of one-off loops.
What a scraped array should look like
Most web scrapers produce an array. Each element should be a record object with stable fields rather than an assortment of selector-specific values:
const raw = [
{ title: " Alpha ", href: "/a", priceText: "$12" },
{ title: "", href: "/missing", priceText: "" },
{ title: "Beta", href: "/b", priceText: "$9" }
];
Keeping one item per object lets every later stage address item.title, item.url and item.price consistently. Normalize missing fields to explicit values such as an empty string or null; do not rely on sparse arrays or empty slots, whose behavior differs across array methods.
How do I manipulate arrays in web scraping?
Use a clear sequence. First reshape raw selector output, then reject bad records, then calculate any aggregate, and only then paginate or export. Each stage has one job:
#1 Best Overall
- Normalize with
map(). It creates a new array populated with the callback result for each element. - Validate with
filter(). Returntruefor rows that meet your quality rules. - Aggregate with
reduce(). Produce a total, grouped object, count, or index. - Copy or edit positions deliberately. Use
slice()ortoSpliced()for non-mutating work; usesplice()only when in-place editing is intentional. - Serialize or paginate. Pass the resulting array to JSON, CSV, a database, or the next crawl stage.
Normalize records with map()
map() is a one-to-one transformation: the output has one element for each input element (including an explicit representation for invalid values if you choose to retain them). Trim text, resolve relative links and parse numeric fields in this stage.
const normalized = raw.map((item) => ({
title: (item.title ?? "").trim(),
url: new URL(item.href ?? "", "https://example.com").href,
price: Number((item.priceText ?? "").replace(/[^0-9.]/g, ""))
}));
Always use the returned array. Calling map() and discarding its result is an anti-pattern; use forEach() or for...of when the purpose is side effects. URL resolution and the exact currency parsing above are implementation choices, so adapt them to the site’s markup and locale.
How do I filter scraped results?
Use filter() with a predicate that states your acceptance rule. It returns a new array and leaves the source untouched.
const valid = normalized.filter((item) =>
item.title.length > 0 &&
item.url.startsWith("https://example.com/") &&
Number.isFinite(item.price)
);
Separate independent rules when debugging:
const withTitles = normalized.filter(item => item.title.length > 0);
const onHost = withTitles.filter(item => {
try { return new URL(item.url).hostname === "example.com"; }
catch { return false; }
});
Filtering after normalization means whitespace, relative URLs and malformed prices have already been handled. If you need to retain rejected rows for an audit, partition into accepted and rejected arrays in one loop rather than silently dropping them.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteHow do I remove duplicates from scraped data?
Choose a key that identifies one logical record, usually a canonical URL. A Map preserves insertion order and makes the policy explicit. The following keeps the first occurrence:
Rank #2
const unique = [...new Map(valid.map(item => [item.url, item])).values()];
To keep the last occurrence, later entries naturally overwrite earlier ones:
const lastSeen = [...new Map(valid.map(item => [item.url, item])).values()];
Normalize the key before deduplicating. Differences such as a trailing slash, fragment, or tracking query can represent the same page, but removing them is a site-specific decision; do not strip query parameters that change content. If URLs are not reliable identifiers, build a composite key such as title plus seller ID, and document collision behavior.
Should I use map(), filter(), or reduce()?
| Method | Purpose | Returns | Mutates source? |
|---|---|---|---|
map() |
One-to-one transformation or field reshaping | New array | No |
filter() |
Keep records matching a predicate | New array | No |
reduce() |
Total, grouping, indexing or another accumulation | Any value | Only if your callback mutates the accumulator or external state |
slice() |
Copy a range or make a shallow copy | New array | No |
splice() |
Insert, replace or delete by position | Removed elements | Yes |
toSpliced() |
Non-mutating positional edit where supported | New array | No |
Chaining these operations keeps intent visible:
const records = raw
.map(item => ({
title: (item.title ?? "").trim(),
url: new URL(item.href ?? "", "https://example.com").href,
price: Number((item.priceText ?? "").replace(/[^0-9.]/g, ""))
}))
.filter(item => item.title && Number.isFinite(item.price));
Use reduce() for totals, groups and indexes
Totals
const total = records.reduce((sum, item) => sum + item.price, 0);
Counts by category
const counts = records.reduce((groups, item) => {
const category = item.category ?? "uncategorized";
groups[category] = (groups[category] ?? 0) + 1;
return groups;
}, {});
Index by URL
const byUrl = records.reduce((index, item) => {
index[item.url] = item;
return index;
}, {});
An initial value such as 0, {} or [] avoids special handling for an empty input array and documents the accumulator’s type. For large crawls, an object or Map index can be more useful than repeatedly scanning the array.
How do I edit an array without changing the original?
JavaScript indexes start at zero, so the first element is index 0. Use slice(start, end) for a non-destructive range or shallow copy:
const firstPage = records.slice(0, 20);
const copy = records.slice();
Where supported by your runtime, toSpliced(start, deleteCount, ...items) performs a splice-like edit and returns a new array:
const withoutFirst = records.toSpliced(0, 1);
const corrected = records.toSpliced(2, 1, replacement);
Use splice() when changing the working array is the desired behavior:
const working = records.slice();
working.splice(0, 1); // delete one
working.splice(1, 0, inserted); // insert
working.splice(2, 1, replacement); // replace
push(), pop(), shift(), unshift(), reverse() and splice() mutate. A shallow copy protects the array container, not nested objects; clone nested data separately if later edits must not affect either version.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Delete by value safely
Find the index and check for -1 before splicing. Otherwise, a failed lookup can accidentally remove the last element when passed to a different indexing expression.
const working = records.slice();
const index = working.findIndex(item => item.url === targetUrl);
if (index !== -1) working.splice(index, 1);
If every matching record should go, use filter() instead:
const remaining = records.filter(item => item.url !== targetUrl);
Pagination and export
After validation and de-duplication, paginate with zero-based offsets:
Rank #4
function page(items, pageNumber, pageSize) {
if (pageNumber < 1 || pageSize < 1) throw new RangeError("Positive page values required");
const start = (pageNumber - 1) * pageSize;
return items.slice(start, start + pageSize);
}
const pageTwo = page(unique, 2, 20);
JSON export preserves types:
const json = JSON.stringify(unique, null, 2);
For CSV, escape quotes, commas and line breaks according to the format’s rules before joining fields. Export only after the final schema and ordering are settled so downstream jobs receive predictable columns.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Troubleshooting array pipelines
- Prices become
NaN: normalize missing text and remove locale-specific currency characters before conversion; retain invalid rows for inspection rather than treatingNaNas zero. - Duplicates remain: deduplicate on a normalized canonical key, not the raw href; decide whether query parameters identify different content.
- The next stage sees changed data: look for mutating methods, especially
splice(),sort()andreverse(); useslice()ortoSpliced()for a copy. - One item disappears after deletion: check the result of
findIndex()before callingsplice(). - Empty input throws in
reduce(): provide an initial accumulator. - Only some fields are missing: normalize every record to the same keys instead of creating sparse arrays.
- URLs point to the wrong host: resolve relative links against the page’s actual base URL and validate the hostname before export.
Performance and reliability practices
Each chained pass is generally linear in the number of records, but several passes allocate several arrays. That trade-off improves readability and makes failures easier to locate. For very large crawls, combine compatible checks in one loop, stream records in batches, or aggregate while parsing so the entire result does not remain in memory. Keep raw input until validation is complete when reproducibility matters, and log counts after normalization, filtering and deduplication to detect selector regressions.
Do not confuse array manipulation with crawling policy. Respect the target site’s access rules, handle request failures separately from empty results, and attach source-page metadata to records when later auditing is important.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
If your workflow mainly needs page images or PDFs rather than DOM records, ScreenshotNeo provides a single screenshot API call. It accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status.
Use the same URL with cURL (see the ScreenshotNeo documentation):
Free tools Windows power users keep installed
One-click scans. No signup required.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients. Its Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Sign up free for ScreenshotNeo.
Best Value
FAQ
Does map() change the objects it receives?
It creates a new array, but your callback can still mutate a nested object. Return a new object when isolation matters.
When is a loop better than a chain?
Use a loop when one pass must classify, count and collect several outputs, or when memory pressure makes intermediate arrays costly. Keep the same explicit stages in your naming and tests.
Can I deduplicate by title?
You can, but titles may legitimately repeat. Prefer a stable URL or source identifier and document the chosen winner.
Recommended Free Tools
Frequently Asked Questions
Does map() change the objects it receives?
It creates a new array, but your callback can still mutate a nested object. Return a new object when isolation matters.
When is a loop better than a chain?
Use a loop when one pass must classify, count and collect several outputs, or when memory pressure makes intermediate arrays costly.
Can I deduplicate by title?
You can, but titles may legitimately repeat. Prefer a stable URL or source identifier and document the chosen winner.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errors




