Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Decode the base64 string into PDF bytes, parse those bytes with a PDF library, then map the extracted text into the JSON shape your application needs. In Node.js, use Buffer.from(value, 'base64'); in a browser, decode to a Uint8Array before passing the data to PDF.js. Extraction reads a PDF’s existing text layer—it does not recognize text in scanned page images.

What the conversion involves

Base64 is an encoding of bytes, not a text format for the PDF itself. A parser expects the decoded PDF bytes. The process is therefore:

  1. Get the base64 value, removing any data-URI prefix if the input has one.
  2. Decode it to binary bytes using the runtime’s base64 decoder.
  3. Pass those bytes to a PDF parser and retrieve text from each page.
  4. Map the page results into your own JSON schema and serialize them if needed.

JSON is not a special output mode shared by PDF parsers. It is an application-defined structure: for example, an array of page numbers and extracted text, or a richer object containing text-item coordinates.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Extract text in Node.js with a base64 string

Node’s Buffer API decodes base64 directly. Its decoder also accepts the URL-safe base64 alphabet and ignores whitespace. The example below uses the documented pdf.js-extract buffer API and returns one text string per page.

#1 Best Overall
WavePad Audio Editing Software - Professional Audio and Music Editor for Anyone [Download]
  • Full-featured professional audio and music editor that lets you record and edit music, voice and other audio recordings
  • Add effects like echo, amplification, noise reduction, normalize, equalizer, envelope, reverb, echo, reverse and more
  • Supports all popular audio formats including, wav, mp3, vox, gsm, wma, real audio, au, aif, flac, ogg and more
  • Sound editing functions include cut, copy, paste, delete, insert, silence, auto-trim and more
  • Integrated VST plugin support gives professionals access to thousands of additional tools and effects
  1. Install the package: npm install pdf.js-extract.
  2. Put the base64 PDF value in base64Pdf (or replace that variable with your input).
  3. Decode it, call extractBuffer, and map the pages into your chosen JSON structure.
import { PDFExtract } from 'pdf.js-extract';

const base64Pdf = process.env.PDF_BASE64;
if (!base64Pdf) throw new Error('Set PDF_BASE64 to the base64-encoded PDF');

const pdfBuffer = Buffer.from(base64Pdf, 'base64');
const extractor = new PDFExtract();

extractor.extractBuffer(pdfBuffer, {}, (err, data) => {
  if (err) throw err;

  const result = {
    pages: data.pages.map((page) => ({
      page: page.info.num,
      text: page.content.map((item) => item.str).join(' '),
    })),
  };

  console.log(JSON.stringify(result, null, 2));
});

The package’s documented callback API is shown here; check the installed package documentation if your version or module setup differs. This example is an implementation pattern, not a claim that every PDF has been tested. In a production service, handle extraction errors at the request boundary rather than letting an unhandled callback error terminate the process.

If the input is a data URI such as data:application/pdf;base64,..., remove the prefix before decoding. A simple boundary check can make that explicit:

const payload = base64Pdf.replace(/^data:application/pdf;base64,/i, '');
const pdfBuffer = Buffer.from(payload, 'base64');

Only strip a prefix your input contract allows; avoid silently changing arbitrary input. Node’s documented decoding behavior is described in the Node.js Buffer documentation. The extraction package documents page text items, buffer extraction, and row-grouping utilities at pdf.js-extract on npm.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Extract page text in a browser with PDF.js

PDF.js accepts binary document data and recommends typed-array data for memory use. If your application already has the PDF as base64, decode it and pass a Uint8Array as the loading task’s data. The following is the core sequence to use after PDF.js has been loaded and exposed as pdfjsLib; how you load the library depends on your application’s bundler or script setup.

async function extractPdfJson(base64Input) {
  const payload = base64Input.replace(/^data:application/pdf;base64,/i, '');
  const binary = atob(payload);
  const bytes = new Uint8Array(binary.length);

  for (let i = 0; i < binary.length; i++) {
    bytes[i] = binary.charCodeAt(i);
  }

  const pdf = await pdfjsLib.getDocument({ data: bytes }).promise;
  const pages = [];

  for (let pageNumber = 1; pageNumber <= pdf.numPages; pageNumber++) {
    const page = await pdf.getPage(pageNumber);
    const content = await page.getTextContent();
    pages.push({
      page: pageNumber,
      text: content.items.map((item) => item.str).join(' '),
    });
  }

  return { pages };
}

const result = await extractPdfJson(base64Pdf);
console.log(JSON.stringify(result));

This maps text items in their returned sequence and joins them with spaces. It is a useful starting point for page text, but it does not reconstruct every layout or semantic relationship. PDF.js’s API describes binary input and document loading at the PDF.js API documentation; its examples and FAQ explain the broader loading workflow at PDF.js examples and PDF.js FAQ.

Mozilla’s FAQ advises decoding base64 before supplying data: “If you have base64 encoded data, please decode it first — not all browsers have atob or data URI scheme support.” If the browser environment does not provide atob, use an available base64 decoder appropriate to that environment rather than assuming the global exists.

Choose a JSON shape that matches the task

A per-page structure is often easier to inspect than a single concatenated string because it preserves page boundaries:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
{
  "pages": [
    { "page": 1, "text": "First page text" },
    { "page": 2, "text": "Second page text" }
  ]
}

If a downstream consumer needs one field, combine the page strings deliberately, for example by joining the page objects’ text values with a newline. If it needs positioning or line grouping, retain the parser’s item coordinates or use the package’s documented grouping utilities instead of flattening early. Neither PDF.js nor pdf.js-extract defines a universal “PDF text JSON” schema.

Handle scanned PDFs, layout, and memory

Scanned or image-only pages

Text extraction only returns text represented in the document’s text layer. A scanned PDF may contain only page images, so ordinary extraction can return little or no text even when the page visibly contains words. The pdf.js-extract documentation explicitly says it does not include OCR. Add a separate OCR stage if image recognition is required; do not treat an empty text result as proof that the PDF page is blank.

Rank #4
MixPad Free Multitrack Recording Studio and Music Mixing Software [Download]
  • Create a mix using audio, music and voice tracks and recordings.
  • Customize your tracks with amazing effects and helpful editing tools.
  • Use tools like the Beat Maker and Midi Creator.
  • Work efficiently by using Bookmarks and tools like Effect Chain, which allow you to apply multiple effects at a time
  • Use one of the many other NCH multimedia applications that are integrated with MixPad.

Tables and reading order

PDF text items and coordinates are not the same as semantic table cells. The Node package provides utilities for grouping lines and rows, but such grouping is a layout aid, not guaranteed table recognition. Validate representative documents, especially when columns, multi-line cells, or unusual reading order matter.

Memory use

Base64 expands the representation relative to the original binary data, and decoding may temporarily leave both representations in memory. PDF.js recommends supplying raw binary data as a typed array rather than converting through base64 when possible. If an upstream service already supplies base64, decode once, avoid unnecessary copies, and consider the size of the document and concurrent jobs in your own application.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Passwords and invalid input

For password-protected PDFs, use the parser’s supported password mechanism. PDF.js exposes a password loading parameter in its API, but compatibility depends on the document and parser; do not assume every encrypted file will open. For malformed or unsupported input, catch parser failures and report a useful error to the caller. The observed behavior can vary by PDF and library, so validate against the kinds of files your application will receive.

Best Value
Corel PDF Fusion Software
  • Save money by using PDF Fusion to view over 100 file formats without having to purchase additional software
  • Merge incompatible files quickly and easily by dragging and dropping in PDF Fusion to create a new PDF documents
  • Save time with PDF Fusion's editing tools to reuse the content from existing documents without starting from scratch
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Which implementation should you use?

Option Best fit Input and output Important boundary
Browser PDF.js Client-side browser workflows Decode to typed-array bytes, load the document, and map page text content Base64 conversion uses memory; OCR inclusion is not established by the cited PDF.js material
Node.js with pdf.js-extract Server-side extraction using a base64 string or buffer Buffer.from(value, 'base64'), then buffer extraction and page text items Package documentation explicitly says no OCR
PDF.js Express Applications using that browser viewer SDK Vendor example converts base64 to a Blob and loads it as a document The cited page concerns document loading, not a general text-extraction recipe; a commercial SDK is not required for the ordinary paths above

PDF.js Express describes the base64-to-Blob viewer workflow at its base64 documentation. No performance comparison is established for these choices; select based on runtime, output structure, and whether OCR or broader viewer functions are needed.

Common problems and fixes

  • “Invalid PDF” or parse failure: Confirm that the input is the PDF’s base64 payload, not a JSON wrapper or full data URI passed unchanged. Check that the source value was not truncated and catch the parser’s error for diagnosis.
  • Empty or sparse extracted text: Check whether the page is scanned or image-only. If so, ordinary extraction will not recognize it; add OCR separately.
  • Unreadable spacing or columns: Joining all item strings with a space flattens layout. Keep item coordinates, evaluate line/row grouping, and inspect output from representative files.
  • Browser says atob is unavailable: The environment may not expose that global. Use a decoder supported by the target runtime, or perform decoding in a suitable server environment.
  • Memory pressure on large files: Avoid repeated base64-to-binary conversions and duplicate buffers. When possible, obtain binary data upstream and use typed-array input in PDF.js.
  • Password prompt or decryption failure: Supply the password through the selected parser’s supported option and handle rejection. The available documentation does not establish compatibility with every encrypted PDF.

Or skip the browser setup

ScreenshotNeo captures a webpage; it does not extract text from an arbitrary base64 PDF buffer. It can be relevant only if the source you need to capture is a publicly reachable webpage. Its one-request API returns a screenshot or PDF, not parsed PDF text. See the ScreenshotNeo site and API documentation for that separate use case.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

For webpage captures, ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000. Sign up for free screenshots.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can I get one text string instead of one JSON entry per page?

Yes. Map the page results to their text values and join them using the separator your application expects, such as a newline. Keep page boundaries if later processing may need them.

Does base64 decoding convert a PDF into JSON by itself?

No. Decoding produces PDF bytes. A PDF parser extracts content; your code then defines and serializes the JSON structure.

Will extracted text always follow the order a person reads the page?

Not necessarily. PDF text items and coordinates can reflect layout rather than semantic reading order. Inspect output from the documents your application handles.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.