Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use the parser that matches the bytes you received. Go has separate, well-tested APIs for JSON (encoding/json), CSV (encoding/csv), XML (encoding/xml), and HTML (golang.org/x/net/html). Map stable data into exported structs; use generic values or token/decoder APIs when the shape is unknown or too large to buffer. Always inspect errors and test the odd cases your source actually produces.

Start by identifying the input and its shape

Before writing a selector or struct, determine the response format, character encoding, and whether the schema is stable. An HTTP Content-Type header is useful evidence, but feeds and scraped pages sometimes mislabel it. Save a representative response, including malformed examples, and decide whether you can hold it in memory.

Source Primary Go API Best first mapping Use streaming when
JSON encoding/json (v1) or encoding/json/v2 Exported struct fields with JSON tags for a known schema; generic values for unknown data Responses are large, records are incremental, or you need token-level control
CSV encoding/csv.Reader Record slices, then explicit column mapping The file is large or you want to process one record at a time
XML encoding/xml Structs with XML tags for a known shape You need namespaces, selective extraction, or incremental records
HTML golang.org/x/net/html Traverse the parsed HTML5 node tree Usually after obtaining a stream; traversal itself is tree-based

Do not use regular expressions as a general HTML parser. HTML5 error recovery can insert implicit nodes and discard explicit malformed tags, so the tree is not always a literal copy of source nesting.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

JSON: deliberate typed mapping or controlled generic data

Decode a known response into a struct

Fields must be exported for decoding. Tags map wire names that do not follow Go’s field-name convention. Fields absent from the destination struct are ignored in the documented tutorial example, which lets a small type select only the data you need.

package main

import (
    "encoding/json"
    "fmt"
    "io"
    "net/http"
)

type User struct {
    ID    int    `json:"id"`
    Name  string `json:"name"`
    Email string `json:"email"`
}

type Response struct {
    Users []User `json:"users"`
}

func main() {
    resp, err := http.Get("https://api.example.com/users")
    if err != nil { panic(err) }
    defer resp.Body.Close()
    if resp.StatusCode < 200 || resp.StatusCode >= 300 {
        panic(fmt.Errorf("HTTP status: %s", resp.Status))
    }
    var result Response
    if err := json.NewDecoder(resp.Body).Decode(&result); err != nil {
        panic(err)
    }
    for _, u := range result.Users {
        fmt.Printf("%d %s <%s>n", u.ID, u.Name, u.Email)
    }
    _, _ = io.Discard, result
}

For an already-buffered payload, json.Unmarshal(data, &value) is straightforward. A decoder reads from an io.Reader and avoids creating a second full copy of the response. Neither approach makes untrusted data safe automatically: validate required fields, ranges, and URLs after decoding.

Handle unknown or mixed JSON

Decode into map[string]any or []any when the shape is genuinely dynamic. Numbers then use the default representation, so preserve precision deliberately if IDs or monetary values can exceed floating-point accuracy. For very large documents or multiple top-level values, use decoder tokens and consume only the members you need instead of recursively materializing everything.

var payload map[string]any
if err := json.NewDecoder(r).Decode(&payload); err != nil {
    return err
}
if v, ok := payload["status"].(string); ok {
    fmt.Println(v)
}

Choose JSON v1 or v2 consciously

Current Go documentation distinguishes encoding/json v1 from encoding/json/v2 and recommends v2 for new usage where it is available in your target toolchain. They are not interchangeable in every edge case. Before migrating, test case matching, duplicate member names, invalid UTF-8, nil slice/map output, and omitempty behavior against your fixtures. Pin the Go version in CI and read the package documentation for that version; the v1 reference viewed for this article identifies Go 1.27.1 and a September 1, 2026 publication date.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

CSV: let the reader handle quoting

A CSV record is not necessarily one source line. A quoted field can contain commas and newlines, so splitting on commas, lines, or strings.Fields corrupts valid RFC 4180-style data. The standard package reads and writes CSV and exposes the format’s documented differences and controls.

Stream records and map columns

package main

import (
    "encoding/csv"
    "fmt"
    "io"
    "os"
)

func main() {
    f, err := os.Open("people.csv")
    if err != nil { panic(err) }
    defer f.Close()

    r := csv.NewReader(f)
    r.FieldsPerRecord = -1 // accept variable-width records; validate below
    r.TrimLeadingSpace = true
    r.Comment = '#'

    header, err := r.Read()
    if err != nil { panic(err) }
    columns := map[string]int{}
    for i, name := range header { columns[name] = i }
    required := []string{"id", "name"}
    for _, name := range required {
        if _, ok := columns[name]; !ok { panic("missing column: " + name) }
    }

    for {
        record, err := r.Read()
        if err == io.EOF { break }
        if err != nil { panic(err) }
        if len(record) <= columns["name"] { panic("short record") }
        fmt.Println(record[columns["id"]], record[columns["name"]])
    }
}

Set Comma for a delimiter such as a tab, and set FieldsPerRecord to the expected width when a strict file contract matters. A value of -1 accepts variable widths, but you must validate each record yourself. Configure Comment only when lines beginning with that rune are comments in your source. Use ReadAll for small files that you intentionally need as a complete matrix; otherwise process Read results incrementally.

When writing CSV, csv.Writer uses LF by default rather than CRLF. Call UseCRLF = true when a consumer requires Windows-style line endings, and always check w.Error() after flushing.

XML: structs for known shapes, tokens for selective extraction

encoding/xml handles XML 1.0 and namespace-aware decoding. Use xml.Unmarshal for a buffered document or xml.Decoder when reading from a stream.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
type Feed struct {
    XMLName xml.Name `xml:"feed"`
    Entries []Entry  `xml:"entry"`
}

type Entry struct {
    ID    string `xml:"id"`
    Title string `xml:"title"`
    Link  string `xml:"link>href,attr"`
}

func decodeFeed(r io.Reader) error {
    var feed Feed
    if err := xml.NewDecoder(r).Decode(&feed); err != nil { return err }
    for _, e := range feed.Entries {
        fmt.Println(e.ID, e.Title, e.Link)
    }
    return nil
}

Attribute and element paths are expressed with XML tags. If namespaces distinguish otherwise identical names, include the namespace-aware form documented by the package rather than matching local names blindly. For a huge feed, inspect decoder.Token(), find each start element, and decode just that subtree; stop on and report syntax errors instead of returning partial data as if it were complete.

HTML: parse the HTML5 tree, then traverse it

The golang.org/x/net/html package implements the HTML5 parsing algorithm. Its parser assumes UTF-8, rejects nesting beyond 512 elements, and may create implied nodes or omit explicit malformed tags. Select elements by walking the resulting tree and examining node types, names, attributes, and text.

doc, err := html.Parse(resp.Body)
if err != nil { return err }

var walk func(*html.Node)
walk = func(n *html.Node) {
    if n.Type == html.ElementNode && n.Data == "article" {
        fmt.Println(strings.TrimSpace(textContent(n)))
    }
    for child := n.FirstChild; child != nil; child = child.NextSibling {
        walk(child)
    }
}
walk(doc)

func textContent(n *html.Node) string {
    if n.Type == html.TextNode { return n.Data }
    var b strings.Builder
    for c := n.FirstChild; c != nil; c = c.NextSibling { b.WriteString(textContent(c)) }
    return b.String()
}

In production, add an attribute helper that checks n.Attr, normalize whitespace, and reject unexpected links before storing them. If a page is rendered by JavaScript, the HTTP body may not contain the data you see in a browser; obtain the rendered HTML through an appropriate browser capture step rather than assuming the parser executes scripts.

Streaming, validation, and error handling

Keep input boundaries explicit

  • Use an io.Reader decoder for network responses and large files.
  • Use byte-slice APIs when the payload is already buffered and small enough for your memory budget.
  • Apply request and read deadlines; a parser cannot recover from a connection that never finishes.
  • Limit accepted body size before parsing when the source is untrusted.

Never silently trust partial data

Return parser errors with the source identifier and record or byte offset when available. For CSV, decide whether one bad record aborts the import or is quarantined with an audit record. For JSON and XML, a syntax error means the document is incomplete; do not publish fields decoded before the error unless your application explicitly supports partial documents.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Validate semantics after syntax

Parsing proves that bytes fit a grammar, not that values are useful. Check required fields, duplicate business IDs, timestamps, enum values, numeric ranges, and relationships. Treat absent, empty, and explicit null as separate states when the source contract distinguishes them; pointers or custom types can preserve that distinction.

Tests that catch real extraction failures

  • JSON with missing fields, unknown fields, null, duplicate names, invalid UTF-8, and numbers at precision limits.
  • CSV with quoted commas, embedded newlines, comments, blank records, a different delimiter, and short or extra fields.
  • XML with namespaces, attributes, empty elements, repeated children, and malformed closing tags.
  • HTML with omitted html/tbody tags, nested elements, duplicate attributes, malformed nesting, non-UTF-8 input handling, and a nesting-depth limit case.

Use table-driven tests and golden fixtures captured from the actual producer. Include a test that verifies your chosen v1 or v2 behavior before changing Go versions or package imports. The official APIs do not provide comparative performance rankings for these formats, so measure your own representative documents if throughput or memory is a requirement.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common failures and fixes

“Unexpected end of JSON input” or XML syntax errors

The body is truncated, an upstream proxy returned an error page, or you attempted to parse a stream before it completed. Log status and content type, enforce a read deadline, and capture a bounded response sample for diagnosis.

CSV columns shift after one long field

Manual splitting ignored quoting. Replace it with csv.Reader, set the delimiter and field-count policy, and add a fixture containing quoted commas and newlines.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

HTML selector finds nothing

The content may be client-rendered, the class may have changed, or HTML5 error recovery produced a different tree. Save the fetched body, parse it, inspect the tree, and use a rendered capture when the data is absent from the response.

Fields remain empty after JSON decoding

Fields are unexported, tags do not match the wire names, or the response shape is different from the struct. Export fields, correct tags, check the decode error, and test the exact payload. Recheck v1/v2 case-matching and other compatibility-sensitive defaults during migration.

Or skip the browser setup

When extraction starts with a rendered webpage, ScreenshotNeo can return a clean PNG, JPEG, WebP, or PDF through one request. It accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

For all request options and parameter names, see the ScreenshotNeo API documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

After downloading the image or PDF, your Go pipeline can archive it, run OCR if needed, or parse the returned HTML from a separate page-info workflow. ScreenshotNeo supports full-page captures with lazy images loaded, CSS-selector element captures, dark mode, 12 device presets plus custom viewports, retina scale, PDF paper size/margins/landscape/page ranges, custom CSS and JavaScript, clicks, selector or network-idle waits, ad/tracker/request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed public image links, asynchronous jobs with signed webhooks, bulk capture of 100 URLs per call, a usage API, and an OpenAPI specification. Existing parameter names used by other screenshot APIs also work.

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; Growth is $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. Create a free ScreenshotNeo account to begin.

FAQ

Should I use a scraper library for every format?

No. Keep format-specific parsers at the boundary and map their output into your own domain types afterward.

Is ReadAll always faster?

Not necessarily. It may simplify code for small inputs, but reader and token APIs avoid whole-document memory costs. Benchmark your real payloads rather than assuming a ranking.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can the HTML package execute JavaScript?

No. It parses HTML that you provide; it does not render a browser page or run scripts.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.