October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk10 min

Go Goroutines vs Java Virtual Threads: Memory Models and Concurrency Overhead

Goroutines and Java virtual threads both multiplex many tasks onto fewer OS threads, but they follow different memory models. Here is what each runtime documents about stacks, shared data and overhead, and why no published figure settles which one uses less memory.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Goroutines and Java virtual threads answer the same practical question: how to keep many blocking tasks in flight without giving each one an operating-system thread. Both are runtime-managed tasks multiplexed over a smaller pool of OS threads. They are not the same mechanism with different syntax, though, and the difference that matters most for correctness is the memory model. Goroutines follow the Go Memory Model. A virtual thread is still a java.lang.Thread, so it follows the Java Memory Model in Chapter 17 of the Java Language Specification. Choosing virtual threads does not change when a write made by one task becomes visible to another.

Memory and throughput are harder to settle. The official documents covered here explain how each runtime stores stacks and schedules work, but none of them is a controlled benchmark that holds runtime versions, hardware, workload, stack depth and allocation pattern constant. This article therefore explains what each runtime stores and guarantees, and what you need to measure before concluding which one uses less memory or delivers more throughput.

Goroutines and virtual threads at a glance

Aspect Go goroutines Java virtual threads
Documents consulted Go FAQ (undated page), Go Memory Model (dated June 6, 2022), Go GC guide JEP 444 (finalized in Java 21, OpenJDK, 2023), Oracle Java SE virtual-thread documentation with versioned pages including Java SE 25 and 26, JLS Chapter 17
Unit you create A goroutine, started with the go statement A java.lang.Thread instance, for example via Thread.ofVirtual()
Mapping to OS threads The runtime multiplexes goroutines onto a set of threads The JDK scheduler maps virtual threads onto platform carrier threads (M:N scheduling)
Behavior when a task blocks The runtime can schedule other goroutines on available threads During supported blocking I/O, the virtual thread is suspended and its carrier is freed
Stack representation Resizable, bounded stacks; the FAQ says a new goroutine starts with a few kilobytes Stack chunks stored as heap objects that grow and shrink, up to the configured platform-thread stack-size limit
Shared-data rules Go Memory Model: channel operations, sync, and sync/atomic; race-free programs have the documented sequential-consistency guarantee JLS Chapter 17 happens-before rules, including monitor unlock/lock, volatile write/read, and Thread.join
Pooling guidance Not stated in the Go FAQ Created per task rather than pooled (JEP 444)
Thread-local values Not addressed in the cited Go documents Need care: virtual threads may be very numerous, and thread-local values can add memory cost (JEP 444)
Documented limits GC behavior can be affected by very large goroutine counts (Go GC guide) Pinning and unsupported or blocking operations can limit scalability, depending on JDK version (Oracle documentation)
Published overhead figures Initial stack of a few kilobytes and roughly three cheap instructions per function call, both high-level FAQ statements No per-thread memory figure stated in JEP 444

How each runtime schedules tasks

Both models avoid the cost of one operating-system thread per task. Neither removes the need for CPU, connections or downstream capacity, and neither behaves identically across every version, so the details below describe documented mechanisms rather than guarantees for every release.

Goroutines

The Go FAQ describes goroutines as independently executing functions multiplexed onto a set of threads. When a goroutine blocks, the runtime can run other goroutines on the threads that are available. The exact scheduling policy is an implementation detail. Do not assume identical ordering or fairness across Go releases, and do not treat the FAQ’s descriptions as a per-task timing guarantee.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Java virtual threads

JEP 444, delivered in Java 21, defines a virtual thread as a java.lang.Thread that runs Java code on an OS-backed platform thread while mounted, without holding that carrier for its whole lifetime. The JDK scheduler maps virtual threads onto platform threads. When supported blocking I/O happens through the relevant Java APIs, the runtime can suspend the virtual thread and free the carrier for other work. The JEP presents this as a way to write thread-per-request code that still reaches high concurrency. It also describes virtual threads as a lightweight implementation of threads that is provided by the JDK rather than the OS, and it names goroutines as another example of user-mode threads.

The two designs share a purpose, not an API or an operational profile. A goroutine is started with go and coordinated with channels and locks. A virtual thread is started through the Thread API and inherits Java’s thread semantics, including its interaction with synchronized blocks and native code.

Stack storage and what the per-task figures mean

Stack storage is where the two runtimes look most alike on paper and differ most in operation. Both grow stacks on demand, but the figures each documentation set publishes describe different things.

Goroutine stacks

The Go FAQ says a newly created goroutine starts with a few kilobytes of stack, and that the runtime grows and shrinks stack memory automatically. It also says the CPU overhead averages about three cheap instructions per function call. Both statements are high-level claims from the FAQ. They are not a cross-language benchmark, a fixed stack size, or a guarantee for every architecture and Go version.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Virtual-thread stacks

JEP 444 states that a virtual thread’s stack is stored in heap-resident stack-chunk objects. The stack grows and shrinks as execution proceeds, up to the configured platform-thread stack-size limit. Because those chunks live on the managed heap, the garbage collector has to account for them along with the application’s own objects. The JEP also says that the heap space and garbage-collector activity for virtual threads are generally difficult to compare with asynchronous code, so it does not offer a per-thread memory number to set beside the Go figure.

Why task counts and stack figures do not measure process memory

A headline such as “a few kilobytes per goroutine” or “a virtual thread per request” describes one component of memory. Total process memory is the sum of several things that vary by workload.

  • Stack depth. Both starting sizes and growth behavior depend on how deep each task’s call chain goes. A deep recursive task costs more than the starting figure suggests.
  • Live heap. Objects reachable from each task, and from shared structures, usually dominate memory in services that hold caches, buffers or request state.
  • Allocation rate and GC. Heap-resident stack chunks and application allocations share the collector. The Go GC guide notes that goroutine stacks are often small relative to the live heap, but very large goroutine populations can affect garbage-collector behavior.
  • Thread-local values. Each virtual thread that stores thread-local data carries that data with it, so the cost scales with the number of tasks.
  • Virtual memory metrics. The Go GC guide cautions against treating virtual memory size (VSS) as a direct measure of a Go program’s useful memory footprint. Measure resident memory and heap use instead.

Because of these factors, a count of goroutines or virtual threads cannot tell you how much memory a process needs. Only measurement under a representative load can.

Memory models: what each language guarantees about shared data

A memory model answers one question: when can a read in one task observe a write made by another? The thread implementation does not change that answer, and the two languages answer it with different rules.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The Go Memory Model

The Go Memory Model specifies when a read in one goroutine can observe a write performed in another. Its advice is to serialize access to shared data using channel operations or the sync and sync/atomic packages. The document states:

Programs that modify data being simultaneously accessed by multiple goroutines must serialize such access.

For programs without data races, the document gives sequential consistency. A data race is a correctness bug in its own right, and it does not become safe because goroutines are cheap.

The Java Memory Model for virtual threads

Java’s rules are defined in Chapter 17 of the Java Language Specification. The central concept is the happens-before relation, formed from program order and synchronization edges. Two examples from the specification are that an unlock of a monitor happens-before a subsequent lock of that monitor, and that a write to a volatile field happens-before subsequent reads of that field.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Virtual threads do not introduce a separate Java memory model. JEP 444 defines them as instances of java.lang.Thread, so the scheduler changes how Java code is multiplexed onto carriers while the Chapter 17 visibility rules still apply unchanged.

A worked example: publishing one result

The following Go program is correct because the channel close is ordered before the receive that returns. The Go Memory Model specifies that the closing of a channel is synchronized before a receive that returns because the channel is closed.

package main

import "fmt"

func main() {
    var data int
    done := make(chan struct{})

    go func() {
        data = 42
        close(done)
    }()

    <-done
    fmt.Println(data) // prints 42
}

Remove the channel operations and the program has a data race. The read of data then has no guaranteed result, and it may observe 0 or 42.

The equivalent Java program uses a virtual thread and Thread.join. Under Chapter 17, all actions in a thread happen-before another thread successfully returns from a join on that thread.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
public static void main(String[] args) throws InterruptedException {
    int[] result = new int[1];
    Thread worker = Thread.ofVirtual().start(() -> result[0] = 42);
    worker.join();
    System.out.println(result[0]); // prints 42
}

Remove the join and the read in the main thread is racy in the same way. The thread type is irrelevant to the outcome; the synchronization edge is what determines it.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Concurrency overhead and operational limits

Lightweight tasks remove one cost, the platform thread held for each blocked task. Several other costs remain, and some of them are specific to one runtime.

Thread-local values

JEP 444 warns that thread-local variables deserve care, because virtual threads may be extremely numerous and thread-local values can add memory costs. A pattern that was harmless with a few hundred pooled platform threads can add up when every request creates a new virtual thread that sets its own thread-local values. The Go documents cited here do not address thread-local storage, so this comparison should be checked against the code in use on each side.

Pinning in Java

Oracle’s virtual-thread documentation discusses pinning, which keeps a virtual thread attached to its carrier. JEP 444 describes the Java 21 design as pinning a virtual thread when it blocks while inside a synchronized block or method, or inside a native frame. Pinned threads cannot release their carrier during the block, so they reduce the concurrency the scheduler can deliver. The exact conditions have changed across later JDK releases, so confirm the behavior in the Oracle documentation for the version you deploy. The Go documents discussed here do not describe an equivalent concept.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

CPU-bound work and downstream limits

Neither runtime makes CPU-bound work cheaper. A virtual thread that computes for a long time still consumes processor capacity, and so does a goroutine doing the same. Concurrency limits also remain in place: connection pools, memory budgets, rate limits and backpressure are application decisions. A service that accepts a million tasks but has a 50-connection database pool still has a 50-connection database.

Can virtual threads replace a thread pool?

For blocking request-per-task code, often yes, but a thread pool typically does two jobs. It reuses expensive platform threads, which virtual threads make unnecessary, and it caps concurrency against a limited resource, which they do not. JEP 444 says virtual threads are created per task rather than pooled, so the first job goes away. The second job must be kept, usually with a semaphore sized to the real resource.

Semaphore dbPermits = new Semaphore(20); // match the database pool's capacity
try (var executor = Executors.newVirtualThreadPerTaskExecutor()) {
    for (var request : requests) {
        executor.submit(() -> {
            dbPermits.acquire();
            try {
                return queryDatabase(request);
            } finally {
                dbPermits.release();
            }
        });
    }
} // close() waits for the submitted tasks to finish

The limit of 20 is illustrative and should match the database’s actual capacity. In Go, the same idea is a buffered channel used as a semaphore, with a sync.WaitGroup to wait for completion.

How to run a fair comparison

A valid comparison between the two runtimes has to specify more than the programming language. Work through these steps before attributing any difference to the runtime.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Pin the versions. Record the exact Go release reported by go version and the exact JDK build reported by java -version. Behavior that differs between Java 21 and a later release, including pinning, can change the result.
  2. Define the workload. Separate blocking I/O from CPU-bound work, and set the blocking pattern, including how long each call waits.
  3. Fix the stack depth and call pattern. Use the same logical depth on both sides, and document it.
  4. Control allocation and live heap. Keep request sizes, object lifetimes and cache behavior comparable, then measure them rather than assuming them.
  5. Record thread-local use. Note every thread-local value on the Java side and any equivalent per-task state on the Go side.
  6. Sweep the concurrency level. Measure at several levels, because the ranking can change as the level rises past the downstream limit.
  7. Include the downstream bottleneck. Use the real connection pool size and rate limits, or the benchmark will mostly measure the bottleneck.
  8. Measure four things. Throughput, tail latency, CPU use, and resident and heap memory under steady load, not only the count of tasks or the virtual memory size.

Choosing between them

The language and team usually decide this before any benchmark runs. The memory-model question is settled by the language, and the performance question is settled by measurement on your workload.

  • Existing Go service. Goroutines, channels and sync are the native tools. Check very large goroutine counts against the GC guidance rather than assuming they are free.
  • Existing Java service with blocking I/O. Virtual threads let you keep thread-per-request code. Audit synchronized blocks and native calls on the blocking path for pinning, and confirm the JDK version you deploy.
  • CPU-bound workload. Neither model adds cores. Tune parallelism to the hardware and keep the work bounded.
  • Shared mutable state. Whichever runtime you use, serialize access with the synchronization primitives its memory model defines. A missing happens-before edge is a bug regardless of the thread type.
  • Memory budget. Compare resident memory and heap under the expected peak load, and include thread-local data, reachable objects and GC behavior in the total.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.