Goroutines and Java virtual threads answer the same practical question: how to keep many blocking tasks in flight without giving each one an operating-system thread. Both are runtime-managed tasks multiplexed over a smaller pool of OS threads. They are not the same mechanism with different syntax, though, and the difference that matters most for correctness is the memory model. Goroutines follow the Go Memory Model. A virtual thread is still a java.lang.Thread, so it follows the Java Memory Model in Chapter 17 of the Java Language Specification. Choosing virtual threads does not change when a write made by one task becomes visible to another.
Memory and throughput are harder to settle. The official documents covered here explain how each runtime stores stacks and schedules work, but none of them is a controlled benchmark that holds runtime versions, hardware, workload, stack depth and allocation pattern constant. This article therefore explains what each runtime stores and guarantees, and what you need to measure before concluding which one uses less memory or delivers more throughput.
Goroutines and virtual threads at a glance
| Aspect | Go goroutines | Java virtual threads |
|---|---|---|
| Documents consulted | Go FAQ (undated page), Go Memory Model (dated June 6, 2022), Go GC guide | JEP 444 (finalized in Java 21, OpenJDK, 2023), Oracle Java SE virtual-thread documentation with versioned pages including Java SE 25 and 26, JLS Chapter 17 |
| Unit you create | A goroutine, started with the go statement |
A java.lang.Thread instance, for example via Thread.ofVirtual() |
| Mapping to OS threads | The runtime multiplexes goroutines onto a set of threads | The JDK scheduler maps virtual threads onto platform carrier threads (M:N scheduling) |
| Behavior when a task blocks | The runtime can schedule other goroutines on available threads | During supported blocking I/O, the virtual thread is suspended and its carrier is freed |
| Stack representation | Resizable, bounded stacks; the FAQ says a new goroutine starts with a few kilobytes | Stack chunks stored as heap objects that grow and shrink, up to the configured platform-thread stack-size limit |
| Shared-data rules | Go Memory Model: channel operations, sync, and sync/atomic; race-free programs have the documented sequential-consistency guarantee |
JLS Chapter 17 happens-before rules, including monitor unlock/lock, volatile write/read, and Thread.join |
| Pooling guidance | Not stated in the Go FAQ | Created per task rather than pooled (JEP 444) |
| Thread-local values | Not addressed in the cited Go documents | Need care: virtual threads may be very numerous, and thread-local values can add memory cost (JEP 444) |
| Documented limits | GC behavior can be affected by very large goroutine counts (Go GC guide) | Pinning and unsupported or blocking operations can limit scalability, depending on JDK version (Oracle documentation) |
| Published overhead figures | Initial stack of a few kilobytes and roughly three cheap instructions per function call, both high-level FAQ statements | No per-thread memory figure stated in JEP 444 |
How each runtime schedules tasks
Both models avoid the cost of one operating-system thread per task. Neither removes the need for CPU, connections or downstream capacity, and neither behaves identically across every version, so the details below describe documented mechanisms rather than guarantees for every release.
Goroutines
The Go FAQ describes goroutines as independently executing functions multiplexed onto a set of threads. When a goroutine blocks, the runtime can run other goroutines on the threads that are available. The exact scheduling policy is an implementation detail. Do not assume identical ordering or fairness across Go releases, and do not treat the FAQ’s descriptions as a per-task timing guarantee.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →#1 Best Overall
Java virtual threads
JEP 444, delivered in Java 21, defines a virtual thread as a java.lang.Thread that runs Java code on an OS-backed platform thread while mounted, without holding that carrier for its whole lifetime. The JDK scheduler maps virtual threads onto platform threads. When supported blocking I/O happens through the relevant Java APIs, the runtime can suspend the virtual thread and free the carrier for other work. The JEP presents this as a way to write thread-per-request code that still reaches high concurrency. It also describes virtual threads as a lightweight implementation of threads that is provided by the JDK rather than the OS, and it names goroutines as another example of user-mode threads.
The two designs share a purpose, not an API or an operational profile. A goroutine is started with go and coordinated with channels and locks. A virtual thread is started through the Thread API and inherits Java’s thread semantics, including its interaction with synchronized blocks and native code.
Stack storage and what the per-task figures mean
Stack storage is where the two runtimes look most alike on paper and differ most in operation. Both grow stacks on demand, but the figures each documentation set publishes describe different things.
Goroutine stacks
The Go FAQ says a newly created goroutine starts with a few kilobytes of stack, and that the runtime grows and shrinks stack memory automatically. It also says the CPU overhead averages about three cheap instructions per function call. Both statements are high-level claims from the FAQ. They are not a cross-language benchmark, a fixed stack size, or a guarantee for every architecture and Go version.
Virtual-thread stacks
JEP 444 states that a virtual thread’s stack is stored in heap-resident stack-chunk objects. The stack grows and shrinks as execution proceeds, up to the configured platform-thread stack-size limit. Because those chunks live on the managed heap, the garbage collector has to account for them along with the application’s own objects. The JEP also says that the heap space and garbage-collector activity for virtual threads are generally difficult to compare with asynchronous code, so it does not offer a per-thread memory number to set beside the Go figure.
Why task counts and stack figures do not measure process memory
A headline such as “a few kilobytes per goroutine” or “a virtual thread per request” describes one component of memory. Total process memory is the sum of several things that vary by workload.
- Stack depth. Both starting sizes and growth behavior depend on how deep each task’s call chain goes. A deep recursive task costs more than the starting figure suggests.
- Live heap. Objects reachable from each task, and from shared structures, usually dominate memory in services that hold caches, buffers or request state.
- Allocation rate and GC. Heap-resident stack chunks and application allocations share the collector. The Go GC guide notes that goroutine stacks are often small relative to the live heap, but very large goroutine populations can affect garbage-collector behavior.
- Thread-local values. Each virtual thread that stores thread-local data carries that data with it, so the cost scales with the number of tasks.
- Virtual memory metrics. The Go GC guide cautions against treating virtual memory size (VSS) as a direct measure of a Go program’s useful memory footprint. Measure resident memory and heap use instead.
Because of these factors, a count of goroutines or virtual threads cannot tell you how much memory a process needs. Only measurement under a representative load can.
Memory models: what each language guarantees about shared data
A memory model answers one question: when can a read in one task observe a write made by another? The thread implementation does not change that answer, and the two languages answer it with different rules.
The Go Memory Model
The Go Memory Model specifies when a read in one goroutine can observe a write performed in another. Its advice is to serialize access to shared data using channel operations or the sync and sync/atomic packages. The document states:
Programs that modify data being simultaneously accessed by multiple goroutines must serialize such access.
For programs without data races, the document gives sequential consistency. A data race is a correctness bug in its own right, and it does not become safe because goroutines are cheap.
The Java Memory Model for virtual threads
Java’s rules are defined in Chapter 17 of the Java Language Specification. The central concept is the happens-before relation, formed from program order and synchronization edges. Two examples from the specification are that an unlock of a monitor happens-before a subsequent lock of that monitor, and that a write to a volatile field happens-before subsequent reads of that field.
Recommended Free Tools
Virtual threads do not introduce a separate Java memory model. JEP 444 defines them as instances of java.lang.Thread, so the scheduler changes how Java code is multiplexed onto carriers while the Chapter 17 visibility rules still apply unchanged.
A worked example: publishing one result
The following Go program is correct because the channel close is ordered before the receive that returns. The Go Memory Model specifies that the closing of a channel is synchronized before a receive that returns because the channel is closed.
package main
import "fmt"
func main() {
var data int
done := make(chan struct{})
go func() {
data = 42
close(done)
}()
<-done
fmt.Println(data) // prints 42
}
Remove the channel operations and the program has a data race. The read of data then has no guaranteed result, and it may observe 0 or 42.
Rank #4
The equivalent Java program uses a virtual thread and Thread.join. Under Chapter 17, all actions in a thread happen-before another thread successfully returns from a join on that thread.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →public static void main(String[] args) throws InterruptedException {
int[] result = new int[1];
Thread worker = Thread.ofVirtual().start(() -> result[0] = 42);
worker.join();
System.out.println(result[0]); // prints 42
}
Remove the join and the read in the main thread is racy in the same way. The thread type is irrelevant to the outcome; the synchronization edge is what determines it.
Concurrency overhead and operational limits
Lightweight tasks remove one cost, the platform thread held for each blocked task. Several other costs remain, and some of them are specific to one runtime.
Thread-local values
JEP 444 warns that thread-local variables deserve care, because virtual threads may be extremely numerous and thread-local values can add memory costs. A pattern that was harmless with a few hundred pooled platform threads can add up when every request creates a new virtual thread that sets its own thread-local values. The Go documents cited here do not address thread-local storage, so this comparison should be checked against the code in use on each side.
Pinning in Java
Oracle’s virtual-thread documentation discusses pinning, which keeps a virtual thread attached to its carrier. JEP 444 describes the Java 21 design as pinning a virtual thread when it blocks while inside a synchronized block or method, or inside a native frame. Pinned threads cannot release their carrier during the block, so they reduce the concurrency the scheduler can deliver. The exact conditions have changed across later JDK releases, so confirm the behavior in the Oracle documentation for the version you deploy. The Go documents discussed here do not describe an equivalent concept.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchBest Value
CPU-bound work and downstream limits
Neither runtime makes CPU-bound work cheaper. A virtual thread that computes for a long time still consumes processor capacity, and so does a goroutine doing the same. Concurrency limits also remain in place: connection pools, memory budgets, rate limits and backpressure are application decisions. A service that accepts a million tasks but has a 50-connection database pool still has a 50-connection database.
Can virtual threads replace a thread pool?
For blocking request-per-task code, often yes, but a thread pool typically does two jobs. It reuses expensive platform threads, which virtual threads make unnecessary, and it caps concurrency against a limited resource, which they do not. JEP 444 says virtual threads are created per task rather than pooled, so the first job goes away. The second job must be kept, usually with a semaphore sized to the real resource.
Semaphore dbPermits = new Semaphore(20); // match the database pool's capacity
try (var executor = Executors.newVirtualThreadPerTaskExecutor()) {
for (var request : requests) {
executor.submit(() -> {
dbPermits.acquire();
try {
return queryDatabase(request);
} finally {
dbPermits.release();
}
});
}
} // close() waits for the submitted tasks to finish
The limit of 20 is illustrative and should match the database’s actual capacity. In Go, the same idea is a buffered channel used as a semaphore, with a sync.WaitGroup to wait for completion.
How to run a fair comparison
A valid comparison between the two runtimes has to specify more than the programming language. Work through these steps before attributing any difference to the runtime.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
- Pin the versions. Record the exact Go release reported by
go versionand the exact JDK build reported byjava -version. Behavior that differs between Java 21 and a later release, including pinning, can change the result. - Define the workload. Separate blocking I/O from CPU-bound work, and set the blocking pattern, including how long each call waits.
- Fix the stack depth and call pattern. Use the same logical depth on both sides, and document it.
- Control allocation and live heap. Keep request sizes, object lifetimes and cache behavior comparable, then measure them rather than assuming them.
- Record thread-local use. Note every thread-local value on the Java side and any equivalent per-task state on the Go side.
- Sweep the concurrency level. Measure at several levels, because the ranking can change as the level rises past the downstream limit.
- Include the downstream bottleneck. Use the real connection pool size and rate limits, or the benchmark will mostly measure the bottleneck.
- Measure four things. Throughput, tail latency, CPU use, and resident and heap memory under steady load, not only the count of tasks or the virtual memory size.
Choosing between them
The language and team usually decide this before any benchmark runs. The memory-model question is settled by the language, and the performance question is settled by measurement on your workload.
Quick Recap
- Existing Go service. Goroutines, channels and
syncare the native tools. Check very large goroutine counts against the GC guidance rather than assuming they are free. - Existing Java service with blocking I/O. Virtual threads let you keep thread-per-request code. Audit synchronized blocks and native calls on the blocking path for pinning, and confirm the JDK version you deploy.
- CPU-bound workload. Neither model adds cores. Tune parallelism to the hardware and keep the work bounded.
- Shared mutable state. Whichever runtime you use, serialize access with the synchronization primitives its memory model defines. A missing happens-before edge is a bug regardless of the thread type.
- Memory budget. Compare resident memory and heap under the expected peak load, and include thread-local data, reachable objects and GC behavior in the total.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




