Skip to content
11 min read

Every Goroutine Needs an Owner That Can Stop It

Go 1.27 made the goroutine leak profile generally available, and it finds a real class of production leaks. The lasting fix is a design rule it cannot enforce for you.

Antonio J. del Águila

Knaisoma

A goroutine leak rarely announces itself as a leak. Consider an illustrative scenario that many operations teams will recognize: a dashboard shows a Go service’s resident memory climbing over a few days, someone raises the container limit or schedules a rolling restart, and the graph settles into a sawtooth that everyone learns to ignore. Memory growth has many causes, and leaked goroutines are only one of them, but they are one that a restart hides very effectively.

Go 1.27, released in August 2026, gives that bug a name the runtime can report. The release notes announce that the goroutineleak profile, an experiment in Go 1.26, is now generally available through runtime/pprof and the /debug/pprof/goroutineleak endpoint. It is a genuinely useful tool and most teams running Go in production should turn it on. Our argument is that it should be treated as the last of three layers rather than the first, because the property that prevents leaks is a design rule no detector can enforce: every goroutine you start needs an owner that can stop it and a place where someone waits for it to finish.

What the runtime can now prove

A leaked goroutine, in the release notes’ definition, is one blocked on a concurrency primitive such as a channel, a sync.Mutex or a sync.Cond that cannot possibly become unblocked. The detection trick is to reuse the garbage collector. As the Go blog’s explanation of the profile describes it, the marking phase starts from goroutines that are not blocked, traces what they can reach, and treats a blocked goroutine as live only if something it is waiting on was reached. Whatever remains blocked and unreached can never wake up, so the runtime reports it.

This matters because the older tools could only guess. Uber’s LeakProf, described in 2022, aggregated ordinary goroutine profiles and flagged source locations with an unusually large number of goroutines blocked on channel operations. Its authors state plainly that it is neither sound nor complete: it misses leaks below its threshold and reports false positives wherever many goroutines are blocked by design. The new profile is built on reachability instead of counts, which is why the Go blog can say it produces “little-to-no false positives.” The underlying research, published at ASPLOS 2025 as Dynamic Partial Deadlock Detection and Recovery via Garbage Collection, came from a collaboration between Aarhus University, Washington University in St. Louis and Uber. Its abstract is a useful calibration: the research prototype detected 94% of partial deadlocks in a set of microbenchmarks but 50% in the test suites of a large industrial codebase. Few false alarms is not the same as full coverage.

Precision about what it finds comes with honesty about what it misses, and the Go documentation is clear on both. Goroutines blocked on file or network I/O, or in direct system calls, are never considered leaked. A primitive that is reachable through a global variable, or through the local variables of a goroutine that is still running, keeps its waiters looking alive even if nothing will ever use it again. And the profile can only report a leak after it has happened, so a leak that occurs on a rare path in production is found in production.

Why this class of bug is common

It is tempting to treat goroutine leaks as a sign of careless code. The evidence says otherwise. The first systematic study of real Go concurrency bugs, Understanding Real-World Concurrency Bugs in Go by researchers at Pennsylvania State University and Purdue University, analyzed 171 bugs across six widely used projects including Docker, Kubernetes and gRPC. Of the 85 blocking bugs, around 58% were caused by message passing rather than shared memory, a result the authors found surprising given how strongly Go encourages channels. The same study ran the reproduced blocking bugs against Go’s built-in deadlock detector, which caught two of them, because it only fires when no goroutine in the whole process can make progress.

The largest industrial account comes from Uber. In a CGO 2024 paper on goroutine leaks in its monorepo, covering about 75 million lines of Go and more than 2,500 microservices, the authors report what happened when they deployed leak detection in tests and in production.

857

Pre-existing leaks found by test-time detection

Saioc et al., CGO 2024

~260

New leaks prevented in one year

Saioc et al., CGO 2024

9.2x

Largest memory reduction after a production fix

Saioc et al., CGO 2024

Two cautions apply to those figures. They come from one company with an unusually large Go estate and a research team invested in the tooling, so they indicate the size of the backlog a mature codebase can carry rather than what your services will show. The 9.2x reduction, and the 34% speedup reported alongside it, are the best cases among the services the authors fixed, not typical outcomes. The shape of the result is still the useful part: a large pool of existing leaks, a steady rate of new ones, and a handful of production fixes with outsized effects on memory.

The shape of a typical leak

The pattern behind the example in the Go blog, and behind many of the channel bugs in the academic studies, is simple: a goroutine sends a result into a channel that its caller has stopped listening to. Here is a lookup with a deadline that looks careful and is not.

func lookup(ctx context.Context, key string) (Value, error) {
    ch := make(chan Value) // unbuffered: a send waits for a receiver

    go func() {
        ch <- fetch(key) // (1) nothing tells fetch to stop
    }()

    select {
    case v := <-ch:
        return v, nil
    case <-ctx.Done():
        return Value{}, ctx.Err() // (2) caller leaves, sender waits
    }
}

Every request that times out leaves one goroutine behind at (2), holding its stack and whatever fetch allocated. Under normal load the timeout rate is low and memory creeps. During a downstream slowdown the timeout rate spikes, and the leak accelerates exactly when the service is already under stress.

There are two common fixes, and they are not equivalent. The minimal one is to make the channel buffered with make(chan Value, 1). The send can then always complete, so the goroutine exits as soon as fetch returns and the abandoned-send leak is gone. It is still a goroutine nobody owns, though. It keeps doing work whose result nobody wants, and if fetch never returns, it never exits either. Under a slowdown you are running a backlog of abandoned fetches against a dependency that is already struggling.

The ownership fix keeps the work inside the caller’s lifetime. For a single lookup, that often means not starting a goroutine at all.

func lookup(ctx context.Context, key string) (Value, error) {
    return fetch(ctx, key) // (3) caller owns and waits for the work
}

func lookupAll(ctx context.Context, keys []string) ([]Value, error) {
    g, ctx := errgroup.WithContext(ctx) // (4) first error cancels rest
    out := make([]Value, len(keys))
    for i, k := range keys {
        g.Go(func() error {
            v, err := fetch(ctx, k)
            out[i] = v
            return err
        })
    }
    return out, g.Wait() // (5) nothing outlives this call
}

At (3) the deadline travels with the call and errors propagate, so there is nothing left behind to leak. When the work really is concurrent, an errgroup.Group from golang.org/x/sync gives each goroutine an owner at (4), which cancels the others on the first error, and a join at (5), which waits for all of them before returning. Neither version is magic: both are only as bounded as fetch is. Passing a context requests cancellation, and the operation still has to honor it, typically by using the context in its network calls or by setting its own deadline.

The first version also shows the profile’s blind spot. There, if fetch returns and the goroutine blocks on the send, the channel is unreachable from any running goroutine and the leak profile will report it. If instead fetch hangs on a network read with no deadline, the goroutine is blocked in I/O, and the goroutineleak profile will never flag it, even though it will never finish either. The runtime can prove that no one will ever receive from a channel. It cannot prove that a remote server will never answer.

Three layers, each catching what the others miss

We recommend treating leak prevention as three layers with different costs and different blind spots.

Design: owner, exit, wait. For every go statement, a reviewer should be able to answer three questions from the surrounding code. Who owns this goroutine, meaning which function or component is responsible for its lifetime? What makes it exit, whether a context, a closed channel or bounded input? Who waits for it, through a sync.WaitGroup, an errgroup.Group or a result channel that is always drained? A goroutine that cannot answer all three is a leak waiting for the right input. This layer costs review attention and catches leaks before they exist, including those blocked on I/O, which the goroutineleak profile does not report. It depends on reviewers applying it consistently, which is why the next two layers exist.

Tests: make leaks fail the build. Uber’s goleak checks that no unexpected goroutines remain when a test or package finishes, and adding it to a package is a few lines.

func TestMain(m *testing.M) {
    goleak.VerifyTestMain(m) // fail if goroutines outlive tests
}

The package-level form matters: the project’s documentation notes that per-test checks cannot tell a leak from a parallel test that has not finished yet. For concurrent code that depends on timeouts, testing/synctest, generally available since Go 1.25, runs tests in a bubble with a fake clock and fails a test when every goroutine in the bubble is durably blocked. Go 1.27 extends it with an in-memory httptest server, which makes it practical for more network-shaped code. Test-time detection is cheap and catches leaks before merge. Because goleak looks at every goroutine still alive, it also catches one stuck in I/O, provided a test drives the cancellation or timeout path that strands it. It only covers paths the tests exercise, however, and the rare error paths that leak are often exactly the ones tests skip.

Production: collect the leak profile on a schedule. The Go blog notes the extra garbage collection work the profile adds and suggests that periodic profiling infrastructure collect it at a modest cadence, “e.g., every 4 hours.” That is a sensible default for a continuous profiling setup: leaks accumulate, so a leak that matters will still be visible hours later. This layer catches what tests never reach, such as untested error paths, real interleavings and real input shapes, but it only reports what already happened, and only for goroutines blocked on concurrency primitives.

Rolling this out without drowning in backlog

The Uber numbers carry a practical warning: when you first switch on leak detection in an established codebase, you may find hundreds of existing leaks. Blocking merges on all of them at once is a good way to get the check disabled within a week. A staged rollout works better.

Start in production, because it is read-only and tells you where the real cost is. Upgrade a representative set of services to Go 1.27, collect the goroutineleak profile on a schedule, and rank the reported stacks by count and by the memory of the services they appear in. Treat the reported stacks as investigation leads, and check them against the services whose memory limits have been raised or whose restarts have been scheduled. Where a heap or memory profile attributes a meaningful share of the growth to the leaked goroutines, you have your first fixes, and they come with a cost argument attached.

Next, add package-level goleak checks to new packages and to packages you are already changing, and record known offenders in an explicit ignore list with an owner and a removal date rather than a silent exception. This stops the backlog from growing while the existing items are worked down. Then write the owner, exit and wait questions into your review checklist, because the leaks blocked on a dependency that stopped answering are invisible to the production profile and are caught in tests only when a test exercises that path. Pair that with a rule that every outbound call has an effective deadline or cancellation path that the operation actually honors, which turns an indefinite I/O hang into a bounded wait.

Finally, stop treating a memory limit increase as a routine operational change. Ask for a leak profile first. If it is clean and memory still grows, you have learned something real about the workload. If it is not, find out how much of the growth the leaked goroutines account for before deciding whether the request pays for capacity or for a bug.

The leak profile is one of the more valuable additions to the Go runtime in years, because it replaces guessing with proof for a large class of bugs. Its best use is to confirm that the ownership discipline in your code is holding, and to show you precisely where it is not.

If your Go services depend on restarts or ever-larger memory limits to stay healthy, we can help. We work with engineering teams to profile production services for goroutine and resource leaks, introduce leak checks into test pipelines without blocking delivery, and review concurrency designs so that every background task has a clear owner and exit. Talk with us about the reliability of your Go services.

Go Software Quality Operations Engineering Leadership
Share:

Stay updated

Get insights on engineering transformation delivered to your inbox.

Newsletter coming soon.