Skip to content
dejan.menges
← Writing

What Go's garbage collector actually charges you

I spent a day arguing about rewriting a Go program in Rust, and a day measuring it instead. Peak memory fell 33 percent, allocations 42 percent, and the whole thing got 3.5 percent faster. Almost none of that came from where the profiler first pointed.

#go#memory#garbage-collection#rust#performance#profiling

A couple of days ago I spent an afternoon seriously considering a rewrite in Rust. Not the whole program, just the part that holds a large graph in memory, because that part was using between three and eight gigabytes and I had convinced myself that a garbage collected language was the reason.

Then I spent a day measuring instead. Peak heap fell 33 percent, allocation count fell 42 percent, and the program got 3.5 percent faster across an 81 repository corpus. Nothing in that day required leaving Go. What it required was giving up four conclusions I was confident about, three of which came from reading a profiler correctly and reasoning about it wrongly.

This post is about memory: who has been responsible for freeing it over the last fifty years, how a collector manages to run while your program is running, what Go specifically charges you for that, and how to find out where your own memory actually goes. The rewrite question is at the end, and the answer is less interesting than the method for getting to it.

Disclosure: the numbers come from enola, a Go tool I maintain that turns repositories into a queryable architecture graph. I mention it once for provenance and then talk about memory. Nothing here is about the tool.

Fifty years of deciding who frees the memory

Every memory management design in history is an answer to one question: at the moment a piece of memory stops being needed, who notices, and how?

You notice. This is malloc and free, and it is still the most common answer by volume of running code. It is also the answer with the best performance characteristics and the worst failure modes, because the question “is anyone still using this?” is genuinely hard, and the language gives you no help answering it. Free too early and you get a use after free. Free twice and you corrupt the allocator. Never free and you leak. Two of those three are exploitable, and the third is just embarrassing.

A counter notices. Reference counting arrived almost immediately, in 1960, and it is still what CPython and Swift and std::shared_ptr do. Every object knows how many references point at it; when the count hits zero, it dies. The memory comes back immediately and predictably, which is the great virtue. It costs an increment and decrement on every reference operation, it needs atomics the moment threads are involved, and it cannot collect cycles at all, which is why CPython also ships a cycle collector and why Rust programs leak when you build a cycle out of Rc.

A scan notices. John McCarthy’s LISP implementation in 1959 introduced garbage collection proper: when you run out, stop, walk everything reachable from the roots, and reclaim the rest. That is mark and sweep, and it answers the hard question by brute force. It also stops your program dead while it does so, which is where a whole subfield came from.

The compiler notices. Cyclone’s region types in the early 2000s, and then Rust, answered it a fourth way: make the lifetime part of the type, and prove at compile time that no reference outlives what it points at. Rust’s ownership model is affine typing plus C++ RAII plus a borrow checker that rejects the programs where the proof fails. A Drop runs at a point the compiler picked and you can predict. There is no collector, no counter on the common path, and no scan.

That is the part everybody knows. Here is the part that gets left out, and it matters for the rewrite question at the end.

Rust does not remove the cost of memory, it removes the collector. Your data still occupies what it occupies. Rc<RefCell<T>> is reference counting with the counting bill attached. Leaking is safe in Rust; mem::forget is not even unsafe. What Rust genuinely gives you, besides safety, is that peak memory is approximately live memory, because nothing is waiting to be collected, and that the cheap thing to write is usually the compact thing. Owned values in a Vec are the path of least resistance. In a garbage collected language, the path of least resistance is a pointer to a heap object, and you pay for that twice: once in bytes and once in scan time.

Hold on to that sentence. It turned out to be the whole answer.

How a collector stopped stopping the world

If you learned about garbage collection from a Java horror story told around 2008, the mental model you have is: the collector runs, everything freezes, and if your heap is big enough the freeze is measured in seconds. That was true. It has not been true for a while, and the mechanism that fixed it is worth understanding because it explains what modern collectors charge you instead.

The progression, compressed:

  • 1959, mark and sweep, in LISP. Stop, trace from the roots, sweep everything unreached. Pause proportional to heap size. Every collector since is a variation on this, and it is still what Go does at heart.
  • 1970, copying collection. Cheney’s algorithm: two halves, copy the live objects from one to the other, and the survivors come out compacted with allocation reduced to bumping a pointer. Pause now proportional to the live set rather than the heap, which is a big win when most objects die. It went into Lisp and ML implementations, and it is still how the young generation of a JVM works today.
  • 1978, the idea that made concurrency possible. Dijkstra, Lamport and colleagues published on-the-fly collection: a collector that runs while the program does, and with it the tri-color abstraction. It was a proof on paper long before hardware made it worth shipping, and it is the direct ancestor of every pauseless collector now in production, Go’s included.
  • 1983 to 1984, generational collection. The observation, from Lieberman and Hewitt and then Ungar’s Generation Scavenging in Berkeley Smalltalk, that most objects die very young. So collect a small nursery frequently and the old generation rarely. This is the single largest practical improvement in the history of the field, and it is why the JVM’s and .NET’s default collectors look the way they do. Go, notably, does not do this, for reasons in the next section.

The 1978 entry is the mechanism, so it is worth being precise. Imagine every object is white, grey, or black. White means not yet reached. Grey means reached but its children have not been scanned. Black means reached and fully scanned. The collector starts with the roots grey and repeatedly takes a grey object, blackens it, and greys its white children. When no grey objects remain, everything still white is garbage.

Now let your program run at the same time as this. It can break the algorithm in exactly one way, and it takes two steps to do it: write a pointer to a white object into a black object, and then destroy the last remaining path to that white object from any grey one. The collector will never revisit black, so it will never find the white object, and it will free memory that is live. This is the only failure mode, and it needs both halves.

So you forbid one half. That is a write barrier: a small piece of code the compiler emits on pointer writes while a collection is in progress.

  • An insertion barrier (Dijkstra) shades the newly stored pointer grey. The white object gets reached after all.
  • A deletion barrier (Yuasa, 1990, also called snapshot at the beginning) shades the overwritten pointer grey. The path that was about to be destroyed is preserved for the duration of this cycle.

Either one makes concurrent marking correct. That is the whole trick, and everything since is engineering on top of it: incremental collectors that interleave in small slices, concurrent collectors that mark on other cores, and the JVM’s ZGC and Shenandoah, which use load barriers instead so they can move objects while the mutator reads them and keep pauses under a millisecond on hundred gigabyte heaps.

Go got a concurrent mark and sweep collector in Go 1.5 in 2015. In Go 1.8 it adopted a hybrid write barrier, combining both forms, which let stacks stay black once scanned and removed the stop the world stack rescan that had been the remaining pause floor. Go 1.14 made goroutine preemption asynchronous, so a tight loop with no function calls could no longer hold up a collection. A modern Go program still stops the world twice per GC cycle, at mark setup and mark termination, and both are typically tens of microseconds.

But the pause did not disappear. It was converted. A concurrent collector pays in three other currencies, and every one of them shows up in the measurements later in this post:

  1. Throughput. The write barrier is real code on a hot path, and marking burns CPU that your program wanted. Go targets 25 percent of GOMAXPROCS for dedicated mark workers when a cycle is running.
  2. Headroom. A concurrent collector must start early enough to finish before the program fills the heap. So it lets the heap grow well past the live set on purpose. This is the space time tradeoff, and it is the one people forget.
  3. Floating garbage. Anything that dies after the marking snapshot survives to the next cycle by construction.

The pause got smaller. The heap got bigger. That was the deal.

What Go actually does with your heap

Go’s specific position in that design space is unusual, and knowing it saves you from a lot of advice that is true of the JVM and false here.

The collector is concurrent, non-generational, and non-moving. No nursery, no compaction, no bump allocation. Go’s allocator descends from TCMalloc: per-P caches, size classes, spans. Small objects go to a size class and fragmentation is managed by segregation rather than by moving things around. This means Go cannot do the generational trick, and it also means a pointer’s address never changes, which is why cgo works as easily as it does.

Escape analysis decides heap versus stack, at compile time. A value that does not outlive its frame stays on the stack and costs the collector nothing. go build -gcflags=-m tells you what escaped and why, and it is the cheapest optimization tool in the language.

GOGC sets a goal, not a limit. This is the single most important sentence about Go memory, so here it is as a formula. After each collection, Go computes:

next heap goal = live heap x (1 + GOGC/100)

With the default GOGC=100, the heap is allowed to grow to twice the live set before the next cycle starts. Live set of 300 MB, goal of 600 MB. GOGC=25 means the goal is 1.25x live. GOGC=off means no goal at all.

Read that formula again and notice what is not in it: how fast you allocate. The goal follows the live set. If you halve your allocation rate, you do not lower the ceiling, you just take twice as long to reach it. This is the trap I walked into later, in full view of a profiler that was telling me the truth.

GOMEMLIMIT, since Go 1.19, is the actual limit. A soft cap on total memory that makes the collector run more aggressively as you approach it, up to running continuously. It is the right tool for a container, and it is designed to be combined with GOGC=off when you know your ceiling. The documented failure mode is the death spiral: if your live set alone approaches the limit, the collector runs forever and reclaims nothing, so a hard backstop still belongs in your deployment.

Go 1.25 made GOMAXPROCS cgroup aware by default, which fixed a long standing class of container bug where a Go process on a 128 core host with a 2 CPU quota spun up 128 worker threads and 128 per-P caches. It also shipped a new collector, Green Tea, as an experiment behind GOEXPERIMENT=greenteagc, aimed at scan efficiency on pointer heavy heaps. All the numbers below are on go1.25.12 with the default collector.

What is good Go, and what is folklore

Good, in rough order of payoff:

  • Reduce the number of live objects, not just live bytes. The collector’s mark cost scales with objects and pointer words, not megabytes. A million element []Thing is one object. A million element []*Thing is a million and one, and each one has to be visited every cycle, forever, for as long as it is live.
  • Do not give every record a map. A map[string]string per object is the most expensive habit in the language. In the program measured here, 1.89 million records each carried a properties map: 8.9 million entries spread over 1.89 million maps. Those maps turned out to be only 211,692 distinct, which meant deduplicating them at publication cut the resident graph from 1,966 MiB to 1,211 MiB and the live object count from 29.3 million to 14.2 million. Same data, same language.
  • Intern repeated strings. 60 MiB of file path strings in that same graph were 1 MiB distinct.
  • Preallocate with a real capacity, and prefer one contiguous buffer with an offset index over N separately allocated slices.
  • sync.Pool for high churn buffers of similar shape, understanding that it is cleared on a two cycle schedule and is not a cache.

Folklore, in the sense of being widely repeated and measurably wrong in Go:

  • “Fewer allocations will lower peak memory.” No. Fewer allocations lower CPU spent collecting. Peak follows the live set and GOGC. I have a very clean measurement of this below.
  • “Check RSS to see how much memory it uses.” On macOS, MADV_FREE_REUSABLE keeps pages resident after they are released, so RSS reads several times the live heap and does not move when memory is genuinely returned. One real run showed 12.9 GB RSS against a 1.2 GB live heap and no leak whatsoever. Measure HeapAlloc, or vmmap --summary footprint on Darwin, never ps rss.
  • “Just lower GOGC.” It is a real lever and it is not free, and the size of the bill depends on a ratio most people never compute. That is a whole section, further down.

A day of profiling, and four wrong conclusions

Here is the workload, because the shape of it matters for everything that follows. Calling it a static analyser would be convenient and not quite true: parsing is one stage of eight, and it is the stage that behaves best. The pipeline, in Go, is fixed and deterministic:

  1. Walk. Enumerate the files under the repository, applying ignore globs so build output, vendored code and generated files never reach a parser.
  2. Hash. Content hash every surviving file and look it up in a cache, so an unchanged file is never re-parsed. This is what makes a warm run different from a cold one, and the warm and cold paths turn out to have completely different memory profiles.
  3. Extract. For each file, parse the text into a syntax tree, tree-sitter for most languages and Go’s own go/ast for Go, walk the tree, and emit facts: small records saying that this function exists at this line, that this file imports that package, that this method calls that one, that this HTTP route is served by that handler. Files run in parallel across the available cores.
  4. Store. Merge every file’s facts into one in-memory store, indexed by kind, file, name and repository.
  5. Link. Resolve the edges no single file could have shown. A call site in one file to a symbol declared in another, a route to the handler that serves it, and across repositories, evidence that one service calls another’s endpoint or consumes its Kafka topic. This runs over the whole assembled fact set, which means the whole fact set has to be assembled first.
  6. Index. Build the bidirectional graph the traversal queries run on.
  7. Analyse. Run deterministic analyses over that graph: cycle detection with Tarjan, layering violations, fan-in and fan-out outliers, dependency depth, complexity outliers. These produce findings, and some of them write derived values back onto the facts.
  8. Write. Render a token budgeted summary, sort everything, hash it, and write the artifacts to disk.

Where the wall clock goes, on the two repositories used throughout this post:

repo walk hash extract link index analyse render
runtime 2.36 s 2.64 s 47.7 s 0.37 s 0.20 s 0.73 s 0.08 s
linux 4.68 s 7.78 s 118.1 s 1.34 s 0.85 s 4.19 s 0.37 s

Extraction is 85 percent of the time, and that is exactly why the memory results below are counterintuitive. Time and memory do not live in the same stage. Those seven stages sum to 54.1 seconds of runtime’s 55.1 second run, which means stage 8, the write, is about the last second of it. The single largest memory win in this whole post is in that last second, and so is the peak heap of both .NET repositories. If you had profiled this program for CPU and then optimized what you found, you would have spent all your time in stage 3, extraction, and never touched the thing that set the ceiling.

Two representative inputs. dotnet/runtime, about 15,700 parsed C# files, produces 397,608 facts. The Linux kernel produces 1,892,479 facts, which index into a graph of 1.83 million nodes and 5.36 million edges. Behind both sits a corpus of 81 public repositories carrying 20 language tags: Ruby, Dart, Scala, PHP, C#, Rust, TypeScript, Python, Go, Java, C and C++, Swift, Kotlin, Vue, Svelte, F#, Razor, XAML and gRPC. That breadth is not decoration. Most of this post is about a change that looked excellent on two of those 81 and was a disaster on seven others.

The starting numbers, taken as the worst of six runs:

baseline final
linux peak heap 7,854 MiB 5,228 MiB -33%
linux allocations per fact 756.7 438.9 -42%
runtime peak heap 1,929 MiB 1,454 MiB -25%
runtime allocations per fact 1305.3 678.0 -48%

Total wall clock across all 81 repositories went from 558s to 539s, so 3.5 percent faster, with no repository more than 6 percent slower. Fact counts stayed identical everywhere, which is the check that none of this changed behaviour.

That is the outcome. The route to it is the useful part.

Instrument first, and know what your instrument measures

The first thing built that day was not an optimization, it was a sampler: read runtime.MemStats on a ticker for the life of the process, keep the high water mark, and write a heap profile at it. Everything else the runtime will hand you is either the survivor of a collection or the operating system’s opinion.

Two things went wrong immediately, and both are worth stealing.

The sampler ran at a fixed 150 ms. A small repository that finished in 374 ms got two samples and reported a peak of “7 MiB”, which is not a measurement, it is a coin flip. The rate is now 10 ms for the first two seconds, and that same repository reports 12 to 13 MiB reproducibly. If your sampling interval is within an order of magnitude of your workload, you are measuring the sampler.

Then, precision. Six runs of the same two repositories gave this:

metric runtime linux spread
allocations per fact 1305.3 / 1305.3 756.7 / 756.7 under 0.001%
peak heap 1,573 to 1,929 MiB 6,537 to 7,854 MiB 2% / 17%

Allocation count is essentially deterministic. Peak heap swings 17 percent run to run on the same binary and the same input, because peak is set by collector pacing as much as by what the code allocates, and the larger the heap the more room pacing has to vary.

This has a direct practical consequence: allocation count is your sensitive instrument and peak heap is your coarse one. A change that improves peak but not churn has probably just been lucky with a collection. When these numbers became a regression gate, the churn metric got an 8 percent tolerance and the peak metric got 25 percent, and anything tighter would have been a flaky test.

Read the profile at the right moment

This is the mistake I would most like to save you, because it wasted hours and it produced a confident, wrong, written down conclusion.

A heap profile in Go describes the last completed mark. HeapAlloc describes now. Those are different moments. I captured a profile at the HeapAlloc high water mark, compared its total against the heap goal implied at a different instant, found a 2.3x gap, and wrote up that pprof was under reporting live memory by more than a factor of two.

It is not. Captured and compared at the same moment, against the peak of NextGC (the collector’s own heap goal, which is where the live set is genuinely largest), the profile totalled 2,984 MB against the 2,995 MB the goal implies. 0.4 percent apart. pprof is accurate. It is just not interchangeable across instants, and a tool that describes a specific moment will happily be compared against a different one without complaining.

So there are two peaks and you need to be explicit about which you are chasing:

  • peak HeapAlloc is the maximum footprint, live plus garbage not yet swept. This is the number that OOM kills you.
  • peak NextGC is the moment the live set is largest. This is the number you can actually do something about by changing your data.

For this workload, peak over steady state was 5.7x on one repository and 5.8x on the other. Generating the graph cost roughly six times what holding it costs, and closing that gap was the entire project.

Attribution is not causation

Two experiments, both clean, both refuting the plan they were meant to confirm.

Experiment one. The profile at maximum live on dotnet/runtime said the serialization buffer was 320 MB, or 45 percent of everything live at that moment. On roslyn, 255 MB and 36 percent. That is the biggest single item on the list, so I rewrote it to stream instead of materializing. It is a genuine reduction in what is retained. Here is what it did to peak heap:

repo before after
roslyn 818 MiB 821 MiB
runtime 882 MiB 886 MiB
linux 3,701 MiB 3,726 MiB

Nothing. On the kernel, because the kernel’s peak is somewhere else entirely. On the .NET repositories, where the buffer genuinely is at the peak, because the heap goal in that window was set by the marshalling passing through rather than by the buffer sitting there. The change was reverted before it shipped.

The reasoning error, stated plainly, because it is the most common one in performance work: “X is 45 percent of live at the peak” is an attribution. “Removing X cuts the peak by 45 percent” is a causal claim. The second requires the peak to be caused by X, and a profile cannot tell you that. Only the experiment can.

Experiment two, same lesson in a different costume. The largest single allocation source in the whole program was a binding that built a fresh Go string on every syntax node type lookup, from a C string the parser had already interned. 180 million objects, 3.2 GB on one repository and 6.3 GB on the other, from 669 call sites. Replacing it with a per grammar lookup table is exactly the kind of change that feels like it must matter:

before after
allocations 516,429,814 271,227,339 (-47%)
total allocated 18,974 MiB 15,691 MiB (-17%)
peak heap 1,160 MiB 1,159 MiB

Halving the allocation count moved the peak by one mebibyte. Which is precisely what the GOGC formula predicts, and I had written that formula down earlier the same day. The heap grows to a goal derived from the live set, and refills to that goal regardless of how fast it is filled. Allocating half as much means collecting half as often to reach the same ceiling. That is a CPU saving, not a memory saving, and it was worth keeping as one.

What actually moved the peak

Every change that worked had the same shape: it removed something that was live, or it removed work entirely. None of them was clever.

Do not build a thing you are about to throw away. A read only mode was serializing the entire fact set to JSON for a cache it had already decided not to persist. On the kernel that is 800 MB of JSON and a 1.5 GB transient, produced and immediately discarded, every run. The fix is one conditional, and it cannot possibly change output.

Do not hold two copies of the same thing. The cache held a marshalled buffer from load until save, alongside the objects decoded from it. Encoding entry by entry straight to the stream removed the retention. Interestingly, this raised allocation count by 0.3 percent, because a few enormous allocations became 1.9 million small ones that die immediately. Peak went down, churn went up, and both being tracked is the reason that trade was visible instead of an argument.

Do not materialize two datasets to produce six integers. This was the big one, and it was invisible for two full phases because nobody had ever profiled a warm run. Every profile until then had been a cold run. When a warm run was finally profiled, 60 percent of the live heap at its peak was one thing: stage 8, the write, was loading the entire previous snapshot back into memory and building a key index over both sides, to produce a summary of six counters and a small map. On the kernel the previous snapshot is 830 MB of JSON. Streaming it and keeping only a key and hash cut the warm peak 47 percent on the kernel and 36 percent on the .NET repository.

That one is also the best example of a bug class worth naming: the expensive part was not doing anything wrong, it was doing something correct at the wrong scale. The computation was right. It just did not need either side materialized.

Do not precompute what is usually never used. The C preprocessor extractor lexed every macro body at collection time into a token slice. On the kernel that table was 671 MB, 22 percent of everything live at maximum, and it stayed live for the entire parse because any file might expand any macro. Storing the body as text and lexing on first expansion: peak down 13 percent, no measurable time cost, output byte identical. Most macros in a kernel are never expanded once.

And a hypothesis that the measurement killed before it cost anything: the lexer sliced its input, so every macro looked like it must be pinning its whole source file alive. Measured over 4,000 kernel sources, the difference between the code as written and a version that cloned every string was 45 MiB against 43 MiB. Two mebibytes. The cost was the tokens themselves, 755,029 of them at 24 bytes each to describe a few characters. Ten minutes of measurement, one wrong fix not written.

The change that worked, except on seventy nine other repositories

The best single number of the day came from one line: lower GOGC for the duration of the work, then restore it. It is the textbook lever, it is exactly what the formula says should work, and it delivered. Peak fell another 22 percent on one repository and 14 to 15 percent on the other, allocation counts and steady state unchanged to the digit, which is the signature of a pure pacing change: same bytes allocated, collected sooner.

Cost, measured on those two repositories: between 1.5 and 4.6 percent wall clock. A third of the peak for 4 percent of the time looked like an obvious trade.

Then the full 81 repository sweep ran, and found seven repositories between 64 and 246 percent slower. All seven were Scala.

repo before after live set allocated
zio 6.0 s 20.6 s (+246%) 18 MB 57 GB
spark 36.1 s 102.2 s (+183%) 161 MB 289 GB
openwhisk 1.7 s 4.7 s (+181%) 12 MB 10 GB
pekko 5.7 s 15.4 s (+171%) 64 MB 43 GB
http4s 2.8 s 6.8 s (+139%) 10 MB 18 GB
pekko-http 1.5 s 3.6 s (+139%) 17 MB 8 GB
lila 4.3 s 7.1 s (+64%) 42 MB 13 GB

That last column needs explaining, because at first sight it is absurd. Spark did not use 289 GB of memory. Nothing on the machine could have. “Allocated” is the cumulative total over the whole run: every byte that was ever handed out, including bytes handed out and reclaimed and handed out again thousands of times over. The “live set” column is what was actually held at any one instant, and for spark that is 161 MB.

The gap between those two columns is the entire subject of this post. Extraction, stage 3, is a conveyor belt: read a file, build a syntax tree, walk it, emit a handful of facts, throw the tree away, next file. The tree for one source file might be a few megabytes and lives for a few milliseconds. Multiply by 200,000 files and you have allocated hundreds of gigabytes through a window that is never more than a couple of hundred megabytes wide. Spark’s 289 GB over 36 seconds is about 8 GB per second of allocation, which sounds enormous and is entirely ordinary for this kind of work, because the same physical memory is being reused continuously.

So the two columns measure genuinely different things: allocated is how hard you work the collector, live is how much memory you need. Every row in that table is a repository that works the collector extremely hard while needing almost nothing, which is exactly the profile that makes lowering GOGC catastrophic.

A same binary A/B settles it: zio takes 20.3 s at GOGC=25 and 5.8 s at GOGC=100, to save 33 MiB on a repository that peaks at 111 MiB.

The predictor is not repository size. It is the ratio of allocation to live set. zio allocates 57 GB against an 18 MB live set. At GOGC=25 the goal is 22.5 MB, so the collector re-triggers roughly every 4.5 MB of allocation and runs thousands of times over the run. Spark has a 161 MB live set and still lost 183 percent, because it allocates 289 GB. The two repositories I had measured on have live sets of 300 MB and 850 MB, where 25 percent headroom is still hundreds of megabytes in absolute terms, which makes them the two repositories in the corpus least able to show this failure.

The whole pacing change was reverted. Peak went back up by about half the total win, and what remains is only the work that was genuinely removed rather than merely re-paced.

Three things I took from that:

  1. A change that alters runtime behaviour rather than work done needs your whole corpus before you believe it. Two data points were never a sample. Every change that removed work generalized perfectly across all 81 repositories. The one that re-tuned the collector generalized to seven catastrophic regressions.
  2. GOGC is not a memory knob, it is an exchange rate, and the rate you get depends on your allocation to live ratio. Compute it before you turn the dial. A high churn, low live workload is the worst possible candidate.
  3. If you want it anyway, it has to be adaptive: pace only once the live heap is large enough for the headroom to be worth anything.

The machine is part of the measurement

Every number above was taken on an Apple M4 Max, 16 cores, 128 GB of RAM. That fact changes how you should read all of them.

The program sets GOMEMLIMIT to 90 percent of system memory, which on this machine is 115 GB. It never engaged, not once, not even on the kernel at 7.8 GB. So these are the collector’s natural peaks, not limited ones. The GOGC formula ran unconstrained the whole time.

On a 16 GB laptop the soft limit would be 14.4 GB and would still not engage for the kernel run. What would happen instead is the operating system: page cache eviction, then swap, then the OOM killer, or on macOS a machine that becomes unusable long before anything reports an error. The failure mode real users hit is the OS, not Go’s soft limit, which is a good argument for setting GOMEMLIMIT to something you chose rather than to a fraction of whatever the machine has.

In a container it gets worse before it gets better. os.freemem and /proc/meminfo report the host’s memory, not the cgroup’s, so a process that sizes anything from “available memory” will size it from a number that has nothing to do with the limit it will be killed at. Go 1.25 fixed the CPU half of this by making GOMAXPROCS respect cgroup quotas; the memory half is still yours to set explicitly.

Two more machine dependent effects worth knowing:

Concurrency is often not the memory variable you think. Halving the worker count from 16 to 8 on the kernel run changed peak heap from 4,328 MiB to 4,331 MiB, and cost 13 percent wall clock. Per worker parse state simply was not the driver here. It very often is, in workloads that buffer per request, and the way you find out is to measure it rather than to reason about it.

Wall clock measured across hours of benchmarking is not controlled. Warm run durations in this corpus climbed steadily across three sweeps, 27.3 to 28.5 to 30.1 seconds, including across a change that removed work. Within a single sweep the spread was 0.1 s. That drift is a laptop over five hours of continuous benchmarking, not the code. If your comparison spans hours of machine time, thermal state is a variable in your experiment.

What the Linux kernel has to do with .NET

The most useful result of the whole day is that the same binary, on two different inputs, has its memory problem in two completely different places.

repo max live when, in the run largest single holder
cpp/linux 2,984 MB t=47s of 141 (33% in) C macro table, 671 MB (22%)
dotnet/runtime 707 MB t=52.7s of 55.1 (96% in) serialization buffer, 320 MB (45%)
dotnet/roslyn 705 MB t=16.1s of 17.7 (91% in) serialization buffer, 255 MB (36%)

The kernel peaks a third of the way in, during extraction, so nothing in the write stage can possibly be its maximum. The .NET repositories parse fast enough relative to their output size that their maximum lands in stage 8, the write, instead, in roughly the last second of the run. Same code, same eight stages, opposite answers, purely from where each input’s peak happens to fall.

This is why the serialization rewrite was a no-op on all three: on the kernel because the buffer is nowhere near the peak, and on the .NET repositories because it is at the peak but is not what causes it.

And then zio, at the other end: 57 GB allocated against an 18 MB live set. Tiny by any size measure and by far the most collector sensitive repository in the corpus. Size is not the axis. The two axes that predicted everything that day were the ratio of allocation to live set, and where in the run the live set peaks. Neither is visible in a line count, and neither is knowable without measuring the specific input you care about.

If you take one operational thing from this post, take this: profile your largest input and your most allocation heavy input, and expect them to disagree. A fix validated on one of them is not validated.

So, Rust?

Back to the afternoon that started this.

The honest case for the rewrite was the peak to steady ratio: 5.7x. Generating this graph cost roughly six times what holding it costs, and a big chunk of that multiple is structurally the collector’s, because a concurrent collector must leave headroom. GOGC=100 is a 2x allowance by definition. That part is not a bug in my code and no amount of tuning deletes it. Rust would.

The honest case against it was already sitting in my notes from an earlier round of this work, and it is the number I keep coming back to. The resident graph at the time was 2,778 MiB across 39.2 million live objects, about 1,540 bytes and 21 objects per record. A prototype of the same data, in the same language, using interned strings, columnar layout and compressed sparse row adjacency over integer node ids, measured 365 MiB and 8,740 objects. About 7.6x smaller, with zero Rust.

So the question is not “would Rust be smaller”. It is: which of my two problems am I actually solving?

  • If your problem is the multiplier between live and peak, a language without a collector removes it structurally. Nothing you do in Go removes it, you can only tune it, and this post is a fairly complete tour of what tuning gets you and what it costs.
  • If your problem is that your live set is several times larger than the data it represents, your language is not the problem. It is layout: pointers where you could have indices, maps where you could have columns, per record allocations where you could have one array. Rust would help only in the sense that it makes the compact version the path of least resistance, which is a real ergonomic advantage and not a capability one.

My live set was 7.6x its own data. That is not a garbage collector problem wearing a disguise, that is a data structure problem, and it was fully available in Go. The measurement said so before I wrote a line of Rust.

The other half of the answer is that not one of the changes that worked that day would have been found by a rewrite. Loading an 830 MB file to produce six integers is the same mistake in every language. Serializing a fact set to a cache you already decided not to write is the same waste in every language. Eagerly lexing 92,841 macro bodies that are never expanded is the same in Rust, and in Rust it is arguably harder to notice because there is no GC to blame and no collector metric to watch.

The general rule I would now apply to any “should we rewrite this in a systems language” question: measure what fraction of your footprint is runtime overhead versus your own layout, before you argue about the runtime. If it is mostly overhead, the rewrite has a real case. If it is mostly layout, you are proposing to spend a year rewriting so that a compiler will nag you into the change you could make on Monday.

There are excellent reasons to choose Rust. “Our Go program uses too much memory” is usually not one of them, because that sentence is usually describing a data structure.

Why you still have to know this, AI or not

Every line of code in that day’s work was written fast, with an AI in the loop. The implementation was never the bottleneck. Here is what the day actually consisted of:

  • a conclusion that a profiler under reported by 2.3x, which was two instruments read at different instants;
  • a fix targeting 45 percent of the live heap, which moved the peak on zero of three repositories;
  • a 47 percent allocation reduction that moved peak memory by one mebibyte, in direct contradiction of a formula written down earlier the same day;
  • a change validated on two repositories that was catastrophic on seven others;
  • a hypothesis about strings pinning source files that was worth two mebibytes.

Not one of those was a coding error. All five were reasoning errors, and every one of them was the kind a language model will commit with total confidence and considerable eloquence, because a plausible mechanism can be generated for any number you put in front of it. Ask why peak memory did not fall after you halved the allocations and you will get a fluent answer. You will get a fluent answer whether or not the real reason is that the heap goal follows the live set.

What broke each of those five was not more reasoning. It was an experiment: revert it, run it, look. The skill that mattered was knowing which number to distrust, which instrument reports which moment, and what the collector’s actual contract with your program is. That knowledge is what converts a fast code generator into leverage instead of into a very rapid way of shipping a confident wrong conclusion.

The AI made the day possible. Fifteen experiments in one day is not a pace I would have hit alone. But the value came from the ones I killed, and killing them required knowing what GOGC does.

The checklist

If you are chasing Go memory on Monday, in order:

  1. Set a goal in the right units. Peak HeapAlloc if you are being OOM killed, max live if you are trying to shrink the data. They are different moments and different projects.
  2. Sample runtime.MemStats over time, at least an order of magnitude faster than the thing you are measuring. Never ps rss, especially on macOS.
  3. Track allocation count and peak separately. Churn is deterministic and will catch a regression; peak swings up to 17 percent run to run and will not.
  4. Take one profile at peak HeapAlloc and one at peak NextGC. Compare each only against numbers from its own moment.
  5. Compute your allocation to live ratio before touching GOGC. High churn over a small live set is the worst possible candidate for lowering it.
  6. Treat every profile line as an attribution, never a cause. Build the fix, measure it, and be prepared to revert. Two of mine were reverted after they worked exactly as designed.
  7. Look for work that should not happen at all before you look for allocations to shave. Everything that generalized across 81 repositories was in this category.
  8. Validate on your largest input and your most allocation heavy input. They will disagree, and the disagreement is the finding.

If you would like more in this direction, I have written about what four different code graph tools spend their memory on, which includes the same layout argument measured across four languages, and about what tradeoffs actually cost when the bill arrives.

One open question I never answered, kept here because a post that only reports wins is not reporting: about 490 MiB stays resident after the output is written, survives a forced collection, and is therefore retained rather than uncollected. I removed the thing I was sure was holding it, and the number moved by 16 MiB. Something else has it, and I have not found it yet.

Comments live on X - or just email me.