A single memory swap operation caused the Go garbage collector to pause for 40 milliseconds, a scenario that developers now see as a cautionary tale about runtime performance in constrained systems. The incident, documented by a Go developer on Hacker News, exposes how operating system swapping can turn a normally microsecond-level garbage collector pause into a significant latency spike.

What You Need to Know

Memory swapping occurs when the operating system moves data between RAM and disk, introducing I/O delays. The Go garbage collector stops the world to manage memory, and if a required page is swapped out, the pause can balloon from microseconds to tens of milliseconds. This matters for latency-sensitive applications running in containers or on systems with limited memory, where swap is a common fallback.

Garbage Collection and Swap Overhead

The Go runtime uses a concurrent garbage collector designed to keep pause times under 500 microseconds in most cases. The collector works by tracing reachable objects while the program continues to run, but it must briefly stop the world at two points: to begin the mark phase and to perform sweep termination. During these stop-the-world windows, the collector accesses memory that may not be in physical RAM if the system is under memory pressure.

The 40ms Incident

The developer reported that a simple workload exhibited a 40ms garbage collector pause. Investigation revealed that the operating system had swapped out a portion of the heap to disk. When the Go runtime attempted to read that memory during the stop-the-world phase, the read triggered a page fault. The system then had to load the data back from the swap file, a process that took tens of milliseconds. The pause was not caused by the garbage collector itself but by the underlying memory subsystem.

Why This Matters

For developers running Go services in cloud environments or containers, the presence of swap can silently degrade performance. Many orchestration platforms, including Kubernetes, configure swap off by default, but custom deployments or developer machines may have it enabled. A single 40ms pause can break latency SLAs for real-time systems, trading platforms or API gateways. The incident also underscores a broader truth: garbage collector behavior cannot be analyzed in isolation from the hardware and operating system layer. Memory pressure, swap usage and page cache states all directly influence runtime guarantees.

Lessons for Developers

Teams can mitigate the risk of swap-induced pauses through several approaches:

  • Disable swap: In production environments, turning off swap entirely ensures the garbage collector always has fast access to memory. This is the most common recommendation.
  • Adjust GOGC: Lowering the GOGC environment variable reduces heap growth, which can cut the volume of memory the collector needs to scan and lower the chance of swap activity.
  • Set memory limits: Use cgroups or container resource limits to prevent the process from exceeding physical RAM, avoiding the need for the OS to swap.

Each option comes with trade-offs. Disabling swap may cause out-of-memory kills under heavy load. Tuning GOGC increases CPU usage. The right choice depends on workload characteristics and tolerance for latency versus throughput.

Broader Performance Implications

The 40ms pause is a reminder that garbage collected languages remain sensitive to memory hierarchy latency. While Go's concurrent collector is a significant improvement over older stop-the-world designs, it cannot overcome physical I/O delays. Developers who profile garbage collection in isolation may miss the real bottleneck. Profiling should include operating system metrics such as major page faults, swap usage and memory pressure to get a complete picture.