⏱️ Reading time: 14 min

Every time a process on Linux calls fork(), the kernel creates a full child process in microseconds without copying memory. The secret behind this is the copy-on-write technique: parent and child share the same physical pages until one of them writes to them.

📑 En este artículo
  1. TL;DR
  2. What is the copy-on-write technique?
  3. Why deferred memory copying matters
  4. How COW (copy-on-write) works internally
    1. The page as the minimum unit
    2. The page fault that triggers the copy
    3. mmap() with MAP_PRIVATE also uses COW
  5. Practical examples
    1. Basic fork: two processes, one starting point
    2. Shared memory until someone writes
    3. The real pattern: workers that inherit data without copying it
  6. Getting started
    1. Cloning a file without copying it (btrfs or xfs)
    2. Checking that a page is still shared
  7. Real-world use cases
  8. Common mistakes and best practices
  9. Comparison with alternatives
  10. Going deeper: what happens at the MMU level
  11. Frequently Asked Questions
    1. What’s the difference between the copy-on-write technique and cloning an entire file?
    2. Why doesn’t Python’s multiprocessing achieve the full savings that deferred memory copying promises?
    3. Does COW (copy-on-write) work on Windows?
    4. Does Redis use COW (copy-on-write) for RDB snapshots?
    5. What happens if the system runs out of memory while copying pages via lazy data duplication?
    6. Which filesystems support block-level lazy data duplication?
  12. References

Git, Docker, and ZFS use the same idea to avoid copying entire files. Redis takes advantage of it to save a snapshot of the database without blocking writes.

TL;DR

  • Copy-on-write lets two processes share the same memory page until one of them modifies it.
  • fork() on Linux creates a full child process without duplicating memory because it inherits the parent’s pages.
  • Docker copies the entire file to the writable layer when a single byte is modified, not just the changed block.
  • Redis uses fork() to take a consistent snapshot of the database without blocking writes.
  • cp –reflink=auto clones a file of any size in milliseconds on btrfs or xfs systems because it doesn’t copy blocks.

What is the copy-on-write technique?

The copy-on-write technique is a mechanism that postpones copying a piece of data until someone tries to modify it. While it’s only being read, multiple consumers share the same physical copy; at the moment of writing, the system creates a private copy of only that modified portion.

The mechanism applies at two different levels. In memory, the unit that gets shared is the page, a 4 KB block managed by the operating system. In storage, the unit is disk blocks or, in Docker’s case, the entire file within a layer.

In both cases the savings are the same. An expensive copy is avoided when it’s most likely that no one will need to modify that data. If no one ever writes to it, the copy never happens.

Why deferred memory copying matters

Without copy-on-write, every fork() would have to duplicate the parent process’s entire memory before starting the child. A server with several gigabytes of resident memory would take a noticeable amount of time to clone itself, and would spend that same amount of RAM for each child, even if the child was only going to run a short command and exit.

With copy-on-write, fork() is nearly instantaneous because it only copies the process’s control structures (page table, file descriptors) and marks the data pages as shared. The real cost is paid later, gradually, only for the pages that actually change.

The same applies at the storage level. Cloning a virtual machine image, taking a database snapshot, or creating a container from an image would be slow and space-expensive operations if each clone involved copying actual bytes. With COW (copy-on-write), these operations are nearly free until something changes.

How COW (copy-on-write) works internally

The page as the minimum unit

The Linux kernel organizes a process’s memory into 4 KB pages. Each page has an entry in the process’s page table that points to a physical address in RAM and carries read, write, or execute permissions.

When a process calls fork(), the kernel doesn’t copy those pages. Instead, it copies the child’s page table so it points to the same physical addresses as the parent’s, but marks all those entries as read-only, even if they originally allowed writing.

A typical fork() on Linux takes microseconds because it copies control structures, not data. Foto de Waypixels en Unsplash
flowchart TD
    A["Parent process in memory"] --> B["fork()"]
    B --> C["Child process: same pages, read-only"]
    C --> D{"Does anyone write to a page?"}
    D -- "No" --> E["Pages remain shared"]
    D -- "Yes" --> F["The kernel copies only that page"]
    F --> G["Parent and child now have distinct private pages"]

The page fault that triggers the copy

When the parent or child tries to write to one of those pages marked as read-only, the MMU (the processor’s memory management unit) detects the violation and generates a page fault. The kernel handles that fault, allocates a new physical page, copies the old content there, and updates the page table of the process that wrote to point to the private copy.

The other process never knows anything happened. It keeps pointing to the original page, which is now exclusively its own. The kernel only copies that 4 KB page, not the process’s entire gigabytes.

sequenceDiagram
    participant P as Process
    participant M as MMU
    participant K as Kernel
    P->>M: writes to a shared page
    M-->>K: generates a page fault
    K->>K: copies the page to a new physical address
    K-->>M: updates the process's page table
    M-->>P: the write completes on the private copy
    Note over P,K: the rest of the pages remain shared

mmap() with MAP_PRIVATE also uses COW

fork() isn’t the only thing that relies on this mechanism. When a program maps a file with mmap(..., MAP_PRIVATE, ...), the kernel shares the file’s pages among all the processes that mapped it that way, but if one of them writes, that page gets copied and becomes private without touching the file on disk. This is how dynamic library loaders share a library’s code segment among dozens of processes while keeping each one’s data segment private.

Practical examples

These three examples in Python go from the simplest case to a real production pattern. All of them require Linux or macOS, because Windows doesn’t implement fork().

Basic fork: two processes, one starting point

import os

pid = os.fork()
if pid == 0:
    print(f'Child: PID {os.getpid()}')
else:
    print(f'Parent: PID {os.getpid()}, child PID {pid}')

This script creates a child process that prints its own PID while the parent prints its own and its child’s. The actual output varies because PIDs are assigned by the system, but it always follows this pattern:

Parent: PID 48213, child PID 48214
Child: PID 48214

Shared memory until someone writes

import os

datos = [0] * 10_000_000  # a large list in memory

pid = os.fork()
if pid == 0:
    print('Child sees the first value:', datos[0])
    os._exit(0)
else:
    os.waitpid(pid, 0)
    print('Parent continues with the same list without duplicating it')

The child only reads the list, never modifies it, so it never triggers a page copy. The output is always this:

Child sees the first value: 0
Parent continues with the same list without duplicating it

The real pattern: workers that inherit data without copying it

import multiprocessing as mp

dataset_grande = {'usuarios': list(range(5_000_000))}

def procesar(worker_id):
    total = len(dataset_grande['usuarios'])
    return f'worker {worker_id} saw {total} users'

if __name__ == '__main__':
    ctx = mp.get_context('fork')
    with ctx.Pool(4) as pool:
        resultados = pool.map(procesar, range(4))
    for r in resultados:
        print(r)

With the fork context, the pool’s four processes inherit dataset_grande via copy-on-write instead of receiving it serialized for each worker. The output:

worker 0 saw 5000000 users
worker 1 saw 5000000 users
worker 2 saw 5000000 users
worker 3 saw 5000000 users

Getting started

Dependencies: a filesystem with COW support (btrfs or xfs with reflink) for the first example, and Python 3 with access to os.fork() for the second. ext4 doesn’t support reflink.

Cloning a file without copying it (btrfs or xfs)

# create a 1 GB test file
fallocate -l 1G original.img

# COW copy: doesn't read the full gigabyte
cp --reflink=auto original.img copia.img

# apparent size vs actual disk space
du -h --apparent-size copia.img
du -h copia.img

The first du shows the file’s logical size. The second shows the actual space it occupies on disk, which barely changes until you modify the copy:

1.0G    copia.img
4.0K    copia.img

Checking that a page is still shared

grep -E 'Shared_Clean|Private_Dirty' /proc/$PID/smaps_rollup

Replace $PID with the process’s actual PID. As long as the pages remain shared, Shared_Clean will be high and Private_Dirty close to zero. An example output:

Shared_Clean:      81920 kB
Private_Dirty:         0 kB

Real-world use cases

Redis and RDB saving. To write a complete snapshot of the database to disk without blocking incoming writes, Redis forks itself. The child process sees a frozen copy of all the memory thanks to COW, while the parent keeps handling commands; if the parent modifies a key, that page gets duplicated, but the rest remains shared.

The layers of a Docker image are read-only; only the top layer can be written to. Foto de Chris Ried en Unsplash

Docker and overlay2. Docker images are made up of stacked read-only layers. When a container modifies a file, overlay2 copies that entire file to the container’s writable layer before applying the change. This operation is called copy-up. Docker also supports a driver on top of Btrfs that does copy at the block level, but overlay2 is the default because it’s simpler to manage.

flowchart TD
    subgraph Image
    L1["Base layer: operating system"]
    L2["Layer 2: dependencies"]
    L3["Layer 3: application"]
    end
    L1 --> L2 --> L3
    L3 --> W["Container's writable layer"]
    W -. "copy-up when a file is modified" .-> L3

Git and local clones. When you clone a local repository with git clone –local, Git links the original repository’s objects via hard links instead of copying them byte by byte. Both repositories share the same inode until one of them runs git gc or rewrites an object.

ZFS and Btrfs. A ZFS or Btrfs snapshot doesn’t copy any blocks when it’s created. The filesystem only freezes the existing pointers; new blocks you write after the snapshot get allocated elsewhere, and the snapshot keeps seeing the old blocks intact. ZFS takes the idea further: its on-disk structure is a copy-on-write block tree, where every write creates new blocks and a new tree root, while the previous root stays intact as long as something references it.

Common mistakes and best practices

Python’s refcounting breaks the savings. CPython keeps a reference count inside every object. Every time a child worker just reads an object, Python still increments that counter, and that write to the counter is enough to trigger a copy of the entire page. In practice, a worker pool that only reads data ends up copying a good chunk of the heap anyway.

💡 Tip: to mitigate this in Python, call gc.freeze() before fork(). It freezes existing objects outside the garbage collector’s tracking and reduces the counter writes that touch them.

Docker copies the whole file, not the block. Unlike Btrfs or ZFS, overlay2 does copy-up at the full-file level. Modifying one byte of a large file inside a container copies the entire file to the writable layer, not just the changed fragment.

⚠️ Heads up: if your container writes frequently to large files, mount a volume instead of relying on the container’s writable layer.

Threads don’t need COW. Threads within the same process already share the same memory space without any extra trick. Copy-on-write solves a different problem: sharing memory between separate processes that used to be a single one.

Memory overcommit can fail late. The kernel allows fork() to promise more memory than is actually available, trusting that not every page will get written to. If many child processes write at once and trigger simultaneous copies, the system can run out of actual RAM. The kernel scores each process with an oom_score based on how much memory it uses and a manual adjustment, and the process with the highest score is the first to go, regardless of whether it was the one that triggered the copies.

Comparison with alternatives

Each mechanism solves the same problem (sharing data without duplicating it) at a different level of the system.

MechanismWhat it sharesWhen it copiesReal example
fork() with COWProcess memory pagesWhen writing to a page (page fault)Redis, preforked servers like gunicorn
ThreadsThe process’s entire memory spaceNever copies; requires locks to avoid collisionsMultithreaded servers in Java or Go
Explicit shared memory (mmap/shm)A hand-chosen memory regionNever copies; synchronization is handled by the appDatabases with a shared buffer pool
Block-level COW (Btrfs/ZFS)Disk blocksWhen writing to a modified blockVolume snapshots
File-level copy-up (overlay2)Entire files in a layerWhen opening the file in write modeDocker containers

Going deeper: what happens at the MMU level

Each process’s page table lives in memory and the MMU queries it on every access. Each entry includes a write-permission bit and, on many architectures, a dirty bit that marks whether the page has already been modified since it was loaded. Copy-on-write turns off the write bit even if the original page had it enabled, and that discrepancy is what triggers the page fault when someone tries to write.

Linux also distinguishes between fork() and vfork(). This second variant, much less commonly used, doesn’t even create a separate page table for the child. It suspends the parent and lets the child use literally the same memory until it calls exec() or exits. It’s faster than fork() with COW, but also riskier, because an error in the child can corrupt the parent’s memory.

The count of how many processes share a physical page lives in a kernel structure called struct page. When that counter reaches one, it means no one else points there anymore, and the kernel can treat that page as private without keeping the read-only permission.

Your next step: run the cp --reflink=auto example on a btrfs or xfs partition and compare the time against cp --reflink=never on the same file to see the difference with your own eyes.

📬 Get new articles by email

We only email about big articles (1-2 a month).

Frequently Asked Questions

What’s the difference between the copy-on-write technique and cloning an entire file?

Cloning an entire file copies all the bytes immediately, regardless of whether they’ll ever be modified. The copy-on-write technique postpones that copy until the exact moment something changes, and in many cases that copy never happens at all.

Why doesn’t Python’s multiprocessing achieve the full savings that deferred memory copying promises?

Because CPython’s garbage collector modifies each object’s reference count even when a worker only reads it, and that write triggers a copy of the page containing that counter.

Does COW (copy-on-write) work on Windows?

Windows doesn’t implement fork() like Linux, but it does use copy-on-write internally in certain process-creation paths, and ReFS supports block-level file-system clones similar to Btrfs’s.

Does Redis use COW (copy-on-write) for RDB snapshots?

Yes. Redis calls fork() before writing the RDB file, and the child process walks through a frozen copy of the database thanks to COW while the parent keeps responding to commands normally.

What happens if the system runs out of memory while copying pages via lazy data duplication?

If the system allowed memory overcommit and suddenly many processes write at once, there might not be enough physical RAM for all the copies. The kernel kills processes to free memory based on their oom_score, and it doesn’t necessarily pick the process that caused the problem.

Which filesystems support block-level lazy data duplication?

Btrfs and ZFS are the most widespread on Linux, both with instant snapshots and clones. XFS added reflink support later, and ReFS offers it on Windows Server. ext4 doesn’t support it.

References

  • Wikipedia: general article on copy-on-write in operating systems and storage.
  • man7.org: man page for fork() on Linux.
  • docs.python.org: official documentation for os.fork() in Python.
  • docs.docker.com: how the overlay2 storage driver works.
  • redis.io: RDB persistence documentation and the use of fork().
  • git-scm.com: documentation for git clone and the –local option.
  • gnu.org: GNU coreutils manual on cp and the –reflink flag.

📱 Enjoy this content? Follow @programacion on Telegram for daily tech content in Spanish: quick summaries, fresh content every day.

Featured image: Foto de Liam Briese en Unsplash

Did it work for you? Got a different error? Say so below: questions get answered and help the next reader.

Leave a comment
Categories: ProgrammingTutorials

Andrés Morales

Developer and AI researcher. Writes about language models, frameworks, developer tooling, and open source releases. Covers ML papers, the tech startup ecosystem, and programming trends.

0 Comments

Leave a Reply

Avatar placeholder

Your email address will not be published. Required fields are marked *

You can include code inside <code>…</code> or, for several lines, <pre><code>…</code></pre>.

This site uses Akismet to reduce spam. Learn how your comment data is processed.