Skip to content

How a box is stored

When you write motor.position = 10, the number has to end up somewhere every process can see, in a form every process understands. This page explains where a box lives in memory, how its values are stored, and how two processes make sure they agree on what the bytes mean. Segment layout has every offset and step in full.

One named mapping per box

Every box lives in one block of memory that the operating system lets several processes map at the same time, its segment. The block has a name, so any process that knows the name can open it, which is how a second process finds your box.1

On Linux it is a POSIX shared memory object, /dev/shm/sharedbox.<name>, created with mode 0600, readable and writable by its owner only.1 On Windows it is a file mapping backed by the page file, Local\sharedbox.<name>, which Windows deletes when the last process closes it.2 The standard library's multiprocessing.shared_memory creates the same kind of object on Windows.3 Creation always asks for a new name, so a box never attaches to a block that another program created first. Names lists every object name.

A fixed layout: header, field table, record

Inside the block, everything has a fixed place. Here is the memory of the tutorial's Motor box, with its int, its bool and its 32-byte str, after the second process set position to 10. The left column is the address, the right one what is stored there; point at a part to read more:

The memory of a Motor box
header 0x000 - 0x07F0x000magic "SHREDBX1", written last0x008layout 2.00x00Cfield_count 30x00Ewaiter_slots 640x010schema_hash0x018record_size 480x01Crecord 0x6C00x020tail 0x0800x024size 0x10000x028create_id, creator start and pid0x03Ctypes_size 00x040seq 2, even: no write running0x048writer_pid 00x04Cwake_word, waiters0x058creator_pidns, then reservedfield table 0x080 - 0x0970x080position +0x00, int, 80x088enabled +0x2C, bool, 10x090label +0x08, str, 32write counts 0x098 - 0x0AF0x098position 10x0A0enabled 00x0A8label 0waiter slots 0x0B0 - 0x6AF0x0B0slot 0 start, pidns, pid, interrupt0x0C8slot 1...slots 2 to 630x6B0padding to a 64-byte boundaryrecord 0x6C0 - 0x6EF0x6C0position = 10 | 0A 00 00 00 00 00 00 000x6C8label length = 6 | 06 00 00 000x6CClabel = "x-axis" | 78 2D 61 78 69 73, then unused0x6ECenabled = False | 000x6EDpadding to 48 bytesunused 0x6F0 - 0xFFF0x6F0rest of the first 4 KiB pageLine 0, up to 0x03F, is written once when the box is created. Line 1, from 0x040, changes with every write and wait. One 8-byte entry per field in declaration order. Each gives the offset in the record and, in one u32, the kind code and the size. A process that opens the box copies this table once. One u64 per field, raised by every write to it, so a watcher can tell which fields changed. 64 slots of 24 bytes, one per thread waiting for changes. Each records the owner's start time, pid namespace and pid, and an interrupt flag. The space up to 0x6C0 is padding. The field values, ordered by alignment, largest first. Every value is little-endian. The mapping is rounded up to whole 4 KiB pages.
header 0x000 - 0x07F0x000magic "SHREDBX1", written last0x008layout 2.00x00Cfield_count 30x00Ewaiter_slots 640x010schema_hash0x018record_size 480x01Crecord 0x6C00x020tail 0x0800x024size 0x10000x028create_id, creator start and pid0x03Ctypes_size 00x040seq 2, even: no write running0x048writer_pid 00x04Cwake_word, waiters0x058creator_pidns, then reservedfield table 0x080 - 0x0970x080position +0x00, int, 80x088enabled +0x2C, bool, 10x090label +0x08, str, 32write counts 0x098 - 0x0AF0x098position 10x0A0enabled 00x0A8label 0waiter slots 0x0B0 - 0x6AF0x0B0slot 0 start, pidns, pid, interrupt0x0C8slot 1...slots 2 to 630x6B0padding to a 64-byte boundaryrecord 0x6C0 - 0x6EF0x6C0position = 10 | 0A 00 00 00 00 00 00 000x6C8label length = 6 | 06 00 00 000x6CClabel = "x-axis" | 78 2D 61 78 69 73, then unused0x6ECenabled = False | 000x6EDpadding to 48 bytesunused 0x6F0 - 0xFFF0x6F0rest of the first 4 KiB pageLine 0, up to 0x03F, is written once when the box is created. Line 1, from 0x040, changes with every write and wait. One 8-byte entry per field in declaration order. Each gives the offset in the record and, in one u32, the kind code and the size. A process that opens the box copies this table once. One u64 per field, raised by every write to it, so a watcher can tell which fields changed. 64 slots of 24 bytes, one per thread waiting for changes. Each records the owner's start time, pid namespace and pid, and an interrupt flag. The space up to 0x6C0 is padding. The field values, ordered by alignment, largest first. Every value is little-endian. The mapping is rounded up to whole 4 KiB pages.

Layout gives every offset. The values you pass to create are in the record before the header's magic word is set, and a process that attaches waits for that word, so it never sees a field before it holds its starting value.

Fixed places might look rigid next to something like a dictionary kept in shared memory, but they buy three things:

  • Reading or writing a field is one copy of a known number of bytes to or from a known address. There is no structure to walk and nothing to rebalance.
  • Values are stored as plain bytes that the native module converts. Nothing is ever unpickled. Unpickling runs code chosen by whoever wrote the bytes,4 so a box whose values were pickles would let any process that can write the block run code in every process that reads it.
  • A process that opens the block with a different version of the class is refused: the schema hash in the header must match the hash the opener computes from its own class.

The record's position is stored as a distance from the start of the block, not as an address, because each process maps the block at a different address.

How values are stored

Every value is turned into a fixed number of bytes: numbers packed the way struct packs them, text as UTF-8. Reading turns the bytes back into a value, and nothing stored is ever unpickled or run as code. The table in SharedBox lists the stored size of each field type, and Layout the bytes of each kind.

Because each value is copied into the record as bytes, a box can only hold the types it knows how to store, listed in Field types, and not any Python object. A field takes the room of its capacity whatever it holds. Capacity goes inside Annotated because it belongs to the stored type, not to the field's options.

Each field is an attribute of the class, and reading or writing that attribute is what reaches into the record. A plain class attribute with the same name in a subclass would hide it, so a subclass can't simply assign a new default to an inherited field; it declares the field again with an annotation instead (see field).

Two box objects on the same segment share one record, so they always hold the same values, and comparing them by value would tell you nothing. So == on two boxes is True only when both are the same Python object, just as is would be.

The schema hash

Imagine one process declares the field at offset 16 as a float and another, running older code, as an int. The second would read the first one's numbers as nonsense, with no error to warn it. The field table can't catch that: it gives each field's offset, capacity and kind, but not which field is called position, and only the classes know what the fields mean. So when a process opens a block, something has to confirm that its class agrees with the creator's about every byte.

That something is the schema hash: a fingerprint of the class's identity and each field's name and type, stored in the header by the process that creates the box. Schema identity has the exact text that is hashed and a worked example. A process that attaches works out the hash from its own class and compares it with the header's before it reads any field.

So anything that changes the meaning of the bytes changes the hash, and attach raises SchemaMismatchError instead of reading nonsense. That covers adding, removing, renaming or reordering a field, changing its type or capacity, and changing the identity (which by default happens when you move or rename the class). Changes that leave the bytes' meaning alone don't affect it: a new method, a docstring, a different default value.

It helps to know what the schema hash doesn't do:

  • It is not a check of who created the block. Any process that can write the block can write any hash into the header. It protects against mistakes, such as an old and a new version of a program running at the same time, not against a hostile process. Reading and writing and the 0600 permissions deal with that.
  • It does not cover changes in how sharedbox itself lays out a record between releases. The header's layout_major and layout_minor do.
  • 8 bytes of SHA-256 give 2^64 possible values, so two different classes sharing a hash by accident is not a practical concern.

Sources


  1. Linux manual page shm_open(3): named shared memory, the name as the way to open it, permission bits. https://man7.org/linux/man-pages/man3/shm_open.3.html ↩↩

  2. Microsoft, CreateFileMappingW: a mapping backed by the paging file, freed when its last handle is closed. https://learn.microsoft.com/en-us/windows/win32/api/winbase/nf-winbase-createfilemappingw ↩

  3. CPython source, Lib/multiprocessing/shared_memory.py (the Windows branch uses a page-file-backed file mapping), and the documentation of SharedMemory.close() and unlink(). https://github.com/python/cpython/blob/main/Lib/multiprocessing/shared_memory.py, https://docs.python.org/3/library/multiprocessing.shared_memory.html ↩

  4. Python documentation, pickle, the warning at the top of the page: unpickling can execute arbitrary code. https://docs.python.org/3/library/pickle.html ↩