Segment layout¶
This page is the full contract for how a box is stored in shared memory:
the names of the objects, every byte of the mapping, and the steps every
process follows to read, write and wait. You need it if you write code that
opens a box without sharedbox.hpp, or if you change sharedbox itself;
to understand the ideas first, read
How a box is stored.
include/sharedbox/sharedbox.hpp implements this contract, and the Python
extension runs on that header.
It describes layout 2.0, which sharedbox 0.4.0 and later write, and
layout 1.0, which the 0.3 releases wrote and later releases still open.
Goal¶
C++ and C code can use a sharedbox box, both as an extension loaded in a Python process and as a standalone program with no Python, without linking a library that sharedbox ships. The layout and protocols below are versioned, and one header-only C++20 library implements them.
Decisions¶
| Topic | Decision | Not taken |
|---|---|---|
| Model | the Arrow PyCapsule Interface; the closest analogue is ArrowArrayStream, a live handle rather than a one-shot transfer |
numpy's __array_struct__ |
| Scope | boxes only; the handle struct is versioned so that queues and streams can get their own dunders later | queues and streams in the same step |
| Contract | a documented memory layout plus a header-only reference implementation | a compiled C library; definitions only |
| Language | the C++20 header sharedbox.hpp, namespace sharedbox |
a C99 header with static inline functions |
| Atomics | std::atomic_ref on plain integer fields |
wrappers over compiler intrinsics; C11 <stdatomic.h> |
| C interface | a minimal sharedbox_c.h, whose functions call sharedbox.hpp, in one .cpp file the consumer compiles |
a static library in the wheel; a C implementation of the protocols |
| Dunder and capsule | __sharedbox_box__, capsule "sharedbox_box" |
_c_ in the name, which reads as "implemented in C" |
| Negotiation | DLPack's max_version; the consumer renames the capsule "used_sharedbox_box"; other keywords raise NotImplementedError, as in Arrow |
none |
| Object names | sharedbox. + box name + suffix, for every object |
the bare box name; names stored in the header |
| Windows scope | Local\ only |
Global\ (see Out of scope) |
| Versioning | layout_major refuses, layout_minor opens and ignores what it does not know |
exact match; flag masks |
| Capsule lifetime | the handle maps the segment again from the box's OS handle | sharing the box's mapping and keeping the box alive |
| Schema identity | identity= class keyword, default module.qualname; it enters the schema hash and names the box; handle::create is available to other languages |
the Python module path written into other languages; open-only for other languages |
| Distribution | headers in the wheel, sharedbox.get_include() and a CMake config |
a vcpkg port or a headers-only wheel |
Names¶
Every OS object of a box is named from the box name:
| Object | Linux | Windows |
|---|---|---|
| mapping | shm_open("/sharedbox.<name>"), seen as /dev/shm/sharedbox.<name> |
CreateFileMappingW(INVALID_HANDLE_VALUE, ..., L"Local\\sharedbox.<name>") |
| wake word | inside the mapping, no name | not used |
waiter slot i |
inside the mapping, no name | auto-reset event Local\sharedbox.<name>.w<i> |
- A box name matches
[A-Za-z0-9_.-]{1,128}. - The default name of a class is 16 hex digits of SHA-256 over its identity
(see Schema identity). Without an
identity=keyword the identity ismodule.qualname, with__mp_main__counted as__main__. - Linux: the mapping is created with
O_CREAT | O_EXCL | O_RDWRand mode0600; a name that exists givesstatus::exists.unlink()callsshm_unlink. - Windows: the mapping is backed by the page file.
CreateFileMappingWreportingERROR_ALREADY_EXISTSmeans the name is taken: the creator closes its handle and givesstatus::exists.ERROR_INVALID_HANDLE, an object of another kind under the name, givesstatus::existstoo.unlink()does nothing; the OS frees the mapping with its last handle.
Layout¶
The mapping holds, from offset 0:
offset 0 header 128 bytes, two cache lines
offset 128 field table field_count x 8 bytes
write counts field_count x 8 bytes
waiter slots waiter_slots x 24 bytes
description table types_size bytes, 8-aligned (layout 2.0)
(zero padding up to a multiple of 64)
record record record_size bytes, 64-byte aligned
(zero padding up to a multiple of 4096)
No structure has implicit padding: gaps are explicit pad and reserved
members, and static_asserts in sharedbox.hpp pin each structure's
sizeof, alignof and every offsetof, and that it is standard-layout
and trivially copyable.
struct header { // sharedbox::header, 128 bytes
// line 0: written once at creation
uint64_t magic; // 0 "SHREDBX1"; written last, with release ordering
uint16_t layout_major; // 8
uint16_t layout_minor; // 10
uint16_t field_count; // 12 1 to 256
uint16_t waiter_slots; // 14 1 to 4096
uint64_t schema_hash; // 16
uint32_t record_size; // 24
uint32_t record; // 28 offset of the record
uint32_t tail; // 32 offset of the field table, always 128
uint32_t size; // 36 mapping size in bytes
uint64_t create_id; // 40 random at create, never 0
uint64_t creator_start; // 48 creator's start time, see Liveness
uint32_t creator_pid; // 56
uint32_t types_size; // 60 bytes of the description table; 0 in layout 1.0
// line 1: changed by writes and waits
uint64_t seq; // 64 sequence lock: even when free, odd while writing
uint32_t writer_pid; // 72 holder of the write lock, 0 if none
uint32_t wake_word; // 76 futex word (Linux)
uint32_t waiters; // 80 occupied waiter slots
uint8_t pad0[4]; // 84 zero
uint64_t creator_pidns; // 88 creator's pid namespace; 0 on Windows
uint8_t reserved1[32]; // 96 zero; minor versions may use it
};
Line 0 is written once at creation and read once by attach, which copies
it. Line 1 holds what writes change, so a write touches one header line.
creator_pidns sits in line 1 only because line 0 has no room left; it is
written once at creation like the other creator fields.
The header is 128 bytes so that minor versions can add fields: a 64-byte
header has no spare byte, and every addition would move the fields after
it. There is one header per box, and readers and writers touch only line 1.
The waiter slots, not the header, are the large part of the tail, and
max_waiters sizes them. Worked example, a box with an int, a float
and a bool and the default 64 waiter slots: header 128, field table 24,
write counts 24, waiter slots 1536, padding 16, record 24 (the Python side
rounds a record up to a multiple of 8): 1752 bytes used, one 4 KiB page. A
box with three fields and 64 slots stays in one page up to about 2.3 KB of
record.
- This is layout
2.0: bytes 8 to 11 are02 00 00 00(layout_major2,layout_minor0). Layout 1.0 is the same except that offset 60 is reserved and zero, there is no description table, and the field kinds are 0 to 5 only. Attach accepts majors 1 and 2. The layout version counts the segment format and is not tied to the package version. - Field table entry, 8 bytes:
u32 offset,u32 capacity_and_kind: top 8 bits the kind code. For kinds below 64 the low 24 bits are the size or capacity; for kinds 64 and up they are the offset of the field's description in the table. - Offsets are the creator's choice. The Python side packs fields by descending alignment; another creator may pack differently. Attachers read offsets from the field table and never compute them, so a box created elsewhere with its own packing reads correctly from Python as long as the field order, names, kinds and capacities match, which the schema hash checks.
- Write counts: one
u64per field, incremented under the write lock. - Waiter slot, 24 bytes:
struct waiter_slot { // sharedbox::waiter_slot
uint64_t owner_start; // process start time, see Liveness
uint64_t owner_pidns; // pid namespace, see Liveness; 0 on Windows
uint32_t owner_pid; // 0 = free
uint32_t interrupt; // set by interrupt(), cleared by the waiter
};
- A
reffield refers to another box, which keeps its own segment. Its capacity is always 144:
struct box_ref { // sharedbox::box_ref
uint64_t create_id; // 0 the box's create_id; 0 = empty
uint64_t schema_hash; // 8 the box's schema_hash
char name[128]; // 16 its name, NUL-padded; ends at the first NUL or at 128
};
An empty reference is 144 zero bytes. A non-empty one always has a
nonzero create_id, since create never draws 0. It is written under the
referring box's sequence lock like any other field. A reader that
attaches the named box compares its create_id with the stored one to
tell it from a box created again under the same name.
Kind codes¶
| code | kind | entry's low 24 bits | size, alignment |
|---|---|---|---|
| 0 | bool | 1 | 1, 1 |
| 1 | int | 8 | 8, 8 |
| 2 | float | 8 | 8, 8 |
| 3 | str | capacity | 4 + capacity, 4 |
| 4 | bytes | capacity | 4 + capacity, 4 |
| 5 | ref | 144 | 144, 8 |
| 6 | complex | 16 | 16, 8 |
| 7 | date | 4 | 4, 4 |
| 8 | time | 16 | 16, 8 |
| 9 | datetime | 16 | 16, 8 |
| 10 | timedelta | 12 | 12, 4 |
| 11 | uuid | 16 | 16, 1 |
| 12 | decimal | capacity | 4 + capacity, 4 |
| 64 | enum | description offset | 2, 2 |
| 65 | flag | description offset | 8, 8 |
| 66 | literal | description offset | 2, 2 |
| 67 | optional | description offset | from the description |
| 68 | union | description offset | from the description |
| 69 | record | description offset | from the description |
| 70 | tuple | description offset | from the description |
| 71 | list | description offset | from the description |
| 72 | set | description offset | from the description |
| 73 | dict | description offset | from the description |
| 74 | array | description offset | data size, 64 |
A field of a kind the reader does not know is opened as opaque bytes: a kind below 64 is sized by its entry, one from 64 by its description's head. Every described size is a multiple of its alignment.
Encodings¶
Little-endian.
bool: 1 byte,0x00or0x01.int: 8 bytes, signed.float: 8 bytes, IEEE 754 double.str,bytes:u32length, then up tocapacitybytes (UTF-8 forstr).ref: 144 bytes, asbox_refabove.complex: twof64, real then imaginary.date:i32date.toordinal(), 1 to 3652059.time:i64microseconds since midnight (wall time),i16UTC offset in minutes,u8flags (bit 0 naive, bit 1 fold), 5 reserved zero bytes.datetime:i64microseconds since 1970-01-01T00:00 (wall time if naive, UTC if aware),i16offset in minutes,u8flags astime, 5 reserved zero bytes.timedelta:i32days,i32seconds,i32microseconds.uuid:UUID.bytes, in network byte order.decimal:u32length, then the text ofstr(value).enum,literal:u16position;flag:u64bits.optional: a presence byte, padding to the member's alignmentA, then the member; sizeround_up(A + member size, A).union: au8tag, padding to the largest member alignmentA, then the member; sizeround_up(A + largest member size, A).record,tuple: members at the offsets their entries give.list,set:u32length, padding tomax(4, A), thencapacityslots ofround_up(element size, element alignment)bytes.dict: aslist, with a key and value pair as the slot: the key at 0, the value atround_up(key size, value alignment).array: the elements in C order.
Descriptions¶
The description table follows the waiter slots. Each description is an
8-byte head, u8 kind, u8 flags (zero), u16 count, u32 size, a
body, and zero padding to a multiple of 8. An entry inside a body has the
shape of a field table entry, with the offset counted from the start of
the value. A name is a u16 length, then UTF-8.
| kind | count | body |
|---|---|---|
| enum | 1 to 65535 | count member names |
| flag | 1 to 64 | count members: a name, then u64 bits |
| literal | 1 to 65535 | count values: a u8 tag, then 0 None (nothing), 1 bool (u8), 2 int (i64), 3 str (a name), 4 bytes (u32 length, then bytes), 5 enum member (a name) |
| optional | 1 | one entry |
| union | 2 to 255 | count entries; member i has tag i |
| record | 1 to 256 | count entries, then count member names |
| tuple | 1 to 256 | count entries |
| list, set | 1 | u32 capacity, then the element's entry |
| dict | 2 | u32 capacity, then the key's entry and the value's |
| array | 0 | u8 DLPack code, u8 bits, u16 lanes (1), u8 ndim (1 to 8), 3 zero bytes, then ndim u64 dimensions |
Attach copies the table and checks it before using it: every description is 8-aligned, inside the table, read once, after its parent and apart from every other; nesting is at most 16 deep; counts are in the ranges above; members lie inside their parent, aligned and apart; every size the description implies is computed without overflow and matches the head; an array's code is one of 0, 1, 2, 4, 5, 6, its bits a multiple of 8 and every dimension at least 1. A failure refuses the segment. A reference may only be a field, and an array only a field or a record member.
Decoding checks¶
A stable sequence number shows a copy is consistent, not that it holds a
value Python can have. Readers check presence bytes and bool (0 or 1),
reserved flag bits (0), union tags and enum and literal positions (below
their count), lengths (at most the capacity), date ordinals (1 to
3652059), times (below a day), offsets (within 1439 minutes either way)
and timedeltas (within Python's range), and treat a failure as a corrupt
value.
Only a bool inside a described kind has to be 0 or 1. A field of kinds
0 to 5 decodes as it did in 0.3, so a bool field whose byte is not 0
reads as True.
Writing a list, set or dict¶
The bytes written for a list, set or dict field are its length and the slots it uses, shorter than the field when it is not full; a reader copies the same part.
Schema identity¶
The schema is the list of fields that fixes a record's shape: names, kinds, capacities and their order. The schema hash shows that a creator and an attacher agree on it and on the class's meaning.
- Hash text: the identity, then one
name:kind:capacityper field in declaration order (base class fields first), joined with|, encoded as UTF-8. Kinds are the wordsbool,int,float,str,bytes; capacity is the payload size in bytes (8 forintandfloat, 1 forbool). Areffield enters asname:ref:<identity>, orname:ref?:<identity>when it may be empty (annotatedX | None), with the identity of the class it refers to in place of the capacity. - A field of a kind from 6 on enters as
name:<type text>. The type text of each kind, withTfor the type text of a member:complex,date,time,datetime,timedelta,uuid: the kind name alone.bool:1,int:8andfloat:8are the texts of those kinds inside a described type.str:n,bytes:nanddecimal:n:nis the capacity in bytes. Abytearrayisbytes:n.enum(A,B): the member names in definition order.flag(R=1,W=2): each member name with its value.literal(v,...): each value in order asNone,TrueorFalse,int:1,str:'x'(the Pythonrepr),bytes:b'x'ormember:NAMEfor an enum member.optional(T), andunion(T,...)with the members in the order stored.record(name:T,...)for a dataclass,NamedTuple,TypedDict, attrs class ormsgspec.Struct, with the members in order. ATypedDictkey that may be missing enters asname?:optional(T).tuple(T,...)for a fixed-length tuple.list[n](T),set[n](T)anddict[n](K,V):nis the capacity in elements. Afrozensetis aset;tuple[T, ...]is alist.array(dtype,d1xd2x...), as inarray(float32,2x3).
A class whose fields are all of kinds 0 to 5 has the hash it had in
layout 1.0.
- schema_hash: the first 8 bytes of SHA-256 over that text, read as a
little-endian u64.
- Identity: the identity= class keyword, a non-empty string; without it,
module.qualname. A subclass does not inherit its base's identity.
- The default box name is 16 hex digits of SHA-256 over the identity, so a
class that sets identity= is found by that identity from any package or
language, with no explicit name.
- A set identity lets another language agree with Python on a string both
sides write down, rather than on a Python module path. Moving or renaming
the Python class then keeps the hash, two packages can each declare a
matching class and share a box, and changing the identity ("motor/2")
refuses processes of the old meaning when no field changed (a change of
units, for instance).
- sharedbox.hpp does not carry SHA-256; C and C++ callers compute the
hash with their own library.
- Test vector:
__main__.Motor|position:int:8|enabled:bool:1|label:str:32
schema_hash = 0x82ce467598596a72
default name = 5b4f7004d44277b7 (16 hex digits of SHA-256 over "__main__.Motor")
__main__.Stage|label:str:4|motor:ref:motor/1
schema_hash = 0xe33ffd91e10fab6d
__main__.Stage|label:str:4|motor:ref?:motor/1
schema_hash = 0xe58153f79aa6fdf6
demo/1|when:date|mode:enum(RED,BLUE)|tags:list[2](str:4)|note:optional(int:8)
schema_hash = 0x2a9e5c5cc2876a4c
The last line is a class with identity="demo/1" and the fields
when: date, mode: Color (members RED and BLUE),
tags: Annotated[list[Annotated[str, Capacity(4)]], Capacity(2)] and
note: int | None.
Versioning rules¶
- A reader refuses a
layout_majorit does not know: this version opens majors 1 and 2. - A reader opens a segment whose
layout_minoris higher than its own, and ignores what it does not know. A handle reports the lower of the segment's minor and its own. - A minor version may only add: fields in bytes that are reserved and zero in older minors, or features that stay off unless the segment says they are on and the reader knows them.
- That includes field kinds:
handle::openandhandle::from_capsuleaccept a field whose kind code they do not know. Its bytes are otherwise opaque, and the field must lie inside the record without overlapping another. A reader opens and reads a field of a kind it does not know; it never writes one, andwritereturnsstatus::rangefor it. The Python extension, which converts every field, refuses such a segment withSchemaMismatchError. - For an unknown kind, the entry's low 24 bits are the field's whole span
in the record (1 byte to 1 MiB), so a future kind with a length prefix
counts the prefix in it, and no alignment is asked. In a layout 2.0
segment this holds for kinds below 64 only: for an unknown kind from 64
on, the low 24 bits are the offset of its description, whose head's
sizeis the span, and no alignment is asked either. - Layout 2.0 adds the description table and kinds 6 to 12 and 64 to 74. A later 2.x minor version may add kinds; a reader opens a field of a kind it does not know as opaque bytes.
- A change to how existing bytes are read or written (the sequence lock,
field encoding, the slot layout, object names) raises
layout_major. - The package takes a semver major step whenever
layout_majorchanges; before 1.0, a minor version step does.
Protocols¶
Every shared word is a plain integer in the mapping, accessed through
std::atomic_ref with the orderings given here.
Create¶
- Create the mapping (see Names) at its full size, rounded up to 4096.
On Linux the size is reserved with
posix_fallocate, retried onEINTR, so a/dev/shmtoo small for the box fails here with an OS error instead of a laterSIGBUS;ftruncateis the fallback where the file system does not support it. New pages read as zero on both platforms. - Write the field table, the creator fields and the rest of line 0
except
magic.create_idcomes from the OS random source (getrandomon Linux, falling back to/dev/urandomwhere it is missing or refused;BCryptGenRandomon Windows) and is drawn again if 0. The creator fields are this process's pid, start time and pid namespace. Every waiter slot is free because the pages are zero. - Write the initial field values into the record.
- Store
magicwith release ordering. Until then no attach succeeds.handle::createdoes this at once.handle::create_unpublishedstops before it, so the creator can write through its handle first (the Python layer runs__post_init__there), andhandle::publish()storesmagiclater, with a compare-exchange from 0; it returnsstatus::rangewhenmagicis already stored, which is always the case for a handle made byopenorfrom_capsule. - On a failure after step 1, remove the name (Linux) and close.
Attach¶
- Open the mapping by name and map all of it. Linux:
fstatgives the size, and a mapping smaller than one page is waited for, since its creator sizes it right after making the name. Windows:VirtualQueryon the view gives the size. - Wait for
magicwith acquire ordering, with backoff (spin, yield, then sleeps doubling up to 1 ms), up to the timeout. A mapping that never getsmagicisstatus::not_found. - Check
layout_major(Versioning rules). - Check every geometry field against the mapping size before reading
through it:
field_countandwaiter_slotsin range,tail == 128, the record 64-byte aligned, after the waiter slots and inside the mapping,sizeequal to the mapping size, and each field table entry (capacity 1 byte to 1 MiB, and the sizes of Kind codes for the fixed kinds, alignment for the known kinds, the field inside the record, no two fields overlapping), then the description table (Descriptions). In layout 2.0record + record_sizeis at most 2^32 - 4096. A field of an unknown kind is opaque (Versioning rules). A failed check isstatus::corrupt. - Copy the field table and use only the copy afterwards.
- Free the waiter slots of dead processes (see Waiter slots).
handle::open does not compare schema hashes: the caller compares
schema_hash() with its own.
Sequence lock¶
- Read: load
seqwith acquire; if odd, back off and retry; copy the field and load its write count; an acquire fence; loadseqagain and retry if it moved. The count is read inside the same window, so a value and its version always match. - Write: compare-and-swap
seqfrom evenstos + 1(acquire), then a release fence; storewriter_pid; write the values; increment each written field's count; unlock; wake. - Unlock: if
seqstill equalss + 1, clearwriter_pid; then compare-and-swapseqfroms + 1tos + 2(release), and do nothing if the swap fails. A failed swap means aforce_unlock()released this writer's lock while it was still running, andseqmay now belong to a later writer's lock. A plain store ofs + 2would release that later writer's lock and moveseqbackwards. The check and the clear are two operations, so aforce_unlock()and a new writer's lock that both land between them leave the new writer'swriter_pidat 0; the lock is not affected, and only a laterLockTimeoutErrornames pid 0 instead of the holder. - Generation: there is no separate counter. Every write adds exactly 2 to
seq, so the generation isseq >> 1(generation()returns it). Aforce_unlock()that storesseq + 1counts as one generation, since the dead writer may have changed data. - A reader or writer gives up after the lock timeout with
status::lock_timeout. force_unlock(): ifseqis odd, compare-and-swap it toseq + 1;writer_pidstays as it was, as a record of who held the lock.
Waiter slots¶
waiter_slotsis set at creation: themax_waitersclass keyword, default 64, 1 to 4096.- Register: first free every slot whose owner is dead (Liveness). Claim a
slot with a compare-and-swap of
owner_pidfrom 0 to this process's pid, then incrementwaiters, clearinterrupt, storeowner_pidnsand lastowner_start(release). A slot with a nonzeroowner_starthas always been counted, so freeing it may decrementwaiters. A slot whoseowner_startis still 0 belongs to a claim in progress; if its pid is no longer running, the claimer died mid-claim and the slot is freed without a decrement. A claimer killed between its compare-and-swap and its increment therefore cannot pushwaitersbelow the true count; one killed after the increment leaves it one too high, which costs a system call on each wake and never loses one. No slot free givesstatus::no_slot. - Free a dead owner's slot: change
owner_startwith a compare-and-swap from the value read toUINT64_MAX - 1(thestart_freeingmarker), which is never a start time on Linux (clock ticks since boot) or Windows (FILETIME). Only the process whose compare-and-swap succeeds goes on, and every scan skips a slot holding the marker. It then checks thatowner_pidstill holds the pid it read; if not, it puts the oldowner_startback and leaves the slot to its new owner. Otherwise it decrementswaitersif the value it read was nonzero, storesowner_pidns = 0, thenowner_pid = 0(release), and last changesowner_startfrom the marker to 0 with a compare-and-swap, which fails harmlessly once a new claimer has stored its own. - A scan reads
owner_pid, thenowner_startandowner_pidns, thenowner_pidagain, and skips the slot if the pid changed: otherwise a slot freed and claimed between the first two reads would pair the dead pid with the new owner's start, and the compare-and-swap would mark a live slot. - Release: store
owner_start = 0(release), decrementwaiters, storeowner_pidns = 0, thenowner_pid = 0(release). A handle frees only a slot it claimed that still records its own(pid, start, pidns); a child created byforkinherits its parent's claims but leaves them to the parent. - A thread waits in a slot it holds. The Python watcher registers one slot
for its lifetime, and a one-off
Segment.wait()without a slot claims one for the call.
Known limits, each after a process is killed at one specific step:
- A freer killed while holding the marker, before it clears
owner_pid, leaves the slot unusable until the segment is created again, andwaitersone too high if it was killed before its decrement. - An owner killed partway through a release leaves a slot with no start,
which is freed without a decrement, so
waitersstays one too high if the kill came before the decrement. On Linux, a kill afterowner_pidns = 0leaves a namespace of 0, which is never freed (see Liveness); the count stays correct. - A claimer killed on Linux before it stores
owner_pidnsleaves the slot stuck the same way, and counted once too many if it had already incrementedwaiters. - The check that
owner_pidstill holds the pid cannot tell an unstamped claimer from another one with the same pid, which needs the pid to be reused within a few instructions.
A count that is too high costs a system call on each write; it never loses a wake-up.
Wait and wake¶
- Wait (slot
i, last seen generationg, timeout): - Load
wake_word, then return at once ifseq >> 1 != gor sloti'sinterruptis set (clearing it with a compare-and-swap). - Linux:
FUTEX_WAIT(shared, not private) onwake_wordwith the value loaded before the check, so a write in between makes the call return at once. - Windows: wait on event
i, which the waiter's process opens on first use and keeps. - Check again after waking; a wake may be spurious. Past the timeout the
result is
status::timeout. - Wake, after every write:
- Increment
wake_wordwith sequentially consistent ordering, so the following load ofwaiterscannot move before it or before the swap ofseq; a waiter that registered just before would otherwise be missed. - If
waitersis 0, stop: no system call on the normal path. - Linux: one
FUTEX_WAKEwithINT_MAXwaiters. - Windows:
SetEventon the event of every occupied slot. interrupt(slot): set the slot'sinterrupt, then wake that slot. On Linux it incrementswake_wordand callsFUTEX_WAKE, and the other waiters check their own flag and sleep again; on Windows it sets that slot's event only. The flag stays set until the waiter sees it, so an interrupt sent before the wait starts still ends it.- The Python watcher waits in steps of at most 1 s. Writes and
interruptwake it at once; after each step it checks that its slot still records its own(pid, start, pidns)and claims a new slot if not. While every slot is taken it checks for changes once a second instead of waiting.
Liveness¶
A process is named by its pid and its start time.
- Start time: Windows, the creation time from
GetProcessTimes; Linux, field 22 of/proc/<pid>/stat, parsed after the last). A real value of 0 counts as 1. - Dead: no process has the pid; or one does and its start time differs
(the pid was reused); or, on Linux, its state is
ZorX. On Windows a process whose handle is signalled has exited. - Alive: the process exists but its start time cannot be read (another
user,
/procmounted withhidepid, access denied on Windows), or the recorded start time was the unreadable marker. A slot is never freed on a guess. - Pid namespaces (Linux): a slot records its owner's pid namespace, the
inode number of
/proc/self/ns/pid. A slot whoseowner_pidnsdiffers from the checking process's own counts as alive, because its pid cannot be checked from there; only its owner, or creating the segment again, frees it. This covers containers that share/dev/shmbut not a pid namespace. A process that cannot read its own namespace records 0, which means unknown: a slot whose namespace is 0, or a checker whose own is 0, is never freed. Windows has no pid namespaces and records 0. - The current pid is cached, and a
pthread_atforkhandler resets the cache in the child afterfork; the start time and namespace are read on first use.
The Python extension uses the same rules for the creator fields: when a
create finds the name taken, SegmentExistsError says whether the creator
still runs, runs in another pid namespace, or has exited. For a box not
published yet it reads the creator fields without waiting for magic,
and a creator that still runs is reported as still creating the box.
Close and unlink¶
- Destroying a handle releases any waiter slot it holds, then unmaps its view and closes its OS handles, including the events it opened.
unlink(name): Linuxshm_unlink("/sharedbox.<name>"); Windows does nothing.
Implementation language¶
The core is C++20, header-only, in namespace sharedbox. Its names are
declared in the inline namespace sharedbox::v2, which changes when the
C++ interface changes incompatibly, so code built against headers with
different inline namespaces can be linked into one program. Names in
sharedbox::detail may change without a new inline namespace, so shared
libraries build with hidden visibility to keep their copies apart. C
programs use it through sharedbox_c.h, whose sbx_* functions are not
versioned this way and are declared with hidden visibility outside
Windows.
- The layout structs are plain standard-layout types with integer members.
- Atomics:
std::atomic_refon those members.static_asserts requirestd::atomic_ref<std::uint32_t>andstd::atomic_ref<std::uint64_t>to be always lock-free, because only lock-free atomics work on memory shared between processes; a fallback that takes a lock would keep that lock in one process.std::atomic_ref::waitandnotify_*are not used, since implementations do not promise they work across processes, andstd::atomicmembers are not used because the standard does not fix their size or layout. Withstd::atomic_refthe compiler emits the barriers each CPU needs, with one code path for every compiler. - 64-bit little-endian targets only; the header refuses a big-endian target at compile time.
- No allocation on the read and write paths. No exceptions: every function
that can fail returns
result<T>, a value or astatus. On C++20result<T>is the header's own type with the part of the interface ofstd::expected<T, status>it needs (has_value,operator bool,operator*,->,value,error,and_then,transform,or_else,value_or);value()on an error callsstd::terminate. Where the standard library hasstd::expectedwith__cpp_lib_expected >= 202211L(C++23),result<T>is an alias forstd::expected<T, status>. Every function returning it is[[nodiscard]], and the header builds with-fno-exceptions. - Every translation unit of a program must see the same
result: build them all as C++20 or all as C++23.
sharedbox.hpp¶
namespace sharedbox {
inline namespace v2 {
inline constexpr std::uint16_t layout_major = 2, layout_minor = 0;
inline constexpr std::uint16_t oldest_layout_major = 1;
// Also: the layout constants handle_version, magic, header_size, name_max, max_fields,
// max_capacity, max_waiter_slots, default_waiter_slots, record_alignment, page_size,
// max_timeout, default_lock_timeout, kind_shift, capacity_mask, max_types_size,
// max_type_depth, max_mapping_size, first_described_kind, max_date_ordinal,
// unix_epoch_ordinal, micros_per_day, max_offset_minutes, literal_none through
// literal_enum, and the kind codes kind_bool, kind_int, kind_float, kind_str,
// kind_bytes, kind_ref, kind_complex through kind_decimal and kind_enum through
// kind_array; the structs header, stored_field, waiter_slot, box_ref and
// type_head (see Layout); dl_dtype and literal_value, which type_view's dtype()
// and literal() return (see Layout 2.0 adds); result<T> and unexpected (see
// Implementation language).
enum class status : int { ok = 0, exists = -1, not_found = -2, layout = -3, schema = -4,
corrupt = -5, lock_timeout = -6, timeout = -7, no_slot = -8,
range = -9, os = -10 };
struct field_spec { std::uint32_t offset; std::uint32_t capacity; std::uint8_t kind; };
struct value { std::uint16_t field; std::span<const std::byte> bytes; };
using seconds = std::chrono::duration<double>;
struct read_value { std::size_t len; std::uint64_t version; };
enum class wake { changed, interrupted }; // a timeout is status::timeout
class handle { // move-only; the destructor releases it
public:
static result<handle> create(std::string_view name, std::span<const field_spec> fields,
std::uint32_t record_size, std::uint64_t schema_hash,
std::uint16_t waiter_slots, std::span<const value> initial,
std::span<const std::byte> types = {});
static result<handle> create_unpublished(std::string_view name, std::span<const field_spec> fields,
std::uint32_t record_size, std::uint64_t schema_hash,
std::uint16_t waiter_slots, std::span<const value> initial,
std::span<const std::byte> types = {});
result<void> publish() noexcept;
static result<handle> open(std::string_view name, seconds timeout);
static result<handle> from_capsule(sbx_handle *capsule);
result<handle> duplicate() const;
sbx_handle *to_capsule() &&;
std::string_view name() const noexcept;
std::uint16_t field_count() const noexcept;
const field_spec &field(std::uint16_t index) const noexcept;
type_view field_type(std::uint16_t index) const noexcept;
std::span<const std::byte> types_table() const noexcept;
std::uint32_t record_size() const noexcept;
std::uint64_t schema_hash() const noexcept;
std::uint64_t create_id() const noexcept;
std::uint16_t waiter_slots() const noexcept;
std::uint16_t minor_version() const noexcept;
std::uint16_t major_version() const noexcept;
void *base() const noexcept;
std::uint64_t size() const noexcept;
result<void> set_lock_timeout(seconds timeout) noexcept;
result<read_value> read(std::uint16_t field, std::span<std::byte> buf) const;
result<std::uint64_t> read_record(std::span<std::byte> buf) const;
result<read_value> read_used(std::uint16_t field, std::span<std::byte> buf) const;
result<read_value> read_large(std::uint16_t field, std::span<std::byte> buf) const;
result<std::uint64_t> read_record_large(std::span<std::byte> buf) const;
result<void> write_large(std::span<const value> values, seconds lock_timeout);
std::span<const std::byte> payload(std::uint16_t field, std::span<const std::byte> record) const noexcept;
result<void> write(std::span<const value> values, seconds lock_timeout);
std::uint64_t generation() const noexcept;
std::uint64_t version(std::uint16_t field) const noexcept;
std::uint32_t writer_pid() const noexcept;
result<void> force_unlock() noexcept;
result<std::uint16_t> register_waiter();
void release_waiter(std::uint16_t slot) noexcept;
bool waiter_held(std::uint16_t slot) const noexcept;
std::uint32_t waiters() const noexcept;
result<wake> wait(std::uint16_t slot, std::uint64_t last_generation, seconds timeout);
result<void> interrupt(std::uint16_t slot);
};
result<void> unlink(std::string_view name) noexcept;
result<header> inspect(std::string_view name) noexcept;
} // namespace v2
} // namespace sharedbox
Every function returning result is [[nodiscard]]. handle also has
lock, unlock and set_wait_hooks, which the tests and the Python
extension use.
Layout 2.0 adds these names to the header:
type_viewandhandle::field_type(index): a field's type, read from its description.handle::types_table()returns the table's bytes.time_value,datetime_valueandtimedelta_value, with anencode_*and adecode_*function for each, and the same pair forcomplex,date,uuid,flagand enum and literal positions (encode_position,decode_position).decode_bool,decode_present,decode_taganddecode_lengthhave no encoder.handle::major_version()andoldest_layout_major, the lowest major versionopenandfrom_capsuleaccept.- The constants
first_described_kind,max_date_ordinal,unix_epoch_ordinal,micros_per_day,max_offset_minutesandliteral_nonethroughliteral_enum, and the structstype_head,dl_dtypeandliteral_value. handle::read_used, which copies only the used part of a list, set or dict field;read_large,read_record_largeandwrite_large, which copy a large value, running the wait hooks once around the copy.- A
typesparameter onhandle::createandhandle::create_unpublished: the description table, empty by default. -
The inline namespace is
v2. -
status:ok, and one code per error the Python side raises, so the extension maps each to its exception class.exists: the name is taken.not_found: no mapping under the name, or one that never becomes a box within the timeout.layout: anotherlayout_major.schemais kept for callers that compare schema hashes.corrupt: a header or field table that fails the attach checks.lock_timeout,timeout,no_slot: see Protocols.range: an argument out of range.os: an OS call failed, witherrnoorGetLastError()left as the call set it. createandwritetake raw bytes in the record encoding, without the length prefix ofstr,bytesandDecimal; converting language values stays in each binding.opentakes no schema hash; callers compareschema_hash()with their own. Itstimeout, in(0, 86400], only bounds the wait for a creator (the Python extension passes at most 1 s). Reads use a lock timeout of 5 s untilset_lock_timeoutchanges it;writetakes its own.readreturnsread_value. When the stored value is longer thanbuf, nothing is copied andlensays how long it is; a buffer of the field's capacity always fits.read_recordcopies every field from one moment and returns the generation of that moment;payloadfinds a field in such a copy.inspect(name)copies a published header without waiting and without the checks ofopen. The Python extension uses it to name the layout of a segment it refuses and the creator inSegmentExistsError.duplicate()makes a second handle with its own mapping from this handle's OS handle (dupandmmapon Linux,DuplicateHandleandMapViewOfFileon Windows), not from the name, so it works afterunlink().minor_version()is the lower of the segment'slayout_minorand thelayout_minorof thesharedbox.hppthe caller was compiled with.- Shared memory that never becomes a box within the timeout is
status::not_found. - Thread safety: every member function may be called from several threads on one handle. Destroying or moving a handle must not overlap another call on it.
- Known limit: on Linux every open handle keeps the mapping's file
descriptor, so about 1000 open boxes reach the default
ulimit -nof 1024.
Capsule handle¶
sbx_handle is the struct a capsule holds. It is plain C data with function
pointers, declared in sharedbox_c.h and included by sharedbox.hpp,
because it passes between extensions compiled separately, possibly with
different compilers and standard libraries. No C++ standard-library type
crosses the capsule.
typedef struct sbx_handle {
uint16_t layout_major, layout_minor;
uint32_t handle_version; /* 1; grows by appending fields */
void *base; /* this handle's own mapping */
uint64_t size;
const char *name; /* the box name, for slot events */
void (*release)(struct sbx_handle *);
void *private_data; /* owned by the build that made the handle */
} sbx_handle;
private_databelongs to the build that made the handle, and only that build'sreleasereads it.- A consumer never reads another build's
private_data. It takes the handle withhandle::from_capsule(orsbx_import). It does not wait formagic: it loads it once and returnsstatus::corruptif it is not set orsizeis below 4096, then runs steps 3 to 5 of Attach on the segment atbase. It keeps its own copy of the struct and of the field table, and checks no schema: the consumer comparesschema_hash(). On success it clears the capsule struct'srelease, and destroying the new handle calls the producer'srelease. The consumer's handle opens slot events bynamewhen it first waits. to_capsule()moves a handle to the heap and returns a newsbx_handlewhosereleasedestroys it; the caller deletes the struct afterwards.- The Python extension's capsule destructor deletes the struct while the
capsule is named
"sharedbox_box"or"used_sharedbox_box", and callsreleaseonly under the first name.
sharedbox_c.h¶
include/sharedbox/sharedbox_c.h is a minimal C99 interface, documented as
"minimal, may be removed in a future major version". Removing it would only
delete code: a semver major step for the headers, with no change to the
layout.
#define SBX_OK 0 /* SBX_E_* have the values of sharedbox::status */
typedef struct sbx_value { uint16_t field; const void *data; size_t len; } sbx_value;
int sbx_open(const char *name, double timeout, sbx_handle *out);
int sbx_import(sbx_handle *capsule, sbx_handle *out);
int sbx_read(const sbx_handle *h, uint16_t field, void *buf, size_t cap,
size_t *len, uint64_t *version);
int sbx_write(sbx_handle *h, const sbx_value *values, size_t n, double lock_timeout);
uint64_t sbx_schema_hash(const sbx_handle *h);
void sbx_release(sbx_handle *h);
- It holds
sbx_handle,sbx_value, the status codes, the six functions above and the typed functions below. Creating, waiting,interrupt,force_unlockandunlinkare in the C++ API only. sbx_field_descreturns a field's kind code and its description's bytes.sbx_read_*andsbx_write_*exist forcomplex,date,time,datetime,timedelta,uuid, enum and literal positions (position), flag bits (flag) and an optional's presence (present, read only). A read givesSBX_E_RANGEfor a field of another kind andSBX_E_CORRUPTfor a stored value no Python value has; a write givesSBX_E_RANGEfor a value out of range.sbx_timeandsbx_datetimeholdint64_t micros,int16_t offset_minutesanduint8_t naive, fold;sbx_timedeltaholdsint32_t days, seconds, microseconds. Asbx_timecounts microseconds since midnight in wall-clock time; asbx_datetimecounts microseconds since 1970-01-01T00:00, wall time when naive and UTC when aware.- Each function converts its arguments and calls
sharedbox.hpp.sbx_openfillsoutwith a handle whoseprivate_dataholds asharedbox::handleand whosereleasedeletes it.sbx_importwraps a capsule's handle throughhandle::from_capsule. sbx_read,sbx_writeandsbx_schema_hashaccept only handles made bysbx_openorsbx_import; for any other,sbx_readandsbx_writereturnSBX_E_RANGEandsbx_schema_hashreturns 0.sbx_writealso returnsSBX_E_RANGEfor a field of a kind the header does not know. When the stored value is longer thancap,sbx_readcopies nothing, sets*lenand returnsSBX_E_RANGE.sbx_releasereleases any handle.- The implementation,
include/sharedbox/sharedbox_c.cpp, needs a C++20 compiler: a C program compiles that one file with its C++ compiler next to its own sources, and the CMake targetsharedbox::cadds it for them. One source file rather than a static library in the wheel, because a static library is tied to the compiler and C runtime that built it.
The PyCapsule interface¶
class Frame(SharedBox, identity="camera/frame/1", max_waiters=64):
exposure: float
count: int
with Frame(0.01, 0) as frame:
camera.run(frame) # an extension that supports sharedbox
The dunder is a protocol for library authors, as __arrow_c_array__ and
__dlpack__ are: Python users pass the box to a consuming library, which
calls it. A capsule is an opaque pointer Python code cannot use, so there is
no Python-level export method.
SharedBox.__sharedbox_box__(self, max_version=None, **kwargs)returns a PyCapsule named"sharedbox_box"holding a heap-allocatedsbx_handlemade byduplicate()andto_capsule(), so it works afterunlink().max_versionis(major, minor); a major the box does not have raisesBufferError, andNonemeans the current version. Any other keyword raisesNotImplementedError.- A consumer that takes the handle renames the capsule
"used_sharedbox_box"and releases the handle itself. releaseunmaps and closes OS handles only: no Python objects, no GIL, so it may run on any thread and after interpreter shutdown.box.close()and the handle never affect each other. On Windows the segment stays alive while a handle is held.sharedbox.get_include()returns the folder holdingsharedbox/sharedbox.hpp,sharedbox/sharedbox_c.handsharedbox/sharedbox_c.cppin the installed package.sharedbox.SupportsSharedBoxis atyping.Protocolwith__sharedbox_box__, for annotations in consuming libraries.sharedbox._native.LAYOUT_VERSIONis(1, 0).
Rust crate (deferred)¶
A Rust crate is not part of this release. It will be a safe layer over
sharedbox.hpp through cxx, so Python, C++ and Rust share one
implementation of the protocols, with schema_hash and the default name
computed in Rust and checked against the test vector above.
Build and packaging¶
include/sharedbox/sharedbox.hpp,sharedbox_c.handsharedbox_c.cppare in the repository; the build installs them into the wheel undersharedbox/include/, andcmake/sharedbox-config.cmakeundersharedbox/share/cmake/sharedbox/, next to asharedbox-config-version.cmakethe build writes from the package version. Before 1.0 it accepts a request for the same minor version only (SameMinorVersion), from 1.0 on the same major version.- CMake targets:
sharedbox::headers, an interface target that requirescxx_std_20(and linksrtandThreads::Threadson Linux,bcrypton Windows), andsharedbox::c, which linkssharedbox::headersand addssharedbox_c.cppto the sources of whatever links it. Both work throughFetchContentand throughfind_package(sharedbox). A project that linkssharedbox::cwithout the CXX language enabled stops at configure time with a message saying so. - A Python extension that uses the header adds
sharedboxto its[build-system] requires, callssharedbox.get_include()and builds as C++20. The sharedbox extension itself builds as C++20.
Out of scope¶
- Queues and streams; they would get
__sharedbox_queue__and__sharedbox_stream__. Global\names on Windows. They would let a service in session 0 share a box with a desktop application, but creating a file mapping there needsSeCreateGlobalPrivilegeand a security descriptor that lets the other session's user open it.- A vcpkg port and a headers-only wheel.
- macOS. Every platform detail lives in
sharedbox.hpp, so macOS would be a third platform there. Known obstacles: POSIX shared memory names are limited to 31 characters, whichsharedbox.<name>with 128-character box names does not fit; there is no futex; there is noposix_fallocate; leftover segments cannot be listed. - A read-only open mode.
- Big-endian hosts.
References¶
- Arrow PyCapsule Interface: https://arrow.apache.org/docs/format/CDataInterface/PyCapsuleInterface.html
- DLPack Python specification (
__dlpack__,max_version, capsule renaming): https://dmlc.github.io/dlpack/latest/python_spec.html - Ulrich Drepper, "Futexes Are Tricky": https://www.akkadia.org/drepper/futex.pdf
- psutil, process start time as identity: https://github.com/giampaolo/psutil
- Microsoft, kernel object namespaces (
Local\,Global\): https://learn.microsoft.com/en-us/windows/win32/termserv/kernel-object-namespaces