Shared Atomic
SharedAtomicU32 / SharedAtomicU64 / SharedAtomicBool
Cross-process atomic counter / flag / bitset cell. The MMF payload
is interpreted directly as AtomicU32, AtomicU64, or AtomicBool;
atomic operations are sound across process boundaries because
hardware cache coherence guarantees the semantics when both
processes map the same physical page (which the OS guarantees for
the same MMF file).
The “atomic-but-shared-across-processes” primitive. Three concrete types ship:
SharedAtomicU32,SharedAtomicU64,SharedAtomicBool. Each is a thin wrapper around the matchingstd::sync::atomic::*type, with the atomic’s storage living in a memory-mapped file. Multiple processes opening the same path see the same counter value via direct atomic ops; no IPC round-trip.
Constraints (read first):
Native sidecar integration: the struct carries a
HandshakeHeader+ObservationRingand implementssubetha_sidecar::AdaptiveInstance. Wrap inSidecarBox::newto register with the global sidecar; rawcreate()/open()return the unregistered type unchanged.MMF-backed: requires a writable backing file path. The file is
ATOMIC_FILE_SIZEbytes (the size of one cache-line-alignedAtomicHeader).Header layout (source lines 27-32):
magic: u32 = 0x4150_5443width: u32+payload_u64: AtomicU64(the payload field is re-cast to the requested atomic width on access).
Three width values (source line 163-164, 184): width=4 →
SharedAtomicU32; width=8 →SharedAtomicU64; width=1 →SharedAtomicBool.Open rejects mismatched width (source lines 83-85, 201-203): opening a u64 file via
SharedAtomicU32::openreturnsSharedAtomicError::LayoutMismatch.Hardware atomic semantics across processes: x86_64 / ARM / RISC-V all guarantee cache-coherent atomic ops on shared memory. No software lock; no IPC overhead.
All standard atomic ops shipped:
load,store,fetch_add,fetch_sub,fetch_or,fetch_and,fetch_xor,swap,compare_exchange(source lines 99-144 for U32/U64).SharedAtomicBoolshipsload,store,swap(source lines 214-216) - no fetch_add semantics for bool.flush/flush_asyncfor durability:flushcallsmsync(POSIX) orFlushViewOfFile + FlushFileBuffers(Windows).flush_asyncskips the file-buffers sync on Windows (page-cache only).Send + Sync unconditional (source lines 53-54, 171-172). Multiple threads in the same process can share an
Arc<SharedAtomicU64>; multiple processes share via the file path.No on-disk persistence guarantee until flush: writes hit the page cache via the atomic op; durable persistence requires
flush.Cross-process backed by MMF.
Table of contents
- What it is
- Three concrete types
- Header layout
- Hardware atomics across processes
- Worked examples
- Bench evidence
- Use case patterns
- Known limitations
- Common pitfalls
- References
What it is
Three concrete types in one source file, each generated by the
shared_atomic_impl! macro (source lines 46-161):
pub struct SharedAtomicU32 { _file: File, mmap: MmapMut }
pub struct SharedAtomicU64 { _file: File, mmap: MmapMut }
pub struct SharedAtomicBool { _file: File, mmap: MmapMut }Each owns a MmapMut over the MMF backing file. The header at
offset 0 carries the magic + width discriminant; the atomic payload
starts at offset_of!(AtomicHeader, payload_u64). The atomic()
accessor casts the payload pointer to the requested
Atomic{U32,U64,Bool} type.
graph LR
F[MMF file]
H[AtomicHeader<br/>magic + width + AtomicU64 payload]
U32[SharedAtomicU32<br/>cast payload as AtomicU32]
U64[SharedAtomicU64<br/>cast payload as AtomicU64]
B[SharedAtomicBool<br/>cast payload as AtomicBool]
F --> H
H --> U32
H --> U64
H --> B
classDef file fill:#1e3a5f,stroke:#5b9bd5,color:#e8f1f5
classDef header fill:#1f4a3a,stroke:#5cb85c,color:#e8f5e8
classDef atomic fill:#5a3a1f,stroke:#d5a05b,color:#f5ede0
class F file
class H header
class U32,U64,B atomic
Three concrete types
| Type | Width | Ops shipped |
|---|---|---|
SharedAtomicU32 | 4 bytes | load, store, fetch_add, fetch_sub, fetch_or, fetch_and, fetch_xor, swap, compare_exchange |
SharedAtomicU64 | 8 bytes | same as U32 |
SharedAtomicBool | 1 byte (in 8-byte payload slot) | load, store, swap |
SharedAtomicBool does not ship fetch_add / fetch_or / etc.
because boolean semantics do not map cleanly to those ops.
Header layout
#[repr(C, align(64))]
struct AtomicHeader {
magic: u32, // 0x4150_5443
width: u32, // 1, 4, or 8
payload_u64: AtomicU64, // re-cast on access
}- 64-byte cache-line aligned to keep the atomic in its own cache line (no false sharing with neighboring file content).
magicvalidates the file is actually a SharedAtomic;openrejects mismatches.widthdiscriminates the atomic type;openrejects width mismatches (e.g., opening a U64 file as U32 returnsLayoutMismatch).
Hardware atomics across processes
The atomic ops on the MMF-backed payload are sound cross-process because:
- The OS maps the same physical page into each process that opens the file.
- CPU cache coherence (MESI / MOESI) operates at the physical-page level, not at the virtual-address level.
- The
lock-prefixed atomic instructions on x86 (LDAR/STLRon ARM) flush the relevant cache line across all cores regardless of which process holds it.
No software synchronization is needed. The fetch_add from process A
is visible to the load in process B as soon as the cache coherence
protocol completes (sub-microsecond on shared L3).
Worked examples
Cross-process counter
use std::sync::atomic::Ordering;
use subetha_cxc::shared_atomic::SharedAtomicU64;
// Process A:
let counter = SharedAtomicU64::create("/tmp/counter.bin", 0).unwrap();
for _ in 0..100 {
counter.fetch_add(1, Ordering::AcqRel);
}
// Process B (or thread B):
let reader = SharedAtomicU64::open("/tmp/counter.bin").unwrap();
let value = reader.load(Ordering::Acquire);
// value is between 0 and 100 depending on when B reads.
Cross-process leader flag
use std::sync::atomic::Ordering;
use subetha_cxc::shared_atomic::SharedAtomicBool;
let leader = SharedAtomicBool::create("/tmp/leader.bin", false).unwrap();
// Atomic test-and-set: only one process wins.
let was_leader = leader.swap(true, Ordering::AcqRel);
if !was_leader {
// This process became the leader.
}Compare-exchange for race-free initialization
use std::sync::atomic::Ordering;
use subetha_cxc::shared_atomic::SharedAtomicU64;
let v = SharedAtomicU64::create("/tmp/version.bin", 0).unwrap();
let init_value = 12345;
// First process to set the value wins; others observe and proceed.
match v.compare_exchange(0, init_value, Ordering::AcqRel, Ordering::Acquire) {
Ok(_) => {
// Won the race; we initialized.
}
Err(observed) => {
// Lost; observed = current value the other process set.
assert_ne!(observed, 0);
}
}Durability
use std::sync::atomic::Ordering;
use subetha_cxc::shared_atomic::SharedAtomicU64;
let v = SharedAtomicU64::create("/tmp/persistent.bin", 12345).unwrap();
v.flush().unwrap(); // sync to disk
// Crash + restart: contents survive.
drop(v);
let v2 = SharedAtomicU64::open("/tmp/persistent.bin").unwrap();
assert_eq!(v2.load(Ordering::Acquire), 12345);Bench evidence
Bench harness: crates/subetha-cxc/benches/shared_atomic.rs.
Captured 2026-06-01 on Windows 11 / Zen+ R7 2700, Criterion with
--sample-size=15 --warm-up-time=1 --measurement-time=2.
Single-thread atomic ops:
| Op | std::sync::atomic::AtomicU64 | SharedAtomicU64 | Delta |
|---|---|---|---|
load(Acquire) | 808.98 ps | 867.11 ps | +58 ps |
fetch_add(1, AcqRel) | 8.69 ns | 8.72 ns | +30 ps |
compare_exchange | 8.81 ns | 8.70 ns | -110 ps (within noise) |
The architectural claim validates: SharedAtomic per-op cost is
within noise of native std::atomic. The only measurable overhead
is the ~60 ps mmap pointer-load on the load path; atomic
operations themselves (lock cmpxchg, lock xadd) run at native
hardware speed.
Rule 3b bench audit
- Fair contender:
std::sync::atomic::AtomicU64is the textbook baseline. - Same ops on both: load, fetch_add, compare_exchange.
- Single-thread workloads: no cross-process contention measured.
- MMF lifecycle managed: bench creates the file, runs ops, deletes the file at end. No leakage across runs.
What the numbers do NOT show
- Cross-process visibility cost: in-process bench cannot measure cross-process round-trip; cache coherence completes in sub-microsecond on shared L3.
- Multi-thread contention: bench is single-threaded. Multi-writer contention on the SAME atomic incurs cache-line bounce; SharedAtomic and std::Atomic share this characteristic.
Use case patterns
Pattern: cross-process global epoch / generation counter
A long-running monolith that wants to bump an epoch counter visible to all workers / consumers. Each worker reads the epoch on each iteration; the master fetch_adds when state changes.
Pattern: leader-election flag
SharedAtomicBool::swap(true, ...) is the test-and-set primitive
for cross-process leader election. The winner sees was_leader = false;
losers see was_leader = true.
Pattern: lock-free statistics counter
Workers fetch_add their per-iteration counts to a shared counter. A monitor process opens the same file and load()s periodically to report current rate.
Pattern: signal flag for graceful shutdown
A controller sets a SharedAtomicBool to true; workers check it
each iteration. No IPC pipe / socket needed.
Known limitations
- Width is fixed at create time: a u32 file cannot be reopened
as u64. The
openAPI enforces this. - No CAS on bool:
SharedAtomicBoolshipsswapinstead ofcompare_exchange. For bool-valued state machines, preferSharedAtomicU32with 0/1 encoding. - MMF size is one cache line per atomic: there is no array form in this primitive. For arrays of cross-process counters use SHARED_VEC.md or a custom MMF layout.
- Durability requires explicit flush: writes are visible
cross-process immediately (cache coherence), but durable to disk
only after
flush. Crash-before-flush loses the latest writes. - No process-crash detection: if a process crashes mid-update (e.g., between two related atomics), the others see the partial state. Use a higher-level coordination primitive (epoch barrier, versioned chain) when atomicity across multiple cells matters.
- The single u64 payload slot is reused for all widths: writes
via the cast atomic touch only the first
widthbytes; the rest of the 8-byte slot is zeroed at create. - MMF backing required: there is no “in-process only” mode.
For single-process atomic ops, use
std::sync::atomicdirectly.
Common pitfalls
Opening a SharedAtomic file with the wrong width. The header rejects mismatches via
LayoutMismatch; check the result rather than.unwrap().Reading a SharedAtomic file that the writer has not yet initialized.
createinitializes the magic / width / value atomically (single page write before returning); a reader that opens during creation may see zeros and reject via LayoutMismatch. Retry with backoff if both processes start at the same time.Forgetting
flushbefore relying on cross-process durability. Cross-process VISIBILITY is immediate via cache coherence; cross-process DURABILITY (survives process crash) requires explicit flush.Wrapping
SharedAtomicU64in a Mutex. Pointless; the atomic is already lock-free and cross-process safe. The Mutex adds latency and serializes call sites the atomic was designed to parallelize.Treating MMF-backed atomics as faster than in-process atomics. They are not. The mmap pointer-deref + atomic-op is the same as in-process atomic-op; the win is the cross-process visibility, not raw speed.
References
- Source:
crates/subetha-cxc/src/shared_atomic.rs(438 lines, 8 unit tests covering U64 round-trip, U32 fetch_add, cross-handle visibility, concurrent fetch_add summation, compare_exchange, bool swap, disk persistence survives reopen, open rejects wrong width). - Sibling primitive: OFFSET_PTR.md - the foundational MMF-backed pointer that other shared types compose on.
- Sibling primitive: SHARED_CELL.md - the non-atomic cross-process cell. Use SharedAtomic when concurrent ops are required; use SharedCell for single-writer scenarios.