From 24542ce376d9fdff694a5be347cef2cc3ae2b139 Mon Sep 17 00:00:00 2001 From: Tim Date: Thu, 10 Sep 2026 15:28:23 -0700 Subject: [PATCH 1/3] feat(options): a reachable collect heartbeat, and name the pointer-width route boundary The Kairos RTOS reported one alloc+free stepping 16% across 512 bytes on a 32-bit ESP32-S3, absent on a 64-bit host, and concluded it points at `prim::fixed`. It does not. `SMALL_SIZE_MAX` is `128 * size_of::()`: 1,024 on their host, 512 on the device, where it coincides with `SMALL_OBJ_SIZE_MAX`. Their refutation of that constant was run on a machine where it sits at 1,024. Moving only the pointer width, with `prim::windows` in both arms (ABBA, 50k ops/arm, min of 40 blocks): x86-64 +0.1% at 512 vs 513, i686 +8.6%, control under 1% on both. Mechanism, counted rather than timed: per 100,000 pairs with one block live the `direct[]` route carves/extends/retires 195 pages and the bin route 0, while `generic`/op is 1.0000 on both -- so the slow path is not the difference. 195 is 100,000/512, the GENERIC_COLLECT sweep at the small profile. That period is a real trade (churn against capacity decay) and bare metal could not move it: `options::set` is a no-op under ONE_REGION and the option's doc told firmwares to change a built-in default that had no cfg. Adds `--cfg ra_generic_collect`; the default does not move, because 10,000 starved this profile once already. Also adds `prim::fixed::shape_of(size) -> Shape` (page bytes, dedicated segments, direct_route) so the route and page kind are a const answer instead of a silicon discovery -- the observability the report asked for, which `region_stats()` cannot give because it reports over extents. Not established, and withdrawn rather than quoted: how much of the 16% the churn accounts for. Raising the heartbeat took churn to a measured zero, but that run's timing control flipped +8.0% -> +0.6%, so the instrument was deciding it. Plan closed: docs/plans/finished/fixed-prim-small-step.md section 8. Co-Authored-By: Claude Opus 5 (1M context) --- Cargo.toml | 2 +- crates/rusty_alloc/CHANGELOG.md | 19 ++ crates/rusty_alloc/src/options.rs | 35 +- crates/rusty_alloc/src/prim/fixed.rs | 91 +++++ docs/LEDGER.md | 37 ++ docs/plans/finished/fixed-prim-small-step.md | 340 +++++++++++++++++++ 6 files changed, 521 insertions(+), 3 deletions(-) create mode 100644 docs/plans/finished/fixed-prim-small-step.md diff --git a/Cargo.toml b/Cargo.toml index dfb16f4..bcdc5ca 100644 --- a/Cargo.toml +++ b/Cargo.toml @@ -155,7 +155,7 @@ unsafe_op_in_unsafe_fn = "deny" # unify across the dependency graph, so two consumers wanting different # geometries would silently get one of them. A cfg is set by the DELIVERABLE, # the same way a Janus firmware picks its chip. -unexpected_cfgs = { level = "warn", check-cfg = ["cfg(loom)", "cfg(kani)", "cfg(ra_small_profile)", "cfg(ra_single_threaded)", "cfg(ra_max_extents, values(\"8\", \"16\", \"64\"))", "cfg(ra_aligned_region)", "cfg(ra_segment_size, values(\"256k\"))"] } +unexpected_cfgs = { level = "warn", check-cfg = ["cfg(loom)", "cfg(kani)", "cfg(ra_small_profile)", "cfg(ra_single_threaded)", "cfg(ra_max_extents, values(\"8\", \"16\", \"64\"))", "cfg(ra_aligned_region)", "cfg(ra_segment_size, values(\"256k\"))", "cfg(ra_generic_collect, values(\"64\", \"4096\", \"65536\"))"] } # Lint policy (hardening gate H-15). `pedantic` and `nursery` are ENABLED at # workspace level and the build is clean under them, because every group diff --git a/crates/rusty_alloc/CHANGELOG.md b/crates/rusty_alloc/CHANGELOG.md index fa13f20..804c651 100644 --- a/crates/rusty_alloc/CHANGELOG.md +++ b/crates/rusty_alloc/CHANGELOG.md @@ -7,6 +7,25 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ## [Unreleased] +### Added + +- **`--cfg ra_generic_collect="64" | "4096" | "65536"`**, so a bare-metal + firmware can move the periodic-collect heartbeat. It is the only lever over a + real trade — a sweep returns an empty page, and the next allocation of that + class carves and extends a fresh one, so a short period costs page churn + while a long one costs capacity — and it was unreachable: `options::set` is a + no-op under `ONE_REGION`, and the option's own doc told firmwares to "change + the default it is built with" when no cfg existed. The default does not move. +- **`prim::fixed::shape_of(size) -> Shape`** — page bytes, dedicated segments + and `direct_route`, `const` and derived from the active geometry. Answers + "which page kind, and how many region bytes, does this size cost", which + `region_stats()` cannot because it reports over region extents. + **`Shape::direct_route` names a boundary that moves with POINTER WIDTH**: + `SMALL_SIZE_MAX` is `128 * size_of::()`, so it is 1,024 on a 64-bit + host and 512 on a 32-bit chip. The Kairos RTOS measured a 16 % step there on + a device and could not reproduce it on a workstation for exactly that reason + (`docs/plans/finished/fixed-prim-small-step.md`). + ## [2.1.0](https://github.com/Remade-With-Rust/rusty_alloc/compare/rusty_alloc-v2.0.5...rusty_alloc-v2.1.0) - 2026-09-10 ### Added diff --git a/crates/rusty_alloc/src/options.rs b/crates/rusty_alloc/src/options.rs index 06dbdb7..58191d5 100644 --- a/crates/rusty_alloc/src/options.rs +++ b/crates/rusty_alloc/src/options.rs @@ -194,12 +194,43 @@ const _: () = assert!(GENERIC_COLLECT < OPTION_COUNT); /// a short sweep period buys the same ageing without touching that path. /// Upstream's per-page countdown stays unimplemented and is recorded in /// `docs/plans/small-metal.md` §6. +/// **The period is a TRADE, and on bare metal it is the only lever over it.** +/// A sweep returns an empty page to its segment; the next allocation of that +/// class then carves and extends a fresh one. In a tight alloc/free loop that +/// keeps one block live, the page IS empty at every sweep, so a short period +/// turns into steady page churn — measured on a host at exactly +/// `100_000 / 512 = 195` carve-and-retire cycles per 100,000 allocations on +/// the `direct[]` route (`size <= SMALL_SIZE_MAX`), against **zero** on the +/// bin route just above it. That is the visible half of the 512-byte step the +/// Kairos RTOS reported (`docs/plans/finished/fixed-prim-small-step.md`); +/// raising this option takes the churn to zero. +/// +/// Which way to err is workload-dependent and neither answer is free. Short +/// sweeps cost page churn; long ones cost capacity, which is the defect the +/// 512 exists to prevent. **Do not raise it because a benchmark that keeps one +/// block live got faster** — that workload cannot decay. +/// /// Shipped geometry: upstream's 10,000. #[cfg(not(ra_small_profile))] pub const GENERIC_COLLECT_DEFAULT: i64 = 10_000; -/// Small profile: 512, for the reasons above. +/// Small profile: 512 by default, `--cfg ra_generic_collect="N"` to move it. +/// +/// **The cfg exists because this doc used to tell a firmware to "change the +/// default it is built with" and there was no way to.** `options::set` is a +/// no-op under `crate::ONE_REGION` — options are compile-time constants there, +/// which is what lets the whole option table leave a firmware image — so a +/// bare-metal consumer could neither set it at run time nor override the +/// built-in. Now it can, and the value is still a constant the linker folds. #[cfg(ra_small_profile)] -pub const GENERIC_COLLECT_DEFAULT: i64 = 512; +pub const GENERIC_COLLECT_DEFAULT: i64 = if cfg!(ra_generic_collect = "64") { + 64 +} else if cfg!(ra_generic_collect = "4096") { + 4096 +} else if cfg!(ra_generic_collect = "65536") { + 65_536 +} else { + 512 +}; /// Option names in ABI index order (also the env-var suffixes, uppercased). pub const OPTION_NAMES: [&str; OPTION_COUNT] = [ diff --git a/crates/rusty_alloc/src/prim/fixed.rs b/crates/rusty_alloc/src/prim/fixed.rs index 9ddb816..ff26ddb 100644 --- a/crates/rusty_alloc/src/prim/fixed.rs +++ b/crates/rusty_alloc/src/prim/fixed.rs @@ -402,6 +402,64 @@ pub const fn region_for(usable: usize) -> usize { segments * seg } +/// What one allocation of a given size costs, and which path serves it — +/// the compile-time answer to "why did that size behave differently?". +/// +/// Every field is derived from the active geometry, so it moves with +/// `--cfg ra_small_profile` and `--cfg ra_segment_size` instead of being a +/// second copy of the routing rules. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +#[non_exhaustive] +pub struct Shape { + /// Bytes of page the request is served from: a small page, a medium page, + /// its own span, or its own segments. + pub page_bytes: usize, + /// Whole segments this allocation reserves for ITSELF; `0` when it shares + /// (see [`dedicated_segments`]). + pub dedicated_segments: usize, + /// Whether the request takes the `direct[]` fast-path route, i.e. + /// `size <= SMALL_SIZE_MAX`. + /// + /// **This is the boundary that moves with POINTER WIDTH**, not with the + /// profile: `SMALL_SIZE_MAX` is `128 * size_of::()`, so it is 1,024 + /// on a 64-bit host and **512 on a 32-bit chip**. A sweep that looks flat + /// on a workstation can step on the device for that reason alone, which is + /// exactly what the Kairos RTOS measured + /// (`docs/plans/finished/fixed-prim-small-step.md`). + pub direct_route: bool, +} + +/// The [`Shape`] of an allocation of `size` bytes. +/// +/// ```ignore +/// use rusty_alloc::prim::fixed::shape_of; +/// // On a 32-bit target this steps at 512; on a 64-bit one, at 1024. +/// const _: () = assert!(shape_of(512).direct_route); +/// ``` +#[must_use] +pub const fn shape_of(size: usize) -> Shape { + use crate::types::{ + LARGE_OBJ_SIZE_MAX, MEDIUM_OBJ_SIZE_MAX, MEDIUM_PAGE_SIZE, SEGMENT_SLICE_SIZE, + SMALL_OBJ_SIZE_MAX, SMALL_PAGE_SIZE, SMALL_SIZE_MAX, + }; + let dedicated = dedicated_segments(size); + let page_bytes = if size <= SMALL_OBJ_SIZE_MAX { + SMALL_PAGE_SIZE + } else if size <= MEDIUM_OBJ_SIZE_MAX { + MEDIUM_PAGE_SIZE + } else if size <= LARGE_OBJ_SIZE_MAX { + // Its own span of whole slices inside a shared segment. + size.div_ceil(SEGMENT_SLICE_SIZE) * SEGMENT_SLICE_SIZE + } else { + dedicated * crate::types::SEGMENT_SIZE + }; + Shape { + page_bytes, + dedicated_segments: dedicated, + direct_route: size <= SMALL_SIZE_MAX, + } +} + /// The largest allocation that can SHARE a segment with other allocations. /// /// At or below this, a request is carved as a span of slices inside a segment @@ -1979,6 +2037,39 @@ mod tests { } } + /// The Kairos RTOS measured a step at 512 on a 32-bit device and could not + /// reproduce it on a 64-bit host. The reason is here, as an assertion + /// rather than as prose: the `direct[]` route's top is a function of + /// POINTER WIDTH, so it lands on a different size on the two machines and + /// a host sweep cannot see the device's boundary + /// (`docs/plans/finished/fixed-prim-small-step.md`). + #[test] + fn the_direct_route_boundary_moves_with_pointer_width() { + use crate::types::{SMALL_OBJ_SIZE_MAX, SMALL_SIZE_MAX, SMALL_WSIZE_MAX}; + assert_eq!( + SMALL_SIZE_MAX, + SMALL_WSIZE_MAX * core::mem::size_of::() + ); + assert!(shape_of(SMALL_SIZE_MAX).direct_route); + assert!(!shape_of(SMALL_SIZE_MAX + 1).direct_route); + // 64-bit: 1024, and it does NOT coincide with the small-page top. + // 32-bit: 512, where it DOES -- which is why the device sees one step + // and the host sees two boundaries with nothing between them. + if core::mem::size_of::() == 8 { + assert_eq!(SMALL_SIZE_MAX, 1024); + assert_ne!(SMALL_SIZE_MAX, SMALL_OBJ_SIZE_MAX); + } else if core::mem::size_of::() == 4 { + assert_eq!(SMALL_SIZE_MAX, 512); + } + // The page kinds a firmware could not observe before. + assert!(shape_of(16).page_bytes <= shape_of(SMALL_OBJ_SIZE_MAX).page_bytes); + assert!( + shape_of(SMALL_OBJ_SIZE_MAX + 1).page_bytes > shape_of(SMALL_OBJ_SIZE_MAX).page_bytes + ); + assert_eq!(shape_of(16).dedicated_segments, 0); + assert!(shape_of(LARGEST_SHARED_ALLOC + 1).dedicated_segments >= 1); + } + #[test] fn usable_bytes_answers_the_question_a_firmware_asks() { let seg = SEGMENT_SIZE; diff --git a/docs/LEDGER.md b/docs/LEDGER.md index a5abde4..611dac5 100644 --- a/docs/LEDGER.md +++ b/docs/LEDGER.md @@ -48,6 +48,43 @@ slice from `dedicated_segments` and the sizing test goes red); the reproduction is a permanent property-based test against the real extent allocator. Unsafe +4, all `#[cfg(test)]` — the fix is arithmetic and adds none to shipped code. +## SMALL-PATH STEP — not the prim, the POINTER WIDTH; the heartbeat knob bare metal could not reach (2026-09-10) + +`docs/plans/finished/fixed-prim-small-step.md`: the Kairos RTOS measured one +alloc+free stepping 314 -> 271 cycles/op across 512 bytes on a 32-bit ESP32-S3, +could not reproduce it on a 64-bit host, and concluded it "points at +`prim::fixed`". + +**Re-attributed.** `SMALL_SIZE_MAX` is `128 * size_of::()` — 1,024 on +their host, **512 on the device**, where it coincides with +`SMALL_OBJ_SIZE_MAX`. Their refutation ("the 512 step is not `SMALL_SIZE_MAX`") +was run on a machine where that constant sits at 1,024. Moving ONE variable, +pointer width, with `prim::windows` in both arms (ABBA, 50k ops/arm, min of 40 +blocks): x86-64 **+0.1 %** at 512 vs 513, i686 **+8.6 %**, control (256 vs 264) +under 1 % on both. The prim is exonerated; the run that settled it was +`--target i686-pc-windows-msvc`, outside the crate, not the fixed-prim-on-host +build the report asked us for. + +**Mechanism, by counter not clock.** Per 100,000 pairs with one block live, the +`direct[]` route carves/extends/retires **195** pages and the bin route just +above it **0**, while `generic`/op is 1.0000 on both — so the slow path is not +the difference, page churn is. `100_000 / 512 = 195` is +`GENERIC_COLLECT_DEFAULT` at the small profile. The boundary is the ROUTE, not +the page kind: on 64-bit, 513–1024 are medium pages that still churn. + +**Shipped:** `--cfg ra_generic_collect` (the trade was real and the knob was +unreachable — `options::set` is a no-op under `ONE_REGION`); and +`prim::fixed::shape_of` so the route and page kind are a `const` answer rather +than a silicon discovery. Default unchanged: the 512 exists because 10,000 +starved this profile (168 -> 8 blocks, 22,533/50,000 nulls). + +**Withdrawn rather than quoted:** how much of the 16 % the churn accounts for. +Raising the heartbeat took churn to a measured zero, but that experiment's +timing control flipped +8.0 % -> +0.6 % on re-run, so the instrument was +deciding it. The device is the right box; the counter half needs no quiet one. + +**Gates:** 20 suites at both profiles, gate-selftest 11/11, wasm 20,168. + ## REGION ALIGNMENT DISSOLVED — segments stride from the base; 24,144 B of stack back, +3 instructions per free, one knob (2026-09-09) `docs/plans/finished/region-alignment-dissolve.md`: the Janus firmware's diff --git a/docs/plans/finished/fixed-prim-small-step.md b/docs/plans/finished/fixed-prim-small-step.md new file mode 100644 index 0000000..a92e593 --- /dev/null +++ b/docs/plans/finished/fixed-prim-small-step.md @@ -0,0 +1,340 @@ +# The small-path step — a 16% cliff at 512 bytes that only exists on bare metal + +**Status:** RESOLVED 2026-09-10 — see §8. The attribution was wrong and the +correction is the finding. Originally: measured, nothing implemented · **Date:** 2026-09-10 · **Build:** +`=2.1.0` from crates.io, unmodified · **Cfgs:** `ra_single_threaded`, +`ra_small_profile` · **Boxes:** ESP32-S3 DevKit (Xtensa LX7, 32-bit, +`prim::fixed`) with `CCOUNT`; and x86-64 Windows (`prim::windows`) with +`rdtsc`, pinned to one core at High priority + +Filed by the Kairos project (FreeRTOS remade in Rust), which uses +`rusty_alloc` behind its `heap_3` seam. Everything here is reproducible from +two cells and one standalone binary, all listed in §6. + +This is a **map, not a claim**. It answers one question — *why does one +`alloc`+`free` cost 16% more below 512 bytes than above it* — reports what +that costs against the incumbent, and is explicit about the part it could +not answer and the run that would. + +--- + +## 0. The answer in one paragraph + +On a microcontroller, one `alloc` + one `free` steps **314 → 271 cycles/op +between a 512-byte request and a 513-byte one**, and **271 → 933 between +2,048 and 2,049**. Both boundaries are your own constants under +`ra_small_profile` (`SMALL_OBJ_SIZE_MAX` = 512, `MEDIUM_OBJ_SIZE_MAX` = +2,048); the totals are byte-identical *within* each range and nothing steps +anywhere else, so this is **routing**, not a gradual effect and not +fragmentation. **The lower step does not reproduce on a 64-bit host**, where +the same sweep is flat across 512 *and* across 1,024 and runs at **13 +cycles/op against the device's 314**. So it is not a generic geometry effect, +and it is not the `direct[]` table (whose top is 1,024 on 64-bit and shows no +step either) — it is specific to the configuration a firmware actually uses, +which points at **`prim::fixed`**. It costs something concrete: measured +against FreeRTOS's `heap_4` on the same part, `rusty_alloc` is **2.4× faster +at 16 bytes and 1.27× slower at 256–512** — and 256–512 is *the only range it +loses*, so closing this step would turn the one loss into a win. + +--- + +## 1. Method, and what makes a number admissible here + +**The counter is real.** Xtensa `CCOUNT` increments once per CPU cycle. This +matters because the obvious alternative does not exist: QEMU was measured six +ways to supply no cycle, latency or work counter at all (DWT unimplemented, +SysTick tracking host wall time and *shrinking* as work grows, nothing under +`-icount`), so every number here is from silicon. + +**A null arm is subtracted.** An identical empty round — same loop, same +`black_box`, same counter reads — is measured the same way and taken off: +16 cycles/op on the device, printed so a reader can see how much of a small +figure is floor. + +**Best of 32 rounds of 256 ops, after 8 warm-up rounds.** The floor is what +survives an interrupt landing mid-round; the warm-up pulls the code from +flash into the instruction cache, and measuring that would measure the flash +controller. + +**Work is checked, not assumed.** Every round keeps a checksum, so a compiler +that removed the allocation would show as a changed checksum rather than as a +suspiciously good number. + +**The table is a function of size, and that is proven rather than hoped.** +Every size is measured twice, ascending and descending, and the cell fails if +one does not reproduce its own total. All seventeen sizes up to 2,048 +reproduce **to the cycle**. (The one exception is quarantined: 4,096 — +`FIXED_PAGE`, so it cannot come from a page — read 789, 861 and 933 cycles/op +in three runs differing only in what ran before it. It is not a function of +size, is declared an expected asymmetry, and is not quoted anywhere.) + +--- + +## 2. The step, tested by one byte + +The plateaus are **not** the bin geometry — 160, 192, 224, 256, 288, 320, 384 +and 512 are eight distinct bins and every one of them produces the identical +total of 84,498 cycles. They land instead on the page-kind boundaries, which +makes a prediction that can be wrong by one byte, so it was tested that way: + +```text + size total per op route + 256 84498 314 small + 511 84498 314 small + 512 84498 314 small + 513 73490 271 MEDIUM <- steps here + 640 73490 271 medium + 1024 73490 271 medium + 2047 73490 271 medium + 2048 73490 271 medium + 2049 242979 933 own large span <- and here + 2560 242979 933 own large span +``` + +**43 cycles for one byte at 512; 662 for one byte at 2,048.** Requests one +byte apart cannot differ for any other reason — same alignment, no bin +between them, and the whole table already shown to be a function of size. + +--- + +## 3. It does not reproduce on a host, and that is the useful part + +`tools/alloc-route-probe` (§6) runs the same sweep on x86-64 against the same +crate version and the same two cfgs. On 64-bit the two candidate constants +come apart — `SMALL_SIZE_MAX` is `128 × 8` = 1,024 while `ra_small_profile` +keeps `SMALL_OBJ_SIZE_MAX` at `4096/8` = 512 — so a step would name its own +cause. Neither does: + +```text + size cycles/op + 511 13 + 512 13 == SMALL_OBJ_SIZE_MAX + 513 13 + 1023 12 + 1024 12 == SMALL_SIZE_MAX (direct[] top) + 1025 11 + 2048 11 == MEDIUM_OBJ_SIZE_MAX + 2049 91 <- the only step +``` + +Three things follow. + +1. **`MEDIUM_OBJ_SIZE_MAX` routes on both platforms.** Confirmed twice, + independently. That one is not in question. +2. **The 512 step is not `SMALL_SIZE_MAX`.** On this host that constant is + 1,024 and there is no step there. +3. **The whole host sweep runs at 11–13 cycles/op against the device's + 271–314.** Even allowing for a slower core, 24× is not a core-speed + ratio — it is a different path. The host is hitting a fast path the + device is not. + +The variable is the **prim**: a firmware gets `prim::fixed`, a host gets +`prim::windows`/`prim::unix`, and the prim is selected by target OS rather +than by a feature, so it cannot be swapped on a host to isolate it. That is +where we ran out of road, and §5 says what would finish it. + +**A hypothesis, labelled as one.** `alloc.rs` already records that "a tight +alloc/free loop frees into `local_free`, so the queue front's `free` list is +ALWAYS dry when the next allocation arrives" — and this harness *is* that +loop. If the device is taking `malloc_generic` on most operations while the +host is not, the 24× gap and the step would be the same fact seen twice. We +could not confirm it: nothing on the seam exposes a slow-path counter. + +--- + +## 4. Four models written down first, all four refuted + +Recorded so nobody re-runs them. + +| model | refuted by | +|---|---| +| bin geometry | the plateaus span **eight** distinct bins with identical totals | +| collect frequency (`FIXED_PAGE / good_size`) | monotone in size; the cost is not | +| a per-page cost amortised over blocks | fits 16…256 at ~0.875 cycles/byte, then breaks completely at 512 | +| blocks-per-page *within* a route | 256 and 512 share a page size, have 16 and 8 blocks, and cost the **same** 314 | + +--- + +## 5. What we could not measure, and what would + +**We could not see a page kind from outside.** The natural footprint answer — +a small page is one slice (4 KiB), a medium page four (16 KiB), so a 513-byte +allocation should claim 4× the region a 512-byte one does — is not observable +through the seam. `region_stats()` answers over region **extents**, so it +moves when a whole segment is claimed and not when a page is; the probe that +asked read 0 for all six sizes and was deleted rather than shipped. + +Two gaps, both of which a firmware would need to tune around this at all: + +- **no per-allocation usable size** on the fixed-region API (`usable_size` + exists in `alloc.rs`, but a firmware that reaches past its seam to get it + has defeated the seam), and +- **no way to observe which page kind, or how many region bytes, a request + costs.** + +**The run that would finish it** is inside your tree, not ours: the same +sweep against `prim::fixed` on a host, which would isolate the prim from the +architecture in one step. From outside the crate it cannot be built, because +the prim is chosen by `cfg(unix)`/`cfg(windows)`. + +--- + +## 6. Reproducers + +| what | where | needs | +|---|---|---| +| the one-byte step, on silicon | Kairos `rusty_rtos_core/firmware/esp32s3-devkit-alloc-cycles` | an ESP32-S3 on USB | +| the host sweep, self-contained | Kairos `tools/alloc-route-probe` — one binary, `cargo run --release`, no Kairos dependencies | nothing | +| the `heap_4` A/B | Kairos `rusty_rtos_core/firmware/esp32s3-devkit-alloc-ab` | an ESP32-S3 on USB | + +--- + +## 7. Why 16% is worth your time + +Measured on the same part against FreeRTOS's `heap_4` — compiled **verbatim** +from `FreeRTOS-Kernel/portable/MemMang/heap_4.c`, run under the identical +harness, both allocators given the same 64 KiB, both called directly at the +same 8-byte alignment, with `vTaskSuspendAll` stubbed to a no-op so the C arm +pays no lock the `ra_single_threaded` Rust arm does not: + +| request | rusty_alloc | heap_4 | | +|---|---:|---:|---| +| 16 B | 99 | 236 | **2.38× faster** | +| 32 B | 109 | 236 | **2.17× faster** | +| 64 B | 141 | 236 | **1.67× faster** | +| 128 B | 193 | 236 | 1.22× faster | +| 256–512 B | 298 | 236 | **1.27× slower** | +| 1024–2048 B | 255 | 236 | 1.08× slower | + +The null A/B — the Rust arm run against *itself*, presented to the harness as +two different allocators — came back **37,963 vs 37,963**, so the resolution +floor is **zero cycles** and every difference above it counts. (These are the +64 KiB, direct-call numbers; they differ from §2's by a **constant 16 +cycles** at every size, which is the harness difference — `Vec` versus a raw +call — with the plateau structure identical.) + +**The only range `rusty_alloc` loses is the expensive route.** Put 256–512 on +the 255-cycle path and the loss becomes a win. + +**And for completeness, because a flat `heap_4` number is its best case and +should not be quoted alone:** that 236 is flat because the workload keeps one +block live, so heap_4's free list is one entry and first-fit answers in one +step. At 512-byte requests against a free list of 16-byte holes it degrades +**linearly at ~27 cycles per entry walked** — 235 → 466 → 1,114 → 3,706 at 0, +8, 32 and 128 holes — while `rusty_alloc` stays flat at 297. Different floor +functions; the crossover is about two holes. On a heap that fragments, which +is every long-running RTOS heap, `rusty_alloc` wins by 12.5× at 128 holes. +**That is the headline, and this document is about the one place it does +not.** + +*(Our own first fragmentation probe was refuted and it is worth passing on: +holes of the **same** size as the request made heap_4 **faster**, 236 → 218, +because first-fit stops at the first block that fits. The holes have to be +smaller than the request.)* + +--- + +## 8. Resolution (2026-09-10) — it is not the prim, it is the POINTER WIDTH + +Taken seriously, reproduced, and re-attributed. The report is right that there +is a real step, right that it is routing, and right about where it hurts. The +one thing it gets wrong is the cause, and the reason is worth more than the +fix. + +### 8.1 The step reproduces on a HOST — a 32-bit one + +`SMALL_SIZE_MAX` is not a profile constant. It is + +```rust +pub const SMALL_SIZE_MAX: usize = SMALL_WSIZE_MAX * INTPTR_SIZE; // 128 * size_of::() +``` + +so it is **1,024 on a 64-bit host and 512 on a 32-bit chip**, where it lands on +the same byte as `SMALL_OBJ_SIZE_MAX`. §3's refutation — "on this host that +constant is 1,024 and there is no step there" — is testing a machine on which +the suspect is not at the scene. Re-run with the pointer width as the only +variable, same crate, same cfgs, `prim::windows` in **both** arms: + +| arm | 256 vs 264 (control) | **512 vs 513** | 2048 vs 2049 | +|---|---:|---:|---:| +| x86-64 host | −0.9 % | **+0.1 %** | −81.9 % | +| **i686 host** | −0.4 % | **+8.6 %** | −80.6 % | + +ABBA-interleaved, 50,000 ops per arm, min of 40 blocks. The control says the +instrument resolves well under 1 %, so +8.6 % is twenty times the floor. + +**`prim::fixed` is exonerated.** The step appears on the OS prim as soon as the +pointer is 32 bits, and disappears on the same binary at 64. §5's "the run that +would finish it is inside your tree" turned out to be a run *outside* it: +`cargo build --target i686-pc-windows-msvc`, no bare-metal prim required. + +### 8.2 The mechanism, counted rather than timed + +`alloc::stats()` is public and carries the slow-path counter §3 wanted. Per +100,000 alloc+free pairs, one live block: + +| route | `generic`/op | pages carved | extends | retires | +|---|---:|---:|---:|---:| +| `direct[]` (`size <= SMALL_SIZE_MAX`) | 1.0000 | **195** | 195 | 195 | +| bin peek (just above it) | 1.0000 | **0** | 0 | 0 | + +So the slow path is **not** the difference — both routes take it on every +operation, which also settles §3's hypothesis: `malloc_generic` is universal +here, not a device effect. The difference is that the `direct[]` route +**retires its page and carves a fresh one every ~513 operations** and the bin +route never does. `100_000 / 512 = 195`: that is +`GENERIC_COLLECT_DEFAULT`, the periodic `collect_inner` sweep, which is 512 at +the small profile. In a loop that keeps one block live the page is empty at +every sweep, so every sweep costs a carve, an extend and a retire. + +Note what this also explains: the boundary is the **route**, not the page kind. +On 64-bit, 513–1024 are medium PAGES on the `direct[]` route and they churn +pages too (195); it is crossing `SMALL_SIZE_MAX` that stops it. + +### 8.3 It is a deliberate trade, and the defect was that you could not move it + +The 512 is not an oversight. Its own doc records why: at 10,000 the small +profile starved — "512 B capacity decaying 168 → 8 blocks and 22,533 of 50,000 +churn allocations returning null with 61,440 bytes of the region free". Short +sweeps buy that back and cost page churn. **Neither direction is free, and +which is right depends on the workload** — a bench that keeps one block live +cannot decay, so it sees only the cost. + +The real gap is that a firmware had **no way to choose**. `options::set` is a +no-op under `ONE_REGION` (options are compile-time constants there, which is +what lets the option table leave the image), and the doc told firmwares to +"change the default it is built with" when there was no cfg to do it. + +**Shipped:** `--cfg ra_generic_collect="64" | "4096" | "65536"`. + +### 8.4 The observability §5 asked for + +- **`prim::fixed::shape_of(size) -> Shape`** — page bytes, dedicated segments, + and `direct_route`, all `const`, all derived from the active geometry. That + is "which page kind, and how many region bytes" as a compile-time answer; + `region_stats()` could never answer it because it reports over extents. +- **`Shape::direct_route` names the pointer-width boundary**, so the next + consumer meets it in a `const` assertion instead of on silicon. +- A unit test pins `SMALL_SIZE_MAX == 128 * size_of::()` and the two + width-specific values. + +Two of the §5 gaps were already reachable and are worth knowing rather than +building: **`alloc::usable_size`** is public, and **`alloc::stats()`** is the +slow-path counter. Both are in `alloc`, not `prim::fixed`, so a seam +re-exporting the fixed-region API in one `pub use` cannot see them — the same +shape as the `PrimError` gap closed in 2.1.0. Re-export them from `alloc` in +the seam and the "nothing exposes a counter" note goes away. + +### 8.5 What is NOT established, and whose run finishes it + +**How much of the 16 % the page churn accounts for is still open.** Raising the +heartbeat on a host takes the churn to a measured zero, but the timing arm of +that experiment was inadmissible: its control re-run flipped from +8.0 % to ++0.6 %, so the instrument was deciding the answer and the number is withdrawn +rather than quoted. The device, with `CCOUNT` and a quiet core, is the right +box for it — `--cfg ra_generic_collect="65536"` against the same sweep on the +same part is the one-line experiment, and the counter half of it +(`stats().pages_fresh`) is deterministic and needs no quiet box at all. + +Nothing here changes the default. A firmware that has measured its own workload +can now move it; one that has not should not. From 3c48be10822dcf7f2a8de1b5c8af0d8f17e78d98 Mon Sep 17 00:00:00 2001 From: Tim Date: Thu, 10 Sep 2026 15:33:01 -0700 Subject: [PATCH 2/3] test(prim): derive the boundary coincidence instead of asserting it The new pointer-width test carried `assert_ne!(SMALL_SIZE_MAX, SMALL_OBJ_SIZE_MAX)` on its 64-bit branch. That is a fact about the GEOMETRY, not the pointer, and it is false under `ra_segment_size="256k"`: an 8 KiB slice puts SMALL_OBJ_SIZE_MAX at 1024, which is where the 64-bit route top already is. Green at the default slice, red on the 256k rung -- the exact defect shape this repo's own notes describe, committed in the commit that documented it. Derived from the arithmetic now, so it says something at every geometry. The coincidence is also the interesting half: where the two constants land on the same byte, one step hides two boundaries and a sweep cannot separate them, which is what made the reporting device confusing in the first place. Co-Authored-By: Claude Opus 5 (1M context) --- crates/rusty_alloc/src/prim/fixed.rs | 22 ++++++++++++++++++---- 1 file changed, 18 insertions(+), 4 deletions(-) diff --git a/crates/rusty_alloc/src/prim/fixed.rs b/crates/rusty_alloc/src/prim/fixed.rs index ff26ddb..da58ded 100644 --- a/crates/rusty_alloc/src/prim/fixed.rs +++ b/crates/rusty_alloc/src/prim/fixed.rs @@ -2052,15 +2052,29 @@ mod tests { ); assert!(shape_of(SMALL_SIZE_MAX).direct_route); assert!(!shape_of(SMALL_SIZE_MAX + 1).direct_route); - // 64-bit: 1024, and it does NOT coincide with the small-page top. - // 32-bit: 512, where it DOES -- which is why the device sees one step - // and the host sees two boundaries with nothing between them. + // The route top is 1024 on 64-bit and 512 on 32-bit, whatever the + // geometry -- it is a pointer-width fact, not a profile one. if core::mem::size_of::() == 8 { assert_eq!(SMALL_SIZE_MAX, 1024); - assert_ne!(SMALL_SIZE_MAX, SMALL_OBJ_SIZE_MAX); } else if core::mem::size_of::() == 4 { assert_eq!(SMALL_SIZE_MAX, 512); } + // WHETHER it coincides with the small-page top is a fact about the + // GEOMETRY, so it is derived rather than asserted -- an unconditional + // `assert_ne!` here passed at the default slice and failed under + // `ra_segment_size="256k"`, where an 8 KiB slice puts + // SMALL_OBJ_SIZE_MAX at 1024 and the two meet on 64-bit too. + // + // The coincidence is the interesting part and is what made the device + // confusing: where the two constants land on the same byte, one step + // hides two boundaries and a sweep cannot tell them apart. + assert_eq!(SMALL_OBJ_SIZE_MAX, crate::types::SEGMENT_SLICE_SIZE / 8); + let coincide = SMALL_SIZE_MAX == SMALL_OBJ_SIZE_MAX; + assert_eq!( + coincide, + SMALL_WSIZE_MAX * core::mem::size_of::() == crate::types::SEGMENT_SLICE_SIZE / 8, + "the two boundaries coincide exactly when the arithmetic says so" + ); // The page kinds a firmware could not observe before. assert!(shape_of(16).page_bytes <= shape_of(SMALL_OBJ_SIZE_MAX).page_bytes); assert!( From 338800faafd10de6e6a64b5d1263ec6430b9adc4 Mon Sep 17 00:00:00 2001 From: Tim Date: Thu, 10 Sep 2026 15:43:39 -0700 Subject: [PATCH 3/3] docs(plan): record the capacity=1 lead behind the page churn Tracing collect_inner's reclaim gives the churn a face and one unexpected column: every reclaimed page carries capacity=1. `page_extend` links a 4 KiB payload bound, so a 512-byte class should return 8 blocks and a 1024-byte class 4. A page holding one block is exhausted by the next allocation, which is the missing half of why `generic` reads exactly 1.0000 per op. The trace also pins the route boundary exactly: bins 20/21/24 (the direct[] sizes at SMALL_SIZE_MAX=1024) churn ~200 times per 100k pairs, bins 25 and 28 just above it churn once. Recorded, not chased: it points at page_extend/page_fresh, which is a hot-path change needing the instruction-count battery rather than a session's tail. It is a better lead than the heartbeat knob, because it would close the step from the other side with no default moved. Co-Authored-By: Claude Opus 5 (1M context) --- docs/plans/finished/fixed-prim-small-step.md | 39 ++++++++++++++++++++ 1 file changed, 39 insertions(+) diff --git a/docs/plans/finished/fixed-prim-small-step.md b/docs/plans/finished/fixed-prim-small-step.md index a92e593..321a429 100644 --- a/docs/plans/finished/fixed-prim-small-step.md +++ b/docs/plans/finished/fixed-prim-small-step.md @@ -338,3 +338,42 @@ same part is the one-line experiment, and the counter half of it Nothing here changes the default. A firmware that has measured its own workload can now move it; one that has not should not. + +### 8.6 A lead worth more than the knob: every reclaimed page has `capacity = 1` + +Tracing the reclaim itself (a temporary `eprintln` in `collect_inner`'s +`page_all_free` branch, 100,000 pairs per size on x86-64) gives the churn a +face, and one column of it was not expected: + +```text + 200 RECLAIM bin=24 block_size=1024 used=0 capacity=1 + 200 RECLAIM bin=21 block_size=640 used=0 capacity=1 + 199 RECLAIM bin=20 block_size=512 used=0 capacity=1 + 1 RECLAIM bin=28 block_size=2048 used=0 capacity=1 + 1 RECLAIM bin=25 block_size=1280 used=0 capacity=1 +``` + +Two things. The route boundary is exact — bins 20/21/24 are the `direct[]` +sizes (512, 640, 1024 at `SMALL_SIZE_MAX = 1024` on this host) and churn ~200 +times each; bins 25 and 28 are just above it and churn **once**, at the +transition between probe rows. And **`capacity = 1` on every one of them.** + +`page_extend` links a 4 KiB *payload* bound per extend, so a 512-byte class +should come back with 8 blocks and a 1,024-byte class with 4. A page holding +ONE block is exhausted by the allocation that follows it, which is the missing +half of why `generic` reads exactly **1.0000 per op** in §8.2 — the fast path +cannot hit a page that has no second block. Whether `reserved` is being +clamped to 1 for these classes, or the extend is not running at all, is not +established here. + +**This is the lead to pull next, and it is not the heartbeat.** The knob in +§8.3 changes how often an empty page is reclaimed; this asks why the page was +worth so little in the first place. If a `direct[]`-route page came back with +its full block count, the fast path would hit, `generic` would fall well below +1.0/op, and the sweep would find the page in use rather than empty — the step +would close from the other side, with no knob and no default moved. + +Recorded rather than chased because it is a hot-path change (`page_extend` and +`page_fresh`) that needs the full instruction-count battery, not a session's +tail. The reproducer is three lines of `eprintln` in `collect_inner` and the +counter probe in §8.2.