Repository navigation
Astro Bot (PPSA21564) #357
Replies: 31 comments 26 replies
Update 2026-10-02: both blockers have a causeSame host and stack 2 as above.
A in detail
B in detail
Standalone reproduction with no AnyPS5 code ( Seen once B is bypassed with the diagnostic
Investigated with AI assistance (Claude Code). |
Update 2026-10-02 (2): runs for minutes with sound, frames still black
The three local diagnostics (none is a proposed fix)
Where it gets to
C in detail
Smaller points
Investigated with AI assistance (Claude Code). |
Update 2026-10-02 (3): first frames, and a correction to the update above
Open, in order
How the correction was checkedThe diagnostic reported shared views as untracked while private memory stayed write-watched. Checked with Dreaming Sarah (PPSA02929) on
The two unconfirmed stops did not appear in the two launches described below. Why the switch avoids BWithout write watching The runs that showed framesStack 2, FFmpeg enabled,
Investigated with AI assistance (Claude Code). |
|
Blocker B (the red zone overwritten when a write fault is resumed on Windows) now has its own issue, since it is on Written with AI assistance (Claude Code). |
|
Great write-up, thanks for digging into the Windows side. My Here is a Linux gameplay capture from last night, through the intro cinematic, the crash site tutorial and the world map into Sky Garden: https://www.youtube.com/watch?v=62wl_LddNnQ. It was recorded before most of the fixes below, so frame rate and visuals are better now. Relevant to your list
Tips for testing
Happy to compare logs if you try the newer head. |
|
If you release this and Gran Turismo I will not need a PS5 PLS add GT7 support in the future after astros playroom. Add them all |
|
I understand this is very difficult and hard and may takes months to years to complete |
Windows: in game, Gorilla Nebula first level reachedThanks to @oneandonlydean: the rendering and gameplay progress below is the work on his
Video: galaxy map, Sky Garden, the dive and the flight into the level, with the log console beside the game window. About 41 s of the white cloud transition are cut and the recording is re-encoded to fit the 10 MB attachment limit, so it is softer than the original 1440p capture. astrobot-windows-8bit.mp4Written with AI assistance (Claude Code). |
|
Thanks, this is great to see, and thanks for the Windows fixes. I've brought #282, #283 and #284 into my The branch has moved on since
Also new: the occlusion dumps no longer drain the GPU (Snowy Canyon went from 1.2 to 2.3 fps on Linux), and on Linux the queue workers are pinned to P-cores on hybrid CPUs. That is off on Windows unless Two things I can't reproduce on Linux and would like to understand:
|
Windows on an AMD GPU (Radeon RX 9060 XT): title card, blocker B in every launch with write watching on
Run 3 on the Radeon: the title card, 272 frames in, at 1.19 FPS. It stayed on this card until I stopped the run at 240 s; the title advances once per presented frame, so the video it had opened was not reached on screen.
B reproduces on every write-watched launch on this machine, so I can test a fix for #361 quickly if that helps. Tested with AI assistance (Claude Code). |
PPSA21567 (the other region's release) on the
|
| Run | Settings | Result |
|---|---|---|
| 1 | default, no input | Logos, title ("press any button") by about 140 s; idle there to 560 s. No error, no GPU hang. |
| 2 | default | Title, NEW GAME save-slot menu, warp. Then it stops: compute shader 0x500622100: AGC graphics: guest snapshot differs from registered memory. That's the same stop SP4C3B4R-8 hit on Windows. |
| 3, 4 | APS5_NO_SNAPSHOT_CHECK=1 |
Both run the full 15 min into the intro cinematic (the mothership, Nebulax's attack, the ship breaking up). No error, no GPU hang, no skipped draws. |
| 5 | as 3 and 4, 35 min | Through the intro cinematic to the crash landing in the desert (Crash Site), then waiting in the tutorial with a controller icon in the corner (no stick input sent). No error, no GPU hang. |
So PPSA21567 v01.018 needs nothing region-specific on top of your branch to get this far. On upstream main it hung the GPU (Xid 109) a few minutes in, inside a heavy compute shader. Your e6ff9656 explanation (two-lane register pressure, spilling) fits everything I'd measured, including why a loop guard that never trips made the hang go away.
What I have that the branch lacks (on my branches built on upstream main, #419 and #425; none merged; each with tests):
- Exact colour compare emulation. For a colour image bound to a comparison sampler, it does the compare in the shader the way the hardware does: per texel, before filtering, including clamp-to-border with black and white borders. Currently the result there is left to the driver (
1ad8fe93). Unsupported formats, clamp modes, the border colour table and aniso still throw. - Single-sample CMASK fast clear. It models the two fill values: CMASK 0 is "cleared to the CB_COLOR_CLEAR_WORD value", and 0xFFFFFFFF is "expanded". It's not rendered uncompressed. Anything else throws. Scoped as MichonGoddijn231849 asked on feat(agc): multisampled color and depth targets with CB resolves #419.
- Depth layout matching / write-back. A colour or storage view of memory a depth surface holds is served from the depth surface in the view's layout: it's written back through an R32 plane view and reloaded. That replaces a stale read. This is the proposal on feat(agc): depth surfaces read as textures and storage images #425; you raised there that the refusal should stay for the case without write-back.
How would you like contributions: PRs against your astrobot branch, or upstream PRs you then pull in? I'd start with whichever of the three you find most useful.
|
@oneandonlydean the exact colour compare emulation, ported onto
I'm away for about a month from today. You're welcome to take it as-is, adapt it, or PR it upstream yourself, whichever suits the branch. The single-sample CMASK model waits until I'm back. |
|
Two measurements from PPSA21567 on Internal resolution. I added a local counter (not pushed) that logs, every 2 s, the colour-target and viewport sizes of the draws and the pixel programs drawing into 3840x2160 targets.
Out of device memory on 8 GB. GPU memory (nvidia-smi, desktop about 0.5 GB included) climbs to about 6.8 GB in the cinematic and peaks at about 7.2 GB at the crash landing. Results of today's long runs (35-minute tests plus one play session), all on this branch, leaving out one real-display run that ended in a GPU hang (Xid 109) instead, 9 in all: 3 unmodified, 4 with my compare commit and 2 with the local counter. Neither of those adds GPU allocations of its own. Both kinds of result happen in each group:
So on an 8 GB card the margin from the intro to the crash landing is a few hundred MB. Each run either fits or doesn't. |
|
@oneandonlydean one region-specific blocker on PPSA21567 (v01.018), relevant to #476. Without my The faulting address is the same offset ( With libc linked Does PPSA21564 bundle |
Update 2026-10-05: blocker B fixed in #637, Crash Site reached on Windows with an Intel CPUSame host as the first post: Windows 11, Core Ultra 9 285K (no AVX-512), RTX 5080, PPSA21564 v01.007.000, relinked with
Upstream
|
| Stack | Stop |
|---|---|
main |
19 s: compute shader 0x50075f000: AGC graphics: unsupported descriptor role Gds |
main + #285 |
13 s: compute shader 0x5005e7b00: AGC driver: guest memory is not readable at 0x30 (range 0x30+0x4) |
Both runs also skip draws with color compression, DCC, endian conversion, nonstandard rounding or color optimization is unsupported (CB_COLOR_INFO 0x00042730). The astrobot branch passes all three.
The launches on astrobot b58df66c
| Build | Settings | Result |
|---|---|---|
| unmodified | default | B at 12 s: 0xc0000005 on SceSndzAudioOutMain, rip eboot+0xec64b5, write |
| unmodified | APS5_NO_WRITE_WATCH=1 |
300 s, no error, 412 frames by 290 s (1.6 to 3 FPS), PlayStation Studios video reached |
| + #637 | default, cold shader cache | 100 s, no error, 60 FPS on the logos, 2094 frames by 95 s |
| + #637 | default, warm shader cache | 52 s: vkWaitForFences present: Vulkan result -4 right after the 3D intro scene starts |
| + #637 | the four settings below, cold cache | 420 s, no error, title and new game reached |
| + #637 | the four settings below, warm cache | 1200 s, no error, 8524 frames, Crash Site reached at about 1050 s |
Settings, taken from @oneandonlydean's post above: APS5_TWO_LANE=500597a00 APS5_INEXACT_SINGLE_LANE=500597a00 APS5_LOOP_GUARD=2000 APS5_TOLERATE_DISPATCH_FAILURES=1. Not established: which of the four avoids the device loss on the RTX 5080; it was seen in one launch without them and in neither launch with them.
Frame rate in the 1200 s launch: about 19 FPS at the title, 2 to 6 FPS in the intro cinematic, 10 FPS on arrival at the Crash Site. Working set about 24 GB, private bytes about 35 GB at the end.
No input device was used: Return was posted to the window every 20 s, which is enough to start a new game and get through the cinematic.
A stall that is not a hang
The first launch after changing those settings recompiles every shader, since they are part of the cache key. After the warp screen the picture stopped for about 80 s while the sound went on: shader-disk-cache reported 192 misses and no hit, at roughly one large wave64 compute program per 10 s. The next launch, with the cache filled, did not stall.
Not tested here: the tutorial with a controller, the save slot without APS5_NO_SNAPSHOT_CHECK=1, and where the 2 to 6 FPS of the cinematic go.
@QDarkRex, since B hits every write-watched launch on your Zen 3, #637 is the change to try there. @SP4C3B4R-8, it also removes the write faults on section views that were suspected for the 0.1 FPS in the first level; a run with it would tell.
Tested and written with AI assistance (Claude Code).
|
@oneandonlydean thanks for the detailed answer, and for #618. Noted on upstream PRs; I'll take memory on 8 GB first. libc on PPSA21567: your data point holds here, with one exception. Relinked with #618: tested and posted on the PR. Part 1 works for libc. This title then stops at Later levels on PPSA21567 (your level select;
Thanks again for the level select; it made this scan possible overnight. |
|
@oneandonlydean on memory with 8 GB: a measurement of what holds device memory when an allocation fails, and a direction I'd like your view on. Setup: PPSA21567 on At the failure, device-local memory is 7.0–7.2 GB of 8 GB (the driver reserves 321 MiB):
The failing requests are small: 5–25 MiB device buffers and textures, and once 60 MiB (a 10240x320 texture). The card is simply full. Why evicting idle cache entries can't help mid-frame. I tried a relief step after an out-of-memory error: evict sampled and storage entries with the exhausted policy, sync, wait for releases, and retry the dispatch once. That policy evicts only entries not used since the current residency frame started. I logged the cache at each relief. One frame touches 3.2–3.8 GB of the ~4 GB sampled cache, so the result depends on when in the frame the failure lands:
In the other run the fatal failures weren't in a dispatch at all. After 9 failed allocations, Direction I'm prototyping locally:
I'm comparing it against unmodified on Snowy Canyon and the intro now (crashes, fps, how much falls back) and will post the numbers. The driver also exposes Do you see a problem with this direction? For example, anything on the branch that assumes device-local placement (mapping, aliasing, or the texture caches' accounting). |
|
Numbers for the system-memory fallback from my note above. Same setup:
So the fallback costs nothing until it triggers, but on its own it only trades the crash for a slideshow. In Snowy Canyon it fired 16,087 and 16,339 times (~8 MiB each, 132–138 GB placed in total, none failed): 9,682 device buffers, 5,905 textures, 464 storage textures and 36 buffers in the first run. fps is 9–10 until the first fallback, drops to 0.1–0.2 within 20 s, and never recovers. Device-local memory stays full of cached textures, so the short-lived device buffers, which are re-created constantly and are hot in compute, all end up in system memory. Next I'll try keeping hot transient allocations in VRAM and pushing cold cache entries out instead. One way is |
Windows on an AMD GPU, update: past the intro with #637, then freezes while the sound keeps playingSame host as my report above (Ryzen 7 5700X, Radeon RX 9060 XT, AMD 26.9.2), PPSA21564 v01.007.000, With #637, B is gone here (no crash in four write-watched launches, two of them over 5 minutes) and the game goes past the intro into gameplay at about 17-19 FPS. I logged FPS, CPU, GPU, RAM and paging once per second over two sessions (200 s and 443 s), the second with the game data on a faster SSD:
astrobot-test-637.mp4astrobot-test-637-2.mp4Not established: what the frame path is waiting on during a freeze. I can rerun with any switch that would show it, and the per-second logs and full stderr of both sessions are available. Tested with AI assistance (Claude Code). |
|
@oneandonlydean results of the memory tracing on 8 GB, with a finding about the sampled texture cache's soft limit that I'd like your view on. Setup: PPSA21567 on What the crash is. The trace is exact (running device-local bytes peak at 7,254 and 7,271 MiB at the first failed allocation). Memory sits at ~2.75 GB for 125–185 frames (the sampled cache holds 0.2 GB then), and then within one frame it jumps by ~3.8 GB. The first failed allocation follows 1–2 frames later in both traces. The driver's own counters say the same: 649 sampled textures created in one 10 s window, the cache going from 226 MiB to 2.7 GB, 0 evicted, 0 rescued. That is the new scene's working set, so evicting idle entries can't free it. It's why my relief prototype freed 0 MiB. What it isn't:
A local test that survives. Skipping the top mip of sampled 2D textures larger than 4 MiB (a hack: the image starts at guest level 1, a quarter of the memory): 2 of 2 runs survive 420 s with no out-of-memory, device memory peaking at 5.9–6.0 GB (nvidia-smi) instead of 7.3–7.6 GB. That build also had pageable-memory extensions enabled; they never triggered (no fallbacks, usage below the budget). Textures are blurrier, I haven't judged by how much. The finding. After the burst that run only manages ~2 fps (fps before the burst: 8–9), and the cache is churning: 29,837 textures created and 30,329 evicted (~137 GiB) in 7 minutes, with the sampled cache held at exactly its soft limit: 2,751 MiB, which is 3/8 of the 7,338 MiB budget. The working set is ~3.5–3.9 GB, so every frame the cache evicts and re-creates. The cache also accounts each texture at its guest size ( Caveats. The trace can't count textures used by replayed draws ("recipe hits" skip the code path where I record use), so "idle after the burst" from the trace is not reliable. I have not run an unmodified build that survives the burst, so I can't say the churn alone explains the 2 fps. Questions.
|
|
Correction to my trace summary above: I wrote that retiring finished batches on every queue ( What the code says about the lifetime: that scratch ( |
|
@oneandonlydean the larger-cache and mip-skip numbers, and a correction to what I said about the frame rate. Setup as before:
Correction. I suggested that the cache churn probably explains the ~2 fps after the burst. It explains only part of it. With the 4,500 MiB target the churn stops after about 270 s (new textures per 40 s: ~2,500, then 500, 10, 0) and fps rises from about 2.1 to 2.9–3.4, not back to the 8–9 fps of the loading phase before the burst. So the level itself runs at about 3 fps on this CPU (Ryzen 5 5600X); I haven't profiled whether that's the queue worker's CPU. The churn costs roughly a third of that. A bigger cache alone doesn't change the crash: it holds more memory, so the burst fails the same way. Only the arms with the mip-skip fit. Your current head |
|
@oneandonlydean follow-up on the scratch memory: what holds it, and one thing I got wrong. Setup: as in my earlier trace note (PPSA21567, Retracted: the release path is not slow. I had suspected head-of-line blocking on the release thread because frees came seconds after batches retired. In 24,256 and 19,738 batches, retire to hand-off took at most 1.1 ms, and destruction a median 0 ms (p99 29–37 ms, max 273 ms). Where the delay is: At the first failed allocation (6,954 and 6,999 MiB of device memory alive): sampled texture images 3,906 / 3,862 MiB; device buffers whose batch had already been destroyed 1,316 / 1,326 MiB; storage images 864 / 691; other 760 / 921; device buffers in unretired batches 108 / 198. Question: would you take a change that flushes the pool's device tier and retries once when a device-local allocation fails? It frees up to ~0.5 GB of the ~3.8 GB burst, with no quality cost; it doesn't fix the crash by itself. |
|
@oneandonlydean follow-up on the scratch memory, with one correction to my last note and a result for flushing the buffer pool. Setup as before: PPSA21567 on Correction. I wrote that ~0.8 GB of scratch whose batch was already destroyed was unexplained. That was my trace's mistake: it didn't record buffers taken back out of the buffer pool, so reused buffers looked parked. With take events added, the split at a failed allocation is: ~0.9 to 1.1 GB of device buffers in use by in-flight work, the pool empty, and 5.75 to 5.85 GB of everything else. So there is no leak to find; that scratch is the burst's live upload working set. Flush the pool and retry once. When a device-local
Ruled out: the poisoned-device bug from #290 (a refused host import of 64 KiB or less makes later device-local allocations return -2). No run log shows the refused-import warning (the logs do show the filter working: "host import ... not attempted: read-only or inaccessible pages"), and in 12 of 12 traces without the flush the first failed allocation comes at 6.8 to 7.3 GB of device memory alive, never near empty. I'm not going to push on freeing memory without lowering quality. Question: you said a quality cap would have to be opt-in. Would you take a setting that drops the top mip of large 2D sampled textures (off by default), with the cache accounting charging real allocation bytes? With the local top-mip skip, Snowy Canyon survives 420 s (2 of 2 runs per arm on your |
|
@oneandonlydean the cache and quality-cap change is up against
With the cap at 4 MiB, Snowy Canyon survives 420 s in 2 of 2 runs (3.1 to 3.2 fps), and Iceberg survives too (it runs out of memory at 142 s without the cap). The intro and Kratos run at the same or better frame rates with 0.5 to 0.8 GB less peak memory. The visual cost I measured is about 4 to 6% less fine detail in the matched frames. Close-up Snowy Canyon gameplay is unmeasured. One limit is stated in the PR and in TechnicalDebt. Shaders that fetch by integer coordinates, query a texture's size or use an explicit LOD see the reduced image. In my shader cache, 11% of the modules fetch from or query a sampled image. Nothing tells the texture lookup how a shader uses an image, so those textures aren't excluded. |
PPSA21567 on Linux with an AMD GPU (RADV): GPU page fault and hang in the introHost: CachyOS (kernel 7.2.8, then 7.2.9), Ryzen 7 7800X3D, RX 9070 XT (RADV GFX1201, Mesa 26.2.3, then 26.2.4; the crash happens on both). Without #943, every host import fails on RADV and the first address-based build stops at the 64 MiB copy limit (details in my comment on #943). SymptomAfter 50 s to 4 min of the intro: an amdgpu page fault, then
Ruled out
RADV hang dump (
|
Windows on an NVIDIA GPU (RTX 4070 Ti SUPER): in-game at about 2.7 FPS, two write-watch costs found on the queue 0 worker
What I see
Measured Level
Patch Both changes are in
diff --git a/core/libs/prx/libSceAgcDriver/Execution/src/GuestMemory.cpp b/core/libs/prx/libSceAgcDriver/Execution/src/GuestMemory.cpp
index a0b54cc6..8c0783b1 100644
--- a/core/libs/prx/libSceAgcDriver/Execution/src/GuestMemory.cpp
+++ b/core/libs/prx/libSceAgcDriver/Execution/src/GuestMemory.cpp
@@ -1026,6 +1026,45 @@ struct ThreadCollectMemo {
};
thread_local ThreadCollectMemo threadCollectMemo;
+// The same memo without the ring's bound: every range the thread walked in its current epoch, kept
+// sorted, disjoint and merged, so a collect inside the walked ranges is found by one binary search.
+// A submission of several thousand draws collects far more than 64 distinct surfaces in one epoch;
+// the ring then forgets ranges it walked and walks them again, and on Windows a walk costs
+// GetWriteWatch time proportional to its pages. Ranges of an older epoch (or an older unwatch serial)
+// are dropped at the first walk of the new one. APS5_COLLECT_MEMO_RING=1 uses the ring as before.
+struct ThreadCollectRanges {
+ std::uint64_t epoch = 0;
+ std::uint64_t unwatched = 0;
+ std::vector<std::pair<std::uint64_t, std::uint64_t>> ranges;
+
+ bool covers(std::uint64_t first, std::uint64_t stop) const {
+ auto next = std::upper_bound(ranges.begin(), ranges.end(), first, [](std::uint64_t value, const auto& range) { return value < range.first; });
+ return next != ranges.begin() && stop <= std::prev(next)->second;
+ }
+
+ void add(std::uint64_t first, std::uint64_t stop) {
+ auto at = std::lower_bound(ranges.begin(), ranges.end(), first, [](const auto& range, std::uint64_t value) { return range.second < value; });
+ auto last = at;
+ while (last != ranges.end() && last->first <= stop) {
+ first = std::min(first, last->first);
+ stop = std::max(stop, last->second);
+ ++last;
+ }
+ if (at == last) {
+ ranges.insert(at, {first, stop});
+ return;
+ }
+ *at = {first, stop};
+ ranges.erase(std::next(at), last);
+ }
+};
+thread_local ThreadCollectRanges threadCollectRanges;
+
+bool collectMemoRing() {
+ static const bool ring = std::getenv("APS5_COLLECT_MEMO_RING") != nullptr;
+ return ring;
+}
+
bool newestCollectFirst() {
static const bool enabled = std::getenv("APS5_SCAN_WHOLE_COLLECT_MEMO") == nullptr;
return enabled;
@@ -1060,10 +1099,21 @@ bool walkWrites(WriteTracker& tracker, std::uint64_t first, std::uint64_t stop,
// reset, so a probe pass first (which doubled the walk of every dirty range) buys nothing.
// APS5_NO_SINGLE_PASS_COLLECT=1 restores the probe pass followed by a resetting pass when dirty.
static const bool singlePass = std::getenv("APS5_NO_SINGLE_PASS_COLLECT") == nullptr;
+ // The result buffer offered to the kernel is bounded by the pages the range can report:
+ // GetWriteWatch's cost grows with the offered count as well as with the range (about 7 us for
+ // 65536 entries whatever the range, against well under 1 us for a small range with a matching
+ // buffer), and most walks cover a few dozen pages. APS5_FULL_COLLECT_BUFFER=1 offers the whole
+ // buffer as before.
+ static const bool fullBuffer = std::getenv("APS5_FULL_COLLECT_BUFFER") != nullptr;
+ const auto offered = [&](std::uint64_t from) {
+ if (fullBuffer) return tracker.pages.size();
+ return static_cast<std::size_t>(std::min<std::uint64_t>(tracker.pages.size(), (stop - (from & ~(page - 1)) + page - 1) / page));
+ };
auto cursor = first;
bool dirty = false;
while (cursor < stop) {
- std::size_t count = tracker.pages.size();
+ const std::size_t capacity = offered(cursor);
+ std::size_t count = capacity;
DWORD granularity = 4096;
if (!GuestArena::GuestArenaCollectWrites_nid_postfix(cursor, static_cast<std::size_t>(stop - cursor), tracker.pages.data(), &count, singlePass)) {
// Uncommitted pages inside the range make the call fail; such ranges are compared instead.
@@ -1071,7 +1121,7 @@ bool walkWrites(WriteTracker& tracker, std::uint64_t first, std::uint64_t stop,
}
if (count != 0) dirty = true;
for (ULONG_PTR i = 0; i < count; ++i) tracker.stamp(tracker.blockOf(reinterpret_cast<std::uintptr_t>(tracker.pages[i])), tracker.generation, kind);
- if (count < tracker.pages.size()) break;
+ if (count < capacity) break;
cursor = reinterpret_cast<std::uintptr_t>(tracker.pages[count - 1]) + (granularity != 0 ? granularity : page);
}
if (dirty) collectDirty.fetch_add(1, std::memory_order_relaxed);
@@ -1079,11 +1129,12 @@ bool walkWrites(WriteTracker& tracker, std::uint64_t first, std::uint64_t stop,
// The resetting pass reports pages again, so a write racing the first pass is stamped too.
cursor = first;
while (cursor < stop) {
- std::size_t count = tracker.pages.size();
+ const std::size_t capacity = offered(cursor);
+ std::size_t count = capacity;
DWORD granularity = 4096;
if (!GuestArena::GuestArenaCollectWrites_nid_postfix(cursor, static_cast<std::size_t>(stop - cursor), tracker.pages.data(), &count, true)) return false;
for (ULONG_PTR i = 0; i < count; ++i) tracker.stamp(tracker.blockOf(reinterpret_cast<std::uintptr_t>(tracker.pages[i])), tracker.generation, kind);
- if (count < tracker.pages.size()) break;
+ if (count < capacity) break;
cursor = reinterpret_cast<std::uintptr_t>(tracker.pages[count - 1]) + (granularity != 0 ? granularity : page);
}
}
@@ -1108,7 +1159,14 @@ std::uint64_t collectWrites(std::uint64_t address, std::size_t bytes, bool memoi
const auto epoch = currentCollectEpoch();
const auto unwatched = unwatchSerial.load(std::memory_order_acquire);
const bool useMemo = memoized && collectMemoEnabled() && bytes != 0;
- if (useMemo && !sharedCollectMemo()) {
+ if (useMemo && !sharedCollectMemo() && !collectMemoRing()) {
+ // As below: a range is remembered only for a completed walk of an in-arena range.
+ const auto& walked = threadCollectRanges;
+ if (walked.epoch == epoch && walked.unwatched == unwatched && walked.covers(first, stop)) {
+ collectMemoHits.fetch_add(1, std::memory_order_relaxed);
+ return tracker.generation.load(std::memory_order_relaxed);
+ }
+ } else if (useMemo && !sharedCollectMemo()) {
// An entry exists only for a completed walk of an in-arena range, so the tracker is
// initialized and watched; nothing below the lock needs asking.
const auto count = threadCollectMemo.entries.size();
@@ -1142,7 +1200,15 @@ std::uint64_t collectWrites(std::uint64_t address, std::size_t bytes, bool memoi
if (collectMemoEnabled()) {
const auto serial = unwatchSerial.load(std::memory_order_relaxed);
if (sharedCollectMemo()) tracker.memo[tracker.nextMemo++ % tracker.memo.size()] = {first, stop, epoch, serial};
- else threadCollectMemo.entries[threadCollectMemo.next++ % threadCollectMemo.entries.size()] = {first, stop, epoch, serial};
+ else if (!collectMemoRing()) {
+ auto& walked = threadCollectRanges;
+ if (walked.epoch != epoch || walked.unwatched != serial) {
+ walked.ranges.clear();
+ walked.epoch = epoch;
+ walked.unwatched = serial;
+ }
+ walked.add(first, stop);
+ } else threadCollectMemo.entries[threadCollectMemo.next++ % threadCollectMemo.entries.size()] = {first, stop, epoch, serial};
}
return tracker.generation;
}
https://github.com/user-attachments/assets/9759997c-18e9-44d7-9288-f2e3851cb0a0
https://github.com/user-attachments/assets/d94b608f-9c22-45e2-8111-1acc3755fa9f
Tested with AI assistance (Claude Code). |
|
I've been working on optimizing Astro Bot and AnyPS5. On windows, Ultra 9 285K and RTX 5080, I'm reaching about 11-12fps on the desert first level. I need to update the project with the major perf improvements i've found. Will update this asap. |
Update 2026-10-10: Crash Site from 3.2 to about 11 FPS on Windows with an NVIDIA GPUSame host as before: Windows 11, Core Ultra 9 285K, RTX 5080, PPSA21564, What worked, one mechanism each:
Measured dead ends, in case they save someone time:
@SP4C3B4R-8, to avoid overlap, these are the areas. The delta copy-back and the VRAM snapshots touch the files #639 and the current ports rewrite, so for now they stay as numbers; happy to share the diff if useful. @QDarkRex, were your 17 to 19 FPS at the Crash Site arrival point? If so, the same branch with no settings runs twice as fast on your Radeon as our stack on NVIDIA, and the imported-memory read cost would be specific to NVIDIA on Windows. Tested and written with AI assistance (Claude Code). |

Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Tested
--windows --to-intel.Status
Not in game on Windows. With the stacks below, the title resolves all its imports, reaches
main(), loads fonts and configuration, finishes PlayGo, creates the Vulkan device, compiles shaders and submits frames. It then stops within about 10 seconds on blocker A or B. The updates below give the cause of both, and the first frames on screen.On Linux, the public
astrobotintegration branch reports the tutorial as playable.Open blockers
shared mapping outside the guest arenaSceSndzAudioOutMainNo open issue or pull request mentioned either of them when this was posted. Not established at that point: whether they are specific to Windows, to Intel hosts or to this version of the title.
A and B as first observed
A. Thrown at
core/libs/prx/libc/src/GuestArena.cpp:177, before any output on stdout; the process aborts. The check is onmain. Not established: whether #286, which touchesGuestArena, plays a part (without it the title stops earlier, so the two could not be compared). Hypothesis: an address space layout that varies between launches, like theVirtualAlloc fixed failedof stack 1.B. Always in the same loop: an audio mixing loop in the title's own code, with no system function in the captured call stack. The loop takes its destination pointers from a table on the stack, and that table does not hold valid buffer addresses. Seen once as a write to
0x7ce1aeb3d7, once as a read at0x4000. Not established: who fills that table. Candidates: the audio libraries the thread calls earlier (libSceAudioOut, libSceAudioPropagation, libSceAcm, libSceAjm) returning wrong buffers, or, unverified, a wrong result from one of the 94EXTRQinstructions the relinker replaced with stubs for--to-intel.Stops on
mainbefore A and B, in the order they are hitFailed to load module libSceJpegDec.prxsceKernelSyncOnAddressWait,sceKernelSyncOnAddressWake, imported by thelibc.prx(SDK 9) shipped with the titlesceKernelMapNamedFlexibleMemoryInternal, samelibc.prxNon-fixed mapping address hints are not implementedStacks tried
main+ feat(libSceJpegDec): implement JPEG decoding #296, feat(libs): declare missing AudioPropagation, AudioIn, NpSessionSignaling and dialog exports #289, feat(libkernel): support mapping address hints and no-overwrite fixed mappings on Windows #286, fix(videoout): read the open param as the service thread's priority and affinity #297, feat(libScePad): implement scePadSetTiltCorrectionState #287, feat(libSceAgcDriver): back the global data share with driver-owned guest memory #285, feat(libkernel): implement sceKernelSyncOnAddressWait and sceKernelSyncOnAddressWake #349, and a local stand-in for the function feat(libkernel): implement sceKernelMapNamedFlexibleMemoryInternal #355 now implements. The title starts drawing. Three runs gave three different stops:sceAudioPropagationSystemQueryMemory not implementedcompute shader ...: AGC driver: guest memory is not readable at 0x30VirtualAlloc fixed failed(core/libs/prx/libkernel/DirectMemory/DirectMemory.cpp:95), seen once, after adding feat(libSceAgcDriver): back the global data share with driver-owned guest memory #285astrobotbranch (eeda258c, 253 commits ahead ofmain) + feat(libSceJpegDec): implement JPEG decoding #296, feat(libkernel): implement sceKernelSyncOnAddressWait and sceKernelSyncOnAddressWake #349, feat(libkernel): implement sceKernelMapNamedFlexibleMemoryInternal #355, feat(libkernel): support mapping address hints and no-overwrite fixed mappings on Windows #286 and feat(libs): declare missing AudioPropagation, AudioIn, NpSessionSignaling and dialog exports #289 without its libSceAudioPropagation file. This goes furthest and gives blockers A and B.All reactions