otimizaçãos round 5

otimizaçãos round 5
This commit is contained in:
Jessica_Natalia
2026-08-16 21:33:01 -03:00
parent 44e4b0183b
commit 3f59905baa
27 changed files with 19362 additions and 1339 deletions
@@ -0,0 +1,235 @@
# VCS Tier-2 SUPERBLOCK V2 COMPLETE — Handoff
Date: 2026-08-16
Base: `PSPRecomp-VCS-TIER2-SUPERBLOCK-V1-PYFIX-2026-08-16`
Target: VCSNative / PSPRecomp / Windows DirectX 12
## Purpose
V1 proved that scheduler-safe second-layer AOT can boot and run, but a single 0154↔0155 trace was too small to move whole-frame performance. V2 replaces that experiment with a profile-guided multi-cluster second layer covering the dominant guest/AOT hotspots measured by `GUESTHOT_UNIT` / `GUESTHOT_PC`.
This is not the rejected GPR cache A/B and not the neutral leaf-inline A/B. The original generated AOT remains the semantic fallback. V2 cold-splits measured hot control-flow closures into smaller host translation units and fuses selected cross-unit direct calls/tail transfers inside those closures while preserving Runtime chain-depth and starvation/scheduler accounting.
## Final V2 coverage
Static generated second-layer corpus:
- 7 profile-guided clusters
- 1,060 selected hot basic blocks
- 17,660 generated cluster source lines (generator-reported)
- 49 static cross-unit/direct call sites fused
- 13 static tail-transfer sites fused
- 25 entry hooks across 10 generated units
- 9 explicit cold exits in the Entity cluster
- all cold/external paths retain the original AOT fallback
Clusters:
| Bit | Mask | Cluster | Units / purpose | Blocks | Fused calls | Fused tails | Hooks |
|---:|---:|---|---|---:|---:|---:|---:|
| 0 | `0x01` | Entity | 0152/0153 accessors + 0154/0155 entity loop | 94 | 13 | 8 | 5 |
| 1 | `0x02` | Geometry | 0084/0085 geometry/collision | 379 | 33 | 0 | 4 |
| 2 | `0x04` | Boundary | 0085/0086 boundary loop | 62 | 0 | 4 | 5 |
| 3 | `0x08` | Matrix | 0044 matrix/VFPU hot function | 73 | 0 | 0 | 1 |
| 4 | `0x10` | Physics | 0129 transform/physics hot function | 15 | 0 | 0 | 1 |
| 5 | `0x20` | World | 0157/0158 world/streaming hot functions | 402 | 1 | 0 | 6 |
| 6 | `0x40` | Edge43 | 0043/0044 boundary | 35 | 2 | 1 | 3 |
| | `0x7F` | **ALL** | | **1,060** | **49** | **13** | **25** |
The 0157 side was added in the final pass because unit 0157 was the sixth-largest unit hotspot in the original profiler. Its measured roots are now in the same World cluster as 0158 rather than leaving half of that hot pair in first-layer AOT.
## Runtime architecture
### Multi-cluster hot CFGs
`profiles/vcs/tools/build_tier2_superblocks.py` parses the generated AOT blocks, computes bounded local control-flow closures from measured PC roots, emits cluster-local labels, and patches only the measured entry PCs in the original units.
Each enabled hook calls a dedicated second-layer function. Local control flow stays in that function. Selected `invoke_chained_direct` calls whose target is inside the same cluster become local second-layer transfers.
### Scheduler / chain-depth preservation
Removed native `invoke_chained_direct` stack frames are retained as logical Runtime frames through the existing scheduler-safe Tier-2 helpers. Fused calls/tails therefore continue to obey chain-depth limits and starvation boundaries. If an external/cold path is required, pending logical frames are completed before returning to normal AOT.
### Original AOT fallback
The original generated unit body remains directly below each hook. Disabling a cluster immediately restores first-layer execution at that PC without rebuilding. Cold exits with known original entry IDs resume the original generated entry function rather than entering a generic PC dispatcher.
## Runtime switches (no rebuild)
Global off:
```bat
set PSPRECOMP_TIER2_SUPERBLOCKS=0
```
Global on/default:
```bat
set PSPRECOMP_TIER2_SUPERBLOCKS=1
```
Cluster mask (default all = `0x7F`):
```bat
set PSPRECOMP_TIER2_CLUSTER_MASK=0x7F
```
Useful masks:
- Entity `0x01`
- Geometry `0x02`
- Boundary `0x04`
- Matrix `0x08`
- Physics `0x10`
- World `0x20`
- Edge43 `0x40`
Masks can be combined with bitwise OR, e.g. Entity+Geometry = `0x03`.
The mask is read once per process. Restart VCSNative after changing it.
## Coverage telemetry
V2 fixes a major blind spot from V1: every PERF interval now emits a `TIER2` line.
Header expected:
```text
stage=tier2-superblock-v2-complete-2026-08-16
perf_telemetry=1 interval_vblanks=60
tier2_superblocks=1 version=2 clusters=7 mask=0x7f hot_blocks=1060 static_fused_calls=49 static_fused_tail=13 hooks=25
```
Per PERF window, expect fields such as:
```text
TIER2 window=60 vblank=... entity_e=... entity_x=... geometry_e=... geometry_x=... ... total_entries=... fused_tail=... fused_calls=... cold=... fallback=... sampled=... sampled_us=...
```
`*_e` = dynamic second-layer entries in that 60-vblank window.
`*_x` = dynamic fused call + tail edges actually executed, not static site count.
`sampled_us` is low-overhead 1/256 sampled inclusive cluster time. It is for coverage/ranking, not an exclusive CPU-time sum.
Counters are thread-local to avoid atomics on hot edges. If a future runtime changes Allegrex execution/reporting to different threads and the TIER2 lines unexpectedly remain zero, the counters must be moved to runtime-owned or atomic aggregation.
## Build integration
`BUILD_VCS_NINJA.bat` now:
1. runs future-timestamp repair;
2. reapplies BOOTFIX-safe transforms only;
3. detects Python 3 using `py -3`, then `python`;
4. regenerates V2 clusters idempotently;
5. invalidates only V2 hook/cluster objects on the first V2 revision build;
6. configures/builds VCSNative and the normal regression/probe targets.
The first V2 revision invalidates units:
`0043 0044 0084 0085 0086 0129 0154 0155 0157 0158`
No manual `.obj` deletion is required.
MSVC cluster flags are kept independent from the giant generated corpus:
- `/O2`
- `/Ob3`
- `/bigobj`
- `/GL-`
A final CMake audit fixed an option-property overwrite: host `/Ob3` is now appended instead of replacing the cluster `/O2 /bigobj /GL-` options.
## Generator correctness / idempotence
Final second execution reports:
```text
host_changed=0
hook_units_changed=0
blocks=1060
lines=17660
fused_calls=49
fused_tail=13
cold_exits=9
```
CFG audit:
- 7/7 cluster files: every `goto SB_*` resolves to a label in the same TU
- zero raw `goto LOCAL_DISPATCH` remains in cluster files
- 25 V2 hooks found across exactly 10 generated units
- zero V1 hook markers remain in generated units
BOOTFIX-safe pass on the final generated corpus with `--check`:
```text
files=0
caps_2048_to_256=0
unaligned_fastview=0
fpu_set_inline=0
fpu_get_inline=0
hot VFPU changed=0
local_cap_2048=0 tail_chain=0 gpr_cache=0 fpr_cache=0
```
So the real Windows build order is idempotent.
## Validation performed
Final tests:
- `psprecomp_tests`: PASS
- `vcs_profile_tests`: PASS
- `vfpu_tier2_tests`: PASS
`vcs_profile_tests` now explicitly tests Tier-2 multi-cluster counter snapshot/reset semantics in addition to the existing scheduler/callback/framebuffer suite.
Actual VCSNative-target object compilation (not only syntax checking):
- 7/7 V2 cluster `.cpp` objects compiled
- 10/10 hooked generated-unit objects compiled: 0043, 0044, 0084, 0085, 0086, 0129, 0154, 0155, 0157, 0158
A full Linux `VCSNative` link was attempted, but `runtime.hpp`/CMake dependency changes made Ninja schedule ~250 generated/host objects and the container execution window interrupted the whole-target build after host compilation had begun. No source/compiler diagnostic was emitted before interruption. Do **not** interpret that as a completed full link. Windows/Ninja remains the authoritative complete build for this profile.
## What is deliberately NOT in V2
The following previously failed or neutral experiments are not reintroduced:
- global/local-cap 2048 change
- EXTREME cross-unit tail-chain semantics
- aggressive GPR/FPR architectural-state promotion
- OPT1 dirty GPR block cache
- GPR read-only A/B
- leaf-inline A/B
- forced GE async / forced vertex-parallel defaults
V2 obtains its structural gain by hot CFG extraction and call/tail fusion while leaving architectural state semantics on the known-good SAFE path.
## Windows test procedure
Apply the V2 overlay over the current V1 PYFIX tree and run:
```bat
profiles\vcs\BUILD_VCS_NINJA.bat
```
Do not delete `.obj` manually.
Then play the same heavy city route for roughly 3–5 minutes and provide the new `VCSNative.log`.
The next analysis should compare both `PERF` and `TIER2` lines. Most important questions:
1. How many dynamic entries/edges does each cluster actually cover?
2. Does `guest_cpu_us_avg` fall for comparable draw/texture/transfer loads?
3. Which cluster has the best sampled cost-to-entry ratio?
4. If performance or correctness regresses, can the offending cluster be isolated immediately by `PSPRECOMP_TIER2_CLUSTER_MASK`?
## Status / progress estimate
- Overall VCS recomp project: ~88%
- Tier-2 performance work: ~86%
- Second-layer architecture/implementation: ~85%
- Performance success against the 160–200 FPS aspiration: **not yet verified**; real Windows gameplay data is still required.
The implementation is complete enough to test as a multi-cluster second layer, but the optimization target is not considered achieved until the runtime logs demonstrate a material reduction in guest CPU time.
@@ -0,0 +1,153 @@
{
"release": "PSPRecomp-VCS-TIER2-SUPERBLOCK-V2-COMPLETE-2026-08-16",
"base": "PSPRecomp-VCS-TIER2-SUPERBLOCK-V1-PYFIX-2026-08-16",
"architecture": {
"clusters": 7,
"hot_blocks": 1060,
"generated_cluster_lines": 17660,
"static_fused_calls": 49,
"static_fused_tail_edges": 13,
"cold_exits": 9,
"hooks": 25,
"hooked_units": [
"0043",
"0044",
"0084",
"0085",
"0086",
"0129",
"0154",
"0155",
"0157",
"0158"
],
"cluster_mask_all": "0x7f",
"global_kill_switch": "PSPRECOMP_TIER2_SUPERBLOCKS=0",
"cluster_mask_env": "PSPRECOMP_TIER2_CLUSTER_MASK"
},
"clusters": {
"entity": {
"mask": "0x01",
"units": [
"0152",
"0153",
"0154",
"0155"
],
"blocks": 94,
"fused_calls": 13,
"fused_tail": 8,
"hooks": 5
},
"geometry": {
"mask": "0x02",
"units": [
"0084",
"0085"
],
"blocks": 379,
"fused_calls": 33,
"fused_tail": 0,
"hooks": 4
},
"boundary": {
"mask": "0x04",
"units": [
"0085",
"0086"
],
"blocks": 62,
"fused_calls": 0,
"fused_tail": 4,
"hooks": 5
},
"matrix": {
"mask": "0x08",
"units": [
"0044"
],
"blocks": 73,
"fused_calls": 0,
"fused_tail": 0,
"hooks": 1
},
"physics": {
"mask": "0x10",
"units": [
"0129"
],
"blocks": 15,
"fused_calls": 0,
"fused_tail": 0,
"hooks": 1
},
"world": {
"mask": "0x20",
"units": [
"0157",
"0158"
],
"blocks": 402,
"fused_calls": 1,
"fused_tail": 0,
"hooks": 6
},
"edge43": {
"mask": "0x40",
"units": [
"0043",
"0044"
],
"blocks": 35,
"fused_calls": 2,
"fused_tail": 1,
"hooks": 3
}
},
"telemetry": {
"perf_interval_vblanks": 60,
"tier2_line_per_perf_window": true,
"entry_timing_sample_stride": 256,
"fields": [
"cluster entries",
"cluster fused edges",
"total_entries",
"fused_tail",
"fused_calls",
"cold",
"fallback",
"sampled",
"sampled_us"
]
},
"validation": {
"psprecomp_tests": "PASS",
"vcs_profile_tests": "PASS incl. Tier2 counter snapshot/reset",
"vfpu_tier2_tests": "PASS",
"cluster_object_compile": "7/7 PASS with actual VCSNative target flags",
"hooked_unit_object_compile": "10/10 PASS with actual VCSNative target flags",
"generator_idempotence": "PASS: host_changed=0, hook_units_changed=0",
"cfg_audit": "PASS: zero unresolved SB gotos, zero raw LOCAL_DISPATCH",
"bootfix_check": "PASS: zero changes",
"full_linux_vcsnative_link": "NOT COMPLETED: whole target rebuild exceeded container execution window; no compiler/source diagnostic before interruption"
},
"msvc": {
"cluster_options": [
"/O2",
"/Ob3",
"/bigobj",
"/GL-"
],
"cmake_property_overwrite_fix": true,
"python_detection": [
"py -3",
"python"
]
},
"performance_status": "UNMEASURED on Windows for V2; runtime log required",
"progress_estimate": {
"overall_project_percent": 88,
"tier2_percent": 86,
"second_layer_architecture_percent": 85
}
}
+29 -11
View File
@@ -151,21 +151,36 @@ if(CMAKE_CXX_COMPILER_ID MATCHES "GNU|Clang")
set_source_files_properties("${VCS_PROFILE_DIR}/host/vcs_profile.cpp" PROPERTIES COMPILE_OPTIONS "-O2")
endif()
# Tier-2 V1 is deliberately a compact hot-region superblock (~1.1k lines),
# not a full-unit mega-TU. Give this host-side trace normal aggressive host
# optimization while keeping the 234 generated AOT units on their measured
# per-source policy above.
# Tier-2 V2 keeps each profile-guided cluster in its own translation unit.
# This is the actual second AOT layer: hot functions/loops are cold-split from
# the giant generated units and selected cross-unit edges are fused locally.
# /GL- on these sources avoids re-inflating them into a giant LTCG partition;
# their hot paths are already visible to the normal /O2+/Ob3 optimizer.
set(VCS_TIER2_CLUSTER_SOURCES
host/vcs_tier2_cluster_entity.cpp
host/vcs_tier2_cluster_geometry.cpp
host/vcs_tier2_cluster_boundary.cpp
host/vcs_tier2_cluster_matrix.cpp
host/vcs_tier2_cluster_physics.cpp
host/vcs_tier2_cluster_world.cpp
host/vcs_tier2_cluster_edge43.cpp
)
if(MSVC)
set_source_files_properties("${VCS_PROFILE_DIR}/host/vcs_tier2_superblocks.cpp" PROPERTIES
COMPILE_OPTIONS "/O2;/Ob3;/bigobj")
set_source_files_properties(
"${VCS_PROFILE_DIR}/host/vcs_tier2_superblocks.cpp"
${VCS_TIER2_CLUSTER_SOURCES}
PROPERTIES COMPILE_OPTIONS "/O2;/Ob3;/bigobj;/GL-")
elseif(CMAKE_CXX_COMPILER_ID MATCHES "GNU|Clang")
set_source_files_properties("${VCS_PROFILE_DIR}/host/vcs_tier2_superblocks.cpp" PROPERTIES
COMPILE_OPTIONS "-O3;-g0")
set_source_files_properties(
"${VCS_PROFILE_DIR}/host/vcs_tier2_superblocks.cpp"
${VCS_TIER2_CLUSTER_SOURCES}
PROPERTIES COMPILE_OPTIONS "-O3;-g0")
endif()
set(VCS_GPU_BACKEND_SOURCE host/ge_gpu_backend_dx12.cpp)
set(VCS_HDR_POST_SOURCE host/vcs_hdr_post_dx12_stub.cpp)
set(VCS_HOST_SOURCES
host/vcs_tier2_superblocks.cpp
host/vcs_bootstrap_paths.cpp
host/vcs_profile.cpp
host/vcs_native_fast_paths.cpp
@@ -234,7 +249,7 @@ endif()
add_executable(VCSNative
${VCS_APP_RESOURCES}
host/main.cpp
host/vcs_tier2_superblocks.cpp
${VCS_TIER2_CLUSTER_SOURCES}
${VCS_HOST_SOURCES}
${VCS_GENERATED}
)
@@ -248,8 +263,11 @@ if(MSVC)
# (MSVC reports this as "D9025: overriding '/Ob0' with '/Ob3'"). Apply it to
# the host sources only.
target_compile_options(VCSNative PRIVATE ${PSPRECOMP_MSVC_MP_FLAG})
set_source_files_properties(host/main.cpp host/vcs_tier2_superblocks.cpp ${VCS_HOST_SOURCES}
DIRECTORY "${VCS_PROFILE_DIR}" PROPERTIES COMPILE_OPTIONS
# APPEND is required here: Tier-2 cluster sources already carry /O2,
# /bigobj and /GL-. Replacing COMPILE_OPTIONS would silently drop those
# flags on MSVC and make the Windows build differ from the validated one.
set_property(SOURCE host/main.cpp ${VCS_TIER2_CLUSTER_SOURCES} ${VCS_HOST_SOURCES}
DIRECTORY "${VCS_PROFILE_DIR}" APPEND PROPERTY COMPILE_OPTIONS
"$<$<CONFIG:Release>:/Ob3>;$<$<CONFIG:RelWithDebInfo>:/Ob3>")
target_link_options(VCSNative PRIVATE /STACK:67108864)
if(PSPRECOMP_LTO)
@@ -1,5 +1,6 @@
#include "psprecomp/runtime.hpp"
#include "generated_units.hpp"
#include "vcs_tier2_superblocks.hpp"
#include <bit>
#include <cmath>
#include <cstdint>
@@ -9480,6 +9481,12 @@ L_088B3FC8:
ctx.gpr[30] = (0u | 0u);
goto L_088B3FCC;
L_088B3FCC:
// TIER2_SUPERBLOCK_V2_HOOK_BEGIN
if (vcs::tier2_cluster_enabled(vcs::Tier2ClusterId::Edge43)) {
vcs::tier2_superblock_edge43(rt, ctx, aot_mem, 0x088B3FCCu);
return;
}
// TIER2_SUPERBLOCK_V2_HOOK_END
ctx.gpr[18] = (aot_mem.aot_load32(ctx.gpr[19] + static_cast<std::uint32_t>(68)));
ctx.gpr[18] = (ctx.gpr[18] + ctx.gpr[30]);
ctx.gpr[5] = (ctx.gpr[29] + static_cast<std::uint32_t>(16));
@@ -1,5 +1,6 @@
#include "psprecomp/runtime.hpp"
#include "generated_units.hpp"
#include "vcs_tier2_superblocks.hpp"
#include <bit>
#include <cmath>
#include <cstdint>
@@ -877,6 +878,12 @@ LOCAL_DISPATCH:
}
}
L_088B4004:
// TIER2_SUPERBLOCK_V2_HOOK_BEGIN
if (vcs::tier2_cluster_enabled(vcs::Tier2ClusterId::Edge43)) {
vcs::tier2_superblock_edge43(rt, ctx, aot_mem, 0x088B4004u);
return;
}
// TIER2_SUPERBLOCK_V2_HOOK_END
ctx.gpr[17] = (static_cast<std::uint32_t>(static_cast<std::int32_t>(static_cast<std::int16_t>(aot_mem.aot_load16(ctx.gpr[18] + static_cast<std::uint32_t>(6))))));
ctx.gpr[4] = (static_cast<std::uint32_t>(static_cast<std::int32_t>(static_cast<std::int16_t>(aot_mem.aot_load16(ctx.gpr[18] + static_cast<std::uint32_t>(14))))));
ctx.gpr[4] = (static_cast<std::int32_t>(ctx.gpr[4]) < static_cast<std::int32_t>(ctx.gpr[17]) ? 1u : 0u);
@@ -979,6 +986,12 @@ L_088B40F4:
ctx.gpr[4] = (aot_mem.aot_load32(ctx.gpr[29] + static_cast<std::uint32_t>(48)));
goto L_088B40F8;
L_088B40F8:
// TIER2_SUPERBLOCK_V2_HOOK_BEGIN
if (vcs::tier2_cluster_enabled(vcs::Tier2ClusterId::Edge43)) {
vcs::tier2_superblock_edge43(rt, ctx, aot_mem, 0x088B40F8u);
return;
}
// TIER2_SUPERBLOCK_V2_HOOK_END
ctx.gpr[4] = (ctx.gpr[4] + static_cast<std::uint32_t>(1));
ctx.gpr[5] = (aot_mem.aot_load32(ctx.gpr[29] + static_cast<std::uint32_t>(52)));
ctx.gpr[30] = (ctx.gpr[30] + static_cast<std::uint32_t>(16));
@@ -1819,6 +1832,12 @@ L_088B4708:
ctx.pc = jump_target;
return;
L_088B4738:
// TIER2_SUPERBLOCK_V2_HOOK_BEGIN
if (vcs::tier2_cluster_enabled(vcs::Tier2ClusterId::Matrix)) {
vcs::tier2_superblock_matrix(rt, ctx, aot_mem, 0x088B4738u);
return;
}
// TIER2_SUPERBLOCK_V2_HOOK_END
ctx.gpr[29] = (ctx.gpr[29] + static_cast<std::uint32_t>(-352));
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(296), std::bit_cast<std::uint32_t>(ctx.fpr[20]));
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(300), ctx.gpr[16]);
@@ -1,5 +1,6 @@
#include "psprecomp/runtime.hpp"
#include "generated_units.hpp"
#include "vcs_tier2_superblocks.hpp"
#include <bit>
#include <cmath>
#include <cstdint>
@@ -3649,6 +3650,12 @@ L_0895543C:
ctx.pc = jump_target;
return;
L_08955444:
// TIER2_SUPERBLOCK_V2_HOOK_BEGIN
if (vcs::tier2_cluster_enabled(vcs::Tier2ClusterId::Geometry)) {
vcs::tier2_superblock_geometry(rt, ctx, aot_mem, 0x08955444u);
return;
}
// TIER2_SUPERBLOCK_V2_HOOK_END
ctx.gpr[9] = (aot_mem.aot_load32(ctx.gpr[4] + static_cast<std::uint32_t>(11768)));
ctx.set_vfpu_scalar_bits_ct<12u>(aot_mem.aot_load32(ctx.gpr[5] + static_cast<std::uint32_t>(4)));
ctx.set_vfpu_scalar_bits_ct<44u>(aot_mem.aot_load32(ctx.gpr[5] + static_cast<std::uint32_t>(8)));
@@ -3716,6 +3723,12 @@ L_089554A0:
ctx.pc = jump_target;
return;
L_089554A8:
// TIER2_SUPERBLOCK_V2_HOOK_BEGIN
if (vcs::tier2_cluster_enabled(vcs::Tier2ClusterId::Geometry)) {
vcs::tier2_superblock_geometry(rt, ctx, aot_mem, 0x089554A8u);
return;
}
// TIER2_SUPERBLOCK_V2_HOOK_END
ctx.gpr[29] = (ctx.gpr[29] + static_cast<std::uint32_t>(-64));
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(12), std::bit_cast<std::uint32_t>(ctx.fpr[20]));
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(16), ctx.gpr[16]);
@@ -1,5 +1,6 @@
#include "psprecomp/runtime.hpp"
#include "generated_units.hpp"
#include "vcs_tier2_superblocks.hpp"
#include <bit>
#include <cmath>
#include <cstdint>
@@ -2460,6 +2461,12 @@ L_08958CF8:
ctx.pc = jump_target;
return;
L_08958D28:
// TIER2_SUPERBLOCK_V2_HOOK_BEGIN
if (vcs::tier2_cluster_enabled(vcs::Tier2ClusterId::Geometry)) {
vcs::tier2_superblock_geometry(rt, ctx, aot_mem, 0x08958D28u);
return;
}
// TIER2_SUPERBLOCK_V2_HOOK_END
ctx.gpr[29] = (ctx.gpr[29] + static_cast<std::uint32_t>(-624));
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(576), std::bit_cast<std::uint32_t>(ctx.fpr[20]));
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(580), ctx.gpr[16]);
@@ -5787,6 +5794,12 @@ L_0895A72C:
ctx.pc = jump_target;
return;
L_0895A760:
// TIER2_SUPERBLOCK_V2_HOOK_BEGIN
if (vcs::tier2_cluster_enabled(vcs::Tier2ClusterId::Geometry)) {
vcs::tier2_superblock_geometry(rt, ctx, aot_mem, 0x0895A760u);
return;
}
// TIER2_SUPERBLOCK_V2_HOOK_END
ctx.gpr[29] = (ctx.gpr[29] + static_cast<std::uint32_t>(-144));
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(72), std::bit_cast<std::uint32_t>(ctx.fpr[20]));
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(76), std::bit_cast<std::uint32_t>(ctx.fpr[22]));
@@ -8395,6 +8408,12 @@ L_0895BF04:
aot_mem.aot_store32(ctx.gpr[17] + static_cast<std::uint32_t>(29552), ctx.gpr[4]);
goto L_0895BF3C;
L_0895BF3C:
// TIER2_SUPERBLOCK_V2_HOOK_BEGIN
if (vcs::tier2_cluster_enabled(vcs::Tier2ClusterId::Boundary)) {
vcs::tier2_superblock_boundary(rt, ctx, aot_mem, 0x0895BF3Cu);
return;
}
// TIER2_SUPERBLOCK_V2_HOOK_END
ctx.gpr[4] = (0u | 0u);
aot_mem.aot_store8(ctx.gpr[29] + static_cast<std::uint32_t>(120), static_cast<std::uint8_t>(ctx.gpr[4]));
ctx.gpr[4] = (aot_mem.aot_load32(ctx.gpr[29] + static_cast<std::uint32_t>(100)));
@@ -8419,6 +8438,12 @@ L_0895BF5C:
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(108), ctx.gpr[4]);
goto L_0895BF7C;
L_0895BF7C:
// TIER2_SUPERBLOCK_V2_HOOK_BEGIN
if (vcs::tier2_cluster_enabled(vcs::Tier2ClusterId::Boundary)) {
vcs::tier2_superblock_boundary(rt, ctx, aot_mem, 0x0895BF7Cu);
return;
}
// TIER2_SUPERBLOCK_V2_HOOK_END
ctx.gpr[18] = (ctx.gpr[20] | 0u);
{ const bool branch_taken = ctx.gpr[30] == 0u;
ctx.gpr[4] = (15u << 16u);
@@ -1,5 +1,6 @@
#include "psprecomp/runtime.hpp"
#include "generated_units.hpp"
#include "vcs_tier2_superblocks.hpp"
#include <bit>
#include <cmath>
#include <cstdint>
@@ -759,6 +760,12 @@ LOCAL_DISPATCH:
}
}
L_0895C000:
// TIER2_SUPERBLOCK_V2_HOOK_BEGIN
if (vcs::tier2_cluster_enabled(vcs::Tier2ClusterId::Boundary)) {
vcs::tier2_superblock_boundary(rt, ctx, aot_mem, 0x0895C000u);
return;
}
// TIER2_SUPERBLOCK_V2_HOOK_END
ctx.gpr[5] = (aot_mem.aot_load32(ctx.gpr[17] + static_cast<std::uint32_t>(29552)));
ctx.gpr[5] = (ctx.gpr[5] + static_cast<std::uint32_t>(4));
aot_mem.aot_store32(ctx.gpr[17] + static_cast<std::uint32_t>(29552), ctx.gpr[5]);
@@ -772,6 +779,12 @@ L_0895C000:
goto L_0895C01C;
}
L_0895C01C:
// TIER2_SUPERBLOCK_V2_HOOK_BEGIN
if (vcs::tier2_cluster_enabled(vcs::Tier2ClusterId::Boundary)) {
vcs::tier2_superblock_boundary(rt, ctx, aot_mem, 0x0895C01Cu);
return;
}
// TIER2_SUPERBLOCK_V2_HOOK_END
ctx.gpr[4] = (aot_mem.aot_load16(ctx.gpr[18] + static_cast<std::uint32_t>(6)));
ctx.gpr[5] = (aot_mem.aot_load32(ctx.gpr[29] + static_cast<std::uint32_t>(92)));
ctx.gpr[19] = (0u | 0u);
@@ -1022,6 +1035,12 @@ L_0895C28C:
ctx.gpr[7] = (aot_mem.aot_load32(ctx.gpr[8] + static_cast<std::uint32_t>(8)));
goto L_0895C2A8;
L_0895C2A8:
// TIER2_SUPERBLOCK_V2_HOOK_BEGIN
if (vcs::tier2_cluster_enabled(vcs::Tier2ClusterId::Boundary)) {
vcs::tier2_superblock_boundary(rt, ctx, aot_mem, 0x0895C2A8u);
return;
}
// TIER2_SUPERBLOCK_V2_HOOK_END
ctx.gpr[4] = (aot_mem.aot_load32(ctx.gpr[29] + static_cast<std::uint32_t>(104)));
ctx.gpr[4] = (ctx.gpr[4] + static_cast<std::uint32_t>(1));
ctx.gpr[20] = (ctx.gpr[20] + static_cast<std::uint32_t>(48));
@@ -1,5 +1,6 @@
#include "psprecomp/runtime.hpp"
#include "generated_units.hpp"
#include "vcs_tier2_superblocks.hpp"
#include <bit>
#include <cmath>
#include <cstdint>
@@ -3859,6 +3860,12 @@ L_08A09AD0:
ctx.pc = jump_target;
return;
L_08A09B2C:
// TIER2_SUPERBLOCK_V2_HOOK_BEGIN
if (vcs::tier2_cluster_enabled(vcs::Tier2ClusterId::Physics)) {
vcs::tier2_superblock_physics(rt, ctx, aot_mem, 0x08A09B2Cu);
return;
}
// TIER2_SUPERBLOCK_V2_HOOK_END
ctx.gpr[29] = (ctx.gpr[29] + static_cast<std::uint32_t>(-16));
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(0), ctx.gpr[16]);
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(4), ctx.gpr[17]);
@@ -6792,12 +6792,12 @@ L_08A6E828:
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(1584), ctx.gpr[6]);
goto L_08A6E894;
L_08A6E894:
// TIER2_SUPERBLOCK_V1_HOOK_BEGIN
if (vcs::tier2_superblocks_enabled()) {
vcs::tier2_superblock_154_155(rt, ctx, aot_mem, 0x08A6E894u);
// TIER2_SUPERBLOCK_V2_HOOK_BEGIN
if (vcs::tier2_cluster_enabled(vcs::Tier2ClusterId::Entity)) {
vcs::tier2_superblock_entity(rt, ctx, aot_mem, 0x08A6E894u);
return;
}
// TIER2_SUPERBLOCK_V1_HOOK_END
// TIER2_SUPERBLOCK_V2_HOOK_END
ctx.gpr[4] = (aot_mem.aot_load32(ctx.gpr[29] + static_cast<std::uint32_t>(1600)));
ctx.gpr[4] = (aot_mem.aot_load32(ctx.gpr[4] + static_cast<std::uint32_t>(0)));
{ const bool branch_taken = ctx.gpr[4] == 0u;
@@ -6808,12 +6808,12 @@ L_08A6E894:
goto L_08A6E8A4;
}
L_08A6E8A4:
// TIER2_SUPERBLOCK_V1_HOOK_BEGIN
if (vcs::tier2_superblocks_enabled()) {
vcs::tier2_superblock_154_155(rt, ctx, aot_mem, 0x08A6E8A4u);
// TIER2_SUPERBLOCK_V2_HOOK_BEGIN
if (vcs::tier2_cluster_enabled(vcs::Tier2ClusterId::Entity)) {
vcs::tier2_superblock_entity(rt, ctx, aot_mem, 0x08A6E8A4u);
return;
}
// TIER2_SUPERBLOCK_V1_HOOK_END
// TIER2_SUPERBLOCK_V2_HOOK_END
ctx.gpr[4] = (aot_mem.aot_load32(ctx.gpr[29] + static_cast<std::uint32_t>(1576)));
ctx.gpr[17] = (aot_mem.aot_load32(ctx.gpr[4] + static_cast<std::uint32_t>(0)));
ctx.gpr[4] = (aot_mem.aot_load32(ctx.gpr[4] + static_cast<std::uint32_t>(8)));
+12 -12
View File
@@ -3417,12 +3417,12 @@ L_08A710F8:
goto L_08A71100;
}
L_08A71100:
// TIER2_SUPERBLOCK_V1_HOOK_BEGIN
if (vcs::tier2_superblocks_enabled()) {
vcs::tier2_superblock_154_155(rt, ctx, aot_mem, 0x08A71100u);
// TIER2_SUPERBLOCK_V2_HOOK_BEGIN
if (vcs::tier2_cluster_enabled(vcs::Tier2ClusterId::Entity)) {
vcs::tier2_superblock_entity(rt, ctx, aot_mem, 0x08A71100u);
return;
}
// TIER2_SUPERBLOCK_V1_HOOK_END
// TIER2_SUPERBLOCK_V2_HOOK_END
{ const bool branch_taken = ctx.gpr[4] != 0u;
// nop
if (branch_taken) {
@@ -3596,12 +3596,12 @@ L_08A711D4:
if (rt.invoke_chained_direct<&recomp_unit_0153_entry, 153u, 104u, 0x08A68CDCu>(ctx, &aot_mem) && ctx.pc == 0x08A711E0u) goto L_08A711E0;
return;
L_08A711E0:
// TIER2_SUPERBLOCK_V1_HOOK_BEGIN
if (vcs::tier2_superblocks_enabled()) {
vcs::tier2_superblock_154_155(rt, ctx, aot_mem, 0x08A711E0u);
// TIER2_SUPERBLOCK_V2_HOOK_BEGIN
if (vcs::tier2_cluster_enabled(vcs::Tier2ClusterId::Entity)) {
vcs::tier2_superblock_entity(rt, ctx, aot_mem, 0x08A711E0u);
return;
}
// TIER2_SUPERBLOCK_V1_HOOK_END
// TIER2_SUPERBLOCK_V2_HOOK_END
ctx.gpr[4] = (aot_mem.aot_load32(ctx.gpr[29] + static_cast<std::uint32_t>(1576)));
{ const bool branch_taken = ctx.gpr[4] != 0u;
// nop
@@ -3611,12 +3611,12 @@ L_08A711E0:
goto L_08A711EC;
}
L_08A711EC:
// TIER2_SUPERBLOCK_V1_HOOK_BEGIN
if (vcs::tier2_superblocks_enabled()) {
vcs::tier2_superblock_154_155(rt, ctx, aot_mem, 0x08A711ECu);
// TIER2_SUPERBLOCK_V2_HOOK_BEGIN
if (vcs::tier2_cluster_enabled(vcs::Tier2ClusterId::Entity)) {
vcs::tier2_superblock_entity(rt, ctx, aot_mem, 0x08A711ECu);
return;
}
// TIER2_SUPERBLOCK_V1_HOOK_END
// TIER2_SUPERBLOCK_V2_HOOK_END
ctx.gpr[4] = (aot_mem.aot_load32(ctx.gpr[29] + static_cast<std::uint32_t>(1600)));
ctx.gpr[4] = (ctx.gpr[4] + static_cast<std::uint32_t>(4));
ctx.gpr[5] = (aot_mem.aot_load32(ctx.gpr[29] + static_cast<std::uint32_t>(1596)));
@@ -1,5 +1,6 @@
#include "psprecomp/runtime.hpp"
#include "generated_units.hpp"
#include "vcs_tier2_superblocks.hpp"
#include <bit>
#include <cmath>
#include <cstdint>
@@ -4358,6 +4359,12 @@ L_08A79DEC:
ctx.pc = jump_target;
return;
L_08A79DF8:
// TIER2_SUPERBLOCK_V2_HOOK_BEGIN
if (vcs::tier2_cluster_enabled(vcs::Tier2ClusterId::World)) {
vcs::tier2_superblock_world(rt, ctx, aot_mem, 0x08A79DF8u);
return;
}
// TIER2_SUPERBLOCK_V2_HOOK_END
ctx.gpr[29] = (ctx.gpr[29] + static_cast<std::uint32_t>(-32));
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(4), ctx.gpr[17]);
ctx.gpr[17] = (2238u << 16u);
@@ -5478,6 +5485,12 @@ L_08A7A658:
ctx.pc = jump_target;
return;
L_08A7A664:
// TIER2_SUPERBLOCK_V2_HOOK_BEGIN
if (vcs::tier2_cluster_enabled(vcs::Tier2ClusterId::World)) {
vcs::tier2_superblock_world(rt, ctx, aot_mem, 0x08A7A664u);
return;
}
// TIER2_SUPERBLOCK_V2_HOOK_END
ctx.gpr[29] = (ctx.gpr[29] + static_cast<std::uint32_t>(-32));
ctx.gpr[5] = (ctx.gpr[5] & 255u);
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(0), ctx.gpr[16]);
@@ -5877,6 +5890,12 @@ L_08A7A928:
ctx.pc = jump_target;
return;
L_08A7A94C:
// TIER2_SUPERBLOCK_V2_HOOK_BEGIN
if (vcs::tier2_cluster_enabled(vcs::Tier2ClusterId::World)) {
vcs::tier2_superblock_world(rt, ctx, aot_mem, 0x08A7A94Cu);
return;
}
// TIER2_SUPERBLOCK_V2_HOOK_END
ctx.gpr[29] = (ctx.gpr[29] + static_cast<std::uint32_t>(-80));
ctx.gpr[4] = (2236u << 16u);
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(56), ctx.gpr[20]);
@@ -7468,6 +7487,12 @@ L_08A7B5A0:
ctx.pc = jump_target;
return;
L_08A7B5A8:
// TIER2_SUPERBLOCK_V2_HOOK_BEGIN
if (vcs::tier2_cluster_enabled(vcs::Tier2ClusterId::World)) {
vcs::tier2_superblock_world(rt, ctx, aot_mem, 0x08A7B5A8u);
return;
}
// TIER2_SUPERBLOCK_V2_HOOK_END
jump_target = ctx.gpr[31];
// nop
local_pc = jump_target;
@@ -1,5 +1,6 @@
#include "psprecomp/runtime.hpp"
#include "generated_units.hpp"
#include "vcs_tier2_superblocks.hpp"
#include <bit>
#include <cmath>
#include <cstdint>
@@ -6336,6 +6337,12 @@ L_08A7EC20:
ctx.pc = jump_target;
return;
L_08A7EC68:
// TIER2_SUPERBLOCK_V2_HOOK_BEGIN
if (vcs::tier2_cluster_enabled(vcs::Tier2ClusterId::World)) {
vcs::tier2_superblock_world(rt, ctx, aot_mem, 0x08A7EC68u);
return;
}
// TIER2_SUPERBLOCK_V2_HOOK_END
ctx.gpr[29] = (ctx.gpr[29] + static_cast<std::uint32_t>(-80));
ctx.gpr[5] = (0u | 12u);
ctx.gpr[6] = (aot_mem.aot_load8(ctx.gpr[28] + static_cast<std::uint32_t>(-7080)));
@@ -6884,6 +6891,12 @@ L_08A7F048:
ctx.pc = jump_target;
return;
L_08A7F080:
// TIER2_SUPERBLOCK_V2_HOOK_BEGIN
if (vcs::tier2_cluster_enabled(vcs::Tier2ClusterId::World)) {
vcs::tier2_superblock_world(rt, ctx, aot_mem, 0x08A7F080u);
return;
}
// TIER2_SUPERBLOCK_V2_HOOK_END
ctx.gpr[29] = (ctx.gpr[29] + static_cast<std::uint32_t>(-80));
ctx.gpr[5] = (0u | 12u);
ctx.gpr[6] = (aot_mem.aot_load8(ctx.gpr[28] + static_cast<std::uint32_t>(-7080)));
+39
View File
@@ -14,6 +14,7 @@
#include "savedata_utility_ui.hpp"
#include "vcs_texture_replacement.hpp"
#include "vcs_runtime_log.hpp"
#include "vcs_tier2_superblocks.hpp"
#include "psprecomp/common.hpp"
#include "psprecomp/deflate.hpp"
@@ -7618,6 +7619,44 @@ void install_profile(psprecomp::Runtime &runtime, std::uint32_t user_arena_start
<< " fence_us=" << (d(r.perf_wait_for_frame_ns, o.perf_wait_for_frame_ns) / 1000u)
<< " finish_us=" << (d(r.perf_finish_frame_ns, o.perf_finish_frame_ns) / 1000u);
runtime_log_line(t.str());
// Tier-2 V2 coverage is reported at the same 60-vblank cadence as
// PERF. Counters are thread-local to the Allegrex execution thread,
// so there are no atomics on the hot superblock edges.
if (tier2_superblocks_enabled()) {
const Tier2CountersSnapshot tier2 = consume_tier2_counters();
std::uint64_t total_entries = 0u;
std::uint64_t total_tail = 0u;
std::uint64_t total_calls = 0u;
std::uint64_t total_cold = 0u;
std::uint64_t total_fallback = 0u;
std::uint64_t total_sample_ns = 0u;
std::uint64_t total_sample_entries = 0u;
std::ostringstream tier2_line;
tier2_line << "TIER2 window=" << n << " vblank=" << display_vblank_index;
for (std::size_t i = 0; i < kTier2ClusterCount; ++i) {
const auto id = static_cast<Tier2ClusterId>(i);
const Tier2ClusterCounters &c = tier2.cluster[i];
total_entries += c.entries;
total_tail += c.fused_tail_edges;
total_calls += c.fused_calls;
total_cold += c.cold_exits;
total_fallback += c.fallbacks;
total_sample_ns += c.sampled_ns;
total_sample_entries += c.sampled_entries;
tier2_line << ' ' << tier2_cluster_name(id) << "_e=" << c.entries
<< ' ' << tier2_cluster_name(id) << "_x="
<< (c.fused_tail_edges + c.fused_calls);
}
tier2_line << " total_entries=" << total_entries
<< " fused_tail=" << total_tail
<< " fused_calls=" << total_calls
<< " cold=" << total_cold
<< " fallback=" << total_fallback
<< " sampled=" << total_sample_entries
<< " sampled_us=" << (total_sample_ns / 1000u);
runtime_log_line(tier2_line.str());
}
if (vcs_configuration().diagnostics.guest_hotspot_profile &&
++guest_hotspot_perf_windows >= 5u) {
report_guest_hotspot_window(display_vblank_index);
+5 -2
View File
@@ -1,5 +1,6 @@
#include "vcs_runtime_log.hpp"
#include "vcs_config.hpp"
#include "vcs_tier2_superblocks.hpp"
#include <chrono>
#include <ctime>
@@ -62,7 +63,7 @@ void runtime_log_initialize(const VcsConfiguration &configuration) {
return;
}
s.file << "VCSNative runtime log\n";
s.file << "stage=tier2-superblock-v1-2026-08-16\n";
s.file << "stage=tier2-superblock-v2-unwind-hotfix-2026-08-16\n";
s.file << "config=" << configuration.source_path.string() << '\n';
s.file << "started=" << timestamp_now() << '\n';
s.file << "perf_telemetry=" << (configuration.diagnostics.perf_telemetry ? 1 : 0)
@@ -70,7 +71,9 @@ void runtime_log_initialize(const VcsConfiguration &configuration) {
<< "\n";
s.file << "guest_hotspot=" << (configuration.diagnostics.guest_hotspot_profile ? 1 : 0)
<< " sample_stride=256 interval_vblanks=300\n";
s.file << "tier2_superblocks=1 cluster=0154+0155 hot_region_edges=8 cold_exits=3\n\n";
s.file << "tier2_superblocks=" << (tier2_superblocks_enabled() ? 1 : 0)
<< " version=2 clusters=7 mask=0x" << std::hex << tier2_cluster_mask() << std::dec
<< " hot_blocks=1060 static_fused_calls=49 static_fused_tail=13 hooks=25 unwind_fix=1 reentry_guard=1\n\n";
if (s.flush_every_line) s.file.flush();
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,894 @@
// AUTO-GENERATED by profiles/vcs/tools/build_tier2_superblocks.py.
// Tier-2 SUPERBLOCK V2 cluster: edge43
#include "vcs_tier2_superblocks.hpp"
#include "psprecomp/runtime.hpp"
#include "generated_units.hpp"
#include <cstdint>
namespace vcs {
using namespace psprecomp;
void tier2_superblock_edge43(psprecomp::Runtime &rt,
psprecomp::AllegrexContext &ctx,
psprecomp::GuestMemory::AotFastView &aot_mem,
std::uint32_t entry_pc) {
tier2_detail::SampleScope tier2_scope(Tier2ClusterId::Edge43);
auto &tier2_stats = tier2_scope.stats();
constexpr std::uint32_t kTier2ReturnCapacity = 32u;
std::uint32_t tier2_pending_transfers = 0u;
std::uint32_t tier2_return_depth = 0u;
std::uint32_t tier2_return_pc[kTier2ReturnCapacity]{};
std::uint32_t tier2_return_unit[kTier2ReturnCapacity]{};
std::uint32_t tier2_return_pending_base[kTier2ReturnCapacity]{};
std::uint32_t tier2_resume_pc = 0u;
std::uint32_t tier2_resume_unit = 0u;
std::uint32_t jump_target = 0u;
std::uint32_t local_pc = 0u;
std::uint32_t local_transfers = 0u;
std::uint32_t entry_id = 0u;
#define TIER2_SB_RETURN() do { tier2_scope.finish(); /* Unwind every logical invoke_chained_direct frame in true LIFO order. Tail frames created inside a fused JAL unwind before that JAL; outer JAL frames are also released after context invalidation. */ while (tier2_return_depth != 0u) { std::uint32_t tier2_base_ = tier2_return_pending_base[tier2_return_depth - 1u]; if (tier2_base_ > tier2_pending_transfers) { ++tier2_stats.fallbacks; tier2_base_ = tier2_pending_transfers; } const std::uint32_t tier2_tail_count_ = tier2_pending_transfers - tier2_base_; tier2_pending_transfers = tier2_base_; if (tier2_tail_count_ != 0u) (void)rt.tier2_complete_fused_transfers(ctx, tier2_tail_count_); --tier2_return_depth; (void)rt.tier2_complete_fused_transfers(ctx, 1u); } if (tier2_pending_transfers != 0u) { (void)rt.tier2_complete_fused_transfers(ctx, tier2_pending_transfers); tier2_pending_transfers = 0u; } return; } while (false)
goto TIER2_ENTRY_DISPATCH;
TIER2_FUSED_RETURN_DISPATCH:
switch (tier2_resume_pc) {
case 0x088B1780u: goto SB_L_088B1780;
case 0x088B17B4u: goto SB_L_088B17B4;
case 0x088B180Cu: goto SB_L_088B180C;
case 0x088B18ECu: goto SB_L_088B18EC;
case 0x088B1940u: goto SB_L_088B1940;
case 0x088B1B74u: goto SB_L_088B1B74;
case 0x088B1BB8u: goto SB_L_088B1BB8;
case 0x088B1BC0u: goto SB_L_088B1BC0;
case 0x088B1BC8u: goto SB_L_088B1BC8;
case 0x088B1BD0u: goto SB_L_088B1BD0;
case 0x088B1C2Cu: goto SB_L_088B1C2C;
case 0x088B1C40u: goto SB_L_088B1C40;
case 0x088B3FCCu: goto SB_L_088B3FCC;
case 0x088B4004u: goto SB_L_088B4004;
case 0x088B4018u: goto SB_L_088B4018;
case 0x088B4020u: goto SB_L_088B4020;
case 0x088B4074u: goto SB_L_088B4074;
case 0x088B407Cu: goto SB_L_088B407C;
case 0x088B40C4u: goto SB_L_088B40C4;
case 0x088B40E0u: goto SB_L_088B40E0;
case 0x088B40F4u: goto SB_L_088B40F4;
case 0x088B40F8u: goto SB_L_088B40F8;
case 0x088B4110u: goto SB_L_088B4110;
case 0x088B4118u: goto SB_L_088B4118;
case 0x088B412Cu: goto SB_L_088B412C;
case 0x088B4134u: goto SB_L_088B4134;
case 0x088B4188u: goto SB_L_088B4188;
case 0x088B4190u: goto SB_L_088B4190;
case 0x088B41D8u: goto SB_L_088B41D8;
case 0x088B41F4u: goto SB_L_088B41F4;
case 0x088B4208u: goto SB_L_088B4208;
case 0x088B420Cu: goto SB_L_088B420C;
case 0x088B4224u: goto SB_L_088B4224;
case 0x088B4234u: goto SB_L_088B4234;
case 0x088B4238u: goto SB_L_088B4238;
default: break;
}
switch (tier2_resume_unit) {
default: break;
}
ctx.pc = tier2_resume_pc;
TIER2_SB_RETURN();
TIER2_LOCAL_DISPATCH_U0043:
if (tier2_return_depth != 0u && local_pc == tier2_return_pc[tier2_return_depth - 1u]) {
tier2_resume_pc = local_pc;
tier2_resume_unit = tier2_return_unit[tier2_return_depth - 1u];
// Match invoke_chained_direct(): when the callee has returned, ctx.pc
// already contains the caller continuation before any starvation
// boundary/accounting can run. A scheduler switch here must never see
// the stale callee PC.
ctx.pc = local_pc;
const std::uint32_t tier2_pending_base =
tier2_return_pending_base[tier2_return_depth - 1u];
bool tier2_same_context = true;
if (tier2_pending_transfers < tier2_pending_base) {
++tier2_stats.fallbacks;
ctx.pc = local_pc;
TIER2_SB_RETURN();
}
const std::uint32_t tier2_nested_tail =
tier2_pending_transfers - tier2_pending_base;
if (tier2_nested_tail != 0u) {
tier2_pending_transfers = tier2_pending_base;
if (!rt.tier2_complete_fused_transfers(ctx, tier2_nested_tail))
tier2_same_context = false;
}
--tier2_return_depth;
if (!rt.tier2_complete_fused_transfers(ctx, 1u))
tier2_same_context = false;
if (!tier2_same_context) TIER2_SB_RETURN();
goto TIER2_FUSED_RETURN_DISPATCH;
}
switch (local_pc) {
case 0x088B1780u: goto SB_L_088B1780;
case 0x088B17B4u: goto SB_L_088B17B4;
case 0x088B180Cu: goto SB_L_088B180C;
case 0x088B18ECu: goto SB_L_088B18EC;
case 0x088B1940u: goto SB_L_088B1940;
case 0x088B1B74u: goto SB_L_088B1B74;
case 0x088B1BB8u: goto SB_L_088B1BB8;
case 0x088B1BC0u: goto SB_L_088B1BC0;
case 0x088B1BC8u: goto SB_L_088B1BC8;
case 0x088B1BD0u: goto SB_L_088B1BD0;
case 0x088B1C2Cu: goto SB_L_088B1C2C;
case 0x088B1C40u: goto SB_L_088B1C40;
case 0x088B3FCCu: goto SB_L_088B3FCC;
default:
ctx.pc = local_pc;
TIER2_SB_RETURN();
}
TIER2_LOCAL_DISPATCH_U0044:
if (tier2_return_depth != 0u && local_pc == tier2_return_pc[tier2_return_depth - 1u]) {
tier2_resume_pc = local_pc;
tier2_resume_unit = tier2_return_unit[tier2_return_depth - 1u];
// Match invoke_chained_direct(): when the callee has returned, ctx.pc
// already contains the caller continuation before any starvation
// boundary/accounting can run. A scheduler switch here must never see
// the stale callee PC.
ctx.pc = local_pc;
const std::uint32_t tier2_pending_base =
tier2_return_pending_base[tier2_return_depth - 1u];
bool tier2_same_context = true;
if (tier2_pending_transfers < tier2_pending_base) {
++tier2_stats.fallbacks;
ctx.pc = local_pc;
TIER2_SB_RETURN();
}
const std::uint32_t tier2_nested_tail =
tier2_pending_transfers - tier2_pending_base;
if (tier2_nested_tail != 0u) {
tier2_pending_transfers = tier2_pending_base;
if (!rt.tier2_complete_fused_transfers(ctx, tier2_nested_tail))
tier2_same_context = false;
}
--tier2_return_depth;
if (!rt.tier2_complete_fused_transfers(ctx, 1u))
tier2_same_context = false;
if (!tier2_same_context) TIER2_SB_RETURN();
goto TIER2_FUSED_RETURN_DISPATCH;
}
switch (local_pc) {
case 0x088B4004u: goto SB_L_088B4004;
case 0x088B4018u: goto SB_L_088B4018;
case 0x088B4020u: goto SB_L_088B4020;
case 0x088B4074u: goto SB_L_088B4074;
case 0x088B407Cu: goto SB_L_088B407C;
case 0x088B40C4u: goto SB_L_088B40C4;
case 0x088B40E0u: goto SB_L_088B40E0;
case 0x088B40F4u: goto SB_L_088B40F4;
case 0x088B40F8u: goto SB_L_088B40F8;
case 0x088B4110u: goto SB_L_088B4110;
case 0x088B4118u: goto SB_L_088B4118;
case 0x088B412Cu: goto SB_L_088B412C;
case 0x088B4134u: goto SB_L_088B4134;
case 0x088B4188u: goto SB_L_088B4188;
case 0x088B4190u: goto SB_L_088B4190;
case 0x088B41D8u: goto SB_L_088B41D8;
case 0x088B41F4u: goto SB_L_088B41F4;
case 0x088B4208u: goto SB_L_088B4208;
case 0x088B420Cu: goto SB_L_088B420C;
case 0x088B4224u: goto SB_L_088B4224;
case 0x088B4234u: goto SB_L_088B4234;
case 0x088B4238u: goto SB_L_088B4238;
default:
ctx.pc = local_pc;
TIER2_SB_RETURN();
}
TIER2_ENTRY_DISPATCH:
switch (entry_pc) {
case 0x088B3FCCu: goto SB_L_088B3FCC;
case 0x088B4004u: goto SB_L_088B4004;
case 0x088B40F8u: goto SB_L_088B40F8;
default:
ctx.pc = entry_pc;
++tier2_stats.fallbacks;
TIER2_SB_RETURN();
}
SB_L_088B1780:
{ const std::uint32_t vfpu_address = ctx.gpr[4] + static_cast<std::uint32_t>(0);
float vfpu_value[4]{
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 0u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 4u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 8u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 12u))};
ctx.write_vfpu_vector_ct<0u, 4u>(vfpu_value); }
{ const std::uint32_t vfpu_address = ctx.gpr[4] + static_cast<std::uint32_t>(16);
float vfpu_value[4]{
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 0u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 4u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 8u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 12u))};
ctx.write_vfpu_vector_ct<1u, 4u>(vfpu_value); }
{ const std::uint32_t vfpu_address = ctx.gpr[5] + static_cast<std::uint32_t>(0);
float vfpu_value[4]{
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 0u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 4u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 8u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 12u))};
ctx.write_vfpu_vector_ct<4u, 4u>(vfpu_value); }
{ const std::uint32_t vfpu_address = ctx.gpr[5] + static_cast<std::uint32_t>(16);
float vfpu_value[4]{
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 0u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 4u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 8u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 12u))};
ctx.write_vfpu_vector_ct<5u, 4u>(vfpu_value); }
ctx.execute_vfpu_vminmax_ct<2u, 0u, 1u, 3u, false>();
ctx.execute_vfpu_vminmax_ct<3u, 0u, 1u, 3u, true>();
ctx.execute_vfpu_compare3_ct<12u, 5u, 2u, 3u, 7u>();
ctx.execute_vfpu_compare3_ct<13u, 3u, 4u, 3u, 7u>();
{ float vfpu_value[4]{};
ctx.write_vfpu_vector_with_destination_prefix_ct<28u, 3u>(vfpu_value); }
{ float vfpu_s[4]{}, vfpu_t[4]{}, vfpu_d[4]{};
ctx.read_vfpu_vector_with_source_prefix_ct<12u, 3u, 0u>(vfpu_s);
ctx.read_vfpu_vector_with_source_prefix_ct<13u, 3u, 1u>(vfpu_t);
for (std::uint32_t i = 0; i < 3u; ++i) vfpu_d[i] = vfpu_s[i] + vfpu_t[i];
ctx.write_vfpu_vector_with_destination_prefix_ct<12u, 3u>(vfpu_d); }
ctx.execute_vfpu_vcmp_ct<12u, 28u, 3u, 7u>();
{ const bool branch_taken = ((ctx.vfpu_ctrl[3] >> 4u) & 1u) != 0u;
{ float vfpu_s[4]{}, vfpu_t[4]{}, vfpu_d[4]{};
ctx.read_vfpu_vector_with_source_prefix_ct<5u, 3u, 0u>(vfpu_s);
ctx.read_vfpu_vector_with_source_prefix_ct<4u, 3u, 1u>(vfpu_t);
for (std::uint32_t i = 0; i < 3u; ++i) vfpu_d[i] = vfpu_s[i] + vfpu_t[i];
ctx.write_vfpu_vector_with_destination_prefix_ct<15u, 3u>(vfpu_d); }
if (branch_taken) {
goto SB_L_088B180C;
}
goto SB_L_088B17B4;
}
SB_L_088B17B4:
{ float vfpu_s[4]{}, vfpu_t[4]{}, vfpu_d[4]{};
ctx.read_vfpu_vector_with_source_prefix_ct<0u, 3u, 0u>(vfpu_s);
ctx.read_vfpu_vector_with_source_prefix_ct<1u, 3u, 1u>(vfpu_t);
for (std::uint32_t i = 0; i < 3u; ++i) vfpu_d[i] = vfpu_s[i] + vfpu_t[i];
ctx.write_vfpu_vector_with_destination_prefix_ct<14u, 3u>(vfpu_d); }
{ float vfpu_s[4]{}, vfpu_t[4]{}, vfpu_d[4]{};
ctx.read_vfpu_vector_with_source_prefix_ct<5u, 3u, 0u>(vfpu_s);
ctx.read_vfpu_vector_with_source_prefix_ct<4u, 3u, 1u>(vfpu_t);
for (std::uint32_t i = 0; i < 3u; ++i) vfpu_d[i] = vfpu_s[i] - vfpu_t[i];
ctx.write_vfpu_vector_with_destination_prefix_ct<8u, 3u>(vfpu_d); }
{ float vfpu_s[4]{}, vfpu_t[4]{}, vfpu_d[4]{};
ctx.read_vfpu_vector_with_source_prefix_ct<1u, 3u, 0u>(vfpu_s);
ctx.read_vfpu_vector_with_source_prefix_ct<0u, 3u, 1u>(vfpu_t);
for (std::uint32_t i = 0; i < 3u; ++i) vfpu_d[i] = vfpu_s[i] - vfpu_t[i];
ctx.write_vfpu_vector_with_destination_prefix_ct<12u, 3u>(vfpu_d); }
{ float vfpu_s[4]{}, vfpu_t[4]{}, vfpu_d[4]{};
ctx.read_vfpu_vector_with_source_prefix_ct<15u, 3u, 0u>(vfpu_s);
ctx.read_vfpu_vector_with_source_prefix_ct<14u, 3u, 1u>(vfpu_t);
for (std::uint32_t i = 0; i < 3u; ++i) vfpu_d[i] = vfpu_s[i] - vfpu_t[i];
ctx.write_vfpu_vector_with_destination_prefix_ct<9u, 3u>(vfpu_d); }
ctx.vfpu_ctrl[0u] = 0x000040C9u;
ctx.vfpu_ctrl[1u] = 0x000043C6u;
ctx.execute_vfpu_vdot_ct<16u, 8u, 12u, 3u>();
ctx.vfpu_ctrl[0u] = 0x000040C2u;
ctx.vfpu_ctrl[1u] = 0x000043C8u;
ctx.execute_vfpu_vdot_ct<48u, 8u, 12u, 3u>();
ctx.vfpu_ctrl[1u] = 0x000003E1u;
ctx.execute_vfpu_vdot_ct<80u, 8u, 12u, 2u>();
ctx.vfpu_ctrl[0u] = 0x000040C9u;
ctx.vfpu_ctrl[1u] = 0x000240C6u;
ctx.execute_vfpu_vdot_ct<17u, 9u, 12u, 3u>();
ctx.vfpu_ctrl[0u] = 0x000040C2u;
ctx.vfpu_ctrl[1u] = 0x000240C8u;
ctx.execute_vfpu_vdot_ct<49u, 9u, 12u, 3u>();
ctx.vfpu_ctrl[1u] = 0x000200E1u;
ctx.execute_vfpu_vdot_ct<81u, 9u, 12u, 2u>();
ctx.vfpu_ctrl[0u] = 0x000007E4u;
ctx.execute_vfpu_vcmp_ct<17u, 16u, 3u, 7u>();
goto SB_L_088B180C;
SB_L_088B180C:
// vflush: architectural no-op that retains VFPU prefixes
ctx.gpr[4] = (ctx.vfpu_scalar_bits_ct<131u>());
ctx.gpr[2] = (ctx.gpr[4] & 16u);
jump_target = ctx.gpr[31];
ctx.gpr[2] = (ctx.gpr[2] < static_cast<std::uint32_t>(1) ? 1u : 0u);
local_pc = jump_target;
goto TIER2_LOCAL_DISPATCH_U0043;
SB_L_088B18EC:
ctx.gpr[29] = (ctx.gpr[29] + static_cast<std::uint32_t>(-128));
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(84), ctx.gpr[16]);
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(88), ctx.gpr[17]);
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(92), ctx.gpr[18]);
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(96), ctx.gpr[19]);
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(100), ctx.gpr[20]);
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(104), ctx.gpr[21]);
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(108), ctx.gpr[22]);
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(112), ctx.gpr[23]);
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(116), ctx.gpr[31]);
ctx.gpr[16] = (ctx.gpr[8] | 0u);
ctx.gpr[17] = (ctx.gpr[7] | 0u);
ctx.gpr[18] = (ctx.gpr[6] | 0u);
ctx.gpr[21] = (ctx.gpr[4] | 0u);
ctx.gpr[22] = (ctx.gpr[29] + static_cast<std::uint32_t>(16));
ctx.gpr[23] = (ctx.gpr[29] + static_cast<std::uint32_t>(32));
ctx.gpr[4] = (ctx.gpr[18] | 0u);
ctx.gpr[6] = (ctx.gpr[29] | 0u);
ctx.gpr[7] = (ctx.gpr[22] | 0u);
ctx.gpr[31] = (0x088B1940u);
ctx.gpr[8] = (ctx.gpr[23] | 0u);
if (rt.invoke_chained_direct<&recomp_unit_0163_entry, 163u, 44u, 0x08A9061Cu>(ctx, &aot_mem) && ctx.pc == 0x088B1940u) goto SB_L_088B1940;
TIER2_SB_RETURN();
SB_L_088B1940:
ctx.gpr[20] = (ctx.gpr[29] + static_cast<std::uint32_t>(48));
ctx.gpr[19] = (ctx.gpr[29] + static_cast<std::uint32_t>(64));
ctx.gpr[10] = (ctx.gpr[29] + static_cast<std::uint32_t>(80));
ctx.gpr[4] = (ctx.gpr[21] | 0u);
ctx.gpr[5] = (ctx.gpr[29] | 0u);
ctx.gpr[6] = (ctx.gpr[22] | 0u);
ctx.gpr[7] = (ctx.gpr[23] | 0u);
ctx.gpr[8] = (ctx.gpr[20] | 0u);
ctx.gpr[31] = (0x088B1968u);
ctx.gpr[9] = (ctx.gpr[19] | 0u);
goto SB_L_088B1B74;
SB_L_088B1B74:
{ const std::uint32_t vfpu_address = ctx.gpr[4] + static_cast<std::uint32_t>(0);
float vfpu_value[4]{
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 0u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 4u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 8u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 12u))};
ctx.write_vfpu_vector_ct<1u, 4u>(vfpu_value); }
{ const std::uint32_t vfpu_address = ctx.gpr[4] + static_cast<std::uint32_t>(16);
float vfpu_value[4]{
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 0u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 4u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 8u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 12u))};
ctx.write_vfpu_vector_ct<2u, 4u>(vfpu_value); }
{ const std::uint32_t vfpu_address = ctx.gpr[5] + static_cast<std::uint32_t>(0);
float vfpu_value[4]{
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 0u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 4u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 8u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 12u))};
ctx.write_vfpu_vector_ct<4u, 4u>(vfpu_value); }
{ const std::uint32_t vfpu_address = ctx.gpr[6] + static_cast<std::uint32_t>(0);
float vfpu_value[4]{
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 0u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 4u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 8u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 12u))};
ctx.write_vfpu_vector_ct<5u, 4u>(vfpu_value); }
{ const std::uint32_t vfpu_address = ctx.gpr[7] + static_cast<std::uint32_t>(0);
float vfpu_value[4]{
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 0u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 4u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 8u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 12u))};
ctx.write_vfpu_vector_ct<6u, 4u>(vfpu_value); }
{ float vfpu_s[4]{}, vfpu_t[4]{}, vfpu_d[4]{};
ctx.read_vfpu_vector_with_source_prefix_ct<5u, 3u, 0u>(vfpu_s);
ctx.read_vfpu_vector_with_source_prefix_ct<4u, 3u, 1u>(vfpu_t);
for (std::uint32_t i = 0; i < 3u; ++i) vfpu_d[i] = vfpu_s[i] - vfpu_t[i];
ctx.write_vfpu_vector_with_destination_prefix_ct<9u, 3u>(vfpu_d); }
{ float vfpu_s[4]{}, vfpu_t[4]{}, vfpu_d[4]{};
ctx.read_vfpu_vector_with_source_prefix_ct<6u, 3u, 0u>(vfpu_s);
ctx.read_vfpu_vector_with_source_prefix_ct<4u, 3u, 1u>(vfpu_t);
for (std::uint32_t i = 0; i < 3u; ++i) vfpu_d[i] = vfpu_s[i] - vfpu_t[i];
ctx.write_vfpu_vector_with_destination_prefix_ct<10u, 3u>(vfpu_d); }
ctx.execute_vfpu_cross_quat_ct<17u, 10u, 9u, 3u>();
ctx.execute_vfpu_vdot_ct<110u, 17u, 17u, 3u>();
{ float vfpu_s[4]{}, vfpu_d[4]{};
ctx.read_vfpu_vector_with_source_prefix_ct<110u, 1u, 0u>(vfpu_s);
for (std::uint32_t i = 0; i < 1u; ++i) vfpu_d[i] = 1.0f / std::sqrt(vfpu_s[i]);
ctx.write_vfpu_vector_with_destination_prefix_ct<110u, 1u>(vfpu_d); }
ctx.execute_vfpu_vscl_ct<7u, 17u, 110u, 3u>();
ctx.execute_vfpu_vdot_ct<103u, 7u, 4u, 3u>();
ctx.execute_vfpu_vdot_ct<24u, 1u, 7u, 3u>();
ctx.execute_vfpu_vdot_ct<56u, 2u, 7u, 3u>();
ctx.execute_vfpu_vcmp_ct<24u, 103u, 1u, 7u>();
{ const bool branch_taken = ((ctx.vfpu_ctrl[3] >> 0u) & 1u) != 0u;
ctx.execute_vfpu_vcmp_ct<56u, 103u, 1u, 7u>();
if (branch_taken) {
goto SB_L_088B1BC8;
}
goto SB_L_088B1BB8;
}
SB_L_088B1BB8:
{ const bool branch_taken = ((ctx.vfpu_ctrl[3] >> 0u) & 1u) == 0u;
// nop
if (branch_taken) {
goto SB_L_088B1C40;
}
goto SB_L_088B1BC0;
}
SB_L_088B1BC0:
{ const bool branch_taken = 0u == 0u;
// nop
if (branch_taken) {
goto SB_L_088B1BD0;
}
goto SB_L_088B1BC8;
}
SB_L_088B1BC8:
{ const bool branch_taken = ((ctx.vfpu_ctrl[3] >> 0u) & 1u) != 0u;
// nop
if (branch_taken) {
goto SB_L_088B1C40;
}
goto SB_L_088B1BD0;
}
SB_L_088B1BD0:
{ float vfpu_s[4]{}, vfpu_t[4]{}, vfpu_d[4]{};
ctx.read_vfpu_vector_with_source_prefix_ct<2u, 3u, 0u>(vfpu_s);
ctx.read_vfpu_vector_with_source_prefix_ct<1u, 3u, 1u>(vfpu_t);
for (std::uint32_t i = 0; i < 3u; ++i) vfpu_d[i] = vfpu_s[i] - vfpu_t[i];
ctx.write_vfpu_vector_with_destination_prefix_ct<3u, 3u>(vfpu_d); }
{ float vfpu_s[4]{}, vfpu_t[4]{}, vfpu_d[4]{};
ctx.read_vfpu_vector_with_source_prefix_ct<103u, 1u, 0u>(vfpu_s);
ctx.read_vfpu_vector_with_source_prefix_ct<24u, 1u, 1u>(vfpu_t);
for (std::uint32_t i = 0; i < 1u; ++i) vfpu_d[i] = vfpu_s[i] - vfpu_t[i];
ctx.write_vfpu_vector_with_destination_prefix_ct<24u, 1u>(vfpu_d); }
ctx.execute_vfpu_vdot_ct<56u, 3u, 7u, 3u>();
{ float vfpu_s[4]{}, vfpu_d[4]{};
ctx.read_vfpu_vector_with_source_prefix_ct<56u, 1u, 0u>(vfpu_s);
for (std::uint32_t i = 0; i < 1u; ++i) vfpu_d[i] = 1.0f / vfpu_s[i];
ctx.write_vfpu_vector_with_destination_prefix_ct<56u, 1u>(vfpu_d); }
{ float vfpu_s[4]{}, vfpu_t[4]{}, vfpu_d[4]{};
ctx.read_vfpu_vector_with_source_prefix_ct<24u, 1u, 0u>(vfpu_s);
ctx.read_vfpu_vector_with_source_prefix_ct<56u, 1u, 1u>(vfpu_t);
for (std::uint32_t i = 0; i < 1u; ++i) vfpu_d[i] = vfpu_s[i] * vfpu_t[i];
ctx.write_vfpu_vector_with_destination_prefix_ct<24u, 1u>(vfpu_d); }
ctx.execute_vfpu_vscl_ct<8u, 3u, 24u, 3u>();
{ float vfpu_s[4]{}, vfpu_t[4]{}, vfpu_d[4]{};
ctx.read_vfpu_vector_with_source_prefix_ct<8u, 3u, 0u>(vfpu_s);
ctx.read_vfpu_vector_with_source_prefix_ct<1u, 3u, 1u>(vfpu_t);
for (std::uint32_t i = 0; i < 3u; ++i) vfpu_d[i] = vfpu_s[i] + vfpu_t[i];
ctx.write_vfpu_vector_with_destination_prefix_ct<1u, 3u>(vfpu_d); }
{ float vfpu_value[4]{};
ctx.write_vfpu_vector_with_destination_prefix_ct<0u, 3u>(vfpu_value); }
{ float vfpu_s[4]{}, vfpu_t[4]{}, vfpu_d[4]{};
ctx.read_vfpu_vector_with_source_prefix_ct<4u, 3u, 0u>(vfpu_s);
ctx.read_vfpu_vector_with_source_prefix_ct<5u, 3u, 1u>(vfpu_t);
for (std::uint32_t i = 0; i < 3u; ++i) vfpu_d[i] = vfpu_s[i] - vfpu_t[i];
ctx.write_vfpu_vector_with_destination_prefix_ct<8u, 3u>(vfpu_d); }
{ float vfpu_s[4]{}, vfpu_t[4]{}, vfpu_d[4]{};
ctx.read_vfpu_vector_with_source_prefix_ct<6u, 3u, 0u>(vfpu_s);
ctx.read_vfpu_vector_with_source_prefix_ct<5u, 3u, 1u>(vfpu_t);
for (std::uint32_t i = 0; i < 3u; ++i) vfpu_d[i] = vfpu_s[i] - vfpu_t[i];
ctx.write_vfpu_vector_with_destination_prefix_ct<11u, 3u>(vfpu_d); }
{ float vfpu_s[4]{}, vfpu_t[4]{}, vfpu_d[4]{};
ctx.read_vfpu_vector_with_source_prefix_ct<1u, 3u, 0u>(vfpu_s);
ctx.read_vfpu_vector_with_source_prefix_ct<4u, 3u, 1u>(vfpu_t);
for (std::uint32_t i = 0; i < 3u; ++i) vfpu_d[i] = vfpu_s[i] - vfpu_t[i];
ctx.write_vfpu_vector_with_destination_prefix_ct<12u, 3u>(vfpu_d); }
{ float vfpu_s[4]{}, vfpu_t[4]{}, vfpu_d[4]{};
ctx.read_vfpu_vector_with_source_prefix_ct<1u, 3u, 0u>(vfpu_s);
ctx.read_vfpu_vector_with_source_prefix_ct<5u, 3u, 1u>(vfpu_t);
for (std::uint32_t i = 0; i < 3u; ++i) vfpu_d[i] = vfpu_s[i] - vfpu_t[i];
ctx.write_vfpu_vector_with_destination_prefix_ct<13u, 3u>(vfpu_d); }
ctx.execute_vfpu_cross_quat_ct<19u, 9u, 10u, 3u>();
ctx.execute_vfpu_cross_quat_ct<15u, 11u, 8u, 3u>();
ctx.execute_vfpu_cross_quat_ct<18u, 9u, 12u, 3u>();
ctx.execute_vfpu_cross_quat_ct<16u, 10u, 12u, 3u>();
ctx.execute_vfpu_cross_quat_ct<14u, 11u, 13u, 3u>();
ctx.execute_vfpu_vdot_ct<102u, 18u, 19u, 3u>();
ctx.execute_vfpu_vdot_ct<101u, 16u, 17u, 3u>();
ctx.execute_vfpu_vdot_ct<100u, 14u, 15u, 3u>();
ctx.execute_vfpu_vcmp_ct<39u, 0u, 3u, 6u>();
{ const bool branch_taken = ((ctx.vfpu_ctrl[3] >> 5u) & 1u) == 0u;
// nop
if (branch_taken) {
goto SB_L_088B1C40;
}
goto SB_L_088B1C2C;
}
SB_L_088B1C2C:
{ float vfpu_value[4]{}; ctx.read_vfpu_vector_ct<1u, 4u>(vfpu_value);
const std::uint32_t vfpu_address = ctx.gpr[8] + static_cast<std::uint32_t>(0);
aot_mem.aot_store32(vfpu_address + 0u, std::bit_cast<std::uint32_t>(vfpu_value[0]));
aot_mem.aot_store32(vfpu_address + 4u, std::bit_cast<std::uint32_t>(vfpu_value[1]));
aot_mem.aot_store32(vfpu_address + 8u, std::bit_cast<std::uint32_t>(vfpu_value[2]));
aot_mem.aot_store32(vfpu_address + 12u, std::bit_cast<std::uint32_t>(vfpu_value[3])); }
{ float vfpu_value[4]{}; ctx.read_vfpu_vector_ct<7u, 4u>(vfpu_value);
const std::uint32_t vfpu_address = ctx.gpr[9] + static_cast<std::uint32_t>(0);
aot_mem.aot_store32(vfpu_address + 0u, std::bit_cast<std::uint32_t>(vfpu_value[0]));
aot_mem.aot_store32(vfpu_address + 4u, std::bit_cast<std::uint32_t>(vfpu_value[1]));
aot_mem.aot_store32(vfpu_address + 8u, std::bit_cast<std::uint32_t>(vfpu_value[2]));
aot_mem.aot_store32(vfpu_address + 12u, std::bit_cast<std::uint32_t>(vfpu_value[3])); }
aot_mem.aot_store32(ctx.gpr[10] + static_cast<std::uint32_t>(0), ctx.vfpu_scalar_bits_ct<24u>());
jump_target = ctx.gpr[31];
ctx.gpr[2] = (0u + static_cast<std::uint32_t>(1));
local_pc = jump_target;
goto TIER2_LOCAL_DISPATCH_U0043;
SB_L_088B1C40:
jump_target = ctx.gpr[31];
ctx.gpr[2] = (0u | 0u);
local_pc = jump_target;
goto TIER2_LOCAL_DISPATCH_U0043;
SB_L_088B3FCC:
ctx.gpr[18] = (aot_mem.aot_load32(ctx.gpr[19] + static_cast<std::uint32_t>(68)));
ctx.gpr[18] = (ctx.gpr[18] + ctx.gpr[30]);
ctx.gpr[5] = (ctx.gpr[29] + static_cast<std::uint32_t>(16));
{ const std::uint32_t vfpu_address = ctx.gpr[18] + static_cast<std::uint32_t>(0);
float vfpu_value[4]{
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 0u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 4u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 8u)),
std::bit_cast<float>(aot_mem.aot_load32(vfpu_address + 12u))};
ctx.write_vfpu_vector_ct<2u, 4u>(vfpu_value); }
ctx.execute_vfpu_vx2i_ct<0u, 2u, 2u, 3u>();
ctx.execute_vfpu_vx2i_ct<1u, 66u, 2u, 3u>();
{ float vfpu_s[4]{}, vfpu_d[4]{};
ctx.read_vfpu_vector_ct<0u, 3u>(vfpu_s);
ctx.apply_vfpu_source_prefix_ct<3u, 0u>(vfpu_s);
const float vfpu_scale = std::ldexp(1.0f, -static_cast<int>(23u));
for (std::uint32_t vfpu_i = 0; vfpu_i < 3u; ++vfpu_i) {
const auto vfpu_integer = static_cast<std::int32_t>(std::bit_cast<std::uint32_t>(vfpu_s[vfpu_i]));
vfpu_d[vfpu_i] = static_cast<float>(vfpu_integer) * vfpu_scale;
}
ctx.write_vfpu_vector_with_destination_prefix_ct<0u, 3u>(vfpu_d); }
{ float vfpu_s[4]{}, vfpu_d[4]{};
ctx.read_vfpu_vector_ct<1u, 3u>(vfpu_s);
ctx.apply_vfpu_source_prefix_ct<3u, 0u>(vfpu_s);
const float vfpu_scale = std::ldexp(1.0f, -static_cast<int>(23u));
for (std::uint32_t vfpu_i = 0; vfpu_i < 3u; ++vfpu_i) {
const auto vfpu_integer = static_cast<std::int32_t>(std::bit_cast<std::uint32_t>(vfpu_s[vfpu_i]));
vfpu_d[vfpu_i] = static_cast<float>(vfpu_integer) * vfpu_scale;
}
ctx.write_vfpu_vector_with_destination_prefix_ct<1u, 3u>(vfpu_d); }
{ float vfpu_value[4]{}; ctx.read_vfpu_vector_ct<0u, 4u>(vfpu_value);
const std::uint32_t vfpu_address = ctx.gpr[5] + static_cast<std::uint32_t>(0);
aot_mem.aot_store32(vfpu_address + 0u, std::bit_cast<std::uint32_t>(vfpu_value[0]));
aot_mem.aot_store32(vfpu_address + 4u, std::bit_cast<std::uint32_t>(vfpu_value[1]));
aot_mem.aot_store32(vfpu_address + 8u, std::bit_cast<std::uint32_t>(vfpu_value[2]));
aot_mem.aot_store32(vfpu_address + 12u, std::bit_cast<std::uint32_t>(vfpu_value[3])); }
{ float vfpu_value[4]{}; ctx.read_vfpu_vector_ct<1u, 4u>(vfpu_value);
const std::uint32_t vfpu_address = ctx.gpr[5] + static_cast<std::uint32_t>(16);
aot_mem.aot_store32(vfpu_address + 0u, std::bit_cast<std::uint32_t>(vfpu_value[0]));
aot_mem.aot_store32(vfpu_address + 4u, std::bit_cast<std::uint32_t>(vfpu_value[1]));
aot_mem.aot_store32(vfpu_address + 8u, std::bit_cast<std::uint32_t>(vfpu_value[2]));
aot_mem.aot_store32(vfpu_address + 12u, std::bit_cast<std::uint32_t>(vfpu_value[3])); }
ctx.gpr[31] = (0x088B3FFCu);
ctx.gpr[4] = (ctx.gpr[20] | 0u);
goto SB_L_088B1780;
SB_L_088B4004:
ctx.gpr[17] = (static_cast<std::uint32_t>(static_cast<std::int32_t>(static_cast<std::int16_t>(aot_mem.aot_load16(ctx.gpr[18] + static_cast<std::uint32_t>(6))))));
ctx.gpr[4] = (static_cast<std::uint32_t>(static_cast<std::int32_t>(static_cast<std::int16_t>(aot_mem.aot_load16(ctx.gpr[18] + static_cast<std::uint32_t>(14))))));
ctx.gpr[4] = (static_cast<std::int32_t>(ctx.gpr[4]) < static_cast<std::int32_t>(ctx.gpr[17]) ? 1u : 0u);
{ const bool branch_taken = ctx.gpr[4] != 0u;
ctx.gpr[16] = (ctx.gpr[17] << 3u);
if (branch_taken) {
goto SB_L_088B40F4;
}
goto SB_L_088B4018;
}
SB_L_088B4018:
{ const bool branch_taken = ctx.gpr[22] == 0u;
// nop
if (branch_taken) {
goto SB_L_088B4074;
}
goto SB_L_088B4020;
}
SB_L_088B4020:
ctx.gpr[4] = (aot_mem.aot_load32(ctx.gpr[19] + static_cast<std::uint32_t>(76)));
ctx.gpr[4] = (ctx.gpr[4] + ctx.gpr[16]);
ctx.gpr[4] = (aot_mem.aot_load8(ctx.gpr[4] + static_cast<std::uint32_t>(6)));
ctx.gpr[4] = (ctx.gpr[4] & 255u);
ctx.gpr[5] = (ctx.gpr[4] ^ 7u);
ctx.gpr[5] = (ctx.gpr[5] < static_cast<std::uint32_t>(1) ? 1u : 0u);
ctx.gpr[6] = (ctx.gpr[4] ^ 8u);
ctx.gpr[6] = (ctx.gpr[6] < static_cast<std::uint32_t>(1) ? 1u : 0u);
ctx.gpr[5] = (ctx.gpr[5] | ctx.gpr[6]);
ctx.gpr[6] = (ctx.gpr[4] ^ 16u);
ctx.gpr[6] = (ctx.gpr[6] < static_cast<std::uint32_t>(1) ? 1u : 0u);
ctx.gpr[5] = (ctx.gpr[5] | ctx.gpr[6]);
ctx.gpr[6] = (ctx.gpr[4] ^ 31u);
ctx.gpr[6] = (ctx.gpr[6] < static_cast<std::uint32_t>(1) ? 1u : 0u);
ctx.gpr[5] = (ctx.gpr[5] | ctx.gpr[6]);
ctx.gpr[4] = (ctx.gpr[4] ^ 12u);
ctx.gpr[4] = (ctx.gpr[4] < static_cast<std::uint32_t>(1) ? 1u : 0u);
ctx.gpr[4] = (ctx.gpr[5] | ctx.gpr[4]);
ctx.gpr[4] = (ctx.gpr[4] & 255u);
{ const bool branch_taken = ctx.gpr[4] != 0u;
// nop
if (branch_taken) {
goto SB_L_088B40E0;
}
goto SB_L_088B4074;
}
SB_L_088B4074:
{ const bool branch_taken = ctx.gpr[21] == 0u;
// nop
if (branch_taken) {
goto SB_L_088B40C4;
}
goto SB_L_088B407C;
}
SB_L_088B407C:
ctx.gpr[4] = (aot_mem.aot_load32(ctx.gpr[19] + static_cast<std::uint32_t>(76)));
ctx.gpr[4] = (ctx.gpr[4] + ctx.gpr[16]);
ctx.gpr[4] = (aot_mem.aot_load8(ctx.gpr[4] + static_cast<std::uint32_t>(6)));
ctx.gpr[4] = (ctx.gpr[4] & 255u);
ctx.gpr[5] = (ctx.gpr[4] ^ 8u);
ctx.gpr[5] = (ctx.gpr[5] < static_cast<std::uint32_t>(1) ? 1u : 0u);
ctx.gpr[6] = (ctx.gpr[4] ^ 16u);
ctx.gpr[6] = (ctx.gpr[6] < static_cast<std::uint32_t>(1) ? 1u : 0u);
ctx.gpr[5] = (ctx.gpr[5] | ctx.gpr[6]);
ctx.gpr[6] = (ctx.gpr[4] ^ 31u);
ctx.gpr[6] = (ctx.gpr[6] < static_cast<std::uint32_t>(1) ? 1u : 0u);
ctx.gpr[5] = (ctx.gpr[5] | ctx.gpr[6]);
ctx.gpr[4] = (ctx.gpr[4] ^ 12u);
ctx.gpr[4] = (ctx.gpr[4] < static_cast<std::uint32_t>(1) ? 1u : 0u);
ctx.gpr[4] = (ctx.gpr[5] | ctx.gpr[4]);
ctx.gpr[4] = (ctx.gpr[4] & 255u);
{ const bool branch_taken = ctx.gpr[4] != 0u;
// nop
if (branch_taken) {
goto SB_L_088B40E0;
}
goto SB_L_088B40C4;
}
SB_L_088B40C4:
ctx.gpr[5] = (aot_mem.aot_load32(ctx.gpr[19] + static_cast<std::uint32_t>(72)));
ctx.gpr[6] = (aot_mem.aot_load32(ctx.gpr[19] + static_cast<std::uint32_t>(76)));
ctx.gpr[6] = (ctx.gpr[6] + ctx.gpr[16]);
ctx.gpr[4] = (ctx.gpr[20] | 0u);
ctx.gpr[7] = (ctx.gpr[23] | 0u);
ctx.gpr[31] = (0x088B40E0u);
ctx.gpr[8] = (ctx.gpr[29] | 0u);
if (tier2_return_depth < kTier2ReturnCapacity) {
if (!rt.tier2_enter_fused_transfer<43u, 0x088B18ECu>(ctx)) {
ctx.pc = 0x088B18ECu;
TIER2_SB_RETURN();
}
tier2_return_pc[tier2_return_depth] = 0x088B40E0u;
tier2_return_unit[tier2_return_depth] = 44u;
tier2_return_pending_base[tier2_return_depth] = tier2_pending_transfers;
++tier2_return_depth;
++tier2_stats.fused_calls;
goto SB_L_088B18EC;
}
++tier2_stats.fallbacks;
if (rt.invoke_chained_direct<&recomp_unit_0043_entry, 43u, 210u, 0x088B18ECu>(ctx, &aot_mem) && ctx.pc == 0x088B40E0u) goto SB_L_088B40E0;
TIER2_SB_RETURN();
SB_L_088B40E0:
ctx.gpr[17] = (ctx.gpr[17] + static_cast<std::uint32_t>(1));
ctx.gpr[4] = (static_cast<std::uint32_t>(static_cast<std::int32_t>(static_cast<std::int16_t>(aot_mem.aot_load16(ctx.gpr[18] + static_cast<std::uint32_t>(14))))));
ctx.gpr[4] = (static_cast<std::int32_t>(ctx.gpr[4]) < static_cast<std::int32_t>(ctx.gpr[17]) ? 1u : 0u);
{ const bool branch_taken = ctx.gpr[4] == 0u;
ctx.gpr[16] = (ctx.gpr[16] + static_cast<std::uint32_t>(8));
if (branch_taken) {
goto SB_L_088B4018;
}
goto SB_L_088B40F4;
}
SB_L_088B40F4:
ctx.gpr[4] = (aot_mem.aot_load32(ctx.gpr[29] + static_cast<std::uint32_t>(48)));
goto SB_L_088B40F8;
SB_L_088B40F8:
ctx.gpr[4] = (ctx.gpr[4] + static_cast<std::uint32_t>(1));
ctx.gpr[5] = (aot_mem.aot_load32(ctx.gpr[29] + static_cast<std::uint32_t>(52)));
ctx.gpr[30] = (ctx.gpr[30] + static_cast<std::uint32_t>(16));
ctx.gpr[5] = (static_cast<std::int32_t>(ctx.gpr[4]) < static_cast<std::int32_t>(ctx.gpr[5]) ? 1u : 0u);
{ const bool branch_taken = ctx.gpr[5] != 0u;
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(48), ctx.gpr[4]);
if (branch_taken) {
if (!rt.tier2_enter_fused_transfer<43u, 0x088B3FCCu>(ctx)) {
ctx.pc = 0x088B3FCCu;
TIER2_SB_RETURN();
}
++tier2_pending_transfers;
++tier2_stats.fused_tail_edges;
goto SB_L_088B3FCC;
}
goto SB_L_088B4110;
}
SB_L_088B4110:
{ const bool branch_taken = 0u == 0u;
// nop
if (branch_taken) {
goto SB_L_088B4208;
}
goto SB_L_088B4118;
}
SB_L_088B4118:
ctx.gpr[16] = (0u | 0u);
ctx.gpr[4] = (aot_mem.aot_load16(ctx.gpr[19] + static_cast<std::uint32_t>(50)));
ctx.gpr[4] = (static_cast<std::int32_t>(ctx.gpr[16]) < static_cast<std::int32_t>(ctx.gpr[4]) ? 1u : 0u);
{ const bool branch_taken = ctx.gpr[4] == 0u;
ctx.gpr[17] = (0u | 0u);
if (branch_taken) {
goto SB_L_088B4208;
}
goto SB_L_088B412C;
}
SB_L_088B412C:
{ const bool branch_taken = ctx.gpr[22] == 0u;
// nop
if (branch_taken) {
goto SB_L_088B4188;
}
goto SB_L_088B4134;
}
SB_L_088B4134:
ctx.gpr[4] = (aot_mem.aot_load32(ctx.gpr[19] + static_cast<std::uint32_t>(76)));
ctx.gpr[4] = (ctx.gpr[4] + ctx.gpr[17]);
ctx.gpr[4] = (aot_mem.aot_load8(ctx.gpr[4] + static_cast<std::uint32_t>(6)));
ctx.gpr[4] = (ctx.gpr[4] & 255u);
ctx.gpr[5] = (ctx.gpr[4] ^ 7u);
ctx.gpr[5] = (ctx.gpr[5] < static_cast<std::uint32_t>(1) ? 1u : 0u);
ctx.gpr[6] = (ctx.gpr[4] ^ 8u);
ctx.gpr[6] = (ctx.gpr[6] < static_cast<std::uint32_t>(1) ? 1u : 0u);
ctx.gpr[5] = (ctx.gpr[5] | ctx.gpr[6]);
ctx.gpr[6] = (ctx.gpr[4] ^ 16u);
ctx.gpr[6] = (ctx.gpr[6] < static_cast<std::uint32_t>(1) ? 1u : 0u);
ctx.gpr[5] = (ctx.gpr[5] | ctx.gpr[6]);
ctx.gpr[6] = (ctx.gpr[4] ^ 31u);
ctx.gpr[6] = (ctx.gpr[6] < static_cast<std::uint32_t>(1) ? 1u : 0u);
ctx.gpr[5] = (ctx.gpr[5] | ctx.gpr[6]);
ctx.gpr[4] = (ctx.gpr[4] ^ 12u);
ctx.gpr[4] = (ctx.gpr[4] < static_cast<std::uint32_t>(1) ? 1u : 0u);
ctx.gpr[4] = (ctx.gpr[5] | ctx.gpr[4]);
ctx.gpr[4] = (ctx.gpr[4] & 255u);
{ const bool branch_taken = ctx.gpr[4] != 0u;
// nop
if (branch_taken) {
goto SB_L_088B41F4;
}
goto SB_L_088B4188;
}
SB_L_088B4188:
{ const bool branch_taken = ctx.gpr[21] == 0u;
// nop
if (branch_taken) {
goto SB_L_088B41D8;
}
goto SB_L_088B4190;
}
SB_L_088B4190:
ctx.gpr[4] = (aot_mem.aot_load32(ctx.gpr[19] + static_cast<std::uint32_t>(76)));
ctx.gpr[4] = (ctx.gpr[4] + ctx.gpr[17]);
ctx.gpr[4] = (aot_mem.aot_load8(ctx.gpr[4] + static_cast<std::uint32_t>(6)));
ctx.gpr[4] = (ctx.gpr[4] & 255u);
ctx.gpr[5] = (ctx.gpr[4] ^ 8u);
ctx.gpr[5] = (ctx.gpr[5] < static_cast<std::uint32_t>(1) ? 1u : 0u);
ctx.gpr[6] = (ctx.gpr[4] ^ 16u);
ctx.gpr[6] = (ctx.gpr[6] < static_cast<std::uint32_t>(1) ? 1u : 0u);
ctx.gpr[5] = (ctx.gpr[5] | ctx.gpr[6]);
ctx.gpr[6] = (ctx.gpr[4] ^ 31u);
ctx.gpr[6] = (ctx.gpr[6] < static_cast<std::uint32_t>(1) ? 1u : 0u);
ctx.gpr[5] = (ctx.gpr[5] | ctx.gpr[6]);
ctx.gpr[4] = (ctx.gpr[4] ^ 12u);
ctx.gpr[4] = (ctx.gpr[4] < static_cast<std::uint32_t>(1) ? 1u : 0u);
ctx.gpr[4] = (ctx.gpr[5] | ctx.gpr[4]);
ctx.gpr[4] = (ctx.gpr[4] & 255u);
{ const bool branch_taken = ctx.gpr[4] != 0u;
// nop
if (branch_taken) {
goto SB_L_088B41F4;
}
goto SB_L_088B41D8;
}
SB_L_088B41D8:
ctx.gpr[5] = (aot_mem.aot_load32(ctx.gpr[19] + static_cast<std::uint32_t>(72)));
ctx.gpr[6] = (aot_mem.aot_load32(ctx.gpr[19] + static_cast<std::uint32_t>(76)));
ctx.gpr[6] = (ctx.gpr[6] + ctx.gpr[17]);
ctx.gpr[4] = (ctx.gpr[20] | 0u);
ctx.gpr[7] = (ctx.gpr[23] | 0u);
ctx.gpr[31] = (0x088B41F4u);
ctx.gpr[8] = (ctx.gpr[29] | 0u);
if (tier2_return_depth < kTier2ReturnCapacity) {
if (!rt.tier2_enter_fused_transfer<43u, 0x088B18ECu>(ctx)) {
ctx.pc = 0x088B18ECu;
TIER2_SB_RETURN();
}
tier2_return_pc[tier2_return_depth] = 0x088B41F4u;
tier2_return_unit[tier2_return_depth] = 44u;
tier2_return_pending_base[tier2_return_depth] = tier2_pending_transfers;
++tier2_return_depth;
++tier2_stats.fused_calls;
goto SB_L_088B18EC;
}
++tier2_stats.fallbacks;
if (rt.invoke_chained_direct<&recomp_unit_0043_entry, 43u, 210u, 0x088B18ECu>(ctx, &aot_mem) && ctx.pc == 0x088B41F4u) goto SB_L_088B41F4;
TIER2_SB_RETURN();
SB_L_088B41F4:
ctx.gpr[16] = (ctx.gpr[16] + static_cast<std::uint32_t>(1));
ctx.gpr[4] = (aot_mem.aot_load16(ctx.gpr[19] + static_cast<std::uint32_t>(50)));
ctx.gpr[4] = (static_cast<std::int32_t>(ctx.gpr[16]) < static_cast<std::int32_t>(ctx.gpr[4]) ? 1u : 0u);
{ const bool branch_taken = ctx.gpr[4] != 0u;
ctx.gpr[17] = (ctx.gpr[17] + static_cast<std::uint32_t>(8));
if (branch_taken) {
goto SB_L_088B412C;
}
goto SB_L_088B4208;
}
SB_L_088B4208:
ctx.gpr[4] = (aot_mem.aot_load32(ctx.gpr[29] + static_cast<std::uint32_t>(56)));
goto SB_L_088B420C;
SB_L_088B420C:
ctx.fpr[12] = std::bit_cast<float>(aot_mem.aot_load32(ctx.gpr[29] + static_cast<std::uint32_t>(0)));
ctx.fpr[13] = std::bit_cast<float>(aot_mem.aot_load32(ctx.gpr[4] + static_cast<std::uint32_t>(0)));
ctx.fcr31 = (ctx.fcr31 & ~0x00800000u) | (((ctx.fpr[12] < ctx.fpr[13])) ? 0x00800000u : 0u);
// nop
{ const bool branch_taken = !((ctx.fcr31 & 0x00800000u) != 0u);
// nop
if (branch_taken) {
goto SB_L_088B4234;
}
goto SB_L_088B4224;
}
SB_L_088B4224:
ctx.fpr[12] = std::bit_cast<float>(aot_mem.aot_load32(ctx.gpr[29] + static_cast<std::uint32_t>(0)));
aot_mem.aot_store32(ctx.gpr[4] + static_cast<std::uint32_t>(0), std::bit_cast<std::uint32_t>(ctx.fpr[12]));
{ const bool branch_taken = 0u == 0u;
ctx.gpr[2] = (0u | 1u);
if (branch_taken) {
goto SB_L_088B4238;
}
goto SB_L_088B4234;
}
SB_L_088B4234:
ctx.gpr[2] = (0u | 0u);
goto SB_L_088B4238;
SB_L_088B4238:
ctx.fpr[20] = std::bit_cast<float>(aot_mem.aot_load32(ctx.gpr[29] + static_cast<std::uint32_t>(60)));
ctx.gpr[16] = (aot_mem.aot_load32(ctx.gpr[29] + static_cast<std::uint32_t>(64)));
ctx.gpr[17] = (aot_mem.aot_load32(ctx.gpr[29] + static_cast<std::uint32_t>(68)));
ctx.gpr[18] = (aot_mem.aot_load32(ctx.gpr[29] + static_cast<std::uint32_t>(72)));
ctx.gpr[19] = (aot_mem.aot_load32(ctx.gpr[29] + static_cast<std::uint32_t>(76)));
ctx.gpr[20] = (aot_mem.aot_load32(ctx.gpr[29] + static_cast<std::uint32_t>(80)));
ctx.gpr[21] = (aot_mem.aot_load32(ctx.gpr[29] + static_cast<std::uint32_t>(84)));
ctx.gpr[22] = (aot_mem.aot_load32(ctx.gpr[29] + static_cast<std::uint32_t>(88)));
ctx.gpr[23] = (aot_mem.aot_load32(ctx.gpr[29] + static_cast<std::uint32_t>(92)));
ctx.gpr[30] = (aot_mem.aot_load32(ctx.gpr[29] + static_cast<std::uint32_t>(96)));
ctx.gpr[31] = (aot_mem.aot_load32(ctx.gpr[29] + static_cast<std::uint32_t>(100)));
jump_target = ctx.gpr[31];
ctx.gpr[29] = (ctx.gpr[29] + static_cast<std::uint32_t>(112));
local_pc = jump_target;
goto TIER2_LOCAL_DISPATCH_U0044;
#undef TIER2_SB_RETURN
}
} // namespace vcs
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,428 @@
// AUTO-GENERATED by profiles/vcs/tools/build_tier2_superblocks.py.
// Tier-2 SUPERBLOCK V2 cluster: physics
#include "vcs_tier2_superblocks.hpp"
#include "psprecomp/runtime.hpp"
#include "generated_units.hpp"
#include <cstdint>
namespace vcs {
using namespace psprecomp;
void tier2_superblock_physics(psprecomp::Runtime &rt,
psprecomp::AllegrexContext &ctx,
psprecomp::GuestMemory::AotFastView &aot_mem,
std::uint32_t entry_pc) {
tier2_detail::SampleScope tier2_scope(Tier2ClusterId::Physics);
auto &tier2_stats = tier2_scope.stats();
constexpr std::uint32_t kTier2ReturnCapacity = 32u;
std::uint32_t tier2_pending_transfers = 0u;
std::uint32_t tier2_return_depth = 0u;
std::uint32_t tier2_return_pc[kTier2ReturnCapacity]{};
std::uint32_t tier2_return_unit[kTier2ReturnCapacity]{};
std::uint32_t tier2_return_pending_base[kTier2ReturnCapacity]{};
std::uint32_t tier2_resume_pc = 0u;
std::uint32_t tier2_resume_unit = 0u;
std::uint32_t jump_target = 0u;
std::uint32_t local_pc = 0u;
std::uint32_t local_transfers = 0u;
std::uint32_t entry_id = 0u;
#define TIER2_SB_RETURN() do { tier2_scope.finish(); /* Unwind every logical invoke_chained_direct frame in true LIFO order. Tail frames created inside a fused JAL unwind before that JAL; outer JAL frames are also released after context invalidation. */ while (tier2_return_depth != 0u) { std::uint32_t tier2_base_ = tier2_return_pending_base[tier2_return_depth - 1u]; if (tier2_base_ > tier2_pending_transfers) { ++tier2_stats.fallbacks; tier2_base_ = tier2_pending_transfers; } const std::uint32_t tier2_tail_count_ = tier2_pending_transfers - tier2_base_; tier2_pending_transfers = tier2_base_; if (tier2_tail_count_ != 0u) (void)rt.tier2_complete_fused_transfers(ctx, tier2_tail_count_); --tier2_return_depth; (void)rt.tier2_complete_fused_transfers(ctx, 1u); } if (tier2_pending_transfers != 0u) { (void)rt.tier2_complete_fused_transfers(ctx, tier2_pending_transfers); tier2_pending_transfers = 0u; } return; } while (false)
goto TIER2_ENTRY_DISPATCH;
TIER2_FUSED_RETURN_DISPATCH:
switch (tier2_resume_pc) {
case 0x08A094D8u: goto SB_L_08A094D8;
case 0x08A094E0u: goto SB_L_08A094E0;
case 0x08A09524u: goto SB_L_08A09524;
case 0x08A09574u: goto SB_L_08A09574;
case 0x08A0958Cu: goto SB_L_08A0958C;
case 0x08A09B2Cu: goto SB_L_08A09B2C;
case 0x08A09BE4u: goto SB_L_08A09BE4;
case 0x08A09C10u: goto SB_L_08A09C10;
case 0x08A09C1Cu: goto SB_L_08A09C1C;
case 0x08A09C44u: goto SB_L_08A09C44;
case 0x08A09C4Cu: goto SB_L_08A09C4C;
case 0x08A09C64u: goto SB_L_08A09C64;
case 0x08A09C9Cu: goto SB_L_08A09C9C;
case 0x08A09CA4u: goto SB_L_08A09CA4;
case 0x08A09CBCu: goto SB_L_08A09CBC;
default: break;
}
switch (tier2_resume_unit) {
default: break;
}
ctx.pc = tier2_resume_pc;
TIER2_SB_RETURN();
TIER2_LOCAL_DISPATCH_U0129:
if (tier2_return_depth != 0u && local_pc == tier2_return_pc[tier2_return_depth - 1u]) {
tier2_resume_pc = local_pc;
tier2_resume_unit = tier2_return_unit[tier2_return_depth - 1u];
// Match invoke_chained_direct(): when the callee has returned, ctx.pc
// already contains the caller continuation before any starvation
// boundary/accounting can run. A scheduler switch here must never see
// the stale callee PC.
ctx.pc = local_pc;
const std::uint32_t tier2_pending_base =
tier2_return_pending_base[tier2_return_depth - 1u];
bool tier2_same_context = true;
if (tier2_pending_transfers < tier2_pending_base) {
++tier2_stats.fallbacks;
ctx.pc = local_pc;
TIER2_SB_RETURN();
}
const std::uint32_t tier2_nested_tail =
tier2_pending_transfers - tier2_pending_base;
if (tier2_nested_tail != 0u) {
tier2_pending_transfers = tier2_pending_base;
if (!rt.tier2_complete_fused_transfers(ctx, tier2_nested_tail))
tier2_same_context = false;
}
--tier2_return_depth;
if (!rt.tier2_complete_fused_transfers(ctx, 1u))
tier2_same_context = false;
if (!tier2_same_context) TIER2_SB_RETURN();
goto TIER2_FUSED_RETURN_DISPATCH;
}
switch (local_pc) {
case 0x08A094D8u: goto SB_L_08A094D8;
case 0x08A094E0u: goto SB_L_08A094E0;
case 0x08A09524u: goto SB_L_08A09524;
case 0x08A09574u: goto SB_L_08A09574;
case 0x08A0958Cu: goto SB_L_08A0958C;
case 0x08A09B2Cu: goto SB_L_08A09B2C;
case 0x08A09BE4u: goto SB_L_08A09BE4;
case 0x08A09C10u: goto SB_L_08A09C10;
case 0x08A09C1Cu: goto SB_L_08A09C1C;
case 0x08A09C44u: goto SB_L_08A09C44;
case 0x08A09C4Cu: goto SB_L_08A09C4C;
case 0x08A09C64u: goto SB_L_08A09C64;
case 0x08A09C9Cu: goto SB_L_08A09C9C;
case 0x08A09CA4u: goto SB_L_08A09CA4;
case 0x08A09CBCu: goto SB_L_08A09CBC;
default:
ctx.pc = local_pc;
TIER2_SB_RETURN();
}
TIER2_ENTRY_DISPATCH:
switch (entry_pc) {
case 0x08A09B2Cu: goto SB_L_08A09B2C;
default:
ctx.pc = entry_pc;
++tier2_stats.fallbacks;
TIER2_SB_RETURN();
}
SB_L_08A094D8:
{ const bool branch_taken = ctx.gpr[5] == 0u;
ctx.gpr[6] = (ctx.gpr[6] & 1u);
if (branch_taken) {
goto SB_L_08A0958C;
}
goto SB_L_08A094E0;
}
SB_L_08A094E0:
ctx.gpr[6] = (ctx.gpr[6] & 255u);
ctx.gpr[7] = (39680u << 16u);
ctx.gpr[6] = (ctx.gpr[6] | ctx.gpr[7]);
ctx.gpr[7] = (2236u << 16u);
ctx.gpr[8] = (aot_mem.aot_load32(ctx.gpr[7] + static_cast<std::uint32_t>(29552)));
aot_mem.aot_store32(ctx.gpr[8] + static_cast<std::uint32_t>(0), ctx.gpr[6]);
ctx.gpr[6] = (aot_mem.aot_load32(ctx.gpr[7] + static_cast<std::uint32_t>(29552)));
ctx.gpr[8] = (4608u << 16u);
ctx.gpr[6] = (ctx.gpr[6] + static_cast<std::uint32_t>(4));
aot_mem.aot_store32(ctx.gpr[7] + static_cast<std::uint32_t>(29552), ctx.gpr[6]);
ctx.gpr[8] = (ctx.gpr[8] + static_cast<std::uint32_t>(277));
aot_mem.aot_store32(ctx.gpr[6] + static_cast<std::uint32_t>(0), ctx.gpr[8]);
ctx.gpr[6] = (aot_mem.aot_load32(ctx.gpr[7] + static_cast<std::uint32_t>(29552)));
ctx.gpr[5] = (ctx.gpr[5] + static_cast<std::uint32_t>(2));
ctx.gpr[6] = (ctx.gpr[6] + static_cast<std::uint32_t>(4));
{ const bool branch_taken = ctx.gpr[4] == 0u;
aot_mem.aot_store32(ctx.gpr[7] + static_cast<std::uint32_t>(29552), ctx.gpr[6]);
if (branch_taken) {
goto SB_L_08A09574;
}
goto SB_L_08A09524;
}
SB_L_08A09524:
ctx.gpr[8] = (ctx.gpr[4] >> 8u);
ctx.gpr[9] = (15u << 16u);
ctx.gpr[8] = (ctx.gpr[8] & ctx.gpr[9]);
ctx.gpr[9] = (ctx.gpr[7] + static_cast<std::uint32_t>(29552));
aot_mem.aot_store32(ctx.gpr[9] + static_cast<std::uint32_t>(20), ctx.gpr[8]);
ctx.gpr[9] = (4096u << 16u);
ctx.gpr[8] = (ctx.gpr[8] | ctx.gpr[9]);
aot_mem.aot_store32(ctx.gpr[6] + static_cast<std::uint32_t>(0), ctx.gpr[8]);
ctx.gpr[8] = (256u << 16u);
ctx.gpr[6] = (aot_mem.aot_load32(ctx.gpr[7] + static_cast<std::uint32_t>(29552)));
ctx.gpr[8] = (ctx.gpr[8] + static_cast<std::uint32_t>(-1));
ctx.gpr[4] = (ctx.gpr[4] & ctx.gpr[8]);
ctx.gpr[6] = (ctx.gpr[6] + static_cast<std::uint32_t>(4));
ctx.gpr[8] = (256u << 16u);
aot_mem.aot_store32(ctx.gpr[7] + static_cast<std::uint32_t>(29552), ctx.gpr[6]);
ctx.gpr[4] = (ctx.gpr[4] | ctx.gpr[8]);
aot_mem.aot_store32(ctx.gpr[6] + static_cast<std::uint32_t>(0), ctx.gpr[4]);
ctx.gpr[6] = (aot_mem.aot_load32(ctx.gpr[7] + static_cast<std::uint32_t>(29552)));
ctx.gpr[6] = (ctx.gpr[6] + static_cast<std::uint32_t>(4));
aot_mem.aot_store32(ctx.gpr[7] + static_cast<std::uint32_t>(29552), ctx.gpr[6]);
goto SB_L_08A09574;
SB_L_08A09574:
ctx.gpr[4] = (1028u << 16u);
ctx.gpr[4] = (ctx.gpr[5] | ctx.gpr[4]);
aot_mem.aot_store32(ctx.gpr[6] + static_cast<std::uint32_t>(0), ctx.gpr[4]);
ctx.gpr[4] = (aot_mem.aot_load32(ctx.gpr[7] + static_cast<std::uint32_t>(29552)));
ctx.gpr[4] = (ctx.gpr[4] + static_cast<std::uint32_t>(4));
aot_mem.aot_store32(ctx.gpr[7] + static_cast<std::uint32_t>(29552), ctx.gpr[4]);
goto SB_L_08A0958C;
SB_L_08A0958C:
jump_target = ctx.gpr[31];
// nop
local_pc = jump_target;
goto TIER2_LOCAL_DISPATCH_U0129;
SB_L_08A09B2C:
ctx.gpr[29] = (ctx.gpr[29] + static_cast<std::uint32_t>(-16));
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(0), ctx.gpr[16]);
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(4), ctx.gpr[17]);
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(8), ctx.gpr[18]);
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(12), ctx.gpr[19]);
// PSP CACHE is a no-op in coherent host memory.
ctx.set_vfpu_scalar_bits_ct<30u>(ctx.gpr[16]);
ctx.set_vfpu_scalar_bits_ct<62u>(ctx.gpr[17]);
ctx.set_vfpu_scalar_bits_ct<94u>(ctx.gpr[18]);
ctx.set_vfpu_scalar_bits_ct<126u>(ctx.gpr[31]);
ctx.gpr[29] = (ctx.gpr[29] + static_cast<std::uint32_t>(-16));
aot_mem.aot_store32(ctx.gpr[29] + static_cast<std::uint32_t>(0), ctx.gpr[19]);
ctx.gpr[19] = (0u | 0u);
ctx.gpr[8] = (static_cast<std::uint32_t>(static_cast<std::int32_t>(static_cast<std::int16_t>(aot_mem.aot_load16(ctx.gpr[4] + static_cast<std::uint32_t>(4))))));
ctx.gpr[9] = (static_cast<std::uint32_t>(static_cast<std::int32_t>(static_cast<std::int16_t>(aot_mem.aot_load16(ctx.gpr[4] + static_cast<std::uint32_t>(6))))));
ctx.gpr[10] = (static_cast<std::uint32_t>(static_cast<std::int32_t>(static_cast<std::int16_t>(aot_mem.aot_load16(ctx.gpr[4] + static_cast<std::uint32_t>(8))))));
ctx.set_vfpu_scalar_bits_ct<29u>(ctx.gpr[8]);
ctx.set_vfpu_scalar_bits_ct<61u>(ctx.gpr[9]);
ctx.set_vfpu_scalar_bits_ct<93u>(ctx.gpr[10]);
{ float vfpu_s[4]{}, vfpu_d[4]{};
ctx.read_vfpu_vector_ct<29u, 3u>(vfpu_s);
ctx.apply_vfpu_source_prefix_ct<3u, 0u>(vfpu_s);
const float vfpu_scale = std::ldexp(1.0f, -static_cast<int>(15u));
for (std::uint32_t vfpu_i = 0; vfpu_i < 3u; ++vfpu_i) {
const auto vfpu_integer = static_cast<std::int32_t>(std::bit_cast<std::uint32_t>(vfpu_s[vfpu_i]));
vfpu_d[vfpu_i] = static_cast<float>(vfpu_integer) * vfpu_scale;
}
ctx.write_vfpu_vector_with_destination_prefix_ct<29u, 3u>(vfpu_d); }
{ float vfpu_matrix[16]{}, vfpu_target_raw[4]{}, vfpu_target[4]{}, vfpu_result[4]{};
ctx.read_vfpu_matrix(vfpu_matrix, 48u, 4u);
ctx.read_vfpu_vector_ct<29u, 4u>(vfpu_target_raw);
constexpr std::uint32_t vfpu_side = 4u;
constexpr std::uint32_t vfpu_input_length = 3u;
for (std::uint32_t i = 0; i < 4u; ++i) vfpu_target[i] = i < vfpu_input_length ? vfpu_target_raw[i] : 0.0f;
if (vfpu_side - 1u >= vfpu_input_length) vfpu_target[vfpu_side - 1u] = 1.0f;
for (std::uint32_t row = 0; row + 1u < vfpu_side; ++row) {
float sum = 0.0f;
for (std::uint32_t column = 0; column < vfpu_side; ++column) sum += vfpu_matrix[row * 4u + column] * vfpu_target[column];
vfpu_result[row] = sum;
}
float vfpu_final_row[4]{vfpu_matrix[(vfpu_side - 1u) * 4u + 0u], vfpu_matrix[(vfpu_side - 1u) * 4u + 1u],
vfpu_matrix[(vfpu_side - 1u) * 4u + 2u], vfpu_matrix[(vfpu_side - 1u) * 4u + 3u]};
ctx.apply_vfpu_source_prefix_ct<4u, 0u>(vfpu_final_row);
ctx.apply_vfpu_source_prefix_ct<4u, 1u>(vfpu_target);
for (std::uint32_t column = 0; column < 4u; ++column) vfpu_result[vfpu_side - 1u] += vfpu_final_row[column] * vfpu_target[column];
const std::uint32_t vfpu_destination_prefix = ctx.vfpu_ctrl[2];
const std::uint32_t vfpu_last_lane = vfpu_side - 1u;
ctx.vfpu_ctrl[2] = ((vfpu_destination_prefix & (1u << 8u)) << vfpu_last_lane) |
((vfpu_destination_prefix & 3u) << (vfpu_last_lane * 2u));
ctx.write_vfpu_vector_with_destination_prefix(vfpu_result, 12u, vfpu_side); }
ctx.execute_vfpu_vcmp_ct<12u, 31u, 4u, 7u>();
// vflush: architectural no-op that retains VFPU prefixes
ctx.gpr[11] = (ctx.vfpu_scalar_bits_ct<131u>());
ctx.gpr[11] = (ctx.gpr[11] & 15u);
ctx.gpr[11] = (ctx.gpr[11] << 16u);
ctx.gpr[8] = (static_cast<std::uint32_t>(static_cast<std::int32_t>(static_cast<std::int16_t>(aot_mem.aot_load16(ctx.gpr[4] + static_cast<std::uint32_t>(14))))));
ctx.gpr[9] = (static_cast<std::uint32_t>(static_cast<std::int32_t>(static_cast<std::int16_t>(aot_mem.aot_load16(ctx.gpr[4] + static_cast<std::uint32_t>(16))))));
ctx.gpr[10] = (static_cast<std::uint32_t>(static_cast<std::int32_t>(static_cast<std::int16_t>(aot_mem.aot_load16(ctx.gpr[4] + static_cast<std::uint32_t>(18))))));
ctx.set_vfpu_scalar_bits_ct<28u>(ctx.gpr[8]);
ctx.set_vfpu_scalar_bits_ct<60u>(ctx.gpr[9]);
ctx.set_vfpu_scalar_bits_ct<92u>(ctx.gpr[10]);
{ float vfpu_s[4]{}, vfpu_d[4]{};
ctx.read_vfpu_vector_ct<28u, 3u>(vfpu_s);
ctx.apply_vfpu_source_prefix_ct<3u, 0u>(vfpu_s);
const float vfpu_scale = std::ldexp(1.0f, -static_cast<int>(15u));
for (std::uint32_t vfpu_i = 0; vfpu_i < 3u; ++vfpu_i) {
const auto vfpu_integer = static_cast<std::int32_t>(std::bit_cast<std::uint32_t>(vfpu_s[vfpu_i]));
vfpu_d[vfpu_i] = static_cast<float>(vfpu_integer) * vfpu_scale;
}
ctx.write_vfpu_vector_with_destination_prefix_ct<28u, 3u>(vfpu_d); }
{ float vfpu_matrix[16]{}, vfpu_target_raw[4]{}, vfpu_target[4]{}, vfpu_result[4]{};
ctx.read_vfpu_matrix(vfpu_matrix, 48u, 4u);
ctx.read_vfpu_vector_ct<28u, 4u>(vfpu_target_raw);
constexpr std::uint32_t vfpu_side = 4u;
constexpr std::uint32_t vfpu_input_length = 3u;
for (std::uint32_t i = 0; i < 4u; ++i) vfpu_target[i] = i < vfpu_input_length ? vfpu_target_raw[i] : 0.0f;
if (vfpu_side - 1u >= vfpu_input_length) vfpu_target[vfpu_side - 1u] = 1.0f;
for (std::uint32_t row = 0; row + 1u < vfpu_side; ++row) {
float sum = 0.0f;
for (std::uint32_t column = 0; column < vfpu_side; ++column) sum += vfpu_matrix[row * 4u + column] * vfpu_target[column];
vfpu_result[row] = sum;
}
float vfpu_final_row[4]{vfpu_matrix[(vfpu_side - 1u) * 4u + 0u], vfpu_matrix[(vfpu_side - 1u) * 4u + 1u],
vfpu_matrix[(vfpu_side - 1u) * 4u + 2u], vfpu_matrix[(vfpu_side - 1u) * 4u + 3u]};
ctx.apply_vfpu_source_prefix_ct<4u, 0u>(vfpu_final_row);
ctx.apply_vfpu_source_prefix_ct<4u, 1u>(vfpu_target);
for (std::uint32_t column = 0; column < 4u; ++column) vfpu_result[vfpu_side - 1u] += vfpu_final_row[column] * vfpu_target[column];
const std::uint32_t vfpu_destination_prefix = ctx.vfpu_ctrl[2];
const std::uint32_t vfpu_last_lane = vfpu_side - 1u;
ctx.vfpu_ctrl[2] = ((vfpu_destination_prefix & (1u << 8u)) << vfpu_last_lane) |
((vfpu_destination_prefix & 3u) << (vfpu_last_lane * 2u));
ctx.write_vfpu_vector_with_destination_prefix(vfpu_result, 12u, vfpu_side); }
ctx.execute_vfpu_vcmp_ct<12u, 31u, 4u, 7u>();
// vflush: architectural no-op that retains VFPU prefixes
ctx.gpr[12] = (ctx.vfpu_scalar_bits_ct<131u>());
ctx.gpr[12] = (ctx.gpr[12] & 15u);
ctx.gpr[12] = (ctx.gpr[12] << 8u);
if (!ctx.execute_signed_add(5u, 5u, 5u)) { rt.arithmetic_overflow(0x08A09BC8u, 0x00A52820u); TIER2_SB_RETURN(); }
ctx.gpr[8] = (ctx.gpr[5] << 2u);
if (!ctx.execute_signed_add(5u, 5u, 8u)) { rt.arithmetic_overflow(0x08A09BD0u, 0x00A82820u); TIER2_SB_RETURN(); }
if (!ctx.execute_signed_add(18u, 4u, 5u)) { rt.arithmetic_overflow(0x08A09BD4u, 0x00859020u); TIER2_SB_RETURN(); }
ctx.gpr[16] = (ctx.gpr[11] | ctx.gpr[12]);
ctx.gpr[17] = (ctx.gpr[4] + static_cast<std::uint32_t>(20));
ctx.gpr[5] = (0u + static_cast<std::uint32_t>(0));
goto SB_L_08A09BE4;
SB_L_08A09BE4:
// PSP CACHE is a no-op in coherent host memory.
ctx.gpr[8] = (static_cast<std::uint32_t>(static_cast<std::int32_t>(static_cast<std::int16_t>(aot_mem.aot_load16(ctx.gpr[17] + static_cast<std::uint32_t>(4))))));
ctx.gpr[9] = (static_cast<std::uint32_t>(static_cast<std::int32_t>(static_cast<std::int16_t>(aot_mem.aot_load16(ctx.gpr[17] + static_cast<std::uint32_t>(6))))));
ctx.gpr[10] = (static_cast<std::uint32_t>(static_cast<std::int32_t>(static_cast<std::int16_t>(aot_mem.aot_load16(ctx.gpr[17] + static_cast<std::uint32_t>(8))))));
ctx.set_vfpu_scalar_bits_ct<15u>(ctx.gpr[8]);
ctx.set_vfpu_scalar_bits_ct<47u>(ctx.gpr[9]);
ctx.set_vfpu_scalar_bits_ct<79u>(ctx.gpr[10]);
{ float vfpu_s[4]{}, vfpu_d[4]{};
ctx.read_vfpu_vector_ct<15u, 3u>(vfpu_s);
ctx.apply_vfpu_source_prefix_ct<3u, 0u>(vfpu_s);
const float vfpu_scale = std::ldexp(1.0f, -static_cast<int>(15u));
for (std::uint32_t vfpu_i = 0; vfpu_i < 3u; ++vfpu_i) {
const auto vfpu_integer = static_cast<std::int32_t>(std::bit_cast<std::uint32_t>(vfpu_s[vfpu_i]));
vfpu_d[vfpu_i] = static_cast<float>(vfpu_integer) * vfpu_scale;
}
ctx.write_vfpu_vector_with_destination_prefix_ct<15u, 3u>(vfpu_d); }
ctx.execute_vfpu_vcmp_ct<15u, 28u, 3u, 1u>();
if (((ctx.vfpu_ctrl[3] >> 5u) & 1u) != 0u) {
ctx.gpr[11] = ((ctx.gpr[16] >> 8u) & 0x000000FFu);
goto SB_L_08A09C9C;
}
goto SB_L_08A09C10;
SB_L_08A09C10:
ctx.execute_vfpu_vcmp_ct<15u, 29u, 3u, 1u>();
if (((ctx.vfpu_ctrl[3] >> 5u) & 1u) != 0u) {
ctx.gpr[11] = ((ctx.gpr[16] >> 16u) & 0x000000FFu);
goto SB_L_08A09C9C;
}
goto SB_L_08A09C1C;
SB_L_08A09C1C:
{ float vfpu_matrix[16]{}, vfpu_target_raw[4]{}, vfpu_target[4]{}, vfpu_result[4]{};
ctx.read_vfpu_matrix(vfpu_matrix, 48u, 4u);
ctx.read_vfpu_vector_ct<15u, 4u>(vfpu_target_raw);
constexpr std::uint32_t vfpu_side = 4u;
constexpr std::uint32_t vfpu_input_length = 3u;
for (std::uint32_t i = 0; i < 4u; ++i) vfpu_target[i] = i < vfpu_input_length ? vfpu_target_raw[i] : 0.0f;
if (vfpu_side - 1u >= vfpu_input_length) vfpu_target[vfpu_side - 1u] = 1.0f;
for (std::uint32_t row = 0; row + 1u < vfpu_side; ++row) {
float sum = 0.0f;
for (std::uint32_t column = 0; column < vfpu_side; ++column) sum += vfpu_matrix[row * 4u + column] * vfpu_target[column];
vfpu_result[row] = sum;
}
float vfpu_final_row[4]{vfpu_matrix[(vfpu_side - 1u) * 4u + 0u], vfpu_matrix[(vfpu_side - 1u) * 4u + 1u],
vfpu_matrix[(vfpu_side - 1u) * 4u + 2u], vfpu_matrix[(vfpu_side - 1u) * 4u + 3u]};
ctx.apply_vfpu_source_prefix_ct<4u, 0u>(vfpu_final_row);
ctx.apply_vfpu_source_prefix_ct<4u, 1u>(vfpu_target);
for (std::uint32_t column = 0; column < 4u; ++column) vfpu_result[vfpu_side - 1u] += vfpu_final_row[column] * vfpu_target[column];
const std::uint32_t vfpu_destination_prefix = ctx.vfpu_ctrl[2];
const std::uint32_t vfpu_last_lane = vfpu_side - 1u;
ctx.vfpu_ctrl[2] = ((vfpu_destination_prefix & (1u << 8u)) << vfpu_last_lane) |
((vfpu_destination_prefix & 3u) << (vfpu_last_lane * 2u));
ctx.write_vfpu_vector_with_destination_prefix(vfpu_result, 12u, vfpu_side); }
ctx.execute_vfpu_vcmp_ct<12u, 31u, 4u, 7u>();
// vflush: architectural no-op that retains VFPU prefixes
ctx.gpr[11] = (ctx.vfpu_scalar_bits_ct<131u>());
ctx.gpr[11] = (ctx.gpr[11] & 15u);
ctx.gpr[16] = (ctx.gpr[16] | ctx.gpr[11]);
ctx.gpr[16] = (ctx.gpr[16] << 8u);
ctx.execute_vfpu_vcmp_ct<28u, 29u, 3u, 1u>();
{ const bool branch_taken = ((ctx.vfpu_ctrl[3] >> 5u) & 1u) != 0u;
// nop
if (branch_taken) {
goto SB_L_08A09CA4;
}
goto SB_L_08A09C44;
}
SB_L_08A09C44:
{ const bool branch_taken = ctx.gpr[16] == 0u;
ctx.gpr[11] = (ctx.gpr[16] | 0u);
if (branch_taken) {
goto SB_L_08A09CA4;
}
goto SB_L_08A09C4C;
}
SB_L_08A09C4C:
ctx.gpr[11] = (ctx.gpr[11] >> 8u);
ctx.gpr[11] = (ctx.gpr[11] & ctx.gpr[16]);
ctx.gpr[11] = (ctx.gpr[11] >> 8u);
ctx.gpr[11] = (ctx.gpr[11] & ctx.gpr[16]);
{ const bool branch_taken = ctx.gpr[11] != 0u;
// nop
if (branch_taken) {
goto SB_L_08A09CA4;
}
goto SB_L_08A09C64;
}
SB_L_08A09C64:
ctx.gpr[31] = (0x08A09C6Cu);
ctx.gpr[6] = (ctx.gpr[19] ^ ctx.gpr[5]);
goto SB_L_08A094D8;
SB_L_08A09C9C:
ctx.gpr[16] = (ctx.gpr[16] | ctx.gpr[11]);
ctx.gpr[16] = (ctx.gpr[16] << 8u);
goto SB_L_08A09CA4;
SB_L_08A09CA4:
ctx.gpr[19] = (ctx.gpr[19] ^ 1u);
ctx.gpr[17] = (ctx.gpr[17] + static_cast<std::uint32_t>(10));
ctx.gpr[5] = (ctx.gpr[5] + static_cast<std::uint32_t>(1));
{ float vfpu_s[4]{}, vfpu_d[4]{};
ctx.read_vfpu_vector_with_source_prefix_ct<28u, 3u, 0u>(vfpu_s);
for (std::uint32_t i = 0; i < 3u; ++i) vfpu_d[i] = vfpu_s[i];
ctx.write_vfpu_vector_with_destination_prefix_ct<29u, 3u>(vfpu_d); }
if (ctx.gpr[17] != ctx.gpr[18]) {
{ float vfpu_s[4]{}, vfpu_d[4]{};
ctx.read_vfpu_vector_with_source_prefix_ct<15u, 3u, 0u>(vfpu_s);
for (std::uint32_t i = 0; i < 3u; ++i) vfpu_d[i] = vfpu_s[i];
ctx.write_vfpu_vector_with_destination_prefix_ct<28u, 3u>(vfpu_d); }
goto SB_L_08A09BE4;
}
goto SB_L_08A09CBC;
SB_L_08A09CBC:
ctx.gpr[31] = (0x08A09CC4u);
ctx.gpr[6] = (ctx.gpr[19] ^ ctx.gpr[5]);
goto SB_L_08A094D8;
#undef TIER2_SB_RETURN
}
} // namespace vcs
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+102 -4
View File
@@ -1,6 +1,10 @@
#pragma once
#include "psprecomp/guest_memory.hpp"
#include <array>
#include <chrono>
#include <cstddef>
#include <cstdint>
namespace psprecomp {
@@ -10,10 +14,104 @@ struct AllegrexContext;
namespace vcs {
enum class Tier2ClusterId : std::uint32_t {
Entity = 0u, // 0152/0153 leaf accessors + 0154/0155 entity loop
Geometry = 1u, // 0084/0085 geometry/collision hot functions
Boundary = 2u, // 0085/0086 unit-boundary loop
Matrix = 3u, // 0044 matrix/VFPU hot function
Physics = 4u, // 0129 hot transform/physics function
World = 5u, // 0157/0158 dominant world/streaming hot functions
Edge43 = 6u, // 0043/0044 cross-unit boundary
Count = 7u,
};
constexpr std::size_t kTier2ClusterCount = static_cast<std::size_t>(Tier2ClusterId::Count);
struct Tier2ClusterCounters {
std::uint64_t entries{};
std::uint64_t fused_tail_edges{};
std::uint64_t fused_calls{};
std::uint64_t cold_exits{};
std::uint64_t fallbacks{};
std::uint64_t sampled_entries{};
std::uint64_t sampled_ns{};
};
struct Tier2CountersSnapshot {
std::array<Tier2ClusterCounters, kTier2ClusterCount> cluster{};
};
namespace tier2_detail {
extern thread_local std::array<Tier2ClusterCounters, kTier2ClusterCount> g_counters;
extern thread_local std::uint32_t g_active_depth;
inline Tier2ClusterCounters &counters(Tier2ClusterId id) noexcept {
return g_counters[static_cast<std::size_t>(id)];
}
// Entry timing is sampled 1/256 so coverage telemetry remains cheap enough to
// leave enabled in performance builds. The destructor sees every exit path,
// including cold fallback and scheduler invalidation.
class SampleScope {
public:
explicit SampleScope(Tier2ClusterId id) noexcept
: counter_(&counters(id)) {
++g_active_depth;
active_ = true;
const std::uint64_t entry = ++counter_->entries;
sample_ = (entry & 0xFFu) == 0u;
if (sample_) {
++counter_->sampled_entries;
start_ = std::chrono::steady_clock::now();
}
}
~SampleScope() noexcept {
finish();
if (active_) {
if (g_active_depth != 0u) --g_active_depth;
active_ = false;
}
}
void finish() noexcept {
if (!sample_ || finished_) return;
const auto ns = std::chrono::duration_cast<std::chrono::nanoseconds>(
std::chrono::steady_clock::now() - start_).count();
if (ns > 0) counter_->sampled_ns += static_cast<std::uint64_t>(ns);
finished_ = true;
}
Tier2ClusterCounters &stats() noexcept { return *counter_; }
private:
Tier2ClusterCounters *counter_{};
std::chrono::steady_clock::time_point start_{};
bool sample_{};
bool finished_{};
bool active_{};
};
} // namespace tier2_detail
bool tier2_superblocks_enabled() noexcept;
void tier2_superblock_154_155(psprecomp::Runtime &rt,
psprecomp::AllegrexContext &ctx,
psprecomp::GuestMemory::AotFastView &aot_mem,
std::uint32_t entry_pc);
bool tier2_cluster_enabled(Tier2ClusterId id) noexcept;
std::uint32_t tier2_cluster_mask() noexcept;
Tier2CountersSnapshot consume_tier2_counters() noexcept;
const char *tier2_cluster_name(Tier2ClusterId id) noexcept;
void tier2_superblock_entity(psprecomp::Runtime &, psprecomp::AllegrexContext &,
psprecomp::GuestMemory::AotFastView &, std::uint32_t entry_pc);
void tier2_superblock_geometry(psprecomp::Runtime &, psprecomp::AllegrexContext &,
psprecomp::GuestMemory::AotFastView &, std::uint32_t entry_pc);
void tier2_superblock_boundary(psprecomp::Runtime &, psprecomp::AllegrexContext &,
psprecomp::GuestMemory::AotFastView &, std::uint32_t entry_pc);
void tier2_superblock_matrix(psprecomp::Runtime &, psprecomp::AllegrexContext &,
psprecomp::GuestMemory::AotFastView &, std::uint32_t entry_pc);
void tier2_superblock_physics(psprecomp::Runtime &, psprecomp::AllegrexContext &,
psprecomp::GuestMemory::AotFastView &, std::uint32_t entry_pc);
void tier2_superblock_world(psprecomp::Runtime &, psprecomp::AllegrexContext &,
psprecomp::GuestMemory::AotFastView &, std::uint32_t entry_pc);
void tier2_superblock_edge43(psprecomp::Runtime &, psprecomp::AllegrexContext &,
psprecomp::GuestMemory::AotFastView &, std::uint32_t entry_pc);
} // namespace vcs
+9 -8
View File
@@ -92,7 +92,7 @@ if not exist "%CTEST_EXE%" set "CTEST_EXE=ctest.exe"
set "NINJA_STATUS=[%%f/%%t %%p ^| %%e elapsed ^| %%r running] "
set "BOOTFIX_STAMP=%BUILD%\.vcs_tier2_bootfix_20260816_v1"
set "SUPERBLOCK_STAMP=%BUILD%\.vcs_tier2_superblock_v1_20260816"
set "SUPERBLOCK_STAMP=%BUILD%\.vcs_tier2_superblock_v2_unwind_hotfix_20260816"
echo ================================================================
echo VCS - NINJA PERFORMANCE INCREMENTAL BUILD
@@ -104,7 +104,7 @@ echo CMake: %CMAKE_EXE%
echo Ninja: %NINJA_EXE%
echo Ninja workers: %JOBS%
echo cl.exe /MP: OFF ^(Ninja owns compile parallelism^)
echo Generated AOT: O3, cold /Ob0, measured hot /Ob3; Tier2 hot superblock host O2 /Ob3
echo Generated AOT: O3, cold /Ob0, measured hot /Ob3; Tier2 V2 clusters O2 /Ob3 /GL-
echo Host/core LTCG: ON
echo AVX2/fast paths: ON
echo ================================================================
@@ -113,7 +113,7 @@ echo [0b/7] Reapplying BOOTFIX-safe Tier-2 transforms (OPT1 semantic transforms
call "%PROFILE%\APPLY_TIER2_EXTREME.bat"
if errorlevel 1 goto :FAIL
echo [0b2/7] Building profile-guided Tier-2 superblock cluster 0154+0155...
echo [0b2/7] Building profile-guided Tier-2 V2 multi-cluster second layer...
set "PYTHON3_CMD="
py -3 -c "import sys; raise SystemExit(0 if sys.version_info.major == 3 else 1)" >nul 2>&1
if not errorlevel 1 set "PYTHON3_CMD=py -3"
@@ -128,11 +128,12 @@ if errorlevel 1 goto :FAIL
if exist "%BUILD%" if not exist "%SUPERBLOCK_STAMP%" (
echo.
echo [0c-super/7] Tier2 SUPERBLOCK V1 revision changed - invalidating fused pair objects once...
del /s /q "%BUILD%\*generated_unit_0154*.obj" >nul 2>&1
del /s /q "%BUILD%\*generated_unit_0155*.obj" >nul 2>&1
echo [0c-super/7] Tier2 SUPERBLOCK V2 UNWIND HOTFIX - invalidating measured hook/cluster objects once...
for %%U in (0043 0044 0084 0085 0086 0129 0154 0155 0157 0158) do del /s /q "%BUILD%\*generated_unit_%%U*.obj" >nul 2>&1
del /s /q "%BUILD%\*vcs_tier2_superblocks*.obj" >nul 2>&1
del /s /q "%BUILD%\*runtime*.obj" >nul 2>&1
del /s /q "%BUILD%\*vcs_tier2_cluster_*.obj" >nul 2>&1
del /s /q "%BUILD%\*vcs_profile*.obj" >nul 2>&1
del /s /q "%BUILD%\*vcs_runtime_log*.obj" >nul 2>&1
)
if exist "%BUILD%" if not exist "%BOOTFIX_STAMP%" (
@@ -169,7 +170,7 @@ echo [2/7] Building VCSNative with Ninja...
"%CMAKE_EXE%" --build "%BUILD%" --parallel %JOBS% --target VCSNative
if errorlevel 1 goto :FAIL
>"%BOOTFIX_STAMP%" echo VCS Tier2 BOOTFIX 2026-08-16 v1
>"%SUPERBLOCK_STAMP%" echo VCS Tier2 SUPERBLOCK V1 2026-08-16
>"%SUPERBLOCK_STAMP%" echo VCS Tier2 SUPERBLOCK V2 UNWIND HOTFIX 2026-08-16
echo.
echo [2b/7] Building tests and DX12 probes...
+59
View File
@@ -1,6 +1,7 @@
#include "framebuffer_capture.hpp"
#include "ge_renderer.hpp"
#include "vcs_profile.hpp"
#include "vcs_tier2_superblocks.hpp"
#include "psprecomp/guest_memory.hpp"
@@ -678,6 +679,63 @@ void test_ge_fixed_mip_level_selection() {
"constant PSP LOD did not select mip level 1");
}
void test_tier2_counter_snapshot_reset() {
using vcs::Tier2ClusterId;
auto &entity = vcs::tier2_detail::counters(Tier2ClusterId::Entity);
auto &world = vcs::tier2_detail::counters(Tier2ClusterId::World);
(void)vcs::consume_tier2_counters();
entity.entries = 3;
entity.fused_tail_edges = 5;
entity.fused_calls = 7;
entity.cold_exits = 11;
entity.fallbacks = 13;
entity.sampled_entries = 17;
entity.sampled_ns = 19000;
world.entries = 23;
world.fused_calls = 29;
const auto snapshot = vcs::consume_tier2_counters();
const auto &entity_snapshot =
snapshot.cluster[static_cast<std::size_t>(Tier2ClusterId::Entity)];
const auto &world_snapshot =
snapshot.cluster[static_cast<std::size_t>(Tier2ClusterId::World)];
require(entity_snapshot.entries == 3, "Tier2 entity entry counter snapshot mismatch");
require(entity_snapshot.fused_tail_edges == 5, "Tier2 fused-tail snapshot mismatch");
require(entity_snapshot.fused_calls == 7, "Tier2 fused-call snapshot mismatch");
require(entity_snapshot.cold_exits == 11, "Tier2 cold-exit snapshot mismatch");
require(entity_snapshot.fallbacks == 13, "Tier2 fallback snapshot mismatch");
require(entity_snapshot.sampled_entries == 17, "Tier2 sampled-entry snapshot mismatch");
require(entity_snapshot.sampled_ns == 19000, "Tier2 sampled-time snapshot mismatch");
require(world_snapshot.entries == 23 && world_snapshot.fused_calls == 29,
"Tier2 multi-cluster snapshot mismatch");
const auto reset_snapshot = vcs::consume_tier2_counters();
for (const auto &cluster : reset_snapshot.cluster) {
require(cluster.entries == 0 && cluster.fused_tail_edges == 0 &&
cluster.fused_calls == 0 && cluster.cold_exits == 0 &&
cluster.fallbacks == 0 && cluster.sampled_entries == 0 &&
cluster.sampled_ns == 0,
"Tier2 counters were not reset after consume");
}
require(vcs::tier2_detail::g_active_depth == 0u,
"Tier2 active-depth started nonzero");
{
vcs::tier2_detail::SampleScope scope(Tier2ClusterId::Entity);
require(vcs::tier2_detail::g_active_depth == 1u,
"Tier2 SampleScope did not arm reentry guard");
require(!vcs::tier2_cluster_enabled(Tier2ClusterId::Geometry),
"Tier2 allowed nested cluster reentry while a superblock was active");
}
require(vcs::tier2_detail::g_active_depth == 0u,
"Tier2 SampleScope leaked active-depth on unwind");
(void)vcs::consume_tier2_counters();
}
void test_framebuffer_formats() {
psprecomp::GuestMemory memory;
constexpr std::uint32_t address = 0x04000000u;
@@ -705,6 +763,7 @@ void test_framebuffer_formats() {
bool rejected = false;
try {
test_tier2_counter_snapshot_reset();
auto invalid = format;
invalid.stride = 0u;
(void)vcs::decode_framebuffer_rgb(memory, invalid);
+553 -165
View File
@@ -1,127 +1,551 @@
#!/usr/bin/env python3
"""Generate the profile-guided VCS Tier-2 hot superblock V1.
"""Generate VCS Tier-2 SUPERBLOCK V2 multi-cluster second-layer AOT.
The second layer deliberately duplicates only the physically hot region, not
whole 16 KiB AOT units. Cold exits resume the original generated unit at
its existing direct-entry id without creating a new logical chain frame.
Logical cross-unit tail frames retain Runtime chain-depth and starvation
accounting through tier2_enter_fused_transfer/tier2_complete_fused_transfers.
V2 is profile-guided and intentionally keeps the original generated corpus as
its semantic fallback. It extracts only measured hot control-flow closures,
fuses selected cross-unit direct calls/tails inside those closures, and leaves
all cold/external paths in the original AOT units.
The generated code preserves Runtime chain-depth and starvation accounting for
removed invoke_chained_direct frames through tier2_enter_fused_transfer() and
tier2_complete_fused_transfers().
"""
from __future__ import annotations
import argparse
import collections
import dataclasses
import pathlib
import re
from typing import Dict, Iterable, List, Mapping, MutableMapping, Sequence, Set, Tuple
U154_START = 0x08A6E894
U154_END = 0x08A6ED98 # cold exit label, excluded
U155_START = 0x08A71100
U155_END = 0x08A71214 # cold epilogue, excluded
HOOK_PCS = {
154: (0x08A6E894, 0x08A6E8A4),
155: (0x08A71100, 0x08A711E0, 0x08A711EC),
}
HOOK_BEGIN = '// TIER2_SUPERBLOCK_V1_HOOK_BEGIN\n'
HOOK_END = '// TIER2_SUPERBLOCK_V1_HOOK_END\n'
HOOK_BEGIN = '// TIER2_SUPERBLOCK_V2_HOOK_BEGIN\n'
HOOK_END = '// TIER2_SUPERBLOCK_V2_HOOK_END\n'
OLD_HOOK_BEGIN = '// TIER2_SUPERBLOCK_V1_HOOK_BEGIN\n'
OLD_HOOK_END = '// TIER2_SUPERBLOCK_V1_HOOK_END\n'
HEADER_INCLUDE = '#include "vcs_tier2_superblocks.hpp"\n'
# Generated unit labels use globally unique guest PCs, so one label namespace
# can safely span several AOT units in a cluster.
LABEL_RE = re.compile(r'^L_([0-9A-F]{8}):\n', re.M)
GOTO_RE = re.compile(r'goto SB_L_([0-9A-F]{8});')
DIRECT_ENTRY_RE = re.compile(r'case (\d+)u: goto L_([0-9A-F]{8});')
DIRECT_CALL_RE = re.compile(
r'(?P<indent>[ \t]*)if \(rt\.invoke_chained_direct<&recomp_unit_(?P<target_name>\d+)_entry, '
r'(?P<target_unit>\d+)u, (?P<entry>\d+)u, (?P<target_pc>0x[0-9A-F]+)u>'
r'\(ctx, &aot_mem\) && ctx\.pc == (?P<cont>0x[0-9A-F]+)u\) '
r'goto SB_L_(?P<cont_label>[0-9A-F]{8});\n(?P=indent)return;')
TAIL_CALL_RE = re.compile(
r'(?P<indent>[ \t]*)\(void\)rt\.invoke_chained_direct<&recomp_unit_(?P<target_name>\d+)_entry, '
r'(?P<target_unit>\d+)u, (?P<entry>\d+)u, (?P<target_pc>0x[0-9A-F]+)u>'
r'\(ctx, &aot_mem\); return;')
LOCAL_DISPATCH_SEQUENCE_RE = re.compile(
r'(?P<indent>[ \t]*)local_pc = (?P<local_expr>[^;]+);\n'
r'(?P=indent)if \(\+\+local_transfers < 256u\) \{ entry_id = 0u; goto LOCAL_DISPATCH; \}\n'
r'(?P=indent)ctx\.pc = (?P<pc_expr>[^;]+);\n'
r'(?P=indent)return;')
def strip_old_hooks(text: str) -> str:
text = re.sub(
re.escape(HOOK_BEGIN) + r'.*?' + re.escape(HOOK_END),
'', text, flags=re.S)
@dataclasses.dataclass(frozen=True)
class UnitSource:
unit: int
raw: str
blocks: Mapping[int, str]
ordered_pcs: Tuple[int, ...]
entries: Mapping[int, int]
@dataclasses.dataclass
class Cluster:
key: str
enum_name: str
function: str
# Unit -> hot seed PCs. Local control-flow closure is extracted from these.
seeds: Dict[int, List[int]]
# Optional explicit PC windows. These are useful for the entity loop where
# unrestricted local closure pulls in unrelated cold functions.
windows: Dict[int, List[Tuple[int, int]]] = dataclasses.field(default_factory=dict)
# Units where cross-unit call targets are recursively pulled into the same
# cluster. This is bounded by the source units in `seeds`/`windows`.
expand_cross_units: bool = False
hooks: Dict[int, List[int]] = dataclasses.field(default_factory=dict)
max_blocks: int = 700
CLUSTERS: List[Cluster] = [
Cluster(
key='entity', enum_name='Entity', function='tier2_superblock_entity',
seeds={},
windows={
152: [(0x08A65EA0, 0x08A65EC8)],
153: [(0x08A68CBC, 0x08A68CE4)],
154: [(0x08A6E894, 0x08A6ED98)],
155: [(0x08A71100, 0x08A71214)],
},
expand_cross_units=False,
hooks={154: [0x08A6E894, 0x08A6E8A4],
155: [0x08A71100, 0x08A711E0, 0x08A711EC]},
),
Cluster(
key='geometry', enum_name='Geometry', function='tier2_superblock_geometry',
seeds={84: [0x08955444, 0x089554A8],
85: [0x08958D28, 0x0895A760]},
expand_cross_units=True,
hooks={84: [0x08955444, 0x089554A8],
85: [0x08958D28, 0x0895A760]},
max_blocks=520,
),
Cluster(
key='boundary', enum_name='Boundary', function='tier2_superblock_boundary',
seeds={85: [0x0895BF3C, 0x0895BF7C],
86: [0x0895C000, 0x0895C01C, 0x0895C2A8]},
expand_cross_units=True,
hooks={85: [0x0895BF3C, 0x0895BF7C],
86: [0x0895C000, 0x0895C01C, 0x0895C2A8]},
max_blocks=180,
),
Cluster(
key='matrix', enum_name='Matrix', function='tier2_superblock_matrix',
seeds={44: [0x088B4738]},
hooks={44: [0x088B4738]},
max_blocks=180,
),
Cluster(
key='physics', enum_name='Physics', function='tier2_superblock_physics',
seeds={129: [0x08A09B2C]},
hooks={129: [0x08A09B2C]},
max_blocks=100,
),
Cluster(
key='world', enum_name='World', function='tier2_superblock_world',
seeds={157: [0x08A79DF8, 0x08A7A94C, 0x08A7A664,
0x08A7B5A8, 0x08A7B5A0, 0x08A7B910, 0x08A7B5B0],
158: [0x08A7EC68, 0x08A7F080]},
expand_cross_units=True,
hooks={157: [0x08A79DF8, 0x08A7A94C, 0x08A7A664, 0x08A7B5A8],
158: [0x08A7EC68, 0x08A7F080]},
max_blocks=700,
),
Cluster(
key='edge43', enum_name='Edge43', function='tier2_superblock_edge43',
seeds={43: [0x088B3FCC], 44: [0x088B4004, 0x088B40F8]},
expand_cross_units=True,
hooks={43: [0x088B3FCC], 44: [0x088B4004, 0x088B40F8]},
max_blocks=100,
),
]
def strip_hooks(text: str) -> str:
for begin, end in ((HOOK_BEGIN, HOOK_END), (OLD_HOOK_BEGIN, OLD_HOOK_END)):
text = re.sub(re.escape(begin) + r'.*?' + re.escape(end), '', text, flags=re.S)
text = text.replace(HEADER_INCLUDE, '')
return text
def region(text: str, start: int, end: int) -> str:
a = text.find(f'L_{start:08X}:')
b = text.find(f'L_{end:08X}:')
if a < 0 or b < 0 or b <= a:
raise RuntimeError(f'cannot extract region {start:08X}-{end:08X}')
return text[a:b]
def parse_unit(path: pathlib.Path, unit: int) -> UnitSource:
raw = strip_hooks(path.read_text(encoding='utf-8'))
matches = list(LABEL_RE.finditer(raw))
blocks: Dict[int, str] = {}
order: List[int] = []
end_markers = [m.start() for m in (
re.search(r'\n}\n\nvoid recomp_unit_\d+\(', raw),
re.search(r'\n}\n\nvoid register_generated_unit_', raw),
) if m is not None]
code_end = min(end_markers) if end_markers else len(raw)
for i, match in enumerate(matches):
pc = int(match.group(1), 16)
end = matches[i + 1].start() if i + 1 < len(matches) else code_end
# Labels belonging to registration metadata do not exist, but clamp the
# final executable block defensively so the generated superblock never
# copies the unit registration function or namespace close.
end = min(end, code_end)
blocks[pc] = raw[match.start():end]
order.append(pc)
entries = {int(pc, 16): int(entry) for entry, pc in DIRECT_ENTRY_RE.findall(raw)}
return UnitSource(unit=unit, raw=raw, blocks=blocks,
ordered_pcs=tuple(order), entries=entries)
def direct_entry_map(full_text: str) -> dict[int, int]:
return {int(pc, 16): int(entry_id) for entry_id, pc in re.findall(
r'case (\d+)u: goto L_([0-9A-F]{8});', full_text)}
def local_closure(source: UnitSource, seeds: Iterable[int], limit: int) -> Set[int]:
queue = collections.deque(seeds)
selected: Set[int] = set()
while queue:
if len(selected) >= limit:
raise RuntimeError(f'unit {source.unit:04d}: hot closure exceeded {limit} blocks')
pc = queue.popleft()
if pc in selected:
continue
block = source.blocks.get(pc)
if block is None:
raise RuntimeError(f'unit {source.unit:04d}: missing hot seed/target 0x{pc:08X}')
selected.add(pc)
for target in re.findall(r'goto L_([0-9A-F]{8});', block):
target_pc = int(target, 16)
if target_pc in source.blocks and target_pc not in selected:
queue.append(target_pc)
return selected
def transform_region(text: str, unit: int, partner: int,
start: int, end: int,
partner_start: int, partner_end: int,
self_entries: dict[int, int]) -> tuple[str, int, int, int]:
# Namespace labels so both source regions can live in one function.
text = re.sub(r'\bL_([0-9A-F]{8})\b', r'SB_L_\1', text)
def window_blocks(source: UnitSource, windows: Sequence[Tuple[int, int]]) -> Set[int]:
out: Set[int] = set()
for start, end in windows:
if start not in source.blocks:
raise RuntimeError(f'unit {source.unit:04d}: window start 0x{start:08X} missing')
for pc in source.ordered_pcs:
if start <= pc < end:
out.add(pc)
return out
fused = 0
# Only tail edges between the two selected regions are fused. JALs to other
# units remain invoke_chained_direct so HLE/thread-switch semantics are exact.
tail = re.compile(
rf' \(void\)rt\.invoke_chained_direct<&recomp_unit_{partner:04d}_entry, {partner}u, '
r'(\d+)u, (0x[0-9A-F]+)u>\(ctx, &aot_mem\); return;'
)
def tail_repl(m: re.Match[str]) -> str:
nonlocal fused
entry_id = int(m.group(1))
target = int(m.group(2), 16)
if not (partner_start <= target < partner_end):
def selected_for_cluster(cluster: Cluster, sources: Mapping[int, UnitSource]) -> Dict[int, Set[int]]:
units = sorted(set(cluster.seeds) | set(cluster.windows))
selected: Dict[int, Set[int]] = {u: set() for u in units}
for unit in units:
src = sources[unit]
if unit in cluster.windows:
selected[unit].update(window_blocks(src, cluster.windows[unit]))
if unit in cluster.seeds:
selected[unit].update(local_closure(src, cluster.seeds[unit], cluster.max_blocks))
if not cluster.expand_cross_units:
return selected
# Recursively pull direct-call targets only when both source and destination
# units are members of this cluster. Local closure of those targets captures
# whole guest functions/loops but remains bounded by max_blocks.
changed = True
while changed:
changed = False
total = sum(len(v) for v in selected.values())
if total > cluster.max_blocks:
raise RuntimeError(f'{cluster.key}: closure exceeded {cluster.max_blocks} blocks')
for unit in units:
src = sources[unit]
for pc in list(selected[unit]):
block = src.blocks[pc]
for target_unit_s, target_pc_s in re.findall(
r'invoke_chained_direct<&recomp_unit_(\d+)_entry, \d+u, \d+u, '
r'(0x[0-9A-F]+)u>', block):
target_unit = int(target_unit_s)
target_pc = int(target_pc_s, 16)
if target_unit not in selected or target_pc in selected[target_unit]:
continue
closure = local_closure(sources[target_unit], [target_pc], cluster.max_blocks)
before = len(selected[target_unit])
selected[target_unit].update(closure)
if len(selected[target_unit]) != before:
changed = True
return selected
def selected_pc_owner(selected: Mapping[int, Set[int]]) -> Dict[int, int]:
owner: Dict[int, int] = {}
for unit, pcs in selected.items():
for pc in pcs:
if pc in owner and owner[pc] != unit:
raise RuntimeError(f'duplicate guest PC 0x{pc:08X}')
owner[pc] = unit
return owner
def transform_block(block: str, source_unit: int, selected: Mapping[int, Set[int]],
sources: Mapping[int, UnitSource], stats: MutableMapping[str, int],
fused_continuations: MutableMapping[Tuple[int, int], int]) -> str:
pc_owner = selected_pc_owner(selected)
selected_all = set(pc_owner)
text = re.sub(r'\bL_([0-9A-F]{8})\b', r'SB_L_\1', block)
# Fuse JAL-style direct calls when the callee target is part of this cluster.
def direct_repl(m: re.Match[str]) -> str:
target_unit = int(m.group('target_unit'))
target_pc = int(m.group('target_pc'), 16)
cont = int(m.group('cont'), 16)
indent = m.group('indent')
if target_pc not in selected_all or pc_owner[target_pc] != target_unit:
return m.group(0)
fused += 1
stats['fused_calls'] += 1
fused_continuations[(source_unit, cont)] = sources[source_unit].entries.get(cont, -1)
original = m.group(0).replace('return;', 'TIER2_SB_RETURN();')
return (
f' if (!rt.tier2_enter_fused_transfer<{partner}u, 0x{target:08X}u>(ctx)) {{\n'
f' ctx.pc = 0x{target:08X}u;\n'
f' TIER2_SB_RETURN();\n'
f' }}\n'
f' ++tier2_pending_transfers;\n'
f' goto SB_L_{target:08X};'
f'{indent}if (tier2_return_depth < kTier2ReturnCapacity) {{\n'
f'{indent} if (!rt.tier2_enter_fused_transfer<{target_unit}u, 0x{target_pc:08X}u>(ctx)) {{\n'
f'{indent} ctx.pc = 0x{target_pc:08X}u;\n'
f'{indent} TIER2_SB_RETURN();\n'
f'{indent} }}\n'
f'{indent} tier2_return_pc[tier2_return_depth] = 0x{cont:08X}u;\n'
f'{indent} tier2_return_unit[tier2_return_depth] = {source_unit}u;\n'
f'{indent} tier2_return_pending_base[tier2_return_depth] = tier2_pending_transfers;\n'
f'{indent} ++tier2_return_depth;\n'
f'{indent} ++tier2_stats.fused_calls;\n'
f'{indent} goto SB_L_{target_pc:08X};\n'
f'{indent}}}\n'
f'{indent}++tier2_stats.fallbacks;\n'
f'{original}'
)
text = tail.sub(tail_repl, text)
text = DIRECT_CALL_RE.sub(direct_repl, text)
# Any ordinary local branch that leaves the selected hot region resumes the
# original generated unit at its existing direct-entry id. This is an
# ordinary native call, *not* invoke_chained_direct: the original local goto
# did not create a logical chain frame or scheduler work item either.
cold_exits: set[int] = set()
cold_dispatch_fallbacks: set[int] = set()
goto_re = re.compile(r'goto SB_L_([0-9A-F]{8});')
# Fuse tail-style direct calls. Their logical chain frames stay pending until
# the superblock leaves, matching nested native tail frames in V1/SAFE AOT.
def tail_repl(m: re.Match[str]) -> str:
target_unit = int(m.group('target_unit'))
target_pc = int(m.group('target_pc'), 16)
indent = m.group('indent')
if target_pc not in selected_all or pc_owner[target_pc] != target_unit:
return m.group(0)
stats['fused_tail'] += 1
return (
f'{indent}if (!rt.tier2_enter_fused_transfer<{target_unit}u, 0x{target_pc:08X}u>(ctx)) {{\n'
f'{indent} ctx.pc = 0x{target_pc:08X}u;\n'
f'{indent} TIER2_SB_RETURN();\n'
f'{indent}}}\n'
f'{indent}++tier2_pending_transfers;\n'
f'{indent}++tier2_stats.fused_tail_edges;\n'
f'{indent}goto SB_L_{target_pc:08X};'
)
text = TAIL_CALL_RE.sub(tail_repl, text)
# Standard generated indirect/local return path. The shared unit-specific
# dispatcher below recognizes a fused-call continuation and performs exactly
# one logical chain unwind before resuming the caller.
def local_dispatch_repl(m: re.Match[str]) -> str:
indent = m.group('indent')
return (f'{indent}local_pc = {m.group("local_expr")};\n'
f'{indent}goto TIER2_LOCAL_DISPATCH_U{source_unit:04d};')
text = LOCAL_DISPATCH_SEQUENCE_RE.sub(local_dispatch_repl, text)
text = text.replace('goto LOCAL_DISPATCH;', f'goto TIER2_LOCAL_DISPATCH_U{source_unit:04d};')
# Any ordinary local branch that leaves this cluster resumes the exact
# original generated entry without creating a new logical chain frame.
def goto_repl(m: re.Match[str]) -> str:
target = int(m.group(1), 16)
in_self = start <= target < end
in_partner = partner_start <= target < partner_end
if in_self or in_partner:
if target in selected_all:
return m.group(0)
cold_exits.add(target)
entry_id = self_entries.get(target)
if entry_id is not None:
return (f'recomp_unit_{unit:04d}_entry(rt, ctx, {entry_id}u, aot_mem); '
entry = sources[source_unit].entries.get(target)
stats['cold_exits'] += 1
if entry is not None:
return (f'++tier2_stats.cold_exits; tier2_scope.finish(); '
f'psprecomp::recomp_unit_{source_unit:04d}_entry(rt, ctx, {entry}u, aot_mem); '
f'TIER2_SB_RETURN();')
cold_dispatch_fallbacks.add(target)
return f'ctx.pc = 0x{target:08X}u; TIER2_SB_RETURN();'
return (f'++tier2_stats.cold_exits; ctx.pc = 0x{target:08X}u; TIER2_SB_RETURN();')
text = GOTO_RE.sub(goto_repl, text)
text = goto_re.sub(goto_repl, text)
# Remaining returns are exits caused by non-fused calls, unsupported paths,
# or chain fallback. Unwind the logical fused tail frames first.
text = text.replace('return;', 'TIER2_SB_RETURN();')
return text, fused, len(cold_exits), len(cold_dispatch_fallbacks)
# External calls / unsupported exits remain the SAFE AOT implementation and
# simply unwind any pending fused tail frames before returning.
text = re.sub(r'(?<!TIER2_SB_)\breturn;', 'TIER2_SB_RETURN();', text)
return text
def patch_hooks(text: str, unit: int) -> str:
text = strip_old_hooks(text)
if '#include "vcs_tier2_superblocks.hpp"' not in text:
def emit_cluster(cluster: Cluster, selected: Mapping[int, Set[int]],
sources: Mapping[int, UnitSource], host: pathlib.Path) -> Tuple[pathlib.Path, Dict[str, int]]:
owner = selected_pc_owner(selected)
selected_all = set(owner)
stats: Dict[str, int] = collections.Counter()
continuations: Dict[Tuple[int, int], int] = {}
chunks: List[str] = []
for unit in sorted(selected):
src = sources[unit]
for pc in src.ordered_pcs:
if pc not in selected[unit]:
continue
chunks.append(transform_block(src.blocks[pc], unit, selected, sources, stats, continuations))
# Entry hooks are intentionally narrower than the selected closure: only
# measured roots pay the extra branch into Tier-2.
entry_cases: List[str] = []
for unit, pcs in cluster.hooks.items():
for pc in pcs:
if pc not in selected.get(unit, set()):
raise RuntimeError(f'{cluster.key}: hook 0x{pc:08X} is outside selected closure')
entry_cases.append(f' case 0x{pc:08X}u: goto SB_L_{pc:08X};')
# Unit-specific local dispatch tables. These handle guest JR $ra and any
# computed local jump that was previously routed through LOCAL_DISPATCH.
unit_dispatch: List[str] = []
for unit in sorted(selected):
local_cases = '\n'.join(
f' case 0x{pc:08X}u: goto SB_L_{pc:08X};'
for pc in sorted(selected[unit]))
unit_dispatch.append(f'''TIER2_LOCAL_DISPATCH_U{unit:04d}:
if (tier2_return_depth != 0u && local_pc == tier2_return_pc[tier2_return_depth - 1u]) {{
tier2_resume_pc = local_pc;
tier2_resume_unit = tier2_return_unit[tier2_return_depth - 1u];
// Match invoke_chained_direct(): when the callee has returned, ctx.pc
// already contains the caller continuation before any starvation
// boundary/accounting can run. A scheduler switch here must never see
// the stale callee PC.
ctx.pc = local_pc;
const std::uint32_t tier2_pending_base =
tier2_return_pending_base[tier2_return_depth - 1u];
bool tier2_same_context = true;
if (tier2_pending_transfers < tier2_pending_base) {{
++tier2_stats.fallbacks;
ctx.pc = local_pc;
TIER2_SB_RETURN();
}}
const std::uint32_t tier2_nested_tail =
tier2_pending_transfers - tier2_pending_base;
if (tier2_nested_tail != 0u) {{
tier2_pending_transfers = tier2_pending_base;
if (!rt.tier2_complete_fused_transfers(ctx, tier2_nested_tail))
tier2_same_context = false;
}}
--tier2_return_depth;
if (!rt.tier2_complete_fused_transfers(ctx, 1u))
tier2_same_context = false;
if (!tier2_same_context) TIER2_SB_RETURN();
goto TIER2_FUSED_RETURN_DISPATCH;
}}
switch (local_pc) {{
{local_cases}
default:
ctx.pc = local_pc;
TIER2_SB_RETURN();
}}
''')
global_resume_cases = '\n'.join(
f' case 0x{pc:08X}u: goto SB_L_{pc:08X};' for pc in sorted(selected_all))
cold_resume_by_unit: Dict[int, List[Tuple[int, int]]] = collections.defaultdict(list)
for (caller_unit, cont), entry in sorted(continuations.items()):
if cont in selected_all:
continue
if entry >= 0:
cold_resume_by_unit[caller_unit].append((cont, entry))
cold_resume_sections: List[str] = []
for unit, entries in sorted(cold_resume_by_unit.items()):
cases = '\n'.join(
f' case 0x{pc:08X}u: ++tier2_stats.cold_exits; tier2_scope.finish(); '
f'psprecomp::recomp_unit_{unit:04d}_entry(rt, ctx, {entry}u, aot_mem); TIER2_SB_RETURN();'
for pc, entry in entries)
cold_resume_sections.append(f''' case {unit}u:
switch (tier2_resume_pc) {{
{cases}
default: break;
}}
break;''')
cpp = f'''// AUTO-GENERATED by profiles/vcs/tools/build_tier2_superblocks.py.
// Tier-2 SUPERBLOCK V2 cluster: {cluster.key}
#include "vcs_tier2_superblocks.hpp"
#include "psprecomp/runtime.hpp"
#include "generated_units.hpp"
#include <cstdint>
namespace vcs {{
using namespace psprecomp;
void {cluster.function}(psprecomp::Runtime &rt,
psprecomp::AllegrexContext &ctx,
psprecomp::GuestMemory::AotFastView &aot_mem,
std::uint32_t entry_pc) {{
tier2_detail::SampleScope tier2_scope(Tier2ClusterId::{cluster.enum_name});
auto &tier2_stats = tier2_scope.stats();
constexpr std::uint32_t kTier2ReturnCapacity = 32u;
std::uint32_t tier2_pending_transfers = 0u;
std::uint32_t tier2_return_depth = 0u;
std::uint32_t tier2_return_pc[kTier2ReturnCapacity]{{}};
std::uint32_t tier2_return_unit[kTier2ReturnCapacity]{{}};
std::uint32_t tier2_return_pending_base[kTier2ReturnCapacity]{{}};
std::uint32_t tier2_resume_pc = 0u;
std::uint32_t tier2_resume_unit = 0u;
std::uint32_t jump_target = 0u;
std::uint32_t local_pc = 0u;
std::uint32_t local_transfers = 0u;
std::uint32_t entry_id = 0u;
#define TIER2_SB_RETURN() do {{ \
tier2_scope.finish(); \
/* Unwind every logical invoke_chained_direct frame in true LIFO \
order. Tail frames created inside a fused JAL unwind before that \
JAL; outer JAL frames are also released after context invalidation. */ \
while (tier2_return_depth != 0u) {{ \
std::uint32_t tier2_base_ = tier2_return_pending_base[tier2_return_depth - 1u]; \
if (tier2_base_ > tier2_pending_transfers) {{ \
++tier2_stats.fallbacks; \
tier2_base_ = tier2_pending_transfers; \
}} \
const std::uint32_t tier2_tail_count_ = tier2_pending_transfers - tier2_base_; \
tier2_pending_transfers = tier2_base_; \
if (tier2_tail_count_ != 0u) \
(void)rt.tier2_complete_fused_transfers(ctx, tier2_tail_count_); \
--tier2_return_depth; \
(void)rt.tier2_complete_fused_transfers(ctx, 1u); \
}} \
if (tier2_pending_transfers != 0u) {{ \
(void)rt.tier2_complete_fused_transfers(ctx, tier2_pending_transfers); \
tier2_pending_transfers = 0u; \
}} \
return; \
}} while (false)
goto TIER2_ENTRY_DISPATCH;
TIER2_FUSED_RETURN_DISPATCH:
switch (tier2_resume_pc) {{
{global_resume_cases}
default: break;
}}
switch (tier2_resume_unit) {{
{chr(10).join(cold_resume_sections)}
default: break;
}}
ctx.pc = tier2_resume_pc;
TIER2_SB_RETURN();
{chr(10).join(unit_dispatch)}
TIER2_ENTRY_DISPATCH:
switch (entry_pc) {{
{chr(10).join(entry_cases)}
default:
ctx.pc = entry_pc;
++tier2_stats.fallbacks;
TIER2_SB_RETURN();
}}
{chr(10).join(chunks)}
#undef TIER2_SB_RETURN
}}
}} // namespace vcs
'''
path = host / f'vcs_tier2_cluster_{cluster.key}.cpp'
old = path.read_text(encoding='utf-8') if path.exists() else None
if old != cpp:
path.write_text(cpp, encoding='utf-8', newline='\n')
stats['host_changed'] = 1
else:
stats['host_changed'] = 0
stats['blocks'] = sum(len(v) for v in selected.values())
stats['lines'] = len(cpp.splitlines())
stats['units'] = len(selected)
stats['hooks'] = sum(len(v) for v in cluster.hooks.values())
return path, stats
def patch_hooks(raw: str, unit: int, hook_specs: Sequence[Tuple[Cluster, int]]) -> str:
text = strip_hooks(raw)
if HEADER_INCLUDE not in text:
text = text.replace('#include "generated_units.hpp"\n',
'#include "generated_units.hpp"\n#include "vcs_tier2_superblocks.hpp"\n', 1)
for pc in HOOK_PCS[unit]:
'#include "generated_units.hpp"\n' + HEADER_INCLUDE, 1)
# Multiple clusters may hook the same unit but never the same PC in V2.
for cluster, pc in hook_specs:
label = f'L_{pc:08X}:\n'
if label not in text:
raise RuntimeError(f'missing hook label {pc:08X} in unit {unit:04d}')
raise RuntimeError(f'unit {unit:04d}: hook label 0x{pc:08X} missing')
hook = (
label + HOOK_BEGIN +
' if (vcs::tier2_superblocks_enabled()) {\n'
f' vcs::tier2_superblock_154_155(rt, ctx, aot_mem, 0x{pc:08X}u);\n'
' return;\n'
' }\n' + HOOK_END
f' if (vcs::tier2_cluster_enabled(vcs::Tier2ClusterId::{cluster.enum_name})) {{\n'
f' vcs::{cluster.function}(rt, ctx, aot_mem, 0x{pc:08X}u);\n'
f' return;\n'
f' }}\n' + HOOK_END
)
text = text.replace(label, hook, 1)
return text
@@ -132,93 +556,57 @@ def main() -> int:
ap.add_argument('profile', nargs='?', type=pathlib.Path,
default=pathlib.Path(__file__).resolve().parents[1])
args = ap.parse_args()
profile = args.profile.resolve()
requested = args.profile
# Convenience: from the repository root, both `vcs` and `profiles/vcs`
# resolve to the VCS profile. The build BAT still passes an absolute path.
if requested == pathlib.Path('vcs'):
profile = pathlib.Path(__file__).resolve().parents[3] / 'profiles' / 'vcs'
else:
profile = requested.resolve()
generated = profile / 'generated'
host = profile / 'host'
raw154 = strip_old_hooks((generated/'generated_unit_0154.cpp').read_text(encoding='utf-8'))
raw155 = strip_old_hooks((generated/'generated_unit_0155.cpp').read_text(encoding='utf-8'))
r154 = region(raw154, U154_START, U154_END)
r155 = region(raw155, U155_START, U155_END)
e154 = direct_entry_map(raw154)
e155 = direct_entry_map(raw155)
t154, f154, c154, d154 = transform_region(
r154, 154, 155, U154_START, U154_END, U155_START, U155_END, e154)
t155, f155, c155, d155 = transform_region(
r155, 155, 154, U155_START, U155_END, U154_START, U154_END, e155)
all_units = sorted({u for c in CLUSTERS for u in (set(c.seeds) | set(c.windows) | set(c.hooks))})
sources = {u: parse_unit(generated / f'generated_unit_{u:04d}.cpp', u) for u in all_units}
cpp = f'''// AUTO-GENERATED by profiles/vcs/tools/build_tier2_superblocks.py.
// Profile-guided second-layer superblock: 0154/0155 dominant entity loop.
#include "vcs_tier2_superblocks.hpp"
#include "vcs_runtime_log.hpp"
#include "psprecomp/runtime.hpp"
#include "generated_units.hpp"
total = collections.Counter()
selected_by_cluster: Dict[str, Dict[int, Set[int]]] = {}
generated_paths: List[pathlib.Path] = []
for cluster in CLUSTERS:
selected = selected_for_cluster(cluster, sources)
selected_by_cluster[cluster.key] = selected
path, stats = emit_cluster(cluster, selected, sources, host)
generated_paths.append(path)
print(f'Tier2 V2 cluster {cluster.key}: ' + ' '.join(f'{k}={v}' for k, v in sorted(stats.items())))
total.update(stats)
#include <cstdlib>
#include <cstring>
hooks_by_unit: Dict[int, List[Tuple[Cluster, int]]] = collections.defaultdict(list)
for cluster in CLUSTERS:
for unit, pcs in cluster.hooks.items():
for pc in pcs:
hooks_by_unit[unit].append((cluster, pc))
namespace vcs {{
hook_units_changed = 0
for unit, specs in sorted(hooks_by_unit.items()):
path = generated / f'generated_unit_{unit:04d}.cpp'
patched = patch_hooks(sources[unit].raw, unit, specs)
if path.read_text(encoding='utf-8') != patched:
path.write_text(patched, encoding='utf-8', newline='\n')
hook_units_changed += 1
// Generated AOT call targets are declared in namespace psprecomp. The copied
// region text intentionally stays byte-close to the original units, so make
// those names visible here instead of rewriting every external helper target.
using namespace psprecomp;
# Remove the V1 generated body so CMake can no longer accidentally compile
# both generations. The common runtime/toggle implementation keeps this name.
legacy = host / 'vcs_tier2_superblocks.cpp'
# common file is handwritten by V2 and must remain; generated cluster files
# carry all hot code.
bool tier2_superblocks_enabled() noexcept {{
static const bool enabled = [] {{
const char *value = std::getenv("PSPRECOMP_TIER2_SUPERBLOCKS");
return value == nullptr || std::strcmp(value, "0") != 0;
}}();
return enabled;
}}
void tier2_superblock_154_155(psprecomp::Runtime &rt,
psprecomp::AllegrexContext &ctx,
psprecomp::GuestMemory::AotFastView &aot_mem,
std::uint32_t entry_pc) {{
std::uint32_t tier2_pending_transfers = 0u;
#define TIER2_SB_RETURN() do {{ \\
if (tier2_pending_transfers != 0u) {{ \\
(void)rt.tier2_complete_fused_transfers(ctx, tier2_pending_transfers); \\
tier2_pending_transfers = 0u; \\
}} \\
return; \\
}} while (false)
switch (entry_pc) {{
case 0x08A6E894u: goto SB_L_08A6E894;
case 0x08A6E8A4u: goto SB_L_08A6E8A4;
case 0x08A71100u: goto SB_L_08A71100;
case 0x08A711E0u: goto SB_L_08A711E0;
case 0x08A711ECu: goto SB_L_08A711EC;
default: ctx.pc = entry_pc; TIER2_SB_RETURN();
}}
{t154}
{t155}
#undef TIER2_SB_RETURN
}}
}} // namespace vcs
'''
out = host/'vcs_tier2_superblocks.cpp'
changed_cpp = not out.exists() or out.read_text(encoding='utf-8') != cpp
if changed_cpp:
out.write_text(cpp, encoding='utf-8', newline='\n')
changes = 0
for unit, raw in ((154, raw154), (155, raw155)):
patched = patch_hooks(raw, unit)
p = generated/f'generated_unit_{unit:04d}.cpp'
if p.read_text(encoding='utf-8') != patched:
p.write_text(patched, encoding='utf-8', newline='\n')
changes += 1
print('Tier2 SUPERBLOCK V1:',
f'host_changed={int(changed_cpp)} hook_units_changed={changes}',
f'fused_tail_edges={f154+f155} cold_exit_labels={c154+c155}',
f'cold_dispatch_fallbacks={d154+d155} lines={len(cpp.splitlines())}')
print('Tier2 SUPERBLOCK V2 COMPLETE:',
f'clusters={len(CLUSTERS)} hook_units_changed={hook_units_changed}',
f'blocks={total["blocks"]} lines={total["lines"]}',
f'fused_calls={total["fused_calls"]} fused_tail={total["fused_tail"]}',
f'cold_exits={total["cold_exits"]}')
return 0
if __name__ == '__main__':
raise SystemExit(main())