VFPU Prefix Semantics and Vector/Matrix Encoding
Type: REFERENCE Status: MAINTAINED Scope: VFPU source/target/destination prefixes, lane transforms, masks, vector/matrix views, and encoding-safe translation. Last reviewed: 2026-08-11
The VFPU is a 128-element 32-bit register file with scalar, vector, and matrix views. A translator must preserve both the physical register mapping and the instruction’s prefix state. Treating a vector name as an independent register, or applying every prefix to every instruction, creates silent numerical corruption.
Architectural prior art
DavidGF’s PSP VFPU documentation describes the 128 scalar registers, eight 4×4 matrix views, control registers, and instruction-specific prefix support. The PSP Developer Wiki’s COP2/VFPU page independently records the reconfigurable scalar/vector/matrix model and its non-fully-IEEE floating-point caveats. PSPSDK’s exact pspvfpu.h reference is the smallest source for the public context and matrix helper surface.
decode instruction -> read PFXS/PFXT/PFXD snapshot<br>source lanes: swizzle -> abs/negate -> constants<br>execute operation at declared width<br>destination: mask -> saturation -> write physical lanes<br>consume/retain prefix state according to the instruction contract
Three prefixes, three responsibilities
- Source (S): lane swizzle, absolute value, negate, and constant selection alter source operands before the operation.
- Target/source (T): the second operand has its own transform and must not inherit S implicitly.
- Destination (D): saturation and write-mask behavior apply when committing result lanes; masked lanes retain their prior physical values.
Prefix lifetime and consumption are architectural state. A rejected or unsupported operation must not accidentally consume a prefix or partially update a masked destination. Conversely, silently ignoring a supported prefix can pass scalar smoke tests while breaking a wide instruction. Build an instruction support table rather than a global “prefixes enabled” switch.
Encoding and physical aliases
Vector and matrix encodings select overlapping physical lanes. The correct identity is the decoded physical register tuple plus the instruction width/orientation, not only the assembly spelling. Loads/stores, matrix rows/columns, and scalar E-field forms need checked bounds. Existing public reference material establishes the register-file concept; it does not make every undocumented wide encoding hardware-proven.
Nakagawa’s translation decisions are implemented in the current tools/codegen.py and related generated-runtime paths. Cite the commit-pinned source for implementation-specific statements, while the public architecture references remain the authority for the guest model.
Measured versus inferred
The site’s VFPU register-addressing page records a bounded hardware measurement, but it is not a survey of all 128 scalar encodings and all possible wide combinations. The transcendentals and denormals article is likewise a measured numeric slice. Neither result licenses a general claim about every prefix, matrix orientation, or invalid encoding.
Safe regression design
For each supported instruction, test identity prefixes, one transform at a time, destination masks, saturation boundaries, width changes, overlapping source/destination tuples, and an intentionally rejected form. Assert physical-lane outputs and prefix state after execution. Keep hardware comparisons to scalar outputs and aggregate results; do not publish private binaries or raw traces.