Skip to main content
Print

VFPU Architecture for Binary Translation

Type: REFERENCE   Status: MAINTAINED   Scope: public VFPU architecture and translation hazards

The VFPU is not merely “128 floats.” It is a physical register file with overlapping scalar, vector, and matrix views, encoded addressing, prefix state, lane masks, and instruction-family-specific behavior. A recompiler can use host SIMD as an optimization, but correctness starts from the physical file and its aliasing rules.

How We Know

physical v[128]
  |-- scalar view: one element
  |-- vector view: 2/3/4 lanes with encoded row/column progression
  |-- matrix view: 2x2 / 3x3 / 4x4 aliases
  |-- prefixes: transform operands and destination writes
  '-- masks/overlap: preserve ordering and untouched lanes
Source-owned conceptual VFPU register/view diagram.

Physical file and views

PRIOR ART: public VFPU references describe 128 physical 32-bit registers conventionally grouped as eight 4×4 matrices. Assembly names, encoded E-fields, physical indices, and matrix/vector helper terminology are related but not interchangeable.

IMPLEMENTATION GUIDANCE: store a physical 128-element file or an exactly equivalent aliasing model. High-level vector objects alone make overlapping writes and mixed scalar/vector access difficult to reason about.

  • Define a single index function per scalar/vector/matrix encoding family.
  • Document row/column progression, transpose/orientation, wrap behavior, and width.
  • Keep textual register naming separate from physical storage indexing.

Prefixes and destination masks

VFPU prefixes can swizzle, negate, take absolute values, inject constants, saturate, and mask destination lanes. Prefix state is architectural state, not a compiler hint. The translator must apply transformations in the documented order and consume or preserve state exactly as the instruction family specifies.

Destination masking is especially easy to get wrong: a masked lane must remain untouched, not be written with a computed value and restored later. Rejected or unsupported instruction forms must not accidentally mutate prefix state.

Instruction families and translation hazards

  • Scalar arithmetic can alias a later vector/matrix view.
  • Vector memory operations combine alignment, lane ordering, and guest-span validation.
  • Conversions and packing are bit-level operations, not host float casts by default.
  • Transcendentals and reciprocal/square-root helpers can diverge across host math, interpreter, JIT, and hardware.
  • Condition codes and control registers can be observed by later instructions.

The practical lowering strategy is to bring up a readable helper/interpreter path first, then add host SIMD only behind tests that compare the complete guest state, including untouched lanes and control state.

Public prior art versus hardware confirmation

The broad architecture above is public prior art. The finite measurements in VFPU Register Addressing are bounded hardware confirmation of selected scalar and wide encodings. The 46-input result comparison in VFPU Transcendentals and Denormals is a separate hardware/differential article. Neither article proves every VFPU encoding or input.

Do Not Infer

  • A 128-register model does not settle every aliasing or prefix rule.
  • One measured E-field or one operation digest does not establish the full VFPU.
  • PPSSPP interpreter/JIT behavior is implementation context, not automatic hardware authority.
  • Host SIMD lane order and exception behavior are not PSP semantics without an equivalence proof.

Open questions

Priority gaps include exhaustive or stratified wide-E coverage, prefix consumption on rejected forms, matrix addressing edge cases, vector load/store boundaries, condition-code behavior, and cross-model/firmware replication. Track these in Open Research Questions rather than silently filling them with emulator behavior.

Primary sources

Table of Contents