AMD’s next-generation RDNA 5 graphics architecture could deliver roughly twice the low-precision matrix throughput of RDNA 4 and Nvidia’s current Blackwell gaming architecture on a per-SIMD basis, according to hardware leaker Kepler_L2. The claim was made in an AnandTech forum discussion about the architecture’s FP8 and FP4 matrix capabilities. AMD has not officially confirmed RDNA 5’s matrix throughput, so the figures should be treated as unverified pre-release information.
Asked what kind of FP8 and FP4 matrix performance could be expected compared with Nvidia’s RTX 50-series architecture, Kepler_L2 said RDNA 5 would operate at half the rate per SIMD of AMD’s gfx1250 architecture. He added that this would place it at approximately double the rate of gfx12 and Blackwell. Gfx12 refers to AMD’s current RDNA 4 generation, while gfx1250 is a newer AMD GPU target already appearing in software tooling and aimed at substantially heavier compute workloads.
The comparison is particularly notable because FP8 and FP4 arithmetic are increasingly important for machine-learning workloads used in modern graphics. Lower-precision matrix operations can accelerate neural networks used for image reconstruction, denoising and other rendering tasks while requiring less compute and memory bandwidth than higher-precision formats. Nvidia has increasingly relied on its Tensor Cores for these workloads, while AMD introduced expanded AI acceleration in RDNA 4 alongside FSR 4.

Kepler_L2’s claim would suggest another significant increase in AMD’s matrix capabilities with RDNA 5. Earlier comments from the leaker indicated that the architecture is expected to support FP4 operations and improve neural rendering performance beyond RDNA 4, while potentially adding substantially more advanced ray-tracing hardware. None of those architectural details have yet been formally disclosed by AMD.
The claim also does not mean an RDNA 5 GPU would automatically deliver twice the overall AI or gaming performance of a comparable Blackwell card. Per-SIMD matrix throughput represents only one portion of the GPU, while total performance will depend on the number of compute units, SIMD configuration, clock speeds, memory bandwidth, software utilization and how effectively games can use the available matrix hardware. Comparisons between AMD and Nvidia execution structures are also not one-to-one, making direct performance conclusions from a single architectural metric unreliable.
Still, stronger FP4 and FP8 support could matter significantly for future ray-traced and path-traced games. Modern neural rendering pipelines increasingly combine traditional graphics processing with machine-learning models for denoising, reconstruction and frame generation. Kepler_L2 has previously said RDNA 5 should accelerate ray denoising considerably compared with RDNA 4, with FP4 support potentially giving AMD more headroom for those workloads.
The leak fits with broader indications that AMD is making extensive changes to its next graphics architecture. LLVM already contains references to gfx13-class AMD GPU targets, while the ongoing AnandTech discussion has pointed to changes involving scheduling, cache design, execution structures and ray-tracing acceleration. Those references confirm AMD is developing newer GPU targets, but they do not independently verify the performance characteristics attributed to RDNA 5.
Timing for the architecture also remains uncertain. A separate Kepler_L2 leak claimed that most RDNA 5 GPUs could arrive in 2028, with an AMD chip known as AT2 potentially appearing earlier in 2027. AMD has not publicly confirmed that schedule or the final lineup.
For now, the new matrix-performance figure offers another indication that AMD may be placing substantially more emphasis on neural graphics with its next generation. Whether that translates into the claimed twofold per-SIMD advantage over RDNA 4 and Blackwell will depend on final silicon and software support, neither of which AMD has detailed publicly.

