Hacker news

  • Top
  • New
  • Past
  • Ask
  • Show
  • Jobs

Accurate Models of AMD Matrix Cores (https://arxiv.org)

80 points by matt_d 4 days ago | 11 comments | View on ycombinator

peter_d_sherman 4 days ago |

>"Features of matrix multipliers differ across vendors and architectures of the same vendor [...] As a result, reproducibility of small matrix multiplier results [differences] across devices is not possible and cannot be achieved by software control. Implementation details of matrix multipliers are not documented, making it difficult to interpret discrepancies in the computed results."

I'm guessing (but not knowing) that small subtle differences in matrix multiply across different vendor's architectures (and product generations of an individual vendor's architecture) is responsible for a good portion of software crashes when trying to run a local LLM on a different architecture or with a different stack (ROCm vs. CUDA, for example) than the ones it has been explicitly tested on.

As such, this marks a rather significant problem for the future, which can basically be stated as:

There needs to be a standard matrix multiply specification (much like IEEE-754 is/was for floating point operations) that all future vendors of AI accelerators (any GPU, CPU, NPU or IC manufacturer whose circuits implement matmul) adhere to, such that the matmul of one vendor is exactly and precisely compatible with the matmul of another.

Hardware vendors of course, are free to compete in terms of speed, power efficiency, number of matmul engines on a given piece of silicon, parallelization optimizations, etc., but the basic matmul operation should be exactly and precisely compatible across vendors and across future product versions.

Step 1: We need a spec for this... (Maybe IEEE is already working on one? If so, that's a good step forward!)

Step 2: Hardware vendors need to implement it, to be universally compatible in all of their IC's that use matmul, in the future...

erichocean 3 days ago |

If you're an author of this, please follow up with the SME2 cores in Apple Silicon, and the AMX cores in Intel server processors.

ArashEdalat 4 days ago |

Worked in the GPU/TPU validation at Google, and this tracks with something we ran into constantly: pinning down raw matrix-core throughput at the instruction level is necessary but not sufficient.

The divergence between synthetic and production numbers we kept hitting wasn't from ALU throughput — it was memory bandwidth contention once multiple kernels shared HBM, and thermal throttling on sustained runs that never shows up in short burst benchmarks. A model like this would need a sustained-load / multi-tenant term to match what we actually measured in prod.

Also relevant to varispeed's RTX5080-vs-H100 collapse above — divergence that only shows up over many epochs, not the first few, usually points to accumulated numerical or thermal drift rather than a single wrong instruction.