|
Download README.md from AethronPhantom/matrix_mult: direct link, hf CLI and curl.
- Browser
- Download file 1.12 kB
-
https://huggingface.co/kernels/AethronPhantom/matrix_mult/resolve/main/README.md
- Command line
-
hf download hf://AethronPhantom/matrix_mult/README.md
-
curl -L -o README.md https://huggingface.co/kernels/AethronPhantom/matrix_mult/resolve/main/README.md
1.12 kB
metadata
license: apache-2.0
tags:
- cuda
- kernel
- gpu-optimization
- hpc
matrix_mult
CUDA matrix multiplication kernel (compile-only baseline).
This repository contains the standalone CUDA source for the matrix_mult lane from
the PyC kernel lab. It is a source artifact for inspection and benchmarking;
it is not a precompiled binary and the result below is not a universal ranking.
Performance
| Kernel | GPU / architecture | Shape | Best recorded result | Evidence |
|---|---|---|---|---|
matrix_mult |
not recorded | not recorded | Not measured in the published campaign | No published performance receipt was found for this lane. |
The result is reported with the original campaign's timing and correctness context. Compare kernels only when GPU, CUDA version, matrix shape, warmup, repeats, and reference/correctness mode match.
Source
kernel.cu— copied fromkernels/prototypes/baseline/matmul/kernel.cu.- Original lane tags:
cuda, matmul.
Build/run contract
{nvcc} -O3 -c {source} -o {build_dir}/{name}.o
(compile-only)