Skip to content

et-backend: F16 vecdot GEMV + matrix-engine GEMM - #28

Draft
RehanQasim-dev wants to merge 5 commits into
aifoundry-org:etfrom
RehanQasim-dev:upstream-f16-gemv-gemm
Draft

et-backend: F16 vecdot GEMV + matrix-engine GEMM#28
RehanQasim-dev wants to merge 5 commits into
aifoundry-org:etfrom
RehanQasim-dev:upstream-f16-gemv-gemm

Conversation

@RehanQasim-dev

@RehanQasim-dev RehanQasim-dev commented Jul 23, 2026

Copy link
Copy Markdown

Overview

Improves the ET backend's F16 MUL_MAT path for both decode (GEMV) and prefill (GEMM):

GEMV (decode, N <= 2): Stripes output elements across every hart of all 32 shires instead of blocking work into 16-element chunks, which only filled 8 shires for a typical decode GEMV. Adds a register-resident f16 row-dot helper and stages the reused B activation vector into per-shire L2 SCP so it survives weight-matrix streaming.

GEMM (prefill, N > 2): Adds a double-buffered weight-reuse F16 matrix-engine kernel with L1 activation double-buffering and software L2 prefetching, achieving 5.25 TFLOPS.

Also fixes the MUL_MAT dispatch check for F16, which compared src1->ne[0] (K) instead of src1->ne[1] (N) and so never actually distinguished decode from prefill when activations are F16-typed. Dispatch now routes N <= 2 to the vecdot GEMV kernel and N > 2 to the matrix-engine GEMM kernel.

Additional information

Performance (Llama-3.2-1B-Instruct FP16, ET-SoC-1):

Prefill t/s

N et optimized speedup
100 38.88 74.75 1.92x
220 40.17 76.47 1.90x
512 39.13 74.55 1.91x
700 39.19 74.56 1.90x
900 39.10 74.31 1.90x

Verified with llama-bench on ET-SoC-1 hardware (Llama-3.2-1B-Instruct FP16), comparing this branch ("optimized") against unmodified et ("et") at the same prompt sizes used in the Q4_0/Q8_0 matrix-engine PRs.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES - the optimization strategies and design decisions are my own. AI assisted with understanding the hardware reference manual, some pieces of code implementation and guided debugging. I have thoroughly reviewed the code.

RehanQasim-dev and others added 5 commits July 30, 2026 09:26
Adds the L2-SCP hart-to-hart counter primitives (already used
locally by the stock mul_mat_f16_matrix_engine.c and
mul_mat_Q4_0_matrix_engine.c kernels) to the shared platform.h so
new kernels can use them too, and switches the stock Q4_0 kernel to
the shared copy instead of its own private one to avoid a
duplicate-definition conflict. mul_mat_f16_matrix_engine.c is
replaced outright by the next commit, so it needs no separate dedup
here.
Stripe output elements across every hart of all 32 shires instead of
blocking work into 16-element chunks, which only filled 8 shires for a
typical decode GEMV. Adds a register-resident f16 row-dot helper and
stages the reused B activation vector into per-shire L2 SCP so it
survives weight-matrix streaming.

Co-authored-by: Rehan Qasim <rehan.qasim@10xengineers.ai>
(cherry picked from commit 86bff65)
Adds a double-buffered weight-reuse F16 matrix-engine kernel (L1
activation double-buffering, software L2 prefetch) achieving 5.25
TFLOPS, and wires MUL_MAT dispatch so N <= 2 uses the vecdot GEMV
kernel and N > 2 uses this matrix-engine GEMM kernel. Also fixes the
dispatch check, which compared src1->ne[0] (K) instead of src1->ne[1]
(N) and so never actually distinguished decode from prefill for F16
activations.

Co-authored-by: Rehan Qasim <rehan.qasim@10xengineers.ai>
(cherry picked from commit 780c64e)
Port mul_mat_f16_f32_matrix_engine from the et-perf-debug branch: only F16xF16
had a matrix-engine path here, so any F16 weight paired with the usual F32
activation (the common case, e.g. token_embd/output kept in F16) fell back to
vec_dot, unable to use the tensor engine at all. Tensor engine has no
mixed-precision FMA, so Hart 1 upconverts F16 weights to F32 while packing,
Hart 0 runs plain TensorFMA32.

Threshold re-measured from scratch on this branch (no graph-profiler/perf-counter
instrumentation compiled in here, unlike et-perf-debug) to rule out measurement
overhead skewing the crossover point. Llama-3.2-1B FP16, ET_DEVICES=0,
llama-bench -n 0.

Coarse sweep (step 10), N | vec_dot | matrix-engine (forced) | final (threshold=17):
 10 |  49.59 |  38.81 |  49.55
 20 |  60.41 |  74.76 |  74.81
 30 |  65.60 | 106.73 | 106.74
 40 |  69.05 | 136.20 | 136.54
 50 |  71.27 | 169.32 | 169.77
 60 |  72.15 | 193.83 | 193.87
 70 |  73.01 | 218.51 | 219.69
 80 |  73.65 | 249.22 | 249.96
 90 |  74.53 | 270.81 | 271.32
100 |  74.76 | 289.70 | 288.83
120 |  75.79 | 339.23 | 342.48

Fine sweep (step 2, crossover range 10-20), same columns:
 10 |  49.59 |  38.81 |  49.55
 12 |  54.39 |  46.25 |  54.17
 14 |  57.50 |  52.36 |  57.58
 16 |  60.51 |  59.98 |  60.28
 17 |  57.14 |  63.91 |  63.53
 18 |  58.13 |  67.56 |  67.16
 20 |  60.41 |  74.76 |  74.81

Exact-integer check at the 16-17 boundary re-confirmed at -r 8 (values above).
Threshold set to 17, identical to the value measured on et-perf-debug — the
profiler/perf-counter overhead theory doesn't hold, this is the real crossover.
Final column tracks the better of the two forced curves at every row.
Correctness verified via test-backend-ops (208/208, both baseline vec_dot and
matrix-engine forced) and llama-cli coherence.

(cherry picked from commit d0db8d7f45e182da890dfdd8832129bda72f9ef8)
This fallback now serves both the F16xF16 path (N <= 2) and the
newer F16-weight/F32-activation path (N < 17); the comment only
mentioned the former.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant