Traceary

Catalog / llama.cpp

b10534

5 days agoaddedchangedOriginal notes

CUDA: adding switch points per HW and quant type to tune the mvq->MMQ decode crossover (#26079)

  • CUDA: runtime GGML_CUDA_MMVQ_MAX to tune the mvq->MMQ decode crossover

Add a runtime override of the mul_mat_vec_q -> MMQ batch crossover (default MMVQ_MAX_BATCH_SIZE). Lowering it routes batches above the threshold from the CUDA-core vector kernel to the int8 MMQ tensor-core path, which is faster once quantized decode becomes compute-bound at B>1 (measured +23-41% at B=8 on RTX 5090 for Q4_K dense, no low-batch loss).

The value is parsed once and clamped to [1, MMVQ_MAX_BATCH_SIZE], since mul_mat_vec_q asserts ncols_dst <= that; invalid input warns and falls back to the default. The override is applied consistently in both the mul_mat_vec_q and MUL_MAT_ID dispatch paths. Default behavior unchanged.

  • Added Blackwell specific switch point, to reduce dependence on runtime env var.

  • Add per-HW switch point values for DGX Spark and removing runtime env var

  • Adding switch points for Ada, tested on RTX 4090

  • Modifying DGX Spark numbers based on latest run and adding some comments and small functional changes relating to MoE

  • Reverting an unnecessary conditional

  • Update ggml/src/ggml-cuda/mmvq.cu


Co-authored-by: praneshgo [email protected] Co-authored-by: Oliver Simons [email protected]

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI: