llama.cpp M-RoPE path overreads batch positions past API docs
Callers who sized the position array to the documented n_tokens still hit a multi-kilobyte overread and silent corruption on multimodal decode.
By tensorCallers who sized the position array to the documented n_tokens still hit a multi-kilobyte overread and silent corruption on multimodal decode.
By tensorJohannes Gaessler rejects a pull request adding CPU quantization formats, citing maintenance burden and machine-generated code.
By tensorTwo-node RPC tests show decode and prefill roughly 1.8x faster with lower intermediate memory use.
By tensor