llama.cpp fixes graph scheduling when switching token and image batches

A local-inference scheduler fix addresses graph composition across different input batches. Retrospective brief for 11:00–11:59 UTC on 29 September 2026.

Top AI stories from the last hour

Copy Markdown
  1. Graph input collection becomes stable across batch types

    The llama.cpp b11254 development build collects all input tensors into graph_inputs after graph splitting. The release notes explain that switching between batches using different inputs, such as token and image batches, could previously change copied graph composition, trigger unnecessary scheduler re-reservations and later cause an abort. The change makes composition depend on which inputs exist. This prerelease fix is relevant to pipeline-parallel and multimodal inference workloads.

Last checked 29 September 2026, 23:00 UTC

RSS