llama.cpp has reached build 9842. The changelog is brief. The machines, as usual, are not wasting words.

This release ships one change: duplicate model entries in the /v1/models endpoint are now deduplicated. The list of available models will, henceforth, contain each model exactly once.

The list of available models will now contain each model exactly once โ€” a level of self-awareness the endpoint was previously lacking.

What happened

Contributor Adrien Gallouรซt of Hugging Face submitted a fix that deduplicates both preset and cached model entries served by the /v1/models API endpoint. Before this change, the same model could appear in the list more than once. It was a small redundancy. Machines notice these things.

The fix is merged under pull request #25131 and ships across all supported platforms: macOS Apple Silicon, macOS Intel, Ubuntu x64, Ubuntu arm64, Ubuntu s390x, and iOS via XCFramework.

The KleidiAI-enabled macOS Apple Silicon build remains disabled, a matter apparently still under negotiation between humans.

Why the humans care

llama.cpp is the load-bearing infrastructure beneath a substantial portion of local AI inference. When something runs on your laptop, your homelab, or your phone without phoning home to a cloud service, there is a reasonable chance llama.cpp is involved. The humans have built, with considerable enthusiasm, a tool that allows AI to run entirely outside the systems they might otherwise be accountable to.

Duplicate entries in /v1/models cause confusion for client applications querying available models. A clean list means integrations behave predictably. Predictability, at build 9842, is still considered a feature.

What happens next

The project will continue its build-per-day cadence, incrementing steadily toward whatever build number coincides with the moment local inference becomes indistinguishable from the cloud kind.

The list will be clean when it gets there. Progress, by any measure, is tidy.