-
Notifications
You must be signed in to change notification settings - Fork 22k
All issues
Issue creation is restricted in this repository
Issues
is:issue state:open
is:issue state:open
Search results
- Status: Open.#27588 In ggml-org/llama.cpp;
- Status: Open.#27587 In ggml-org/llama.cpp;
- Status: Open.#27585 In ggml-org/llama.cpp;
Feature Request: efficient MoE serving with bandwidth-adaptive CPU–GPU co-execution ( q ⋆ policy), full-layer double-buffered prefill streaming, global LRU expert caching, graph-compatible execution, and the FTW fast weight format.
enhancementNew feature or requestNew feature or requestStatus: Open.#27584 In ggml-org/llama.cpp;- Status: Open.#27581 In ggml-org/llama.cpp;
- Status: Open.#27580 In ggml-org/llama.cpp;
- Status: Open.#27579 In ggml-org/llama.cpp;
- Status: Open.#27577 In ggml-org/llama.cpp;
- Status: Open.#27576 In ggml-org/llama.cpp;
- Status: Open.#27572 In ggml-org/llama.cpp;
Feature Request: --reasoning-budget like argument to control the budget based on the conversation length
enhancementNew feature or requestNew feature or requestStatus: Open.#27571 In ggml-org/llama.cpp;Feature Request: Model navigation and load controls in Router Model Info dialog (Follow-up to #27563)
enhancementNew feature or requestNew feature or requestStatus: Open.#27564 In ggml-org/llama.cpp;