#
ik-llama-cpp
Here are 5 public repositories matching this topic...
-
Updated
May 27, 2026 - TypeScript
Turboquant Q4/Q3 with IQK FA
-
Updated
Jul 22, 2026 - C++
Turboquant Q4/Q3 with IQK FA
-
Updated
Apr 19, 2026 - Python
A one-command inference server and benchmark harness for running large language models that don't fit in your GPU's VRAM. Optimized for RTX PRO 6000 and Deepseek v4 flash
cuda nvidia rtx llm llama-cpp deepseek mxfp4 rtx-pro-6000 rtx-6000 deepseek-v4 ik-llama-cpp deepseek-v4-flash
-
Updated
Aug 21, 2026 - Shell
Improve this page
Add a description, image, and links to the ik-llama-cpp topic page so that developers can more easily learn about it.
Add this topic to your repo
To associate your repository with the ik-llama-cpp topic, visit your repo's landing page and select "manage topics."