Skip to content
#

sm120a

Here are 3 public repositories matching this topic...

Language: All
Filter by language

Rust + CUDA inference engine for NVIDIA RTX PRO 6000 Blackwell and RTX 5090. Serves safetensors and GGUF over an OpenAI-compatible API, with per-device tuned defaults and speculative decode gated byte-identical to plain decode. Hosted instance: inference.tiyuvta.ai

  • Updated Aug 23, 2026
  • Rust

Improve this page

Add a description, image, and links to the sm120a topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the sm120a topic, visit your repo's landing page and select "manage topics."

Learn more