-
Notifications
You must be signed in to change notification settings - Fork 2.7k
All issues
Issue creation is restricted in this repository
- #15044 · laikhtewari opened
on Jun 6, 2026 1 - #3148 · juney-nvidia opened
on Mar 29, 2025 5 - #3124 · juney-nvidia opened
on Mar 27, 2025 11
Issues
is:issue state:open
is:issue state:open
Search results
[RFC] DFlash2 for Qwen3.8 on consumer Blackwell: integration boundary before we propose a PR
Speculative Decoding<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafter<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafterStatus: Open.#18085 In NVIDIA/TensorRT-LLM;MODEL_TYPE_TO_TOOL_PARSER maps qwen3_5/qwen3_5_moe to a parser whose format the template does not emit
LLM API<NV>High-level LLM Python API & tools (e.g., trtllm-llmapi-launch) for TRTLLM inference/workflows.<NV>High-level LLM Python API & tools (e.g., trtllm-llmapi-launch) for TRTLLM inference/workflows.Status: Open.#18084 In NVIDIA/TensorRT-LLM;Reasoning-parser auto-selection misclassifies reasoning-at-start templates, silently emptying reasoning_content
LLM API<NV>High-level LLM Python API & tools (e.g., trtllm-llmapi-launch) for TRTLLM inference/workflows.<NV>High-level LLM Python API & tools (e.g., trtllm-llmapi-launch) for TRTLLM inference/workflows.Status: Open.#18083 In NVIDIA/TensorRT-LLM;[Performance]: Deduplicate shared RoPE SMEM staging in fused DiT QK norm
General perf<NV>Broad performance issues not specific to a particular component<NV>Broad performance issues not specific to a particular componentStatus: Open.#18035 In NVIDIA/TensorRT-LLM;[Bug]: ClusterStorage.get return annotations exclude valid None results
LLM API<NV>High-level LLM Python API & tools (e.g., trtllm-llmapi-launch) for TRTLLM inference/workflows.<NV>High-level LLM Python API & tools (e.g., trtllm-llmapi-launch) for TRTLLM inference/workflows.Status: Open.#17977 In NVIDIA/TensorRT-LLM;[New Model]: Add TeleChat4 support with Multi-Head Hyper-Connection (mHC)
Model customization<NV>Adding support for new model architectures or variants<NV>Adding support for new model architectures or variantsnew modelRequest to add a new modelRequest to add a new modelSpeculative Decoding<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafter<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafterStatus: Open.#17965 In NVIDIA/TensorRT-LLM;[Bug]: Qwen3_5ForCausalLM (Qwen3.5-2B) incorrect output with batch_size > 1
bugSomething isn't workingSomething isn't workingCustomized kernels<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.Pytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesStatus: Open.#17964 In NVIDIA/TensorRT-LLM;Parallelize unfinished beam finalization across beams
Customized kernels<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.Status: Open.#17928 In NVIDIA/TensorRT-LLM;Optimize beam-search cache indirection for short generation
Customized kernels<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.Status: Open.#17927 In NVIDIA/TensorRT-LLM;KvCacheManagerV2 segfault on ~256k prompts with Nemotron-3-Ultra when max_seq_len=1M and avg_seq_len unset
KV-Cache Managementkv-cache management for efficient LLM inferencekv-cache management for efficient LLM inferencePytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesStatus: Open.#17926 In NVIDIA/TensorRT-LLM;A tool parser that matches its marker but extracts no calls returns 200 with no diagnostic
Inference runtime<NV>General operational aspects of TRTLLM execution not in other categories.<NV>General operational aspects of TRTLLM execution not in other categories.Status: Open.#17917 In NVIDIA/TensorRT-LLM;DeepSeekR1Parser ignores a per-request enable_thinking=False, returning empty content
Decoding/Sampling<NV>Token sampling algorithms in TRTLLM for text gen (top-k, top-p, beam).<NV>Token sampling algorithms in TRTLLM for text gen (top-k, top-p, beam).Status: Open.#17916 In NVIDIA/TensorRT-LLM;