Efficient LLM inference on Slurm clusters.
-
Updated
Aug 17, 2026 - Python
Efficient LLM inference on Slurm clusters.
PipelineLLM 是一个系统性的大语言模型(LLM)后训练学习项目,涵盖从监督微调(SFT)到偏好优化(DPO)、强化学习(RLHF/PPO/GRPO)再到持续学习(Continual Learning)的完整技术栈。
A practical, multi-layered JSON repair library for Elixir that intelligently fixes malformed JSON strings commonly produced by LLMs, legacy systems, and data pipelines.
Phi-Bench (Φ-Bench): 85 open-source LLM-infrastructure engineering tasks for frontier LLMs & coding agents (KFC/LH/E2E). Self-contained public Dockerfile + offline scoring + reference solution per task. Website: llminfrabench.com
Multi-model AI agent runtime. Define agents in YAML, route each role to a model, orchestrate with 7 patterns (ReAct, Plan & Execute, Fan-Out, Pipeline, Supervisor, Swarm, Glyph), and deploy as a REST/WebSocket API with RAG, memory, MCP tools, guardrails and OpenTelemetry observability.
AI Workload Control Layer for routing deterministic, reusable, retrieval-needed, tool-needed, and provider-needed work before model invocation.
A browser-based UI for launching, monitoring, clustering and managing multiple llama.cpp server instances from inside a Docker container. Includes an Ollama-compatible API proxy
Reliability control for the Anthropic Python SDK
Persistent Cognitive Memory Infrastructure — durable, multi-tenant memory layer for AI agents. HTTP + gRPC, hybrid retrieval, background workers, observability. Go.
nvProbe — Open-source NVIDIA GPU benchmark suite for CUDA workload automation, Slurm HPC cluster profiling, and MLPerf reporting
A production-grade, schema-aware PostgreSQL MCP server for enterprise AI. Features Zero-Trust SQL validation, multi-tier permissions, and real-time schema introspection for secure, autonomous database operations.
Krako 2.0 – Energy-efficient, triadic multi-tier inference infrastructure enabling adaptive routing across heterogeneous edge–cloud nodes.
Joule is a budget-controlled AI agent runtime that minimizes energy and token usage through hierarchical routing and deterministic tool execution.
One command. Full LLM stack. Zero config.
An intelligent gateway for Claude APIs that dynamically routes requests to the most cost-efficient model, caches responses, and escalates based on confidence signals — reducing LLM spend without sacrificing quality.
secrets.wtf: defensive AI infrastructure exposure index for Ollama and LM Studio hosts, local LLM APIs, model observations, remediation, and takedown requests.
A humble exploratory PoC for a hardware-native, optical timing-frozen control plane engine. Fuses inline CUDA PTX, RAII memory tunnels, and 4D JAX shard_map structures to investigate 0ns-overhead fault-tolerant routing for hyperscale distributed AI.
A Branchless, Zero-Jitter Ingress Router for 32-GPU Distributed Mesh Networks utilizing JAX/XLA and NCCL.
High-performance Triton kernels for NVIDIA H100. Implements fused FP8 LayerNorm, tiled FlashAttention, and SRAM-optimized memory primitives for Hopper architecture.
Add a description, image, and links to the llm-infrastructure topic page so that developers can more easily learn about it.
To associate your repository with the llm-infrastructure topic, visit your repo's landing page and select "manage topics."