Strix Halo guide for AMD Ryzen AI MAX+ 395 / Radeon 8060S local LLM setup and benchmarks: Ollama, llama.cpp, Vulkan/RADV, ROCm, GGUF, and raw evidence.
-
Updated
Aug 22, 2026 - Python
Strix Halo guide for AMD Ryzen AI MAX+ 395 / Radeon 8060S local LLM setup and benchmarks: Ollama, llama.cpp, Vulkan/RADV, ROCm, GGUF, and raw evidence.
Operator-grade GPU monitor for NVIDIA GPUs with native GB10 / DGX Spark coherent UMA support — PSI pressure, clock detection, ConnectX-7 network layer
Local inference server for Apple Silicon — hot-swaps MLX models (LLM, vision, embeddings, TTS, STT) via OpenAI API
A CUDA implementation of the transpose-free Quasi-Minimal Residual method
Unified Memory Abstraction Layer for AI Inference on AMD APUs and Intel iGPUs
A turnkey, fully-local AI workstation engineered for the AMD Ryzen AI Max+ 395. LLM inference, voice, document parsing, browser automation, agents — all on-device.
gpu thrashingNVIDIA GPU Unified Memory diagnostic tool — architecture-aware, measurement-based, PCIe/coherent transport detection
Apple Silicon Unified Memory for GPU-Accelerated Analytics — TPC-H benchmarks across DuckDB, NumPy, and MLX
Talos-O (Omni): A sovereign, embodied agentic organism forged on AMD Strix Halo. Integrating the Chimera Kernel (Linux 7.0), Zero-Copy Introspection, and the Phronesis Engine. Built from First Principles.
Fundamentals of Accelerated Computing C/C++ is a course provided by NVIDIA.
NVML unified memory shim for NVIDIA DGX Spark Grace Blackwell GB10 - enables MAX Engine, PyTorch, and GPU monitoring
Unlock fast, local LLM inference on AMD-powered mini PCs delivering 65-87 t/s for large models without cloud or subscription costs
Honest local LLM deployment planning and benchmarking for high unified-memory Macs and future Linux/NVIDIA rigs.
Run LLMs larger than your RAM — native GGUF inference engine with SSD streaming, no GPU required
Empirical kernel scheduling characterization for NVIDIA GB10 (SM121a). Sweeps GEMM tile configurations, classifies PTX instruction paths, captures hardware telemetry
Research into CUDA Unified Memory as a VRAM extension for LLM inference
Performance comparison of two different forms of memory management in CUDA
The real-time coordination layer for teams of developers running Claude Code agents. Git coordinates code at rest; Datum coordinates agents in motion.
3D U-Net with tf.keras for Large-Model-Support or Unified Memory
GB10-aware CUPTI Activity collector — runtime kind detection, phase management, and JSON output for hardware-coherent UMA platforms
Add a description, image, and links to the unified-memory topic page so that developers can more easily learn about it.
To associate your repository with the unified-memory topic, visit your repo's landing page and select "manage topics."