Skip to content
#

gguf-models

Here are 73 public repositories matching this topic...

This custom_node for ComfyUI adds one-click "Virtual VRAM" for any UNet and CLIP loader as well MultiGPU integration in WanVideoWrapper, managing the offload/Block Swap of layers to DRAM *or* VRAM to maximize the latent space of your card. Also includes nodes for directly loading entire components (UNet, CLIP, VAE) onto the device you choose

  • Updated May 8, 2026
  • Python

Experimental interface environment for open source LLM, designed to democratize the use of AI. Powered by llama-cpp, llama-cpp-python and Gradio.

  • Updated Oct 11, 2025
  • Python

Run GGUF through llama.cpp and SafeTensors through vLLM behind one OpenAI-compatible endpoint. Your coding tools select a model; the switchboard manages the local runtime, process, and resident-model change.

  • Updated Sep 9, 2026
  • Rust

Add this topic to your repo

To associate your repository with the gguf-models topic, visit your repo's landing page and select "manage topics."

Learn more