Pinned Loading
-
llama-cpp-pascal-cuda-windows
llama-cpp-pascal-cuda-windows PublicCUDA compile guides for llama-cpp-python on legacy NVIDIA Pascal GPUs (GTX 1050/1060/1070/1080) — Windows, VS2019/CUDA 12.1 and VS2022/CUDA 12.6 paths
-
SovereignKernel
SovereignKernel PublicCustom C++/CUDA LLM inference runtime built from scratch. Standalone GGUF loading, RoPE, KV cache management, and custom kernels without cuBLAS or external high-level frameworks.
Cuda
-
turbovec-quantization-benchmark
turbovec-quantization-benchmark PublicDoes TurboVec quantization hold up on exact-fact retrieval, or only on semantic similarity? Tests recall@k across bit_width 4/2 vs full precision on a synthetic report with 104 ground-truth facts. …
Python
-
qwen-local-code-eval
qwen-local-code-eval PublicEmpirical benchmark of Qwen quantization tiers (IQ1, Q6, IQ4) for local code generation under strict 8GB VRAM limits.
-
local-whisper-transcription-pipeline
local-whisper-transcription-pipeline PublicLocal video/audio to subtitles pipeline. Zero cloud, VRAM-aware model selection, runs locally on any NVIDIA GPU. Built on faster-whisper and ffmpeg. Designed for readable orchestration logic, not d…
Python
-
sdxl-lora-gguf-blueprint
sdxl-lora-gguf-blueprint PublicBlueprint for training custom SDXL LoRAs and running local GGUF inference on 8GB VRAM — no cloud, no ComfyUI
If the problem persists, check the GitHub status page or contact support.