awq-quantization
Here are 5 public repositories matching this topic...
scaleDown AI: Enterprise Model Quantization Platform Slash inference costs by 70%. Deploy LLMs anywhere. A microservices-based orchestration engine for isolating and automating incompatible AI optimization workflows.
-
Updated
Jan 28, 2026 - Python
Embeddable Python Engine for VibeVoice TTS with AWQ-INT4 quantization (50% VRAM, 2x Faster RTF) of VibeVoice Large 7B model
-
Updated
May 7, 2026 - Python
Fine-tuning pipeline for Qwen2.5-Coder-1.5B using QLoRA, DPO alignment, and AWQ quantization on free-tier Colab GPU.
-
Updated
Aug 19, 2026 - Jupyter Notebook
A from-scratch implementation of Llama-3.2-1B in PyTorch, decode-latency benchmarks on three GPUs (T4, L4, A100), three weight-only quantization methods (RTN, GPTQ, AWQ) measured against both, and a packed int4 format with a fused Triton GEMV so the quantized weights are actually 4 bits in HBM.
-
Updated
Sep 3, 2026 - Jupyter Notebook
Add this topic to your repo
To associate your repository with the awq-quantization topic, visit your repo's landing page and select "manage topics."