A high-throughput and memory-efficient inference and serving engine for LLMs
-
Updated
Oct 4, 2026 - Python
A high-throughput and memory-efficient inference and serving engine for LLMs
๐ฅ MaxKB is an open-source platform for building enterprise-grade agents. ๅผบๅคงๆ็จ็ๅผๆบไผไธ็บงๆบ่ฝไฝๅนณๅฐใ
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.8, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM5.3, Gemma4, Llava, Phi4, ...) (AAAI 2025).
Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training. Reinforcement learning for Qwen3.6, GPT-OSS, Llama, and more!
Jev-like family of decision models built on top of Qwen3.5/3.8 you can train and run on your own
Open Source Deep Research Alternative to Reason and Search on Private Data. Written in Python.
๐A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.๐
A native macOS app that allows users to chat with a local LLM that can respond with information from files, folders and websites on your Mac without installing any other software. Powered by llama.cpp.
ASR/STT subtitle generator. Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD. Noise-robust for JAV
Serve large Qwen models fast on the GPUs you actually own. Qwen3.8-27B on a single 24 GB card with vLLM: 127 tok/s single-user (381 when the answer quotes the prompt), ~1,035 tok/s at 64 concurrent, 150k-262k context. vLLM patches, requant pipeline, benchmarks.
Chat2API enables zero-cost access to leading AI models by leveraging official web UIs. It supports providers such as DeepSeek, GLM, Kimi, MiniMax, Qwen, and Z.ai, and seamlessly integrates with tools like openlcaw, Cline, and Roo-Code.
AI-powered tool for efficient abstract and PDF screening in systematic reviews.
Fully Open Framework for Democratized Multimodal Training
๐ฌ Generate images from any camera viewpoint via 3D interactive control. Drag the camera in 3D space or use sliders to set azimuth/elevation/distance, then generate. Built with Three.js + Gradio, bilingual ZH/EN UI.
๐ Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support
The fastest way to run Qwen3.8-Flash-Next on Strix Halo (gfx1151)
๐๐๐A collection of some awesome public projects about Large Language Model(LLM), Vision Language Model(VLM), Vision Language Action(VLA), AI Generated Content(AIGC), the related Datasets and Applications.
MimikaStudio - A local-first application for macOS (Apple Silicon) + Agentic MCP Support
To associate your repository with the qwen3 topic, visit your repo's landing page and select "manage topics."