Repositories list
27 repositories
sglang
Publicruninfra-sdk
PublicOfficial RunInfra SDK (TypeScript + Python) | optimized inference deploymentslocal-kimi
PublicOptimized local serving engine for Kimi-Linear-48B: INT4 quantizer, fused decode kernels for a measured 3.18x, and an OpenAI-compatible server. Ships with k3, a…inference-cost-truth
PublicWhat LLM inference actually costs. 874 verified price rows across 24 providers, 378 GPU rental rates, 395 cited throughput datapoints, and a self-hosting break-…TIDE
PublicDynamic per-token early exit for LLM inference. Skip layers tokens don't needMemoir
PublicWeights that never freeze. A continual-learning architecture, and the pre-registered experiment that falsified its central claim. https://arxiv.org/abs/2607.207…inkling-turbo
PublicFaster attention kernels for serving TML's Inkling model on vLLM. 2.7x over the shipping path on H100, and the only implementation that runs on A100.hclsm
PublicHierarchical Causal Latent State Machines for Object-Centric World Modelingautotree
PublicTree execution engine for LLM inference: fork, merge, prune KV cache at token granularityAutoMegaKernel
PublicAn agent harness that compiles a model into one provably-correct, self-retargeting CUDA megakernel and self-tunes it past cuBLAS at batch-1 LLM decode, paper: h…autoevolve
PublicAgent-native evolutionary optimization. Say the goal in english, evolve code toward a measured target with a population of coding agents.runinfra-cli
Publicbonsai-turbo
PublicSingle-launch batch-1 decode engine for PrismML Bonsai 27B (ternary and 1-bit) on NVIDIA GPUs. 1.76x the vendor llama.cpp fork on H100, same outputs.auto
Publicthe agi compiler: records llm agent behavior, proves what repeats, and compiles it into verified, sandboxed wasm binaries that run for microdollars. nothing fig…openfang
PublicOpen-source Agent Operating SystemStreamIndex
PublicMemory-bounded compressed sparse attention via streaming top-k. Triton kernels for the DeepSeek-V4 lightning indexer. 32x regime extension on a single H200 | by…ouroboros
PublicDynamic weight generation for recursive transformers via input-conditioned LoRA modulationautokernel
Publicqwen3.5-triton
Publicpicolm
Publictiny-tpu
PublicMinimal TPU implementation with 8x8 systolic array and PyTorch integrationgpuci
PublicGPU CI/CD tool that tests CUDA kernels across multiple GPUs in parallel - Part of RightNowRightNow-GPU-Database
PublicRightNow-Tile
PublicOpen-source transpiler for CUDA Tile (13.1) migrationrightnow-cli
Publicgpu-profiler
Public- RightNow Arabic LLM Corpus - One of the largest high-quality Arabic text datasets for LLM training
ProTip! When viewing an organization's repositories, you can use the
props. filter to filter by custom property.