RiftStack researches and builds tools for inference optimization, focusing on compiler technology and serving techniques. Its open-source ML compiler Emmy transforms Torch IR through multiple representations down to optimized CUDA kernels. The lab publishes tutorials and benchmarks on GPU performance, memory virtualization for long-context inference, and kernel optimizations for RTX GPUs.
In the AI & ML space, RiftStack takes a focused approach. It focuses on running large language models efficiently and cost-effectively on consumer and commodity GPUs rather than only high-end datacenter hardware. RiftStack is an open-source project aimed at AI researchers and inference engineers. RiftStack is open source under the Apache-2.0 license. RiftStack is available on the web and API, and it can be self-hosted.
RiftStack builds and maintains RiftStack, and it first shipped in 2025. The project is developed in the open on GitHub with 67 stars and 332 commits in the last 90 days. Among its 5 catalogued features are ML Compiler, Kernel Optimization, and GPU Matmul Tuning.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Same category — not a similarity match