RunAnywhere is an open-source platform focused on AI model inference optimized for consumer hardware, such as Apple GPUs and Qualcomm NPUs. It provides hand-written GPU and NPU kernels, bypassing generic abstraction layers to maximize performance on specific silicon. The platform enables running models for tasks including large language models (LLM), vision language models (VLM), speech-to-text (STT), text-to-speech (TTS), and embeddings directly on a wide range of devices.
Central to RunAnywhere's approach is a single C++ core that handles inference, model management, routing, and telemetry, serving as the foundation for all supported platforms. This core is exposed through six open-source SDKs, offering thin bindings for Swift, Kotlin, React Native, Flutter, TypeScript, and C++. These SDKs ensure that applications across iOS, Android, macOS, Windows, Linux, web, and embedded environments can access the same models, API, and performance characteristics. The hosted console complements the SDKs by providing capabilities for over-the-air model updates and fleet management.
The platform includes specialized inference engines such as MetalRT for Apple GPUs and QHexRT for Qualcomm Hexagon NPUs. MetalRT features custom-written Metal kernels and supports tasks like speech-to-speech and vision language model inference, with benchmarks demonstrating notable speed improvements on Apple silicon. QHexRT enables running LLM, VLM, STT, TTS, and embeddings entirely on Qualcomm NPUs, with performance metrics published for transparency.
RunAnywhere targets developers and organizations seeking to deploy on-device AI across diverse hardware, emphasizing open-source accessibility and hardware-native performance. The entire stack above the kernel, including the core and SDKs, is open source, with resources available for integration and fleet management. This positions RunAnywhere as a research-driven, cross-platform AI inference solution for applications requiring speed, privacy, and hardware efficiency.
RunAnywhere is an Inference & model serving project. It enables developers to run AI inference efficiently across devices, cloud, and edge environments. RunAnywhere is an open-source project aimed at developers. RunAnywhere is open source under the Open Source license. RunAnywhere is available on the web, iOS, and Android, and it can be self-hosted.
It is developed by RunAnywhere (United States), and it first shipped in 2025. The project is developed in the open on GitHub with 10.3k stars and 1.6k commits in the last 90 days. Key capabilities include AI inference engines, Cross-platform SDKs, and On-device AI.
What PulseGate has recorded for this listing
Same category — not a similarity match