Petals is a system for running large language models at home in a decentralized manner. It allows users to generate text and fine-tune models such as Llama 3.1 up to 405 billion parameters, Mixtral 8x22B, Falcon 40B and larger, and BLOOM 176B. The approach follows a BitTorrent-style peer-to-peer network where each participant loads a portion of the model and contributes by serving other parts to the network.
Inference performance reaches up to six tokens per second for Llama 2 70B in single-batch mode and up to four tokens per second for Falcon 180B. This speed supports use in chatbots and interactive applications. The system extends beyond standard API interfaces by permitting any fine-tuning and sampling methods, custom paths through the model, and access to hidden states. It combines the convenience of an API with the flexibility of PyTorch and Hugging Face Transformers.
Petals is offered as a tool that can run on consumer-grade GPUs or within Google Colab notebooks. It forms part of the BigScience research workshop. Documentation and source code are available on GitHub, with development discussion taking place in Discord. A network status display shows current top contributors who are providing GPU resources.
In the AI & ML space, Petals takes a focused approach. It focuses on running and fine-tuning massive LLMs that exceed single-GPU memory limits without relying on expensive centralized APIs. It is built as an open-source project for developers and researchers. Petals is open source under the MIT license. It runs on the web and API, and it can be self-hosted.
It is developed by BigScience, and it first shipped in 2022. The project is developed in the open on GitHub with 10.3k stars. Key capabilities include Distributed Inference, Model Fine-tuning, and BitTorrent Network. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Same category — not a similarity match