Vitpose Plus Large is a Vision Transformer model for keypoint detection hosted on Hugging Face by the University of Sydney. It forms part of the ViTPose family and addresses the task of locating human body keypoints in images.
The model is provided as a set of weights and associated code that can be loaded through the Transformers library. Users instantiate an AutoImageProcessor and a VitPoseForPoseEstimation class directly from the repository identifier usyd-community/vitpose-plus-large. It supports inference on local hardware or through compatible inference providers and is distributed with both model card documentation and example notebooks for Google Colab and Kaggle.
The repository lists two associated arXiv papers (2204.12484 and 2212.04246) that describe the underlying ViTPose approach. It carries an Apache-2.0 license, making the weights and code openly available for modification and redistribution. The page classifies the work under the keypoint detection task and the Transformers architecture while indicating use of the Safetensors format for the model files.
Vitpose Plus Large sits in PulseGate's Other AI category. It focuses on detecting human body keypoints and estimating poses from images. It is built as an open-source project for computer vision researchers and developers. The project is open source (Apache-2.0). It runs on the web and API.
Behind Vitpose Plus Large is University of Sydney, and it first shipped in 2022. The project is developed in the open on GitHub with 2.1k stars. Key capabilities include Pose Estimation, Keypoint Detection, and Vision Transformer. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do