Skywork-Reward-V2-Qwen3-0.6B is a compact 0.6 billion parameter reward model built on the Qwen3 architecture. It is designed to evaluate and score text outputs for use in reinforcement learning from human feedback (RLHF) and Direct Preference Optimization (DPO) pipelines. The model is openly available on Hugging Face for researchers and developers fine-tuning smaller language models.
Skywork Reward V2 Qwen3 0.6B sits in PulseGate's Other AI category. It focuses on scoring and ranking model outputs to align language models with human preferences at small scale. Skywork Reward V2 Qwen3 0.6B is an open-source project aimed at developers. The project is open source (Open Source). It runs on the web and API.
Skywork builds and maintains Skywork Reward V2 Qwen3 0.6B, and it first shipped in 2025. The project is developed in the open on GitHub with 152 stars.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do