bert-large-japanese-v2 is a large BERT model pretrained on a combination of CC-100 and jawiki data using whole word masking and unidic-lite tokenization. It serves as a foundation model for various Japanese natural language processing tasks. Available through Hugging Face Transformers, it supports direct loading for fine-tuning or feature extraction in research and production environments.
Bert Large Japanese sits in PulseGate's Other AI category. It focuses on providing a strong pretrained Japanese language model for downstream NLP tasks. Bert Large Japanese is an open-source project aimed at Japanese NLP researchers and developers. The project is open source (Apache-2.0). Bert Large Japanese is available on the web and API.
Behind Bert Large Japanese is Tohoku NLP, based in Japan, and the product first shipped in 2018. The GitHub repository has been archived. Among its 3 catalogued features are Pretrained BERT, Japanese Tokenization, and Whole Word Masking. It exposes integrations via a public API.
Latest indexed changes and source events
tohoku-nlp/bert-large-japanese-v2 verified by the PulseGate indexer
⚠ Archived