Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
orchestra-research avatar

Sentencepiece

  • 400 installs
  • 11.2k repo stars
  • Updated June 16, 2026
  • orchestra-research/ai-research-skills

sentencepiece is an agent skill that trains and compares BPE versus Unigram SentencePiece tokenizers for developers who need correct vocabulary and subword regularization before fine-tuning or serving custom models.

About

sentencepiece is a tokenizer training skill from orchestra-research/ai-research-skills focused on the SentencePiece library's BPE and Unigram modes. It explains merge-based BPE training with worked corpus iterations—such as merging 'e'+'s' then 'es'+'t'—and contrasts Unigram probabilistic segmentation plus subword regularization tradeoffs. Developers reach for sentencepiece when building language-agnostic vocabularies, choosing BPE vs Unigram for a domain corpus, or generating `.model` files before Hugging Face or custom training. The guide includes Python snippets using `import sentencepiece as spm` and `spm.SentencePieceTrainer` patterns for reproducible vocabulary creation.

  • Compares BPE (merge-by-frequency) vs Unigram (probabilistic pruning) with concrete corpus walkthrough
  • Includes SentencePieceTrainer snippet for BPE with vocab_size control
  • Explains deterministic BPE splits vs Unigram sampling and subword regularization behavior
  • Covers when compression and training speed favor BPE vs when probabilistic tokenization helps
  • Grounds choices in implementation via the sentencepiece Python API

Sentencepiece by the numbers

  • 400 all-time installs (skills.sh)
  • +37 installs in the week ending Jul 18, 2026 (Skillselion tracking)
  • Ranked #492 of 2,066 Data Science & ML skills by installs in the Skillselion catalog
  • Security screen: HIGH risk (skills.sh audit)
  • Data as of Jul 28, 2026 (Skillselion catalog sync)
npx skills add https://github.com/orchestra-research/ai-research-skills --skill sentencepiece

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs400
repo stars11.2k
Security audit1 / 3 scanners passed
Last updatedJune 16, 2026
Repositoryorchestra-research/ai-research-skills

How do you train SentencePiece BPE vs Unigram tokenizers?

Train and choose BPE vs Unigram SentencePiece tokenizers with correct tradeoffs before fine-tuning or serving a custom vocabulary.

Who is it for?

NLP engineers standardizing tokenization with SentencePiece before fine-tuning multilingual or domain-specific LLMs.

Skip if: Projects already locked to a pretrained Hugging Face tokenizer with no custom vocabulary training step.

When should I use this skill?

An agent must train, compare, or configure SentencePiece BPE vs Unigram tokenizers for a new corpus.

What you get

Trained SentencePiece `.model` vocabulary, BPE-vs-Unigram tradeoff notes, and subword regularization configuration.

  • SentencePiece model file
  • BPE vs Unigram selection notes

By the numbers

  • Covers 2 SentencePiece algorithms: BPE and Unigram

Files

SKILL.mdMarkdownGitHub ↗

SentencePiece - Language-Independent Tokenization

Unsupervised tokenizer that works on raw text without language-specific preprocessing.

When to use SentencePiece

Use SentencePiece when:

  • Building multilingual models (no language-specific rules)
  • Working with CJK languages (Chinese, Japanese, Korean)
  • Need reproducible tokenization (deterministic vocabulary)
  • Want to train on raw text (no pre-tokenization needed)
  • Require lightweight deployment (6MB memory, 50k sentences/sec)

Performance:

  • Speed: 50,000 sentences/sec
  • Memory: ~6MB for loaded model
  • Languages: All (language-independent)

Use alternatives instead:

  • HuggingFace Tokenizers: Faster training, more flexibility
  • tiktoken: OpenAI models (GPT-3.5/4)
  • BERT WordPiece: English-centric tasks

Quick start

Installation

# Python
pip install sentencepiece

# C++ (requires CMake)
git clone https://github.com/google/sentencepiece.git
cd sentencepiece
mkdir build && cd build
cmake .. && make -j $(nproc)
sudo make install

Train model

# Command-line (BPE with 8000 vocab)
spm_train --input=data.txt --model_prefix=m --vocab_size=8000 --model_type=bpe

# Python API
import sentencepiece as spm

spm.SentencePieceTrainer.train(
    input='data.txt',
    model_prefix='m',
    vocab_size=8000,
    model_type='bpe'
)

Training time: ~1-2 minutes for 100MB corpus

Encode and decode

import sentencepiece as spm

# Load model
sp = spm.SentencePieceProcessor(model_file='m.model')

# Encode to pieces
pieces = sp.encode('This is a test', out_type=str)
print(pieces)  # ['▁This', '▁is', '▁a', '▁test']

# Encode to IDs
ids = sp.encode('This is a test', out_type=int)
print(ids)  # [284, 47, 11, 1243]

# Decode
text = sp.decode(ids)
print(text)  # "This is a test"

Language-independent design

Whitespace as symbol (▁)

text = "Hello world"
pieces = sp.encode(text, out_type=str)
print(pieces)  # ['▁Hello', '▁world']

# Decode preserves spaces
decoded = sp.decode_pieces(pieces)
print(decoded)  # "Hello world"

Key principle: Treat text as raw Unicode, whitespace = ▁ (meta symbol)

Tokenization algorithms

BPE (Byte-Pair Encoding)

spm.SentencePieceTrainer.train(
    input='data.txt',
    model_prefix='bpe_model',
    vocab_size=16000,
    model_type='bpe'
)

Used by: mBART

Unigram (default)

spm.SentencePieceTrainer.train(
    input='data.txt',
    model_prefix='unigram_model',
    vocab_size=8000,
    model_type='unigram'
)

Used by: T5, ALBERT, XLNet

Training configuration

Essential parameters

spm.SentencePieceTrainer.train(
    input='corpus.txt',
    model_prefix='m',
    vocab_size=32000,
    model_type='unigram',
    character_coverage=0.9995,  # 1.0 for CJK
    user_defined_symbols=['[SEP]', '[CLS]'],
    unk_piece='<unk>',
    num_threads=16
)

Character coverage

Language TypeCoverageRationale
English0.9995Most common chars
CJK (Chinese)1.0All characters needed
Multilingual0.9995Balance

Encoding options

Subword regularization

# Sample different tokenizations
for _ in range(3):
    pieces = sp.encode('tokenization', out_type=str, enable_sampling=True, alpha=0.1)
    print(pieces)

# Output (different each time):
# ['▁token', 'ization']
# ['▁tok', 'en', 'ization']

Use case: Data augmentation for robustness.

Common patterns

T5-style training

spm.SentencePieceTrainer.train(
    input='c4_corpus.txt',
    model_prefix='t5',
    vocab_size=32000,
    model_type='unigram',
    user_defined_symbols=[f'<extra_id_{i}>' for i in range(100)],
    unk_id=2,
    eos_id=1,
    pad_id=0
)

Integration with transformers

from transformers import T5Tokenizer

# T5 uses SentencePiece internally
tokenizer = T5Tokenizer.from_pretrained('t5-base')
inputs = tokenizer('translate English to French: Hello', return_tensors='pt')

Performance benchmarks

Training speed

CorpusBPE (16k)Unigram (8k)
100 MB1-2 min3-4 min
1 GB10-15 min30-40 min

Tokenization speed

  • SentencePiece: 50,000 sentences/sec
  • HF Tokenizers: 200,000 sentences/sec (4× faster)

Supported models

T5 family: t5-base, t5-large (32k vocab, Unigram) ALBERT: albert-base-v2 (30k vocab, Unigram) XLNet: xlnet-base-cased (32k vocab, Unigram) mBART: facebook/mbart-large-50 (250k vocab, BPE)

References

  • [Training Guide](references/training.md) - Detailed options, corpus preparation
  • [Algorithms](references/algorithms.md) - BPE vs Unigram, subword regularization

Resources

  • GitHub: https://github.com/google/sentencepiece ⭐ 10,000+
  • Paper: https://arxiv.org/abs/1808.06226 (EMNLP 2018)
  • Version: 0.2.0+

Related skills

How it compares

Pick sentencepiece for training `.model` files with the SentencePiece library; use huggingface-tokenizers when comparing WordPiece and Hugging Face tokenizer internals.

FAQ

What algorithms does sentencepiece compare?

sentencepiece contrasts BPE merge training and Unigram probabilistic segmentation in SentencePiece, including subword regularization effects and Python trainer setup with worked merge examples on sample corpora.

How do you train a SentencePiece model in Python?

sentencepiece shows patterns with `import sentencepiece as spm` and `spm.SentencePieceTrainer`, walking through vocabulary initialization, iterative merges for BPE, and trainer options before exporting a `.model` file.

Is Sentencepiece safe to install?

skills.sh reports 1 of 3 security scanners passed. Review the Security Audits panel on this page before installing in production.

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.