OpenAI / GPT-5: local raw-text counts using o200k_base, the GPT-5 encoding mapped by OpenAI’s tiktoken. Compare text with the official OpenAI Tokenizer. Chat messages, tool definitions and media can add tokens.
Claude: This local reference uses Anthropic’s legacy tokenizer, which is not accurate for Claude 3 and later. Claude 4.7 and later use a newer tokenizer, so counts from older versions should not be reused. The visible pieces and IDs belong to that legacy tokenizer, not your current Claude model. Use the official token-counting API to check your model.
Gemini: Google’s official Gemma 4 vocabulary. Google’s local SDK maps Gemini 3.5 Flash, 3.1 Flash-Lite and 3.1 Pro Preview to Gemma 4. Local counts cover raw text; they are not verified counts for every Gemini version or complete API requests. Use the official countTokens API for model-specific verification.
DeepSeek: DeepSeek’s official V4 Pro vocabulary, pinned to a specific release. Counts cover raw text, excluding chat templates and media. Other versions may use different tokenizers. Check the API usage for billing.
Grok: rough budgeting estimate: UTF-8 bytes divided by 4, rounded up. This heuristic is used in xAI’s open-source Grok Build for budgeting; it is not a model tokenizer and may differ substantially for language, code, punctuation or model versions. Token pieces and IDs are unavailable. For a chosen model, use the official xAI Tokenize API.
No text is sent to any provider. All vocabulary files are served locally. First use loads the local tokenizer; future model versions require an explicit update. Reference review: 29 September 2026.