Can language models work without word tokens?
Yes, at research scale. A peer-reviewed Nature study converted existing 1B–8B language models to operate on UTF-8 bytes using 49.1 billion continued-training tokens. The resulting models improved character-level tasks and approached their source models elsewhere, but the evidence is benchmark-based and does not yet prove lower deployment cost or better real-world products.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1The researchers byteified four model variants derived from OLMo, Qwen and Llama-family base models at roughly 1B–8B parameters, rather than training every byte-level model from scratch.
- 2Training used 9.8 billion tokens in a frozen-global-model distillation stage and 39.3 billion tokens in end-to-end training: 49.1 billion continued-training tokens in total, described as less than 1% of a typical pretraining budget.
- 3Bolmo 7B strongly improved character benchmarks but remained slightly behind its OLMo 3 source on several general benchmark groups. The paper did not test production reliability, user outcomes or actual energy savings.
Research topic
Whether existing subword language models can be retrofitted to operate directly on UTF-8 bytes without losing most of their benchmark capability

The direct answer: yes, but the model still builds internal patches
The Nature paper shows that a competitive language model does not need an external vocabulary of words or word fragments at its input and output. The researchers converted established subword models into systems that read and generate raw UTF-8 bytes. Their strongest open model, Bolmo 7B, substantially improved tasks that require manipulating individual characters and approached the source OLMo 3 model on many broader benchmarks. The team also applied the method to Qwen3 8B and Llama 3 8B checkpoints, suggesting that the conversion is not tied to one model family.
That does not mean the architecture performs every calculation one byte at a time. A lightweight local encoder groups contextualized bytes into variable-length latent patches before a large transformer processes them; a local decoder then expands the patch representations back to bytes. The important distinction is that those groupings are learned inside the model rather than imposed by a fixed external tokenizer. This can preserve character information while keeping the global sequence short enough to be computationally practical.[1][2]
Why ordinary word-fragment tokenization creates blind spots
Most current large language models first split text into a finite vocabulary of whole words, fragments and symbols. That compression is efficient, but it can obscure the characters inside a token. The paper connects this to familiar weaknesses in spelling, string manipulation, code and biological sequences, where one character can change the meaning. Vocabulary construction can also allocate capacity unevenly across writing systems: common English fragments may receive compact representations while other languages are broken into longer or less useful sequences.
Byte input avoids an external vocabulary because any UTF-8 text can be expressed with the same small set of byte values. The trade-off is length: a sentence represented as bytes produces many more input units than a tokenized sentence. Earlier byte-level systems often spent too much computation on those long sequences or lagged subword models trained with more mature recipes. The new study addresses that practical gap by preserving the deep transformer from an existing model and learning new byte-level input, patching and output components around it.[1]
The conversion used two training stages and 49.1 billion tokens
In stage one, the researchers froze the source model’s global transformer and trained the new local encoder, local decoder, boundary predictor and byte prediction head on 9.8 billion tokens, approximately 43 billion bytes. The objective was a form of self-distillation: learn patch boundaries and representations that reproduce the behaviour of the original subword system before asking the converted model to exploit finer byte information. The boundary predictor exceeded 99% accuracy against the source tokenizer’s boundaries during this stage.
In stage two, the whole model was trained end to end on another 39.3 billion tokens, approximately 173 billion bytes. Total continued training was therefore 49.1 billion tokens. The paper describes that as less than 1% of a typical pretraining budget, which is meaningful relative to frontier-scale training, but it is not a small experiment in ordinary terms. The available training mix contained about 172 billion tokens from Dolma 3 plus 75 million character-focused tokens; the models trained for less than one epoch on that mixture.[1][2]
The benchmark result is a trade-off, not a universal win
The clearest gain was character understanding. The released benchmark summary reports Bolmo 7B scoring 78.6 on CUTE compared with 56.9 for its OLMo 3 source, and 71.6 versus 55.1 on the multilingual EXECUTE character benchmark. Code performance in the summary was 41.0 for Bolmo and 40.1 for OLMo. Against the earlier BLT 7B byte model, Bolmo’s absolute gain on the paper’s STEM grouping was 16.5 percentage points. These comparisons support the claim that byteification overcame a substantial performance gap for byte-level models of similar size.
The converted model did not dominate its source everywhere. Bolmo 7B scored 48.9 versus 55.3 on the reported mathematics grouping, 65.5 versus 66.3 on multiple-choice STEM, 75.8 versus 77.7 on multiple-choice non-STEM and 70.9 versus 72.4 on generative question answering. The paper also reports different behaviour between pass-at-one and pass-at-sixteen code measures. A product team would therefore need task-specific evaluation rather than assuming that removing a tokenizer improves every capability.[1][2]
Multilingual promise is plausible but not yet demonstrated broadly
A shared byte vocabulary removes one structural source of language imbalance: the model no longer needs a hand-built inventory that represents some scripts more compactly than others. The study’s multilingual character benchmark improved even though the added character-focused training examples were purely in English. That is an encouraging sign that learning to inspect characters can transfer across scripts, and the model can represent any valid UTF-8 sequence without an out-of-vocabulary token.
It is not yet evidence of equal language quality. The study did not train on multilingual character exercises, evaluate conversational usefulness across communities or measure whether byte sequences consume comparable compute across languages. Better spelling and string manipulation do not guarantee better factual knowledge, cultural competence or safety. A strong follow-up would report per-language generation quality, latency, sequence compression and error patterns, especially for scripts poorly served by current tokenizers.[1]
What this could change for developers and users
For developers, retrofitting creates a cheaper research path than rebuilding a competitive byte model from random initialization. The paper produced Bolmo variants from fully open OLMo checkpoints and showed conversions of Qwen3 and Llama 3 base models. The public repository includes code, training scripts, model weights, data-processing instructions and evaluation guidance. That openness makes the central result more testable and lowers the barrier for teams studying code, genomics, multilingual text or other domains where individual symbols matter.
For users, any benefit remains indirect. A byte-level model could eventually make spelling, unusual names, mixed scripts, source code and biological strings more reliable. Flexible patch sizes could also support faster inference at a chosen quality trade-off. The study did not run a production service, compare user satisfaction, test adversarial byte sequences or measure total energy consumption. Its claims about lower deployment cost are potential consequences of the architecture, not observed outcomes from a live system.[1][2]
What would change the assessment
The next evidence should compare matched subword and byteified models under the same end-to-end compute, memory and latency budgets on real applications. Evaluators should include multilingual generation, secure code handling, long contexts, malformed UTF-8, prompt attacks and domains such as genomics where a single-character error can be consequential. Prospective deployments would show whether improved character benchmarks translate into fewer user-visible mistakes without degrading reasoning or increasing serving cost.
The work used computing resources from the US Department of Energy’s Oak Ridge facility, the National Artificial Intelligence Research Resource pilot, Microsoft Azure and Google’s TPU Research Cloud. Named authors disclosed support from the UK EPSRC, the US National Science Foundation and the European Research Council; the paper reports no competing interests. Those resources and disclosures matter when judging reproducibility. The most defensible conclusion today is that byteification removes an important technical barrier and opens a credible alternative model-design path—not that fixed tokenizers are already obsolete.[1][2]
What this means for people
- More reliable character handling could help users working with names, mixed scripts, code and scientific sequences.
- A universal byte interface may reduce one source of language imbalance, but it does not guarantee equal model quality across languages.
- Open weights and training code let more researchers test the claim, although reproducing 49.1 billion-token continued training still requires substantial compute.
Global context
Tokenizer design affects every language and technical domain that a model handles. The research spans US, UK and German institutions and releases open artefacts, but the training resources remain concentrated in large research organisations. Broad value will depend on independent tests across languages, hardware and applications rather than benchmark leadership alone.
What the evidence does not yet show
- The study evaluated base models at roughly 1B–8B parameters, not frontier-scale production assistants.
- Results are dominated by benchmark comparisons; no live product, user outcome, reliability study or downstream scientific workflow was evaluated.
- Bolmo 7B improved character tasks but remained slightly behind its source model on several mathematics, multiple-choice and question-answering groups.
- The added character-focused training data were English-only, so broader multilingual benefits remain uncertain.
- Potential energy and deployment savings were discussed but not established through a matched real-world cost or lifecycle assessment.
What to watch next
- Independent reproduction using the released code, checkpoints and training mixture.
- Matched latency, memory, energy and quality comparisons in deployed applications.
- Per-language evaluation across scripts that current tokenizers represent inefficiently.
- Safety testing for malformed encodings, invisible characters, code and biological sequences.
Living evidence record
Impact record IAI-1TGJLMF
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
7 October 2026
Source trail
2 direct sources across 2 source types.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Related-source reporting disclosure
This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.
Evidence trail
Sources used for this report
Links checked 7 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Technology
Is Gemini 4 Argon available now, and what do Google's benchmark claims show?
Google has announced a frontier model for coding, professional work and cyber defence, but access is initially limited to trusted defenders. Its 19-row comparison is broad and often strong; it is still a vendor evaluation, not an independent test or a public release.
7 min · 3 sources
Technology
Does GPT-6.1 Sol make AI work cheaper? Count successful tasks, not just tokens
OpenAI's 29 September release reports stronger performance at lower task costs in selected evaluations. The useful comparison for buyers is the cost of checked, completed work in their own setting.
4 min · 2 sources
Technology
Can a lightweight AI read Bangladesh’s road signs?
In a 720-image test, a 4.6-million-parameter classifier reached 99.34% accuracy on a new 48-class Bangladeshi dataset. It classified prepared sign images; it did not detect signs in moving traffic or prove safe vehicle use.
8 min · 2 sources
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.