快讯
Allen AI在Nature发表论文:以不足1%预算将预训练模型改造为字节级
Allen AI在Nature发表论文,提出名为byteification的两阶段蒸馏方法,无需从头训练,即可将Olmo、Llama 3、Qwen等现有预训练模型改造为字节级模型,成本不足原预训练预算的1%。所得Bolmo、Blama、Bwen模型推理速度达到实用水平,并在字符级推理任务上表现出色。
事件来源
查看原文
RT Turing Post (Ksenia Se) An interesting paper was just published by @allen_ai in Nature on ditching subword tokenization without burning millions of dollars in compute: Retrofitting language models to operate over bytes. The classic headache with byte-level models has always been speed and the sheer cost of training them from scratch. Instead of starting from zero, the authors introduce "byteification" – a neat two-stage distillation trick to retrofit existing pretrained models (like Olmo, Llama 3, and Qwen) into byte-level systems using under 1% of the original pretraining budget. The key architectural insight addresses a subtle flaw in earlier latent tokenizer setups. Subword tokenizers look ahead at future bytes when drawing boundaries, whereas standard byte-level patch predictors were strictly causal. By adding just 1 byte of lookahead during prefill and predicting fused boundary tokens during decode, they bridge that expressivity gap. The resulting models (Bolmo, Blama, Bwen) run at practical inference speeds, crush character-level reasoning tasks where standard tokenizers trip up, and you can even transfer instruction tuning across using simple task arithmetic. Definitely worth a close read if you work on token-free architectures.