H-Net (dynamic chunking)
End-to-end hierarchical byte-level model that learns its own chunking; a two-stage H-Net matches a token-based Transformer twice its size.
Replaces fixed BPE with a learned dynamic-chunking module inside the network, so segmentation is trained jointly. Strong on languages with weak tokenisation heuristics and DNA (nearly 4x data efficiency).
- Date
- Thursday, 10 July 2025
- Lab
- Hwang, Wang, Gu (academic)
- Kind
- paper
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Two-stage H-Net vs Transformer | matches 2x larger byte-level vs BPE, matched compute and data | authors |
Authors Sukjun Hwang, Brandon Wang, Albert Gu; arXiv v1 2025-07-10. Institutions and weight availability not verified from the abstract page.
Sources
This record was checked against its sources on 6 October 2026. How we check