RiLM: Parameter-Efficient Language Modeling via Geodesic Decoding
This is novel geometry applied to a real constraint: tiny models waste capacity on the output matrix. HypRiLM beats baselines on WikiText-2, which is promising. If you're deploying edge models or exploring parameter-efficient architectures, the manifold-based decoding is worth testing. The gains are solid but models this small are niche.