arXiv cs.CLPaper
Lost in the bf16 Cast: Exporting Ternary Language Models Can Revert Most Low-Learning-Rate Code Changes
This is a hard finding: deployed ternary models (BitNet, Falcon-E, BitCPM) lose most of their ability on standard benchmarks because of an accidental rounding bug in the export step. The labs documented the pipelines correctly, but the actual released weights disagree with the code. If you're using or planning to use ternary models, verify the export step or you're getting crippled performance.