arXiv cs.LGPaper
Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090
The real story is cost, not capability: a usable small model trained for under $7,000 versus the $700K to $1.5M price tags cited for comparable open efforts. If the recipe holds up under scrutiny, it lowers the bar for academic labs and indie teams to pretrain rather than just fine-tune, which is a meaningful shift in who gets to build foundation models.