arXiv cs.LGPaper
Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers
The practical problem here is real: VLM-as-policy is slow and unreliable at scale. SAGE tackles this by treating the VLM as a fallible guide rather than ground truth, weighting its advice by environment feedback. If you're building vision-based agents, this distillation pattern—use expensive models for training signal only—should become standard in your pipeline.