Fine-tuning a 450M VLM on 50K browser screenshots boosts score from 1/100 to 44/100

A 450M vision-language model improved from 1/100 to 44/100 on a browser task after fine-tuning on 50K screenshots. The result was shared on Reddit's r/LocalLLaMA.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- AutoResearchClaw generates research papers from chat prompts
- 63% of Religious Books on Amazon Likely AI-Written, Study Finds
- Why real-time AI at scale is so hard
- OpenAI users report ongoing usage limit issues
- Claude user shares simple skill to track long sessions