Reddit user trains LLM on 1800's texts with 160GB dataset

A Reddit user pre-trained language models exclusively on 1800's London data, culminating in a 40B-token (160GB) dataset of 1800-1875 English text. They plan to train a 2B parameter model on it.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Robotic systems struggle with physical navigation in recent demonstrations
- ChatGPT Visualize skill turns notes into interactive interfaces
- Twitch streamers sue Twitch, Amazon over AI training data use
- Fastino launches GLiNER2.5 with boundary-prediction architecture
- NVIDIA Sol Engine cuts MiniMax H3 latency to 14.93s