They apparently pre-train with all data up to 1900 and then fine-tune with 1900-...

mmooss · 2025-12-19T00:25:50 1766103950

They pre-train with all data up to 1900 and then fine-tune with 1900-1913 data.

Where does it say that? I tried to find more detail. Thanks.

tootyskooty · 2025-12-19T00:54:33 1766105673

See pretraining section of the prerelease_notes.md:

https://github.com/DGoettlich/history-llms/blob/main/ranke-4...

pests · 2025-12-19T01:41:24 1766108484

I was curious, they train a 1900 base model, then fine tune to the exact year:

"To keep training expenses down, we train one checkpoint on data up to 1900, then continuously pretrain further checkpoints on 20B tokens of data 1900-${cutoff}$. "