I was curious, they train a 1900 base model, then fine tune to the exact year: "...

		pests 1 day ago \| parent \| context \| favorite \| on: History LLMs: Models trained exclusively on pre-19... I was curious, they train a 1900 base model, then fine tune to the exact year: "To keep training expenses down, we train one checkpoint on data up to 1900, then continuously pretrain further checkpoints on 20B tokens of data 1900-${cutoff}$. "