Models pre-trained on various obfuscated versions of a sample of 10B tokens to analyse whether such obfuscated models are still performant.