Submitted by Jingfeng Yao 105 Towards Scalable Pre-training of Visual Tokenizers for Generation MiniMax 441 4