Readings & resources: Pre- and post-processing
- Dodge, J., Sap, M., Marasović, A., Agnew, W., Ilharco, G., Groeneveld, D., Mitchell, M., Gardner, M. “Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus.” EMNLP 2021.
- Soldaini, L. et al. “Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research.” ACL 2024.
- “Interactive Byte-Pair Encoding Visualizer.” Cornell CS4782.