From the source
Replit describes its experience deploying transformer-based language models ranging from ~100 million to over 100 billion parameters, including its GhostWriter code autocomplete product.
The post covers techniques for reducing nonsense output, maintaining quality via benchmarks and A/B testing, and improving inference speed through FasterTransformer, knowledge distillation, and quantization.





