From the source
LG AI Research introduced BeamCLIP, a method for transferring representations from large pre-trained multimodal models (e.g., CLIP) to smaller models (e.g., ResNet-18), using Cross-modal Similarity Matching (CSM) and Context-based Prompt Augmentation (CPA), achieving 66.2% linear probe accuracy on ImageNet-1K.





