From the source
The paper proposes a domain generalization method that uses natural language supervision to ground visual representations, introducing two modules: Visual and Textual Joint Embedder and Textual Explanation Generator, and achieves state-of-the-art results on CUB-DG and DomainBed benchmarks.





