From the source
Lead story
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
Strengthening language discrimination reduces the multilingual performance gap in speech models.
Apple researchers found that enhancing language discrimination during pretraining of multilingual speech models (using an auxiliary language classifier and per-language k-means targets) reduces the performance gap compared to monolingual models.
In a controlled English/French HuBERT setting, phone discrimination error decreased from 11.6% (bilingual baseline) to 10.4%, while lexical and prosodic performance improved, in some cases matching or exceeding monolingual baselines.
The strongest gains occurred when language discrimination was introduced in the first training iteration.
From the source
We show that strengthening the model's ability to discriminate languages during pretraining reduces and, on some measures, closes this multilingual gap on continuous phonetic and higher-level linguistic measures, while preserving substantial cross-language sharing.
machinelearning.apple.com