From the source
Lead story
Ask your AI
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
Method for compressing on-device speech encoders via distillation
Apple researchers present a method for compressing streaming neural audio encoders used in on-device dictation through latent-space distillation.
The technique trains a student encoder to regress the teacher's pre-quantizer latent representation, achieving 2.8× compression while maintaining within 1.9% relative word error rate of the teacher on most model pairs without fine-tuning.
From the source
At 2.8× compression the distilled student stays within 1.9% relative WER of its teacher on five of six teacher–student pairs without any fine-tuning, and improves on an independently trained tokenizer of identical capacity by 3.9% relative.
machinelearning.apple.com