Rasmus Aagaard @rasgaard.com · Jan 16

LiteASR (arxiv.org/pdf/2502.20583) finds that latency caused by the Whisper encoder and decoder varies significantly at different settings. Compressing the decoder is common (distill/turbo variants) but the encoder is largely ignored. The authors look into encoder compression - super cool work!

6 likes 1 replies

?

Replies

Vladimir Salnikov · Jan 17

Love your recent posts on Model optimization and Edge AI/SLM/TTS related things 🙏 Although not my domain, it is interesting to learn new stuff and in general I think it is helpful for AI research diversity on this platform