Where an embedding stack bottlenecks: Perplexity's case
Because embedding models are small, the bottleneck is not the model but the work around it: kernel launching, tokenisation, an idle CPU. The transferable lessons from the stack Perplexity published.
Research