Perplexity AI announced two new versions of its pplx‑embed‑v2‑late model. The smaller variant has 0.6 billion parameters and is designed to run on edge hardware like phones or IoT gadgets. The larger variant has 9 billion parameters and is aimed at building high‑quality vector indexes for search or recommendation systems. Both models are released under the MIT license, so anyone can download and host them on their own servers.
How the models perform
In Perplexity’s own tests, the 9 billion‑parameter model achieved a 92.4% score on the MADQA benchmark, a standard test for question‑answering accuracy. The same model scored 61.2% on the ViDoRe v3 Markdown benchmark, which measures handling of markdown‑formatted text. The company says the 0.6 billion model is optimized for speed and low memory use, while the 9 billion model focuses on accuracy.
Why it matters
Developers can now run a lightweight embedding model directly on phones or small devices without needing a cloud service. Larger teams that need precise search or recommendation features can use the higher‑capacity model while keeping full control of their data, because both versions are open‑source.