Perplexity Photon Speeds Up Search, but Fast Search Trades Relevance for Cost
marktechpost.com

Perplexity Photon Speeds Up Search, but Fast Search Trades Relevance for Cost

Tech News
4 min read

Published by AINave Editorial

TL;DRPerplexity says its new Rust-based Photon engine cut production retrieval-and-ranking p99 latency from about 800 ms to 65 ms. Its Fast Search API mode costs $1 per 1,000 requests, but Perplexity’s own results show a measurable drop in long-tail relevance.

Perplexity’s Photon retrieval engine is now handling its production search traffic, and the company reports that it cut p99 retrieval-and-ranking latency from about 800 milliseconds to 65 milliseconds. That figure covers Photon’s internal stages, not the full API call. The separate Fast Search API mode reports 160 ms p50 and 230 ms p95 single-call latency, so the measurements describe different parts of the system. Perplexity’s reported latency figures

Photon’s speedup comes from changing the retrieval path

Photon is an in-house engine written in Rust that replaces a forked open-source system. Perplexity describes a broker that sends each request to a group of shards, where retrieval and two ranking stages run before the broker combines candidates and fetches document fields. The engine now handles all production traffic.

The implementation is designed to avoid paying the full cost of reading and scoring every candidate. Photon uses different posting-list formats for sparse and dense data, and a WAND-like traversal checks inexpensive bounds before reading exact term frequencies. Compact document records store ranking information, while batched asynchronous disk reads fetch records at known offsets. In practical terms, the design tries to spend disk and ranking work on candidates that might actually make the cut. Perplexity’s account of Photon’s retrieval and ranking design

The change also targets operational problems beyond slow queries. The old system’s p99 reportedly reached about 1.2 seconds during index merges, and deploying and syncing an extra cluster could take more than a week. Perplexity says Photon uses about 20% fewer serving machines than its old content nodes and can build a full web index in a single-digit number of hours. The reported production and indexing changes

Fast Search is a separate API choice, with a quality cost

Developers can select Fast Search by setting search_type: "fast" on POST /search; the reported price is $1 per 1,000 requests. Perplexity also reports results across six benchmarks and 3,554 tasks: Fast scored 64.3% at an estimated $59.73 in combined model-plus-search cost, compared with 64.0% and $187.60 for the default preset. Those task-cost figures are not the API’s per-request price, and they reflect Perplexity’s reported evaluation rather than an independent comparison. Fast Search pricing and benchmark results

Preset Benchmark score Estimated model-plus-search cost Long-tail DCG Answer availability
Fast 64.3% $59.73 2.21 0.567
Default 64.0% $187.60 2.45 0.596

The trade-off is visible in Perplexity’s internal long-tail measures: Fast scored 2.21 on DCG, against 2.45 for default, and answer availability was 0.567 versus 0.596. The company recommends Fast for routine agent loops and default search for hard, ambiguous queries. That distinction matters because a similar score across a benchmark set does not mean the presets retrieve equally well on every kind of request. Perplexity’s long-tail results and stated guidance

Photon itself is not open source, so the reported route for developers is a hosted API, not self-hosting. Fast Search makes the faster, lower-cost option accessible, but its usefulness depends on whether a workflow can tolerate the relevance reduction Perplexity reports.

FAQs

Photon is Perplexity’s in-house Rust-based retrieval and ranking engine. The company says it replaced its forked open-source engine and now handles production traffic; Photon itself is not open source. Perplexity’s description of Photon

Sources

Latest Tech News