
Perplexity Photon Speeds Up Search, but Fast Search Trades Relevance for Cost
Published by AINave Editorial
Perplexity’s Photon retrieval engine is now handling its production search traffic, and the company reports that it cut p99 retrieval-and-ranking latency from about 800 milliseconds to 65 milliseconds. That figure covers Photon’s internal stages, not the full API call. The separate Fast Search API mode reports 160 ms p50 and 230 ms p95 single-call latency, so the measurements describe different parts of the system. Perplexity’s reported latency figures
Photon’s speedup comes from changing the retrieval path
Photon is an in-house engine written in Rust that replaces a forked open-source system. Perplexity describes a broker that sends each request to a group of shards, where retrieval and two ranking stages run before the broker combines candidates and fetches document fields. The engine now handles all production traffic.
The implementation is designed to avoid paying the full cost of reading and scoring every candidate. Photon uses different posting-list formats for sparse and dense data, and a WAND-like traversal checks inexpensive bounds before reading exact term frequencies. Compact document records store ranking information, while batched asynchronous disk reads fetch records at known offsets. In practical terms, the design tries to spend disk and ranking work on candidates that might actually make the cut. Perplexity’s account of Photon’s retrieval and ranking design
The change also targets operational problems beyond slow queries. The old system’s p99 reportedly reached about 1.2 seconds during index merges, and deploying and syncing an extra cluster could take more than a week. Perplexity says Photon uses about 20% fewer serving machines than its old content nodes and can build a full web index in a single-digit number of hours. The reported production and indexing changes
Fast Search is a separate API choice, with a quality cost
Developers can select Fast Search by setting search_type: "fast" on POST /search; the reported price is $1 per 1,000 requests. Perplexity also reports results across six benchmarks and 3,554 tasks: Fast scored 64.3% at an estimated $59.73 in combined model-plus-search cost, compared with 64.0% and $187.60 for the default preset. Those task-cost figures are not the API’s per-request price, and they reflect Perplexity’s reported evaluation rather than an independent comparison. Fast Search pricing and benchmark results
| Preset | Benchmark score | Estimated model-plus-search cost | Long-tail DCG | Answer availability |
|---|---|---|---|---|
| Fast | 64.3% | $59.73 | 2.21 | 0.567 |
| Default | 64.0% | $187.60 | 2.45 | 0.596 |
The trade-off is visible in Perplexity’s internal long-tail measures: Fast scored 2.21 on DCG, against 2.45 for default, and answer availability was 0.567 versus 0.596. The company recommends Fast for routine agent loops and default search for hard, ambiguous queries. That distinction matters because a similar score across a benchmark set does not mean the presets retrieve equally well on every kind of request. Perplexity’s long-tail results and stated guidance
Photon itself is not open source, so the reported route for developers is a hosted API, not self-hosting. Fast Search makes the faster, lower-cost option accessible, but its usefulness depends on whether a workflow can tolerate the relevance reduction Perplexity reports.
FAQs
search_type to fast on POST /search. Perplexity also describes using extra_body={"search_type": "fast"} with Python SDK versions 0.43.4 and 0.43.5. The Fast Search API instructions


















