Perplexity has released Photon, an in-house retrieval and ranking engine written in Rust. It replaces an open-source engine Perplexity had forked for its AI-native search stack. Photon now handles retrieval and ranking for all production traffic. It also powers a new Fast Search mode in the Perplexity Search API. Perplexity reports single-call latency of 160 ms at p50 and 230 ms at p95.
Is it deployable? Yes, as a hosted API. Set search_type: "fast" on POST /search and pay $1 per 1,000 requests. Photon itself is not open source, so the engine cannot be self-hosted.
Why Perplexity Replaced its Old Engine
The old engine hit 3 limits as the index grew:
- Tail latency: Production p99 sat near 800 ms. The dataset exceeded RAM, so
mlockwas not an option. Cold reads triggered major page faults that stalled queries. - Merge spikes: During disk index fusion, p99 climbed to about 1.2 s for 10 to 15 minutes.
- Slow recovery: Deploying and syncing an extra cluster could take more than a week. Recovery also raised the share of partial responses.
Perplexity team concluded that building from scratch was simpler and cheaper than maintaining its fork.
How Photon Works
A load balancer routes each request to a Photon broker. The broker fans out to a shard group and watches for timeouts. Each shard runs retrieval, initial ranking, and second-stage ranking. The broker then merges candidates and fetches key document fields.
- Adaptive posting lists: Short lists sit inline within a single page. Longer lists split into blocks of fixed document ID ranges. Sparse blocks store sorted offset arrays and use galloping search. Dense blocks use bitmaps, so membership becomes a single bit lookup.
- Budgeted traversal: A WAND-like algorithm splits lists into driving lists and probe lists. Cheap presence checks bound each candidate’s maximum score first. Exact term frequencies are read only when a candidate can clear the threshold.
- Docblob records: Each document gets a compact record of frequencies, field masks, and positions. Terms use Elias-Fano encoding, so ranking decodes only the matched terms. Ranking a candidate needs just 1 lookup per document.
- Batched async reads: Record offsets are known upfront, so disk reads go out in batches through
io_uring. The cache checks the whole batch first. Readers take no locks, and eviction uses CLOCK instead of a shared LRU list. - Separate build and serve: Indexers build versioned shard indexes from YTsaurus tables on dedicated nodes. A controller rotates serving groups one at a time and warms caches with replayed search-log queries.
A full web index now builds in a single-digit number of hours.
Interactive Explainer: Inside Photon
/* PANE 1: posting lists */
function lenFromSlider(v){return Math.round(Math.pow(10,0.7+v/100*6));}
function fmt(n){return n.toLocaleString(‘en-US’);}
var seed=7;function rnd(){seed=(seed*9301+49297)%233280;return seed/233280;}
function drawPL(){
var n=lenFromSlider(+$(‘px-len’).value),den=+$(‘px-den’).value/100;
$(‘px-lenv’).textContent=fmt(n);$(‘px-denv’).textContent=Math.round(den*100)+’%’;
var box=$(‘px-plist’);box.innerHTML=”;seed=7;
if(n<=64){
var ids=[];var cur=0;for(var i=0;i
‘;
resize();return;
}
var nb=Math.min(12,Math.max(3,Math.round(Math.log10(n)*2)));
var g=document.createElement(‘div’);g.className=”plist”;
for(var b=0;b
var el=document.createElement(‘div’);el.className=”blk “+(dense?’dense’:’sparse’);el.style.animationDelay=(b*45)+’ms’;
var h=”
Block “+(b+1)+’: ‘+(dense?’bitmap’:’offset array’)+’
‘;
if(dense){h+=’
‘;for(var k=0;k<32;k++)h+=’‘;h+=’
‘;}
else{var c=Math.max(2,Math.round(bd*12)),o=0,arr=[];for(var k3=0;k3