Architect Launches Liquid Inference, a Real-Time Auction for LLM Inference






Architect Financial Technologies has launched Liquid Inference, an LLM router that runs a live auction for every request. Liquid Inference is an LLM inference marketplace from Architect where providers bid to serve each prompt. The buyer pays the lowest offer that meets its rules. For developers, it is quite simple message: swap a base URL, keep your code, and let providers compete on price.

What is Liquid Inference?

Liquid Inference is an exchange-style router for LLM inference. According to Architect, providers post offers to serve specific models. Each request is auctioned across every provider quoting the named model. The lowest-priced offer that meets the buyer’s rules wins.

The product comes from a trading firm, not an AI lab. Architect runs the AX perpetual futures exchange. In May 2026 it acquired a US Designated Contract Market to list GPU compute futures, pending regulatory review. The team used its experience building financial exchanges to create two-sided price discovery for inference.

How does the inference auction work?

The flow has 4 steps:

  1. Request: A client sends a standard OpenAI or Anthropic API call.
  2. Rules: The buyer’s constraints filter eligible offers.
  3. Auction: Providers quoting that model compete. The lowest qualifying offer wins.
  4. Receipt: The max price is locked before generation. Billing covers metered usage only.

Buyers can set per-job cost caps, time to first token limits and minimum throughput. They can also require approved regions, zero data retention and provider or model allow lists. An Auto mode can pick the model for a given unit of work. Harrison’s LinkedIn post adds routing rule presets and full multi-modal support.

Account holders can view live order books, per-provider and per-model quotes, and cleared transactions. That level of market data is unusual for an LLM API.

What do buyers get?

  • Drop-in compatibility with agentic coding tools. The post lists Claude Code, Codex, OpenCode, Cursor, Pi and Cline.
  • Free email signup. The first 500 users get $20 of free inference, per Harrison.
  • A referral program: 20% of referred fees as free inference, plus 10% on second-level referrals.

What do inference providers get?

Providers onboard through the Liquid Inference app. Harrison says new providers are verified “in minutes, not weeks.” All prompts use the OpenAI API standard. A REST and WebSocket API registers models and quotes.

Providers can update quotes based on their own costs. That lets them sell spare GPU capacity only when they want. Payouts run through Stripe, with itemized records of every job.

How does Liquid Inference compare with OpenRouter and Hugging Face?

FeatureLiquid InferenceOpenRouterHugging Face Inference Providers
Routing modelPer-request auction across quoting providersPrice-weighted load balancing, inverse square of priceFastest provider by default; :cheapest suffix optional
Model / provider count“Hundreds” of models; providers not disclosed500+ models, 80+ providers18 listed partners
API compatibilityOpenAI and Anthropic-compatibleOpenAI-compatibleOpenAI-compatible, chat only
Price capMax price locked before first tokenmax_price parameterNot disclosed
Data controlsZDR, regions, allow listszdr, data_collection, only/ignoreProvider preference order
Platform feeNot disclosed5.5% card credit fee, $0.80 minimumNo markup
Free credits$20 for first 500 usersNot disclosed$0.10/month free, $2.00 PRO
Public market dataLive order books and cleared tradesNot disclosedNot disclosed

‘}).join(”);setTimeout(function(){el.querySelectorAll(‘i’).forEach(function(i){i.style.width=(state.bars?i.dataset.w:0)+’%’})},30)}
var step=-1,book=document.getElementById(‘li-book’),msg=document.getElementById(‘li-msg’);
var msgs=[‘A coding agent sends a standard chat request. Every provider quoting that model is invited.’,’Buyer rules filter the book. Provider B misses the TTFT rule; Provider E is not ZDR.’,’Remaining offers compete. Provider C wins at $0.36, the lowest qualifying quote.’,’Max price is locked before generation. Receipt shows winner, cap, final charge and competition depth.’];
function render(){R.querySelectorAll(‘.st’).forEach(function(s,i){s.classList.toggle(‘on’,i===step)});var elig=offers.filter(function(o){return o.ok});var win=elig.slice().sort(function(a,b){return a.p-b.p})[0];rows(book,offers,{bars:step>=0,out:step>=1?function(o){return !o.ok}:null,win:step>=2?win:null});msg.innerHTML=step<0?’Press Next step to walk one request through the auction.’:msgs[step];ping()}
document.getElementById(‘li-next’).addEventListener(‘click’,function(){step=step>=3?0:step+1;render()});
document.getElementById(‘li-reset’).addEventListener(‘click’,function(){step=-1;render()});
var act={},book2=document.getElementById(‘li-book2’),msg2=document.getElementById(‘li-msg2’);
R.querySelectorAll(‘.rl’).forEach(function(r){r.addEventListener(‘click’,function(){r.classList.toggle(‘on’);act[r.dataset.r]=r.classList.contains(‘on’);draw2()})});
function draw2(){var on=Object.keys(act).filter(function(k){return act[k]});var out=function(o){return on.some(function(k){return !o.f[k]})};var el=offers.filter(function(o){return !out(o)});var win=el.slice().sort(function(a,b){return a.p-b.p})[0];rows(book2,offers,{bars:true,out:out,win:win});msg2.innerHTML=win?el.length+’ of 5 offers eligible. Winner: ‘+win.n+’ at $’+win.p.toFixed(2)+’.’:’No offer meets every rule. The request is not filled.’;ping()}
R.querySelectorAll(‘.cd’).forEach(function(c){c.addEventListener(‘click’,function(){c.classList.toggle(‘open’);setTimeout(ping,380)})});
function ping(){try{parent.postMessage({mtpLiH:R.offsetHeight+40},’*’)}catch(e){}}
render();draw2();window.addEventListener(‘load’,ping);window.addEventListener(‘resize’,ping);setTimeout(ping,400);
})();



Source link

  • Related Posts

    NVIDIA PivotOPD Teaches Multi-Turn AI Agents to Recover From Pivotal Mistakes

    NVIDIA researchers, with Princeton University and the University of Maryland, have introduced PivotOPD, an on-policy distillation method for multi-turn LLM agents. PivotOPD on-policy distillation trains an agent to avoid its…

    Perplexity AI Releases pplx-embed-v2-late: A 0.6B Edge Model and a 9B Model Scoring 92.4% on MADQA

    Perplexity has released pplx-embed-v2-late, a pair of ColBERT-style multimodal embedding models. They come in 2 sizes: 0.6B for fast, cheap queries and 9B for maximum quality. Both models retrieve text,…

    Leave a Reply

    Your email address will not be published. Required fields are marked *