Perplexity API

Perplexity API pricing charges you twice for the same call.

Pricing Model:

Pricing Model:

Usage, on two axes at once

Usage, on two axes at once

Usage, on two axes at once

Packaging Model:

Packaging Model:

Good / Better / Best (GBB)

Good / Better / Best (GBB)

Good / Better / Best (GBB)

Credit Model:

Credit Model:

Prepaid, Stripe-billed, auto reload on a balance threshold

Prepaid, Stripe-billed, auto reload on a balance threshold

Prepaid, Stripe-billed, auto reload on a balance threshold

Updated on:

Perplexity Sonar API pricing: two meters running on every call

Perplexity API pricing charges you twice for the same call. You pay per million tokens like any model API, and you pay a separate request fee for the web search that grounds the answer. That fee isn't flat: it climbs with the search_context_size you ask for. On a short Sonar query the search fee is the bill and the tokens round to nothing. Perplexity documents all of it.

Key takeaways

  • Sonar bills $1 per million tokens in and out, then adds $5, $8 or $12 per 1,000 requests depending on search context size.

  • Perplexity's own Sonar worked example totals $0.005420 for one query, and $0.005 of that is the request fee rather than tokens.

  • Sonar Deep Research carries no request fee. It bills citation tokens at $2 per million, reasoning tokens at $3 per million, and search queries at $5 per 1,000.

  • Sonar Chat Completions is deprecated. Perplexity supports it until 27 September 2026 and points new builds at the Agent API, which prices tools per invocation.

Perplexity API pricing in 2026

Model

Input

Output

 

Sonar

$1

$1

Sonar Pro

$3

$15

Sonar Reasoning Pro

$2

$8

Sonar Deep Research

$2

$8

Model

Low context

Medium context

---

---

---

Sonar

$5

$8

Sonar Pro

$6

$10

Sonar Reasoning Pro

$6

$10

Sonar Pro with Pro Search (search_type: pro)

$14

$18




Source: docs.perplexity.ai/docs/getting-started/pricing, read 22 Sep 2026.

What Perplexity actually meters

Perplexity meters tokens and requests separately, and the docs state the formula outright: total cost per query equals token costs plus a request fee. That fee applies to Sonar, Sonar Pro and Sonar Reasoning Pro, and not to Sonar Deep Research.

The fee tracks search_context_size, set inside web_search_options. Low is the default and the cheapest, high buys maximum search depth. Moving Sonar from low to high more than doubles the fee, from $5 to $12 per 1,000, without touching the token rate.

That split matters. Perplexity's published Sonar example uses 9 input and 411 output tokens at low context: tokens cost $0.000009 and $0.000411, the request fee costs $0.005. So 92% of that line is search, not the model. Budget from token counts alone and you'll be out by an order of magnitude.

Sonar Deep Research inverts the shape, dropping the request fee and adding three meters instead. Perplexity's worked example for one call totals $0.816123, of which $0.581841 is reasoning tokens. Every response carries a usage.cost object splitting request_cost from token costs, so you reconcile per call.

How credits work

Perplexity runs prepaid credits, not postpaid invoicing. You buy credits in the API console, Stripe takes the payment, and calls draw the balance down. Auto reload adds credits when the balance falls below a threshold you set, and the docs recommend enabling it.

Cumulative purchases also drive your usage tier, which sets your rate limits. Tier 0 starts at $0 and Tier 5 at $5,000 of lifetime purchases, with $50, $250, $500 and $1,000 in between. Tiers count lifetime spend rather than current balance and never downgrade: on the Agent API that's 1 query per second at Tier 0 against 33 at Tier 5, and on Sonar 50 requests a minute against 4,000.

Credit expiry, rollover, pooling across projects and deduction order are all not documented, and neither is any refund path for unused balance. Enterprises can buy credits through AWS Marketplace.

What happens when you hit the limit

You get blocked, not billed. Perplexity's docs are blunt about it: run out of credits and your API keys are blocked until you top the balance up. Calls then return a 401, which the FAQ lists alongside invalid and deleted keys as a credential failure rather than a billing one. There's no grace allowance and no arrears billing, which makes auto reload worth checking before any launch.

Rate limiting behaves differently. Exceed your tier's QPS or per-minute limit and you get a 429 with a Retry-After header, and Router requests rejected with a 429 aren't billed.

How Perplexity's pricing has changed across all these years

Date

Milestone

Source

 

27 Sep 2026

Sonar support ends

docs.perplexity.ai/docs/resources/changelog

Jul 2026

Sonar folded into the Agent API. The new Router API bills per token with no request fee

docs.perplexity.ai/docs/resources/changelog

Apr 2026

API credits become purchasable through AWS Marketplace

docs.perplexity.ai/docs/resources/changelog

Nov 2025

Pro Search reaches GA on Sonar Pro, carrying a higher request-fee band

docs.perplexity.ai/docs/resources/changelog

Jul 2025

Responses return a cost object with request_cost split from token costs

docs.perplexity.ai/docs/resources/changelog

Mar 2025

Search context modes introduced and citation tokens stop billing except on Deep Research. Default from 18 April 2025

docs.perplexity.ai/docs/resources/changelog

Jan 2025

Sonar and Sonar Pro launch, replacing the llama-3.1-sonar family

docs.perplexity.ai/docs/resources/changelog

Source: docs.perplexity.ai/docs/resources/changelog, read 22 September 2026. Wayback's CDX index

Flexprice’s Take

Sonar's two-axis meter is the honest way to price a grounded model API, and Perplexity documents it better than most vendors document a single meter.

Retrieval costs real money and it doesn't scale with token count, so Perplexity meters it separately. Folding search into a token rate would subsidise heavy searchers or overcharge light ones. Instead search_context_size is its own dial: Sonar moves from $5 to $12 per 1,000 requests between low and high context, and the token rate never budges.

The per-response cost object proves the bill line by line, splitting request_cost from token costs on every call. Perplexity publishes worked arithmetic on every model page too, a standard nothing else in this index meets.

Predicting the bill is the hard part. Four models, three context tiers, a separate Pro Search band reaching $22 per 1,000, an auto mode whose fee varies by classification, and Deep Research running four meters at once. It's also on a clock: Sonar ends on 27 September 2026.

Best For

Teams that pin search_context_size deliberately and reconcile against the response cost object.

Watch Out For

The 27 September 2026 end date, and search_type: auto, which makes the request fee unpredictable by design.

Manish Choudhary

CEO & Co-founder, Flexprice

Reselling search calls and need per-model cost and margin per customer?

Flexprice tracks both per model per customer, so you can see which accounts pay for themselves,

Flexprice’s Take

Sonar's two-axis meter is the honest way to price a grounded model API, and Perplexity documents it better than most vendors document a single meter.

Retrieval costs real money and it doesn't scale with token count, so Perplexity meters it separately. Folding search into a token rate would subsidise heavy searchers or overcharge light ones. Instead search_context_size is its own dial: Sonar moves from $5 to $12 per 1,000 requests between low and high context, and the token rate never budges.

The per-response cost object proves the bill line by line, splitting request_cost from token costs on every call. Perplexity publishes worked arithmetic on every model page too, a standard nothing else in this index meets.

Predicting the bill is the hard part. Four models, three context tiers, a separate Pro Search band reaching $22 per 1,000, an auto mode whose fee varies by classification, and Deep Research running four meters at once. It's also on a clock: Sonar ends on 27 September 2026.

Best For

Teams that pin search_context_size deliberately and reconcile against the response cost object.

Watch Out For

The 27 September 2026 end date, and search_type: auto, which makes the request fee unpredictable by design.

Manish Choudhary

CEO & Co-founder, Flexprice

Reselling search calls and need per-model cost and margin per customer?

Flexprice tracks both per model per customer, so you can see which accounts pay for themselves,

Customer
Sentiment Highlights

Frequently Asked Questions

Frequently Asked Questions

How much does the Perplexity Sonar API cost?

What is the Perplexity request fee?

Does the Perplexity API have a free tier?

What happens if my Perplexity credits run out?

Launch usage-based billing this week, not next quarter

Launch usage-based billing this week, not next quarter

Get Instant Feedback on Your Pricing | Join the Flexprice Community with 400+ Builders on Slack

Join the Flexprice Community on Slack