AIVAX

Pricing

Service usage prices are listed below in USD. M means one million tokens; 1k means one thousand units. Approximate prices (~) vary with the model used and the work performed.

See subscription pricing for monthly plan prices and Plans and limits for quotas. Usage rates are subject to the plan multiplier:

  • Free: +25% on inference taxes;
  • Pro: +5% on inference taxes;
  • Max: 0% on inference taxes.

BYOK are not affected by inference taxes.

Free, Pro, and Max include separate daily allowances for eligible RAG embeddings, Reflex reranking, Julia-1 semantic decisions, and Fetch/OCR extraction. The rates below apply when a metered item is not covered. Coverage is all-or-nothing per item, not necessarily per complete request: an item that cannot fit within the remaining allowance and its permitted margin is billed in full. Compare the allowances and check exclusions in Plans and limits. LLM subscription coverage is currently disabled.

Inference and Moderation #

Inference rates depend on the selected model, provider, input size, and media type. Moderation is charged separately in Processing Units (PUs), covering input, cached input, and output usage; its PU price varies with the model and provider used.

DescriptionPricing
AI model and AI Gateway inferenceSelected model and provider rates
Input moderationVariable price per PU; separate from the main inference charge

Semantic decisions #

The decision-model rates below are base USD prices per million input tokens, before account and plan adjustments. Output tokens have no charge in the current decision-model catalog. Julia-1 is eligible for the daily allowance described in Plans and limits; other decision models are billed normally.

ModelInput price / million tokens
@supersonic-labs/julia-1$0.008
@typesafe/jev-1.13$0.042
@respan/span-01$0.020
@respan/span-01-lite$0.000
@jaredpalmer/kev-4b$0.042
@upstage/solar-decide$0.050
@cloudflare/clef$0.240
@cloudflare/clef-flash$0.090
@liquid/d1$0.040
@perplexity/pplx-decider-v1-27b$0.040
@openai/gpt-6-luna-decisions$0.100

See Semantic decisions for model selection and how input usage is measured.

Agentic Tests #

Each test includes the selected model or AI Gateway’s inference charges, plus simulated-user and judge usage at the selected profile’s rates.

DescriptionPricing
Model or AI Gateway under testRegular inference rates
Low profile - simulated userInput $0.25/M tokens; cache $0.025/M tokens; output $1.50/M tokens
Low profile - judgeInput $0.30/M tokens; cache $0.03/M tokens; output $2.50/M tokens
Medium profile - simulated userInput $0.75/M tokens; cache $0.075/M tokens; output $3.75/M tokens
Medium profile - judgeInput $0.75/M tokens; cache $0.075/M tokens; output $3.75/M tokens
High profile - simulated userInput $0.75/M tokens; cache $0.075/M tokens; output $3.75/M tokens
High profile - judgeInput $1.25/M tokens; cache $0.15/M tokens; output $4.25/M tokens

RAG and Collections #

Indexing and search are billed by token usage. Generated RAG responses are charged separately from query embedding, and their price varies with the summarization model.

DescriptionPricing
Collection text embedding$0.10/M tokens
Semantic search - query cache miss$0.10/M tokens
Semantic search - query cache hitZero
RAG response generation~$0.50/M tokens, excluding query rates
Reflex - cache miss$0.015/M tokens
Reflex - cache hit$0.003/M tokens

Media Injector #

Converting media into RAG documents is billed for input, cached input, output, and media usage. The source file, optional context, and generated content affect the total. Rates depend on media type and input-token volume.

DescriptionPricing
PDFs and images - up to 272K input tokensInput $0.30/M tokens; cache $0.03/M tokens; output $1.80/M tokens
PDFs and images - above 272K input tokensInput $0.60/M tokens; cache $0.06/M tokens; output $3.60/M tokens
Audio - up to 256K input tokensInput/media $0.60/M tokens; cache $0.12/M tokens; output $3.00/M tokens
Audio - above 256K input tokensInput/media $1.20/M tokens; cache $0.24/M tokens; output $6.00/M tokens
VideoInput/media $0.45/M tokens; cache $0.045/M tokens; output $3.75/M tokens

Text Tools #

Text segmentation and classification are billed by token usage.

DescriptionPricing
Text segmentation$0.30/M tokens
Text classification$0.10/M tokens

Voice and Media #

Generation and transcription rates depend on the selected model. Media description pricing is approximate and depends on the available processing model.

DescriptionPricing
Voice SessionsSelected realtime model rates
Speech-to-textVaries by model
Text-to-speechVaries by model
Image generationFixed output and reference-image tariffs by model
Media descriptions~$1.50/M tokens

Image generation charges each delivered output at the selected model’s fixed output price, plus its per-reference price for every reference sent with that output. Prompt processing is included. Token- and megapixel-priced providers use rounded-up estimates, not exact provider-cost pass-through. No additional AIVAX image-generation markup or account and plan multiplier applies. Current tariffs are listed in the Models catalog; see Image generation.

Web Search, OCR and Fetch #

Web and X searches are billed per search. Advanced web search is billed by token usage and varies with the model and number of interactions. Fetch and OCR extraction use Processing Units (PUs), with a daily free allowance by plan. Optional schema-guided JSON conversion is charged separately. The extraction allowances and PU rates do not apply to JSON conversion or moderation.

DescriptionPricing
Web search$5/1k searches
X (Twitter) search$5/1k searches
Advanced web search~$0.75/M tokens
Fetch and OCR extraction - FreeBase daily allowance; uncovered items $0.15/1k PUs
Fetch and OCR extraction - Pro10× Free daily allowance; uncovered items $0.05/1k PUs
Fetch and OCR extraction - Max5× Pro daily allowance; uncovered items $0.02/1k PUs
Fetch JSON conversion (responseSchema)Variable inference-based price per PU; charged separately, with no daily extraction allowance

For the Fetch API, processingUnits reports text/OCR extraction usage and jsonProcessingUnits reports the additional schema-guided JSON conversion usage. JSON PUs account for input, cached input, and output token usage at the processing model and provider’s rates; they are not priced at the plan’s OCR rate. The plan’s inference multiplier applies to JSON conversion. Omitting responseSchema or setting it to null disables conversion, reports jsonProcessingUnits: 0, and incurs no JSON conversion charge.

Storage #

Each plan includes storage. Pro and Max overages are billed hourly at the monthly rates below; Free storage cannot be expanded.

DescriptionPricing
Free storage30 MB included; no expansion
Pro storage2 GB included; excess $0.50/GB/month
Max storage20 GB included; excess $0.20/GB/month

Data Collection Discount #

Accounts that enable Data collecting receive a 10% discount on eligible RAG query-embedding usage and eligible reranking operations performed while collection is enabled. Other services keep their regular prices.

Other Tools #

The following tools have no separate tool charge. Model inference used to invoke them is still billed at its regular rate.

DescriptionPricing
Memory and calendarNo separate charge
Advanced requestsNo separate charge
Document generationNo separate charge
Web page generationNo separate charge

Type to search the documentation.