Skip to main content

Models & Pricing

HPP Router exposes models from multiple providers behind one OpenAI-compatible API. Billing is token-based at the resolved model's rates. You can pay with prepaid quota (default) or, for paid models, settle on-chain via the x402 wallet rail (commonly USDC.e; the challenged accepts[0].asset is authoritative).

Listing models​

Use the OpenAI-compatible models endpoint to discover what's available, including the virtual hpprouter/auto model. No API key is required for this endpoint.

curl https://router.hpp.io/llm/v1/models

The response is an OpenAI-style list with pricing and capability metadata. pricing.* values are USD per token (multiply by 1_000_000 for the familiar $/1M figure):

{
"object": "list",
"data": [
{
"id": "hpprouter/auto",
"object": "model",
"owned_by": "hpprouter",
"name": "Auto",
"description": "Smart routing picks a provider and model based on your request.",
"pricing": null,
"context": null,
"max_output": null,
"tool": null,
"structured": null,
"knowledge_cutoff": null,
"input_modalities": [],
"output_modalities": []
},
{
"id": "openai/gpt-5",
"object": "model",
"owned_by": "openai",
"name": "GPT-5",
"description": null,
"pricing": {
"input": 0.00000125,
"output": 0.00001,
"cache_write": null,
"cache_read": 1.25e-7
},
"context": 272000,
"max_output": 128000,
"tool": true,
"structured": true,
"knowledge_cutoff": null,
"input_modalities": ["text", "image"],
"output_modalities": ["text"]
}
]
}
FieldMeaning
idModel identifier — use this as the request model.
owned_byProvider that owns the model.
name / descriptionDisplay name and short description (may be null).
pricingnull for hpprouter/auto (billed at the resolved model).
pricing.inputUSD per input token.
pricing.outputUSD per output token.
pricing.cache_write / cache_readUSD per cache token, when applicable (may be null).
context / max_outputContext window and max output tokens, when known.
tool / structuredTool calling / structured output support (may be null).
knowledge_cutoffKnowledge cutoff date string, when known.
input_modalities / output_modalitiesAccepted / produced modalities (e.g. text, image).
note

The models list is the source of truth for what is currently enabled. The example rates below are illustrative and may change.

Model identifiers​

Specify a model as provider/model, or use the virtual smart-routing model:

Example modelDescription
hpprouter/autoSmart routing — the gateway resolves an actual model per request.
openai/gpt-5OpenAI GPT-5.
openai/gpt-4oOpenAI GPT-4o (vision-capable).
anthropic/claude-sonnet-4Anthropic Claude Sonnet 4.
moonshotai/kimi-k2.6Moonshot Kimi.
ollama/gpt-oss:120bLocal Ollama model (billed at $0).
ollama/solidity-master:2Solidity finetune model on a dedicated Ollama backend (billed at $0).

How billing works​

  • API pricing.* values are USD per token. Human-readable $/1M ≈ pricing.input × 1_000_000.
  • When using $/1M rates from the table below, cost ≈ (prompt_tokens × input_rate + completion_tokens × output_rate) / 1,000,000.
  • When plugging API pricing.* values directly, omit the / 1,000,000 and also include cache tokens when present:
    (prompt_tokens × input) + (completion_tokens × output) + (cache_write_tokens × cache_write) + (cache_read_tokens × cache_read).
  • The usage block in each response reports the token counts used for billing.
  • Local models (e.g. ollama/*) are tracked at $0 cost, but token usage is still logged.

Example rates​

ModelInput (per 1M)Output (per 1M)
openai/gpt-5$1.25$10
openai/gpt-image-1 (image generation)$5$40

See Image Generation for how image quality affects output-token usage.

Billing with smart routing​

When you request hpprouter/auto, billing uses the resolved model's pricing — not a price for auto itself (pricing is null on the models list). The resolved model is returned in the X-HPP-Router-Resolved-Model response header and recorded in your usage logs. See Smart Routing for details.

Checking your usage​

  • GET /api/quota-check — check prepaid quota remaining for the authenticated consumer.
  • GET /api/usage — usage summary (requests, total tokens, total cost); use ?rail=wallet for on-chain settlement fields.

See Usage & Settlement.