Skip to content
PALEPALE / DOCUMENTATION

Choose a model, know the rate

Compare all eight offers and understand availability, savings, and token prices.

Compare the three prices

Open the full model table to compare input, cached input, and output rates per million tokens. Sort any price column or search for a model. Each access link keeps the model you chose.

Input is the context you send. Output is the answer a model generates. Eligible cached input uses its own rate instead of the regular input rate. There is no cached output price.

What the savings mean

The displayed percentage is the smallest saving across input, cached input, and output, rounded down. The comparison rates were recorded on 2026-10-08 and are not independently verified provider list prices. A discount does not describe a plan's token allowance.

Read the active catalog

bash
curl "$PALEPALE_BASE_URL/models"

Use an exact id from the data array in your requests. context_window is the total context capacity; max_output_tokens is the response ceiling. Inspect capabilities.streaming, capabilities.tool_calling, capabilities.reasoning, capabilities.temperature and capabilities.caching before enabling those features. Missing capabilities are not a promise of support.

The pricing object contains input_per_1m_usd, output_per_1m_usd and optional cached_input_per_1m_usd. Compare the same units: USD per one million tokens.

Availability and limits

The public catalog is a planned lineup. Confirm model access on Discord before activation. After activation, use your gateway's GET /v1/models for approved IDs, context windows, output ceilings, and capabilities. A published offer does not make a model callable.

Pick for your workload

For long conversations, compare input and cached-input prices. For long answers, compare output prices. Check context and output limits against the work you expect. Learn the cost calculation or compare plans.

Need a hand?

Open Palepale Discord for activation or account support. Compare API offers or choose a plan.

API tokens. Up to 98% less.

Choose your model and pay by token usage. Palepale rates in green; comparison rates crossed out. USD per 1 million tokens.

API access
GPT 6 AstraOpenAI · Plannedgpt-6-astra
Palepale: $0.45675Reference: $10.00
Palepale: $0.03654Reference: $1.00
Palepale: $1.827Reference: $50.00
−95%
128000
Get API ↗
GPT 6 SolOpenAI · Plannedgpt-6-sol
Palepale: $0.16Reference: $2.00
Palepale: $0.016Reference: $0.20
Palepale: $0.80Reference: $10.00
−92%
128000
Get API ↗
Kimi K3Moonshot · Plannedkimi-k3
Palepale: $0.27Reference: $3.00
Palepale: $0.027Reference: $0.30
Palepale: $1.35Reference: $15.00
−91%
128000
Get API ↗
GPT 5.6 SolBest dealOpenAI · Plannedgpt-5.6-sol
Palepale: $0.08Reference: $4.00
Palepale: $0.008Reference: $0.40
Palepale: $0.40Reference: $20.00
−98%
128000
Get API ↗
Gemini 3.8 FlashGoogle · Plannedgemini-3.8-flash
Palepale: $0.0675Reference: $0.75
Palepale: $0.0067Reference: $0.075
Palepale: $0.3375Reference: $3.75
−91%
128000
Get API ↗
Claude Fable 5.1Anthropic · Plannedclaude-fable-5.1
Palepale: $1.50Reference: $10.00
Palepale: $0.0375Reference: $0.25
Palepale: $7.50Reference: $50.00
−85%
128000
Get API ↗
DeepSeek V4.1 FlashDeepSeek · Planneddeepseek-v4.1-flash
Palepale: $0.045Reference: $0.30
Palepale: $0.0009Reference: $0.006
Palepale: $0.18Reference: $1.20
−85%
128000
Get API ↗
Claude Opus 5.5Anthropic · Plannedclaude-opus-5.5
Palepale: $0.60Reference: $4.00
Palepale: $0.03Reference: $0.20
Palepale: $3.00Reference: $20.00
−85%
128000
Get API ↗
GLM 5.3Z.ai · Plannedglm-5.3
Palepale: $0.756Reference: $1.40
Palepale: $0.140401Reference: $0.26
Palepale: $2.376Reference: $4.40
−45%
128000
Get API ↗