- Documentation
- Choose a model, know the rate
Choose a model, know the rate
Compare all eight offers and understand availability, savings, and token prices.
Compare the three prices
Open the full model table to compare input, cached input, and output rates per million tokens. Sort any price column or search for a model. Each access link keeps the model you chose.
Input is the context you send. Output is the answer a model generates. Eligible cached input uses its own rate instead of the regular input rate. There is no cached output price.
What the savings mean
The displayed percentage is the smallest saving across input, cached input, and output, rounded down. The comparison rates were recorded on 2026-10-08 and are not independently verified provider list prices. A discount does not describe a plan's token allowance.
Read the active catalog
curl "$PALEPALE_BASE_URL/models"Use an exact id from the data array in your requests. context_window is the total context capacity; max_output_tokens is the response ceiling. Inspect capabilities.streaming, capabilities.tool_calling, capabilities.reasoning, capabilities.temperature and capabilities.caching before enabling those features. Missing capabilities are not a promise of support.
The pricing object contains input_per_1m_usd, output_per_1m_usd and optional cached_input_per_1m_usd. Compare the same units: USD per one million tokens.
Availability and limits
The public catalog is a planned lineup. Confirm model access on Discord before activation. After activation, use your gateway's GET /v1/models for approved IDs, context windows, output ceilings, and capabilities. A published offer does not make a model callable.
Pick for your workload
For long conversations, compare input and cached-input prices. For long answers, compare output prices. Check context and output limits against the work you expect. Learn the cost calculation or compare plans.
Need a hand?
Open Palepale Discord for activation or account support. Compare API offers or choose a plan.
API tokens. Up to 98% less.
Choose your model and pay by token usage. Palepale rates in green; comparison rates crossed out. USD per 1 million tokens.
| API access | ||||||
|---|---|---|---|---|---|---|
GPT 6 AstraOpenAI · Planned gpt-6-astra | Palepale: $0.45675 | Palepale: $0.03654 | Palepale: $1.827 | −95% | 128000 | Get API ↗ |
GPT 6 SolOpenAI · Planned gpt-6-sol | Palepale: $0.16 | Palepale: $0.016 | Palepale: $0.80 | −92% | 128000 | Get API ↗ |
Kimi K3Moonshot · Planned kimi-k3 | Palepale: $0.27 | Palepale: $0.027 | Palepale: $1.35 | −91% | 128000 | Get API ↗ |
GPT 5.6 SolBest dealOpenAI · Planned gpt-5.6-sol | Palepale: $0.08 | Palepale: $0.008 | Palepale: $0.40 | −98% | 128000 | Get API ↗ |
Gemini 3.8 FlashGoogle · Planned gemini-3.8-flash | Palepale: $0.0675 | Palepale: $0.0067 | Palepale: $0.3375 | −91% | 128000 | Get API ↗ |
Claude Fable 5.1Anthropic · Planned claude-fable-5.1 | Palepale: $1.50 | Palepale: $0.0375 | Palepale: $7.50 | −85% | 128000 | Get API ↗ |
DeepSeek V4.1 FlashDeepSeek · Planned deepseek-v4.1-flash | Palepale: $0.045 | Palepale: $0.0009 | Palepale: $0.18 | −85% | 128000 | Get API ↗ |
Claude Opus 5.5Anthropic · Planned claude-opus-5.5 | Palepale: $0.60 | Palepale: $0.03 | Palepale: $3.00 | −85% | 128000 | Get API ↗ |
GLM 5.3Z.ai · Planned glm-5.3 | Palepale: $0.756 | Palepale: $0.140401 | Palepale: $2.376 | −45% | 128000 | Get API ↗ |