gpt-5.6-luna

GPT-5.6 Luna API — price per token

GPT-5.6 Luna is the fast, economical tier of the GPT-5.6 line — built for high-volume, latency-sensitive work at one fifth of the flagship price.

Pricing per 1M tokens

RateOfficial OpenAIHere, from (−60%)Here, best (−70%)
Input$1$0.4$0.3
Cached input$0.1$0.04$0.03
Cache write$1.25$0.5$0.375
Output$6$2.4$1.8

Input

Official OpenAI
$1
Here, from (−60%)
$0.4
Here, best (−70%)
$0.3

Cached input

Official OpenAI
$0.1
Here, from (−60%)
$0.04
Here, best (−70%)
$0.03

Cache write

Official OpenAI
$1.25
Here, from (−60%)
$0.5
Here, best (−70%)
$0.375

Output

Official OpenAI
$6
Here, from (−60%)
$2.4
Here, best (−70%)
$1.8

Every request is metered at the official rate first, then your progressive B2C discount (60% at the start, up to 70% as cumulative top-ups grow) is subtracted before it touches your prepaid balance. Context window: 272K tokens. Max output: 32K tokens. Reasoning efforts: none, low, medium, high, xhigh, max.

Best for

  • Classification, extraction and summarization at scale.
  • Latency-sensitive chat and routing layers.
  • Cheap pre-processing before a Sol or Terra call.

Good to know

  • Same reasoning-effort range as the flagship, including max.
  • Text and image input, text output. SSE streaming on both Responses and Chat Completions.
  • Requests above 272K input tokens bill at OpenAI long-context rates: 2× input and 1.5× output on the whole request.

How to use GPT-5.6 Luna

Create a free account, generate one key, and point any OpenAI-compatible tool at https://openai.api.apitoken.sale/v1 with model ID gpt-5.6-luna — Responses and Chat Completions both work, authenticated with Authorization: Bearer. New accounts include $10 of API usage at official prices — enough to test the model before topping up.

Frequently asked questions

How much does the GPT-5.6 Luna API cost?

Officially $1 per 1M input tokens and $6 per 1M output tokens, with cached input at $0.10. With the apiToken.sale discount that starts at $0.40/$2.40 and reaches $0.30/$1.80 — the cheapest way to run GPT-5.6.

What is Luna good for?

High-volume, low-latency work: classification, extraction, summarization, routing and simple chat. For complex reasoning, step up to Terra or Sol.

What is the model ID?

gpt-5.6-luna. It works on the same apiToken.sale key, balance and OpenAI-compatible endpoint as every other GPT model.

Run GPT-5.6 Luna on the OpenAI-compatible API at up to 70% off — instant key, prepaid balance, card or crypto.