GPT-5.6 Sol, Terra & Luna:
Full Review, Benchmarks, Pricing & Access Guide (2026)

If you are a developer or AI team lead running Codex, Cursor, or custom agent workflows, OpenAI's June 26, 2026 launch of GPT-5.6 Sol, Terra, and Luna reshapes coding and security research choices: Sol tops TerminalBench 2.1 at 91.9% with a 96.7% CTF hit rate. The catch: only about 20 vetted partner organizations can access the models today. This guide covers pricing, Max/Ultra reasoning modes, every benchmark number, the government restriction story, the Claude Mythos 5 comparison, July Cerebras 750 token/s, and a six-step access playbook. Node tiers are on the NOVAKVM pricing page.

GPT-5.6 family at a glance (June 27, 2026)
Model Tier Input Output Highlight
GPT-5.6 Sol Flagship $5 / 1M tokens $30 / 1M tokens TerminalBench 2.1 #1 at 91.9%
GPT-5.6 Terra Balanced $2.50 / 1M tokens $15 / 1M tokens GPT-5.5-level performance, 50% lower cost
GPT-5.6 Luna Lightweight $1 / 1M tokens $6 / 1M tokens High-frequency tasks, ~80% cheaper than Sol

OpenAI introduced celestial naming for the first time: Sol (Sun), Terra (Earth), Luna (Moon) for flagship, balanced, and lightweight tiers. This is OpenAI's most significant release since GPT-5.5 — and the first family where every tier, including entry-level Luna, crossed OpenAI's internal High cybersecurity risk rating.

  • Limited access: Following Trump's June 2 executive order, GPT-5.6 is limited to about 20 approved preview partners. General ChatGPT users cannot access it yet; broad availability is expected within weeks.
  • A new precedent: This is the first time the U.S. government formally required an AI company to restrict a frontier model release. OpenAI complied while publicly opposing permanent government gatekeeping.
  • June got blocked: OpenAI GPT-5.6 (limited), Anthropic Claude Fable 5 / Mythos 5 (offline June 12 under export control), and Google Gemini 3.5 Pro (delayed to July) all stalled in the same month.
  • Context window: All three tiers report roughly 1.5M tokens (~50% above GPT-5.5's 1M). Confirm with the full system card at general release.
  • Short reign at #1: Claude Mythos 5 held TerminalBench's top spot for only 17 days (since June 9) before Sol's 91.9% dethroned it.

"We don't believe this kind of government access process should become the long-term default. It keeps the best tools from users, developers, enterprises, cyber defenders, and global partners who need them." — OpenAI CEO Sam Altman

GPT-5.6 Sol is OpenAI's most capable model to date — built for hard coding, long-horizon cybersecurity research, and multi-step autonomous agents. Pricing matches GPT-5.5 ($5 / $30 per MTok) with a major capability jump.

  • Max mode: Sol takes additional reasoning time before responding — trading latency for accuracy when the answer must be right, not just fast.
  • Ultra mode: A multi-agent architecture that spawns parallel subagents, splits work, and merges results. This is why Sol hit 91.9% on TerminalBench (standard mode: 88.8%). Ultra consumes significantly more tokens — reserve it for genuinely complex tasks.

GPT-5.6 Terra is the daily enterprise workhorse for high-volume support, internal tools, and document analysis. Performance sits near GPT-5.5 at half the API cost ($2.50 / $15 per MTok).

GPT-5.6 Luna optimizes for high-frequency, low-latency workloads: summarization, drafting, routine automation. Luna is OpenAI's first non-flagship model to earn High ratings in both cybersecurity and biology ($1 / $6 per MTok).

Which GPT-5.6 model should you use?
Your need Recommended model
Complex coding agents and multi-step workflows Sol (Ultra mode)
High-volume business docs, support, API at scale Terra
Summarization, drafting, routine automation Luna
GPT-5.5-level capability on a tight budget Terra (same tier performance, 50% lower cost)
Latency-critical real-time apps (from July) Sol on Cerebras (up to 750 token/s)

TerminalBench 2.1 tests multi-step command-line planning across 89 complex challenges — closer to real agent work than single-shot code completion:

TerminalBench 2.1 coding agent leaderboard
Model Score Mode
GPT-5.6 Sol 91.9% Ultra (multi-agent)
GPT-5.6 Sol 88.8% Standard
Claude Mythos 5 88.0% Standard
GPT-5.5 83.4% Standard
Gemini 3.1 Pro Preview 70.7% Standard

Agent's Last Exam (long-horizon professional tasks, code mode): Sol at 50.9% — the only model to cross 50%. Luna sits slightly above GPT-5.5.

CTF hit rates: Sol 96.7%, Terra 91.84%, Luna 85.19%. GPT-5.6 is the first OpenAI family where all three tiers hit the High cybersecurity classification.

ExploitBench: Sol matches Anthropic's Mythos Preview while using only ~1/3 of the output tokens — same security research capability at dramatically lower cost. OpenAI red-teaming confirmed Sol can identify vulnerabilities and exploit primitives in Chromium and Firefox codebases but cannot autonomously engineer a complete, functional exploit chain, staying below the Cyber Critical threshold.

Life sciences: On GeneBench v1, Sol matches or exceeds GPT-5.5 with fewer tokens. HealthBench Professional: 60.5+8.7 points above GPT-5.5.

GPT-5.6 Sol vs Claude Mythos 5 head-to-head
Category GPT-5.6 Sol Claude Mythos 5
TerminalBench 2.1 91.9% (Ultra) / 88.8% 88.0%
ExploitBench Near-identical, ~3× cheaper on tokens Strong (restricted access)
Input pricing $5 / M $10 / M (currently offline)
Availability Limited preview → general release soon Offline (U.S. export control)
Context window ~1.5M tokens 200K tokens

Sol beats Mythos 5 on TerminalBench and offers comparable security research at half the price. Claude Fable 5 may still lead on SWE-Bench Pro; we'll update once OpenAI publishes the complete GPT-5.6 system card.

Government restriction: On June 2, 2026, President Trump signed an executive order allowing U.S. agencies up to 30 days of pre-release access to review frontier models. On June 26, following a White House request coordinated by OSTP and ONCD, OpenAI limited GPT-5.6 to approximately 20 pre-approved trusted partner organizations.

June 2026: all three flagship releases got blocked
Company Model Status
OpenAI GPT-5.6 Sol/Terra/Luna Limited preview (~20 orgs)
Anthropic Claude Fable 5 / Mythos 5 Forced offline June 12 via export control
Google Gemini 3.5 Pro Delayed to July

Cerebras acceleration (July): GPT-5.6 Sol on Cerebras hardware targets up to 750 tokens per second. Most frontier models today run 50–150 token/s — a 5× to 15× speed jump. A 10-second response could complete in under one second for real-time coding assistants and live customer-facing AI.

Right now (June 2026): About 20 approved partner organizations via API and Codex only. General ChatGPT users are still waiting.

Coming in July 2026: General ChatGPT availability (Plus and Pro first), public API access, and GPT-5.6 Sol on Cerebras for select enterprise customers (up to 750 token/s). Polymarket traders assign an 87% probability that GPT-5.6 will be broadly released by July 31, 2026.

Safety built into GPT-5.6: Real-time misuse classifiers on every output, account-level review for sensitive workflows, 700,000 A100-equivalent GPU hours of automated red-teaming, universal jailbreak testing, a specialized large reasoning model as a backstop filter, and external security organization testing before launch.

  1. Verify your access tier: Check platform.openai.com and ChatGPT settings for partner whitelist status. If not listed, keep GPT-5.5 or Gemini 3.5 Pro as production fallback.
  2. Externalize model routing: Manage gpt-5.6-sol, gpt-5.6-terra, and gpt-5.6-luna via LiteLLM or env vars — never hard-code unreleased slugs.
  3. Match model to workload: Sol Ultra for complex agents and coding; Terra for high-volume business APIs; Luna for summarization and classification — avoid routing everything through the flagship.
  4. Budget for Ultra tokens: Multi-agent Ultra mode burns significantly more tokens. Set daily Sol caps and automatic Terra/Luna downgrade rules in your cost dashboard.
  5. Build a launch-day checklist: ChatGPT availability does not equal API availability. Plan a 24–48 hour API buffer and watch openai.com/blog plus the Deployment Safety System Card.
  6. Fix your agent runtime: Move Codex CLI, Cursor Cloud Agent, and multi-model trials to a stable 24/7 Mac host so sleep, OAuth expiry, and disk pressure do not eat your preview window.

  • TerminalBench 2.1 record: GPT-5.6 Sol Ultra 91.9%, standard 88.8%; Claude Mythos 5 held #1 only 17 days (June 9 through June 26).
  • CTF triple High: Sol 96.7%, Terra 91.84%, Luna 85.19% — OpenAI's first all-tier High cybersecurity product line.
  • Cerebras throughput: Up to 750 token/s from July — roughly 5–15× today's 50–150 token/s frontier range.
  • HealthBench Professional: Sol 60.5, +8.7 vs GPT-5.5; ExploitBench token use about one-third of comparable rivals.

FAQ:

  • Is GPT-5.6 available on ChatGPT now? Not for the general public. Limited to ~20 trusted partners; full rollout expected in July 2026.
  • Is GPT-5.6 Sol better than Claude Fable 5 for coding? Sol leads TerminalBench (91.9% vs Mythos 5's 88%). Fable 5 still leads SWE-Bench Pro; wait for OpenAI's full benchmark report.
  • What is Ultra mode? Parallel subagents split work and merge results — major gains on complex tasks, higher token cost.
  • Why is GPT-5.6 restricted? White House / OSTP / ONCD requested limited release under Trump's June 2 executive order. OpenAI complied but opposes making this permanent.
  • How fast on Cerebras? Up to 750 token/s for select enterprise customers starting July 2026.
  • Context window size? Approximately 1.5M tokens reported; confirm when the full system card ships.

GPT-5.6 marks breakthroughs in capability (Ultra multi-agent #1 on TerminalBench), efficiency (security research at one-third the tokens), and speed (Cerebras 750 token/s) — while setting a precedent for U.S. government involvement in frontier model releases.

Primary references — re-open after policy or release updates:

OpenAI Official: Previewing GPT-5.6 Sol

OpenAI Deployment Safety System Card

VentureBeat: GPT-5.6 Launch Coverage

SiliconAngle: GPT-5.6 vs Claude Mythos 5

TechTimes: Government Lock Analysis

Running Codex and Ultra multi-agent trials on a laptop that sleeps will waste a 91.9% TerminalBench score to interrupted jobs, expired OAuth, and full disks. Cloud APIs alone cannot replace a stable 24/7 execution surface with on-machine Xcode validation.

If you need 24/7 agent trials, fixed SSH, and predictable Apple Silicon capacity before Sol goes GA, moving Cursor Cloud Agent, Codex batch jobs, and LiteLLM routing to dedicated bare-metal Mac hosts is usually cheaper than firefighting unstable environments: NOVAKVM offers multi-region Mac Mini M4 / M4 Pro flexible terms with fixed bandwidth and default SSH — built for iOS CI/CD and AI agent automation on one machine. See the pricing page, order page, and help center.