OpenAI shipped two new models on the same day this week, GPT-6 Sol and GPT-6 Luna, and quietly cut its own API prices in half while doing it. If you have been sending every prompt to the flagship model because you never bothered to check whether a cheaper tier could do the job just as well, this is the week that habit starts costing you real money.
What Are GPT-6 Sol and Luna?
GPT-6 Sol and GPT-6 Luna are two new OpenAI models released on 22 September 2026, trained with the same methods behind the flagship GPT-6 Astra but priced far lower. Sol is built for coding and complex professional work; Luna is built for high-volume clerical tasks like summarising, extracting data and answering quick questions.
Both models arrived alongside a permanent (not promotional) price cut of roughly 50% compared to the previous GPT-5.6 generation, according to VentureBeat's coverage of the release. OpenAI says Sol now makes roughly half as many mistakes as its predecessor, closing most of the gap to Astra-level reliability at a fraction of the cost.
The naming is worth understanding once, because it will keep coming up. Astra is the flagship line, the model OpenAI ships when it wants to show off its best reasoning and coding ability regardless of cost. Sol and Luna are the "everyday" tiers built from the same underlying research, tuned to trade a small amount of peak capability for a much larger drop in price. This is not a new idea; Anthropic runs the same playbook with Opus, Sonnet and Haiku, and Google does it with Gemini Pro and Flash. What changed this week is how close the cheap tiers have gotten to flagship quality.
How Much Do Sol and Luna Actually Cost?
Sol costs $2 per million input tokens and $10 per million output tokens. Luna costs $0.10 per million input tokens and $0.50 per million output tokens. Astra, the flagship, sits far above both. Here is the breakdown that matters if you or your team touch the API directly, or use a third-party tool built on it:
--- GPT-6 Luna: $0.10 input / $0.50 output per million tokens, built for quick clerical tasks
--- GPT-6 Sol: $2 input / $10 output per million tokens, built for coding and complex reasoning
--- GPT-6 Astra: roughly 2.5 times Sol's price and about 47 times Luna's at fixed usage, reserved for the highest-stakes work
Even if you never touch an API key yourself, this matters. Many of the writing, research and automation tools practitioners already use are built on OpenAI's API underneath, so a price cut like this usually shows up as a "fast" versus "pro" toggle inside tools you already pay for, or as those tools quietly getting cheaper.
Put the numbers into a real scenario. A marketer running 200 short social captions a day through Luna, each roughly 300 input tokens and 150 output tokens, spends well under a dollar a day at Luna's pricing. Run the same volume through Astra and the bill climbs into a very different range, for output the marketer likely could not tell apart in a blind test. That gap is the entire argument for learning to route tasks deliberately instead of defaulting to whichever tier feels safest.
Sol vs Luna vs Astra: Which One Should You Actually Use?
Use Luna for quick, high-volume, low-stakes tasks: sorting a spreadsheet of leads, drafting a one-line email subject, pulling a date out of a contract. Use Sol for genuinely complex or professional work, coding included, where you need near-flagship judgment without flagship pricing. Reserve Astra for the handful of tasks where being wrong is expensive and cost is genuinely secondary.
The reliability numbers back this up. Luna's hallucination rate on OpenAI's toughest internal test dropped from 93% to 77%, a real improvement but still noticeably higher than Sol's rate on the same benchmark, according to reporting from The New Stack. That gap is exactly why Luna belongs on low-stakes, high-volume work and Sol belongs on anything where a wrong answer would actually cost you time to catch and fix.
One counterintuitive point worth remembering: token price alone does not tell you the real cost of a task. Astra's higher per-token price sometimes produces a lower total cost per successful outcome, because it needs fewer retries to get a hard task right the first time. Measure cost per finished task on your own work, not just the sticker price per million tokens.
Concretely, a content operations lead in Hong Kong running a small team might route a day's work like this: Luna handles the morning inbox triage and pulls key dates out of vendor contracts; Sol drafts the first version of a client report and writes a script to reformat a messy spreadsheet; Astra only gets called in for the one client-facing strategy document where a factual slip would be genuinely expensive to fix after the fact. None of that requires new tools, only a different default.
How to Build a Model-Routing Habit in Your Own Workflow
A model-routing habit means classifying each task by stakes and complexity before you open a chat window, then deliberately sending it to the cheapest model that can still get it right, instead of defaulting to whichever tab happens to be open. This is the single highest-leverage habit change for anyone running dozens of AI tasks a day.
The easiest way to build this habit is to make an AI assistant do the classifying for you. Paste this into any model you have access to, then describe your actual task underneath it:
Try This Prompt:
You are my AI task router. I will describe a task. Classify it using this exact format:
TASK: [one-line description]
STAKES: [low / medium / high, how costly is a wrong or sloppy answer?]
COMPLEXITY: [low / medium / high, does it need multi-step reasoning, coding, or judgment calls?]
RECOMMENDED TIER: [fast/cheap model / mid-tier model / flagship model]
REASON: [one sentence]
Here is my task: [describe your actual task here]
This works no matter which vendor's tiers you use: Sol, Luna and Astra from OpenAI, Haiku, Sonnet and Opus from Anthropic, or Gemini's Flash and Pro tiers. The classification logic is identical even though the model names change every few months.
Common Mistakes When Switching to a Cheaper Model
The most common mistake is routing everything to the cheapest tier and only noticing the quality drop after a client or manager flags a sloppy output. Cheaper is not free of consequence; it is a trade you are making deliberately, one task at a time.
The second mistake is assuming the sticker price per million tokens is the whole story. A cheaper model that needs three retries to get a task right can end up costing more in your own time than a pricier model that gets it right the first time.
The third mistake is switching an entire existing workflow to a new model without re-testing your actual prompts first. Prompts tuned for one model's quirks do not always transfer cleanly, even between models from the same company.
The fourth mistake is treating this as a one-time decision. Model pricing and reliability shift every few months, as this week's release proves, so a routing choice that made sense in July can be outdated by September. Revisit your routing habit roughly once a quarter, not once and never again.
Try It Now
Pull up your last five AI tasks from today. Run each one through the router prompt above. Count how many were actually high-stakes, complex work that justified a flagship model, and how many could have gone to a cheaper tier without you noticing the difference. Most practitioners find the answer is fewer than they expected.
If you manage a small team, turn this into a five-minute exercise at your next check-in rather than something you do alone. Ask each person to classify their own last three tasks the same way. The gaps between how different people route similar work are usually where the easiest cost savings and quality fixes are hiding.
The Real Takeaway
This week's release is also a reminder that "best model" is the wrong question to be asking day to day. The right question is "best model for this specific task, at this specific stakes level, today", and that question has a different answer for a quick email reply than it does for a client-facing report. Model releases like Sol and Luna will keep happening every few months, and the names will keep changing. The skill that does not expire is knowing how to classify your own tasks by stakes and complexity, and routing them accordingly, instead of defaulting to habit. That is the difference between using AI and running an efficient AI workflow.
We understand AI. We understand you better. With UD by your side, AI doesn't feel cold, whether that means picking the right model for a task or building the right AI team around your business.
Ready to Put the Right AI to Work?
Model routing is one small habit. Building a full AI workforce that knows which task goes where is a bigger project, and UD's AI Employee Hub is built for exactly that. We'll walk you through every step, from choosing the right roles to deploying them safely inside your own workflow.
Reviewed by the UD AI team.