Four new models landed within two days.
Anthropic released Claude Opus 5.5. OpenAI followed with GPT-6 Sol and GPT-6 Luna. SpaceXAI introduced Grok 4.7 just ahead of them.
Each company brought the usual launch material: benchmark wins, selected customer stories, safety claims and charts built to make its model look like the sensible choice. Taken together, though, the releases say something more useful about where the market is going.
The labs are starting to compete on the cost of completed work.
That is a better metric than price per token, and a harder one to advertise cleanly. A cheap model can take longer, retry the task or produce work that needs heavy review. A more expensive model may finish in one pass. The invoice only tells part of the story.
We have seen this measurement problem before. In Token Spend Is Not a Trophy, we wrote about companies confusing AI consumption with AI value. Model pricing creates the same trap from the other direction. Paying less for activity tells you very little about the result.
This week gave us an unusually clear view of the tradeoffs.
Four models, three different bets
Anthropic is making its premium model easier to use every day. Opus 5.5 is cheaper and faster than Opus 5, with strong early results in coding and long professional tasks. Artificial Analysis placed it at the top of its initial rankings for agentic knowledge work. The appeal is straightforward: more difficult work completed with less waste.
OpenAI is pushing much harder on price. GPT-6 Sol targets coding and professional agents, while Luna brings capable inference into a range where high-volume use becomes practical. Both cost half as much as their GPT-5.6 equivalents. Sol looks competitive with more expensive models; Luna may matter more simply because so many routine tasks can now run cheaply.
Grok 4.7 concentrates its gains in coding and professional deliverables. Independent testing shows a meaningful jump over Grok 4.6 on those tasks, although its broader benchmark improvement is modest. It also uses far more output to get there. Grok has reached the frontier through heavier reasoning, which makes the low sticker price less conclusive than it first appears.
Each company can claim a win under the right conditions. Anthropic has the strongest early case for demanding agent work. OpenAI is setting the price pressure. SpaceXAI has closed a meaningful gap. None of those conclusions produces a universal ranking.
The API bill is only one part of the cost
Benchmark percentages and API prices are useful. They also leave out retries, latency and the time a person spends checking or repairing the result.
That gap becomes more visible as agents move into longer jobs. A code migration or research assignment can run for hours, call tools and revisit earlier decisions. The cheapest model on the rate card can easily become the expensive choice once an engineer has to clean up after it.
The latest releases reflect that reality. Anthropic is reducing the work required to get a strong result. OpenAI is lowering the entry price. SpaceXAI is spending more inference where longer reasoning pays off. All three are trying to make autonomous work economically useful.
What we know so far
Launch-week certainty is cheap.
The models have been public for less than two days. Most companies have not tested them on their own workflows. Production data on reliability, latency and review time does not exist at meaningful scale yet. Claims of a universal winner are running ahead of the evidence.
The useful evaluations will come from repeatable work inside real systems: support cases resolved without escalation, pull requests accepted, reports that survive fact-checking, workflows completed without a person repairing the last step.
Those results will vary by company. A model that excels at software migrations may be the wrong choice for thousands of short commerce tasks. The best option can even change within one workflow.
This is why we keep AI systems model-flexible at REBL. The leading model changes too often, and the economics move even faster. Business logic, data and evaluation should live outside the provider. When a cheaper model reaches the required quality, switching should be a configuration decision rather than a rebuild.
We reached the same conclusion when Kimi K3 arrived at a third of the price. Two months later, the names and numbers have already changed. The architecture lesson has held up.
The release cycle this week strengthens that case. Every lab improved something meaningful. None earned the right to become permanent infrastructure.
The question worth asking
Benchmark tables will keep moving. Token prices will keep falling. Teams deciding what to use need a metric tied to the work itself.
How much does it cost to produce an outcome we are willing to ship?
Answering that requires more than an API price. Measure retries, latency and human review alongside usage. Track failures that make it into production. Use the smallest model that clears the quality bar, then keep testing because that answer will change.
Claude Opus 5.5, GPT-6 Sol and Luna, and Grok 4.7 made the frontier cheaper this week. The companies that benefit most will be the ones that know exactly what they need from it.





