live
10 labs · 30 models
All articles
9 min readChatPlus Team

Claude Opus 5.5, GPT-6 Sol and Grok 4.7: three releases in two days and which model to choose

In two days at the end of September, three major releases arrived: Grok 4.7 from xAI, Claude Opus 5.5 from Anthropic, and OpenAI's GPT-6 Sol and Luna. Here is a straightforward look at how they differ, where each one truly excels, and which to choose for your task. All four are already available in ChatPlus.

ModelsAnthropicOpenAIxAINews
Max mascot floats in outer space above the edge of a planet, reaching for three holographic cards with the logos and labels "Anthropic — Claude Opus 5.5", "OpenAI — GPT-6 Sol · Luna", and "xAI — Grok 4.7". Cover image for an article about new models
Listen to the article00:00

On September 21, xAI released Grok 4.7. The next day, Anthropic unveiled Claude Opus 5.5, while OpenAI released GPT-6 Sol and GPT-6 Luna. Three labs, four models, two days. It is hard to call that a coincidence: the race is so close that no one wants to leave a competitor alone in the headlines for even a week.

The short version: Opus 5.5 is the strongest of the new releases and, according to independent measurements, currently the best model on the market overall. GPT-6 Sol has barely improved in raw intelligence, but costs half as much and makes fewer factual mistakes. Luna is lightweight and very affordable. Grok 4.7 is a major step forward for xAI, but still sits in the second tier for now, albeit at an attractive price. Here are the details and the honest caveats for each.

Four models at a glance

For reference, here are the core specs. Prices are listed per million API tokens, so you can see which models are expensive; ChatPlus does not charge separately for tokens, as everything is included in the subscription.

ModelDeveloperInput / output, $ContextBest at
Claude Opus 5.5Anthropic4 / 201MLong-form coding, documents, clear writing
GPT-6 SolOpenAI2 / 101.05MCode, agents, factual accuracy
GPT-6 LunaOpenAI0.10 / 0.501.05MFast, simple tasks
Grok 4.7xAI2 / 6500KCode, engineering, law

Claude Opus 5.5: Fable-level performance without the flagship price

Opus 5.5 launches the new Claude 5.5 family. Anthropic sums up its main advantage in one sentence: the model performs at the level of Claude Fable 5.1 — the company's most capable and most expensive model — while costing 40% less than Opus 5. Token prices are down 20%, and the rest comes from efficiency: the model uses fewer steps and fewer tokens for the same task. It also generates text more than 30% faster than its predecessor.

Independent testing confirms this. In the Artificial Analysis aggregate index, Opus 5.5 scored 58 points at maximum settings — the best result ever measured there, ahead by several points. It leads almost everywhere: agentic programming (66.4% on Terminal-Bench 4.0 versus 55.8% for Fable 5.1 and 57.9% for GPT-6 Astra), office work, and complex tool-use questions.

But benchmark numbers say little about what a model feels like in real work. Early tester stories are more revealing. One migrated a project with 680,000 lines of code in less than a day, a task that would have taken an engineering team weeks. Another audited and repaired a 200,000-line codebase in three hours, while Opus 5 needed more than twenty for the same work. At another company, Opus 5.5 completed a complex task that had previously required 38 prompts and four days in 11 prompts and three hours. Anthropic also ran an internal experiment: it asked Opus 5.5 and Fable 5.1 to rewrite HAProxy, a popular load balancer, from C to Rust. Both succeeded, but Opus 5.5 finished in 9.5 hours instead of 12 and cost half as much.

The second major change is how the model writes. Opus 5 was often criticized for verbosity: long introductions, corporate phrasing, and the point buried in the third paragraph. Opus 5.5 puts the main point first, drifts into jargon less often, and follows the style rules you give it more reliably. One tester put it simply: "It writes the way I write." You notice this immediately when working with text: there is less cleanup after the model.

There is one characteristic worth knowing in advance: you cannot turn off reasoning in Opus 5.5. It always thinks first, then responds. That is a benefit on complex tasks. For a question like "what is the English word for a window sill?" it means extra seconds of waiting. For small tasks like that, a simpler model is a better choice.

On ChatPlus, Claude Opus 5.5 is already available in chat — on the Max plan, alongside other top-tier models.

GPT-6 Sol: the same bar, at half the price

At the beginning of September, OpenAI released its GPT-6 Astra flagship, and the GPT-6 family now looks like this: Astra at the top, Sol in the middle, Luna at the bottom. Sol is designed as a workhorse: it takes Astra's precision and ability to carry a task through to completion, but costs considerably less.

OpenAI is selling Sol on reliability rather than records. The company took real conversations where users had flagged model errors and tested the new version on them: Sol makes mistakes roughly half as often as GPT-5.6 Sol and reaches Astra-level reliability. An independent hallucination test confirms this, though with a caveat. The share of fabricated answers fell from 92% to 60%, largely because the model refuses to answer more often when it is uncertain. The overall share of correct answers even dropped slightly, from 59% to 54%. For everyday work, that is more good news than bad: an honest "I don't know" is better than a confident invention.

In programming, Sol keeps pace with much more expensive models. On DeepSWE, a test of long engineering tasks in real repositories, it scores 68.8%. Claude Fable 5 has the best score at 69.9%, but costs roughly five times more per task. On Terminal-Bench 4.0, Sol improved from 37% to 43%.

Now for what the press release does not mention. In the aggregate intelligence index, Sol is almost unchanged from its predecessor: 48 points versus 47. It has even regressed in office-document work, and trails both Astra and the previous Sol in computer control. So this is not a leap forward, but the same capability bar at half the price. For most tasks, that is exactly what you need: coding, research, correspondence, and analytics. If you need a model that clicks through interfaces on its own, choose Astra.

GPT-6 Sol is available on our PRO subscription, and because it costs half as much as the previous Sol, your allowance goes noticeably further in long conversations. Open a chat with GPT-6 Sol.

GPT-6 Luna: the smaller model with the freshest knowledge

OpenAI describes Luna simply: a model for high-volume tasks with a clear goal. Summarize a document, extract numbers and names from text, answer a question quickly. There is nothing modest about that: these tasks make up most of what people do with AI every day.

Its price is almost symbolic: ten cents per million input tokens and fifty cents per million output tokens. That is half the input price of GPT-5.6 Luna and almost two and a half times cheaper on output. In the aggregate intelligence index, the model remains at the level of its predecessor. It slipped slightly in agentic coding, but on DeepSWE at maximum settings it delivers 66.6% for just 22 cents per task. According to OpenAI, that is comparable to Claude Opus 5 and Fable 5 at medium settings, while costing 93–96% less.

One interesting detail: the smallest model in the family has the freshest knowledge. Luna was trained on data through May 2026, while Sol and Astra stop at April.

In ChatPlus, GPT-6 Luna is available on the PRO subscription and does not use your allowance — the 5-hour limit from which requests to expensive models are deducted. You can use it as much as you like without watching the counter. It is not the best choice for complex multi-file code, but you can switch to Sol or Opus in the same conversation.

Grok 4.7: large, late, and unexpectedly affordable

Grok 4.7 arrived before the others, despite having been anticipated the longest: Elon Musk postponed the release at least five times, starting in late July. The delay is understandable. This is not a cosmetic update but a new, substantially larger base model: according to media reports, around 2.1 trillion parameters. It received longer reinforcement learning training, with an emphasis on tasks that take humans hours. xAI also improved self-checking: the model reviews its own answers more carefully and handles long context better. Incidentally, the company now calls itself SpaceXAI in official materials, though the models are still called Grok.

The improvement over Grok 4.6 is substantial. On CursorBench 4.0, a test of long programming tasks, the score rose from 40.4% to 46.3%. On Terminal-Bench 4.0, it nearly doubled, from 20.3% to 37.6%. In multi-hour office work — reports, spreadsheets, presentations — the model gained more than a hundred Elo points. Its legal performance is particularly surprising: on the Harvey legal benchmark, Grok 4.7 scores 19.6%, almost three times Claude Fable 5.1's 6.7%. On the electrical engineering benchmark EEBench, it gets 64% versus Fable's 56.4%.

Its place in the overall rankings is more modest. In most comparisons, Grok 4.7 is second: it trails Fable 5.1 in office work, GPT-6 Astra in engineering tasks, and remains far behind Opus 5.5. Independent tests put it at 46 points on the intelligence index. That was enough for xAI to enter the top four labs, but no further. Musk himself promises "Astra and Fable class" only with Grok 4.9. One more caveat: at maximum settings, the model uses more than twice as many tokens per task as Grok 4.6, so the same token price does not mean the same cost per result.

Grok 4.7's key advantage is its price-to-quality ratio. It costs the same as Grok 4.6 — two dollars per million input tokens and six per million output tokens — and is noticeably cheaper than Anthropic and OpenAI flagships. Grok 4.7 is available on ChatPlus with the PRO subscription.

So which one should you choose?

We would break it down like this.

If the task is large and important — complex code, a project audit, an analytical report, or a long document that needs clear thinking — choose Claude Opus 5.5. It is currently the strongest available model, and it writes noticeably better than previous Claude models.

For everyday work where you need balance, GPT-6 Sol is a great fit: code, research, emails, and spreadsheets. It is careful, hallucinates less often, and uses less of your allowance.

For small tasks such as summaries, translations, short reference answers, and drafts, choose GPT-6 Luna. It is fast and does not use your allowance at all.

Grok 4.7 is worth trying for legal and engineering questions, where it is unexpectedly strong, and as a second opinion when you want to compare another model's answer.

The easiest way to compare is on your own task. In ChatPlus, you can switch models mid-conversation: ask Opus a question, switch to Sol or Grok, send the same prompt again, and immediately see who handles it best. Detailed model pages are available for Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and Grok 4.7.

In short

Four new models arrived in two days. Claude Opus 5.5 is the new leader: Fable 5.1-level capability at 40% less than Opus 5, best-in-class agentic coding, and noticeably clearer writing. GPT-6 Sol has barely improved in intelligence, but costs half as much and makes factual errors half as often. GPT-6 Luna is a very affordable model for simple tasks with the freshest knowledge in the family. Grok 4.7 made a major leap forward, especially in code, law, and engineering, but remains in the second tier for now — at an accessible price. All four are already available in ChatPlus.

Try the best AI models on ChatPlus

GPT, Claude, Gemini, GLM, and other AI models in one window.