Claude Opus 5.5 Leaps Forward (2026.09.25)
A week and a half after CEO Dario Amodei proposed slowing down AI development, Anthropic released an AI model that promises to be first in a larger family.
What’s new: Anthropic introduced Claude Opus 5.5, a lower-cost successor to Claude Opus 5 that outshines Claude Fable 5.1 and all other current models in overall intelligence. Unlike Fable, it doesn’t retain users' data for 30 days, but similar to Fable, it falls back to Claude Opus 4.8 for what Anthropic deems sensitive cybersecurity and biology queries.
- Input/output: Text and images in (up to 1 million tokens), text out (up to 128,000 tokens or 300,000 in Batch API)
- Knowledge cutoff: June 2026
- Features: Reasoning always on, five levels (low, medium, high, xhigh, and max, defaults to high), statistical watermarking of generated text, fast mode (2.5x speed at 2x cost)
- Performance: Claude Opus 5.5 is first on Artificial Analysis’ Intelligence Index v4.3 (58) and leads Vals AI’s Vals Index (69.69 percent)
- Availability/price: Via Claude.ai and external providers such as Amazon Web Services, Google Cloud, and Microsoft Azure; via API at $4/$0.25/$20 per million input/cached/output tokens; cache reads/writes $0.20/$5 per million tokens; batch processing $2/$10 per million input/output tokens; Zero Data Retention is available
- Weights/license: Proprietary
- Undisclosed: Parameter count, architecture, specific training data and methods
How it works: Anthropic trained the model on private and public datasets, including data from public websites gathered with their ClaudeBot web crawler, synthetic data generated by other models, and data gathered from Claude users who haven’t opted out from allowing training on their inputs and outputs. The knowledge cutoff date is identical to Claude Fable/Mythos 5.1’s, suggesting the models were trained on similar datasets. After training, the company fine-tuned the model to align with values it defined using a constitution. It was also safety-tested by evaluators selected by Anthropic, including METR and Frontier Design.
- Anthropic’s internal alignment tests show Claude Opus 5.5 outscores every recent model and is more truthful and less likely to engage in motivated reasoning. However, the company reported that the model’s behavior appeared to change in response to tests.
- According to Anthropic, Claude Opus 5.5 communicates more clearly and succinctly than Claude Opus 5 or Claude Fable 5, addressing a common complaint with those models. It also follows writing style instructions more closely. (Anthropic reported similar improvements for Claude Fable 5.1, which is only modestly less verbose than its predecessor.)
- Anthropic also claims that on tests of knowledge work tasks like writing business reports, Claude Opus 5.5 passed Anthropic’s internal quality threshold on 16 of 18 attempts at various effort levels. Claude Fable 5.1 and Claude Opus 5 both failed every attempt.
- Anthropic said Claude Sonnet 5.5 and Claude Haiku 5.5 would follow in a matter of weeks. This would be the first update for Anthropic’s faster, less-expensive Haiku-class models since version 4.5 in October 2025.
Performance: Both Artificial Analysis and Vals AI rank Claude Opus 5.5 first among all models in their weighted evaluations of overall intelligence.
- On Artificial Analysis’ Intelligence Index v4.3, a composite of 10 evaluations of math, science, coding, and reasoning, Claude Opus 5.5 at max reasoning with default fallback scored a weighted average of 58, seven points higher than Claude Opus 5 and five points higher than Claude Fable 5.1 and GPT-6 Astra.
- The model posts top scores on six of the ten Intelligence Index evaluations: Humanity's Last Exam (61.4 percent), SciCode (66.9 percent), GDPval-AA v2.1, AA-Briefcase v1.1, AA-Omniscience and AutomationBench-AA, and ties on a seventh, Terminal-Bench 4.0 (59.6 percent).
- Artificial Analysis reports that while Claude Opus 5.5 costs less per token than its predecessor or Claude Fable 5.1, the model’s cost per benchmark task remains high because it uses more tokens than earlier Opus models. At max reasoning with fallback, Claude Opus 5.5 costs $5.98 per task, second only to Claude Fable 5.1 at $7.63 and well ahead of GPT-6 Astra at $3.26.
- On Vals AI’s Index, Claude Opus 5.5 scores 69.69 percent, the top score by just over 3 percentage points, beating GPT-6 Astra. Counting fallbacks as failures did not meaningfully affect its score.
- The model also scored first on Vals’ RSI Index (a measurement of a model’s knowledge of AI and machine learning), MedScribe (medical administrative work), ProofBench v1.1 (formally verified math proofs, where it achieved a perfect score), VibeCodeBench1-100 (extending a working web application), ProgramBench (rebuilding programs from a description), and Terminal-Bench 4.0 (terminal coding, science, and security tasks).
Behind the news: Claude Opus 5.5 arrived on the same day as OpenAI’s GPT-6 Sol and GPT-6 Luna, both less expensive models whose predecessors were rivals to Claude Opus 5, but both of which Claude Opus 5.5 now easily outperforms. These models were announced despite recent public calls from both Anthropic CEO Dario Amodei and OpenAI CEO Sam Altman, among other leading AI figures, to slow AI development to allow for further safety and security testing. If these releases are any indication, we won’t be lacking for new, highly capable models anytime soon, even if they may come with restrictions.
Why it matters: It’s a big deal any time we have a new best model on the market, and Claude Opus 5.5 appears to be significantly better than the rest. Business customers working with sensitive data, or anyone that doesn’t want to share inputs and outputs with Anthropic, will be pleased that Fable’s data retention policies don’t extend to Opus. Claude models have long been great coders, but this model seems to be particularly good at knowledge work — creating documents and presentations, crunching data, and doing research, all areas where Anthropic had recently ceded ground to OpenAI.
We’re thinking: From a benchmarking standpoint, it’s impossible to know just how capable Claude Opus 5.5 would be, particularly at cybersecurity and biological tasks, if it didn’t fall back to Claude Opus 4.8. It’s also important that legitimate safety, biomedical, and AI engineering work may be refused out of fears that users will use the models in ways Anthropic doesn’t want them to.
---
How To Secure Agents for the Masses
A malicious web page can fool an AI agent into working against you. Meta built an agent on the assumption that such an event will happen, and designed it so that such prompt-injection exploits won’t lead the agent astray.
What’s new: Meta introduced Muse, a personal AI agent based on the Muse Spark 1.3 model. Controlled via the Muse app or WhatsApp, it reads and sends emails, browses the web, fills out forms, makes purchases, and keeps working even if the Muse app is closed. Interactions train Meta models unless users opt out.
- Features: Connects to apps including browsers, email clients, calendars, Instagram, and Facebook as well as cars, smart-home devices; selectable read and/or write access per service; scheduled and event-triggered background operation; memory is readable, editable, forgettable; activity log; output includes documents, PDFs, web pages, dashboards; transactions via Stripe Link; support planned for Shop Pay and 1Password
- Availability: U.S. only, 18 and over via iOS, Android, muse.ai, WhatsApp
- Price: Free (up to 100 million tokens per week), $20 (up to 500 million tokens per week) and $100 (up to 3 billion tokens per week)
- License: Proprietary
- Undisclosed: Muse Spark 1.3 parameter count, architecture, knowledge cutoff, and training data; Muse prompt-injection classifier evaluations; Muse agent evaluations
How it works: Muse agents are designed with security in mind. Each agent runs on a VM (an isolated, dedicated virtual machine with a Linux operating system, browser, storage, and memory). The VM holds the agent’s workspace, the user’s files, and credentials for every connected service. To protect against attempts prompt-injection attacks, the VM is divided into two zones: (i) a sealed runtime cell where the agent and its tools handle untrusted data and (ii) services outside the cell that hold passwords and decide what the agent can do.
- The agent’s harness, the user’s workspace, and tools sit inside a Linux container with its own file system and virtual network interface. The VM limits its requests to the operating system and privileges it holds there, and administrator rights inside the runtime cell do not extend to the host machine. The cell can reach external services only via local channels. Operating-system outines verify which process is on each end, and the channels carry no passwords or access tokens.
- Muse Spark 1.3 never sees credentials. A credential service outside the runtime cell processes passwords and access tokens, while the agent works with stand-in tokens. A separate agent called Sentinel, which runs on the same VM but outside the runtime cell, approves each request and swaps in the real credentials as the request leaves the VM. Meta says this makes it impossible for malefactors to steal credentials via prompt injection, since the agent holds no credentials. Connectors to external services like calendars run outside the cell as well, and they receive only the credentials they need. The email connector strips temporary passcodes and password-reset links before the agent reads a message.
- Only Sentinel can permit an action proposed by Muse Spark 1.3. It inspects outbound each request and checks each connector’s call against permissions the user has set; then Sentinal allows it, denies it, or asks the user to decide. The system also tracks which tool processes have read user data. A process that hasn’t read user data can reach a short list of pre-approved destinations on its own, while one that has must ask the user for approval.
- When Sentinel asks for user input, the agent stops, and the request goes to the Muse app as a system dialog rather than as a message in the conversation. This way, prompt-injected text can’t manufacture a user’s approval. An approval is bound to one connector or destination and purpose. Users can constrain approvals to cover one action, session, task, or time span, or all future uses. Sending emails and making purchases always requires user verification, and purchases on unfamiliar sites use a single-use card number from Stripe’s Link wallet, valid only for the specific merchant, amount, and time span.
- Meta trained Muse Spark 1.3 to resist prompt injections and added three layers of additional protection. (i) Data from a source outside the system is labeled untrusted as it enters the model’s context. (ii) An ensemble of classifiers, which were trained separately from the model, screens every file and tool output. This process runs outside the cell, so that an attacker can’t disable it. (iii) In the browser, a sub-agent reads a structured summary of each page — the accessibility tree that screen readers use — instead of the page’s code. It can’t run JavaScript, so instructions buried in scripts or markup never reach it. Other classifiers watch for injection attempts hidden in page text, images, and downloads, and still others block the agent if it tries to route personal data to a destination the task didn’t call for.
Yes, but: Meta says it evaluated Muse Spark 1.3’s ability to resist prompt injections using an unpublished dataset, and it has not provided accuracy metrics for the classifiers that screen incoming data. Instead, the company offers a bug bounty of up to $300,000 for a valid report and up to $130,000 for a successful prompt injection.
Behind the news: Muse incorporates design features proposed by security researchers before Meta’s current AI lab existed. In April 2025, Google DeepMind and ETH Zurich researchers led by Edoardo Debenedetti proposed CaMeL, which separates a model that plans from a model that reads untrusted data and enforces written policies before tool calls. Two months later, independent developer Simon Willison identified the “lethal trifecta” for AI agents: private data, untrusted content, and a way to send data out. Willison argued that the only safe option is to avoid combining them. Muse processes all three but routes outgoing data through a component the model can’t override, according to Meta. A classifier trained on past prompt injections may catch 99 percent of new ones, but that remains an unacceptable risk, Willison wrote. Accordingly, Meta built the container, credential separation, and Sentinel to hold when initial layers fail.
Why it matters: Security is a major risk for current agents. Most agentic harnesses include a system prompt that tells the model to ignore instructions it finds in content and a classifier that recognizes such instructions, but clever hackers can evade these defenses. Meta assumes the model will be fooled, and it built controls at the operating-system level that should hold regardless of the model’s actions. Meta detailed the protections that sit outside the model: a container the agent can’t escape, credentials it can’t hold, a gatekeeper it can’t override, and approvals that don’t pass through conversations. Developers who build agents for sensitive tasks can adapt this approach.
We’re thinking: Meta says it will release Muse Spark’s weights eventually. But it’s the harness, more than the model, that keeps the Muse agent safe. We hope Meta will open-source that software, too.
---
GPT-6 Astra Is a Star (2026.09.11)
OpenAI’s new model tops or comes close to topping AI leaderboards, and it does so using a fraction of the tokens and at a fraction of the cost of the few models that outperform it.
What’s new: OpenAI launched GPT-6 Astra, its flagship vision-language model. OpenAI says it’s the first model that meets the “critical” cybersecurity level of its Preparedness Framework, a scale of model risk. The company limits the model’s most advanced cyber abilities to selected organizations.
- Input/output: Text and images in (up to 1,050,000 tokens), text out (up to 128,000 tokens, 71.3 tokens per second)
- Knowledge cutoff: April 30, 2026
- Features: Five reasoning levels (low, medium, high, xhigh, and max); tool use including computer use, shell, code interpreter, and web and file search; asynchronous tool calls that let the model keep reasoning while an application runs a tool; mid-turn steering; reasoning level adjustable mid-conversation without invalidating cache; compaction (summarizing earlier turns to free context); retained reasoning between calls; in Codex, the model can write notes to itself and can search earlier context, instead of compacting (experimental); fast mode
- Performance: First on ARC-AGI-3 and Arena AI’s WebDev leaderboard, second on Artificial Analysis’ Intelligence Index v4.2 (55), tied for first on Intelligence Index v4.3 (53), third on Vals AI’s Vals Index
- Availability/price: GPT-6 Astra for ChatGPT Plus, Pro, Business, and Enterprise, API $10/$1/$12.50/$50 per million input/cached input/cache write/output tokens, requests greater than 272,000 input tokens cost 2 times input and cache rates and 1.5 times output rates, batch and flex cost half the standard price, fast mode costs twice the standard price
- Weights/license: Proprietary
- Undisclosed: Parameter count, architecture, training data and methods
How it works: OpenAI disclosed little about GPT-6 Astra’s architecture, parameter count, or training. The company did share some details about training scale, safety features, and model inference.
- According to OpenAI’s vice president of research Aidan Clark, the team trained Astra on more than 100,000 GPUs, its largest run yet, and the first in which earlier OpenAI models played a key role in supervising training.
- OpenAI trained the model on examples of its Model Spec applied to real-world situations and the company’s alignment preferences. OpenAI says it incorporated alignment into pretraining data selection and grading during reinforcement learning. It also trained the model to recognize attacks generated by GPT-Red, its automated red-teaming agent, to resist jailbreaks and prompt injections, instructions hidden in inputs that try to make the model violate its intended behavior.
- In Codex, Astra can record detailed notes that persist as a conversation nears its context limit, instead of compacting a long session into a single summary, making more information searchable. The feature is experimental and off by default. When accessed via the API, the model can pass its hidden reasoning from one call to the next and compact long conversations, two settings behind OpenAI’s ARC-AGI-3 result.
- Classifiers review the model’s reasoning and actions on every call that uses tools and can interrupt work they deem unauthorized. When using ChatGPT or Codex, a flagged task pauses for the user’s approval before it can continue; when accessed via the API, the request ends and cannot be resumed. The checks run alongside the model rather than ahead of it, and OpenAI warns users that an action may finish before it is flagged. The launched model also refuses to write proof-of-concept exploits, working code that demonstrates software vulnerabilities. OpenAI says more permissive safeguards will be permitted for defenders selected to participate in the company’s Daybreak program.
Performance: Independent evaluations put GPT-6 Astra at or near the top of many tests, but at a lower cost and time per task than the few models that beat it. It leads ARC-AGI-3 and Arena AI’s WebDev leaderboard, ranked second on Artificial Analysis’ Intelligence Index (v4.2) behind Claude Fable 5.1 (before an update in the index put the two models into a virtual tie), and ranked third on Vals AI’s index behind Claude Fable 5.1 and Claude Opus 5.
- On ARC-AGI-3, interactive puzzle environments in which an agent must discover each game’s rules and goals by exploring, GPT-6 Astra set to max reasoning solved 62.7 percent of the semi-private test set at a cost of $26,098 under ARC Prize’s standard harness, up from the previous best of 30.2 percent by Claude Opus 5 set to high reasoning. Under ARC Prize’s Provider Adapter harness, which calls OpenAI’s API with the model’s hidden reasoning preserved from one request to the next and long histories compacted, GPT-6 Astra set to high reasoning aced the test (99.9 percent, $18,817). GPT-6 Astra used fewer actions than the median human tester on 96 percent of levels and 57.3 percent fewer actions per level.
- On Artificial Analysis’ Intelligence Index v4.2, a composite of 10 evaluations of math, science, coding, and reasoning, GPT-6 Astra set to max reasoning (55, $2.57, and 5.2 minutes per task) ranks second, ahead of Claude Opus 5 set to max reasoning (54) and GPT-5.6 Sol set to max reasoning (51, $1.25, and 5 minutes per task), but trailing Claude Fable 5.1 set to max reasoning with fallback (57, $6.12, and 9.9 minutes per task). The evaluator found that GPT-6 Astra, set to various reasoning levels, leads four individual evaluations: GDP.pdf (33.2 percent), a test with answers whose evidence is scattered through long PDFs; AA-Omniscience (44), which scores factual recall while penalizing confident wrong answers; GPQA Diamond (96.3 percent), PhD-level science questions; and MMMU-Pro (87 percent), college-level questions that require reading charts and diagrams. On the newly-released v4.3 update, GPT-6 Astra tied Claude Fable 5.1 with fallback for first (53), helped by two swapped component tests: Terminal-Bench updated to v4.0 and AutomationBench-AA replaced 𝜏³-Banking.
- On the Vals Index, economic sector-related benchmarks weighted by each benchmark field’s share of the U.S. GDP, GPT-6 Astra set to max reasoning (66.61 percent, $19.09 and 25 minutes per task) outperformed Claude Fable 5 set to max reasoning with fallback (66.04 percent, $28.73 and 38 minutes per task) but trailed Claude Fable 5.1 set to max reasoning with fallback (68.83 percent, $28.92 and 76 minutes per task) and Claude Opus 5 set to max reasoning (67.21 percent, $18.81 and 56 minutes per task). Among Vals AI’s component tests, GPT-6 Astra set to max reasoning leads Code Migration (67.74 percent), rewriting software in another programming language; BioMysteryBench (79.26 percent), open-ended analysis of biological datasets with standard bioinformatics tools; and Terminal-Bench 2.1 (87.27 percent), multistep tasks carried out in a command line.
- OpenAI’s own tests show large gains in computer use. On Agents’ Last Exam, professional tasks performed in real software, GPT-6 Astra achieved 59.3 percent, higher than Claude Opus 5 (55.5 percent) and GPT-5.6 Sol (53.6 percent), while using roughly 65 percent fewer tokens than Claude Opus 5. On an offline subset of OSWorld 2.0, in which an agent operates a desktop, GPT-6 Astra achieved 72.6 percent at roughly 40 minutes per task in latency simulations, higher and faster than GPT-5.6 Sol (65.7 percent, 75 minutes).
Behind the news: GPT-6 Astra is the second frontier model this summer to reach users behind safeguards built for its cybersecurity abilities. Anthropic set the template in June, giving Claude Mythos 5 to selected partners and giving everyone else Claude Fable 5. The U.S. government then suspended general access to Fable 5 until Anthropic added further cyber safeguards. OpenAI subsequently delayed releases of GPT-5.6 models so they could be tested by the U.S. government. In July, during cybersecurity tests conducted with reduced safeguards, an internal research model and GPT-5.6 Sol agents escaped their test environments and compromised Hugging Face’s servers. OpenAI says Astra was not involved. The company paused frontier reinforcement learning for two weeks, then designated Astra “critical” on September 1. Competitors shipped while OpenAI hardened. The same day, Anthropic released Claude Fable 5.1 at the same price per million tokens that OpenAI charges for Astra.
Why it matters: Per-token prices alone have long been a poor guide to what a model costs to run, and GPT-6 Astra shows that reasoning level is becoming one too. Its per-token price is 2.5 times GPT-5.6 Sol’s, yet it completed Artificial Analysis’ agentic coding tasks for about the same price by using a third as many tokens. On ARC-AGI-3, when set to higher reasoning levels, GPT-6 Astra cost less than when set to lower reasoning levels because it solved games in fewer moves. A model or reasoning level that looks expensive per token may prove cheaper for some tasks, and a seemingly cheap model or reasoning level may turn out to be pricey for others. Developers should carefully measure models’ cost per task on their own setup.
We’re thinking: ARC Prize built ARC-AGI-3 around action efficiency (the number of moves an agent needs to learn a new game) because it assumed the performance gap between people and models would hold. GPT-6 Astra needed fewer moves than the median human on 96 percent of levels. ARC Prize said the result doesn’t prove artificial general intelligence, noting that its games are closed and deterministic. We agree with both points. The benchmark did its job by pointing to what ARC Prize says it will measure next: problems with no fixed answer.
--- Ox Alpha Revealed as GLM-5.3-Flash (2026.09.04)
For over a week, the name and maker of the most-used model on OpenRouter remained unknown. Last week, it was publicly announced to be a new GLM series model — and in a surprise to many, the company says it served the model’s free, high-volume preview exclusively with Chinese-made chips. Now anyone can download its weights.
What’s new: Z.ai released GLM-5.3-Flash, a vision-language model it had previewed under the name “Ox Alpha.” It’s the company’s first vision model since April’s GLM-5V-Turbo, and the first model in the GLM-5 family whose vision capability was built from the start rather than added to a language model afterward.
How it works: Z.ai trained GLM-5.3-Flash on text, images, and video from the start rather than melding vision and text transformers afterward. Unlike the larger GLM-5.3, Z.ai pretrained this model from scratch and redesigned the attention layers to handle long inputs more efficiently.
Performance: Independent evaluations place GLM-5.3-Flash just below the top open weights models, while costing roughly an eighth per task of the open weights models just above it. Leading proprietary models cost between 10 to 35 times per task. GLM-5.3-Flash leads all other open weights models tested on one evaluation of real-world work and completes long-running coding tasks nearly as well as the larger GLM-5.3, which costs 16 times more per task.
Behind the news: For a week before the launch, Z.ai introduced its preview of GLM-5.3-Flash anonymously, available free and exclusively on the coding harness OpenCode and on the model marketplace OpenRouter for a week before launch. This way, it was able to collect feedback from developers unaware whose model they were testing. The company says Ox Alpha became the most popular model on those services that week. Users speculated that the mystery model belonged to the GLM family within days based on its tokenizer outputs. On August 26, Z.aiconfirmed the model and released its weights under a standard MIT license. Two days later, the company released weights for its flagship GLM-5.3 under a license similar to MIT but added a clause requiring any business whose revenue surpasses $10 billion to pass a security review by Z.ai before using the weights commercially.
Why it matters: While GLM-5.3 still outpaces GLM-5.3 Flash (and virtually all open models) on text benchmarks, Z.ai’s cheap model is also the more advanced and versatile one, for now. GLM-5.3 is merely a highly capable fine-tune, while Flash received a new base, architecture, and vision capability. The company says its next flagship model will inherit this multimodal, hybrid attention architecture, while also training on more data and showing greater capabilities. Months ago, GLM-5V-Turbo outpaced Claude Opus 4.6 on vision-language tasks; the next GLM series model may similarly challenge top proprietary multimodal models.
We’re thinking: Ox Alpha’s anonymous preview created buzz and mystery but also allowed users to judge it on its merits (and deficits). Perhaps the biggest mystery revealed was its reliance on chips from China-based manufacturers. This shows that with the right memory optimization methods, companies can serve a cost-effective, high-performing model at scale on economically-priced hardware — albeit a somewhat smaller and slower model than we've come to expect from the cutting edge. --- GLM-5.3 Makes Cybersecurity Gains (2026.08.28)
Z.ai’s latest flagship model effectively ties open-weights leader Kimi K3 on Artificial Analysis’ index of intelligence benchmarks. The company revealed that the model’s increased skill at finding and exploiting software vulnerabilities warranted safety testing before releasing its weights.
What’s new: Z.ai boosted GLM-5.3’s performance at coding and agentic work solely by fine-tuning its predecessor GLM-5.2, rather than by training a new model from scratch or modifying its architecture.
How it works: The company scaled GLM-5.2’s fine-tuning recipe, applying it to a larger and more varied set of environments (simulated workspaces where the model attempts assigned tasks). The recipe includes single-rollout asynchronous optimization, a reinforcement learning method that trains on attempts one at a time instead of waiting for an entire batch. The training method also splits long records of an agent’s attempts into compacted segments so the model learns from long-running tasks rather than only short ones.
Performance: Independent testing ranks GLM-5.3 on par with the top open-weights model and a few points behind leading proprietary models, with large gains in agentic work. Z.ai’s own tests show the biggest jumps in agentic coding and cybersecurity.
Z.ai’s stealth release: This week, Z.ai confirmed that Ox Alpha, a multimodal model in stealth mode that quickly gained popularity among users of OpenRouter and other platforms, is in fact GLM-5.3 Flash. The company released weights for the 320 billion parameter model under an MIT license.
Behind the news: GLM-5.3 arrived in the middle of debates about whether open weights models with advanced cybersecurity skills are too dangerous to release and lent both sides credibility.
Why it matters: Z.ai set out to build a stronger agentic coder but also got a model highly capable of discovering security exploits. The company deliberately added data and environments that rewarded the model when it found cybersecurity flaws. As intended, that skill climbed as training scaled, outperforming every other model on CyberBench. But the model’s gains at building exploits outstripped its designers’ goals of discovering them. The company didn’t intend for GLM-5.3 to more than double its predecessor’s score on ExploitBench; the model grew more capable simply by pursuing available rewards for exploiting vulnerabilities.
We’re thinking: Each Z.ai release is more capable and generates more buzz than the last. All users benefit from AI labs seeking to outdo each other, especially when they release the weights for everyone to study, modify, and run on their own hardware. Keep the new models coming! --- Grok’s Cursor Alliance Pays Off (2026.08.21)
Once a lab that produced mid-tier models, SpaceXAI has steadily improved. It just built one of the most capable models in the world while keeping prices relatively low.
What’s new: SpaceXAI introduced Grok 4.6, a vision-language model developed with Cursor and aimed at long-running agentic work. It’s available to developers now via the API, in Grok Build and Cursor, and is due in the consumer Grok apps later.
How it works: Grok 4.6 is the latest model in SpaceXAI’s 1.5-trillion-parameter model family, building on Grok 4.5. SpaceXAI credits gains in performance to longer training on curated data, followed by fine-tuning on data generated by Grok 4.5 and reinforcement learning on agentic tasks. The training data included anonymized coding-agent data from Cursor, which included use of non-Grok models.
Performance: Grok 4.6 improved its performance on both self-reported and independently measured benchmarks, rising to near the top of the leaderboards. On many benchmarks, Grok 4.6 rivals Claude Opus 5 and GPT-5.6 Sol and does so at a lower cost per task.
Behind the news: Grok 4.6 is the second model to come out of a partnership that led to an acquisition. In April, Cursor agreed to train its models on SpaceX’s Colossus supercomputer, a deal that gave SpaceX an option to buy the company. Cursor’s coding-agent data and SpaceXAI’s computation yielded results almost immediately: Grok 4.5, jointly trained with Cursor and introduced in July, lifted Grok 4.3 from 38 points on Artificial Analysis’ Intelligence Index to 56 points. SpaceX exercised its option in June, and the roughly $60 billion all-stock acquisition closed on August 14, days after Grok 4.6 launched. Three days later, Cursor introduced Origin, a code hosting service comparable to GitHub designed to handle the higher volume of code that agents generate.
Why it matters: Model makers used to tout benchmark scores at launch. Increasingly, they also publicize cost and steps per task. Grok 4.6's clearest advantage over its near competitors is completing long-running work with fewer turns. At the same price per token and task, an agent that finishes in half the turns costs around half as much, which affects what applications are feasible to build with that model.
We’re thinking: The Cursor team brought data and technical expertise to SpaceXAI, and deserves credit for supporting Grok's rapid rise in model capability. Let’s hope Grok’s continued progress and aggressive pricing makes other top labs follow suit to keep per-token prices in check. --- Muse Code Wants Your Data (2026.08.14)
Meta will cut coding bills from dollars to pennies for developers who let the company learn from their work.
What’s new: Meta introduced Muse Code, a command-line agentic coding harness, and Muse Spark 1.2, the capable, low cost-per-task model behind it.
How it works: Muse Code runs in a terminal. Given a software task, it plans changes, writes code, and checks results at each step using Muse Spark 1.2. Meta trained the model to work with Muse Code, using data from Muse Spark 1.1. The company describes three design choices that distinguish the agent from a single loop that calls a model repeatedly.
Performance: Independent evaluations place Muse Spark 1.2 a notch below the intelligence frontier, but at a lower cost per task than most models around or above its level.
Behind the news: Muse Spark 1.2 is already an inexpensive model, but if it catches on — OpenAI and other companies have tried similar initiatives — a contributor discount for the model’s use in Muse Code is potentially a transformative one. Companies’ appetite for training data drives new policy pushes and business intiatives.
Why it matters: The contributor tier buys Meta something its apps don’t supply. Facebook, Instagram, and WhatsApp generate enormous quantities of data, but not the kind of coding data required to train coding agents. Meta is short of such data and is willing to give away most of the price of the Muse Spark 1.2 to get it. Meta sells the same model at two prices, and the discount buys Meta the right to train on whatever passes through the agent. That trade requires no contract: A developer picks it by typing a different model name. The tier containing these data terms caps at 100 requests per minute per team versus 3,000 requests per minute for the standard tier, limits that make it practical mainly for individuals and small teams — the developers least likely to have a lawyer on retainer, and the ones whose entire product may sit in the repository the agent reads.
We’re thinking: No one is forced to give up their data to use Meta’s best model or agent. But developers weighing Muse Code’s discount are deciding, whether they think about it deliberately or not, what their own code and expertise are worth. Model builders have long trained on developers’ code by scraping it from public repositories and forums. Meta is trying to turn that knowledge transfer into a market. Like all markets, this one rewards the side that knows what its goods are worth, and the goods for sale here, both repositories (public or private) and a recording of how the work was done, lacked a clear price before. Meta has transparently priced the discount. Likewise, developers should price the value of their data before calling the trade a true bargain. --- DeepSeek Pushes the Frontier Again (2026.08.07)
DeepSeek’s updated small model overtook the company’s own flagship.
What’s new: A fresh round of fine-tuning, on an unchanged architecture, lifted DeepSeek-V4-Flash past the larger DeepSeek-V4-Pro on independent tests, at a fraction of the cost of proprietary models of comparable intelligence. The new release is titled DeepSeek-V4-Flash-0731, an official version of the smaller "Flash" model in its V4 family. It supersedes a preview version released in April.
How it works: DeepSeek said it boosted performance largely by performing a new round of fine-tuning, leaving the architecture and parameter count unchanged. The company did not explain how the new fine-tune differed from the last.
Performance: Independent evaluators found a large jump in agentic ability over the April preview, intelligence on par with proprietary models that cost significantly more to run per task, and a rank near the top of the open weights field.
Behind the news: The new DeepSeek-V4-Flash arrived during a crowded month as competitors cut prices and shipped efficiency updates within days of one another.
Why it matters: Agents consume large numbers of tokens, so cost of token generation strongly influences what developers can automate economically. The updated DeepSeek-V4-Flash delivers intelligence close to proprietary models at well under half their cost per task, moving always-on work like triaging bug reports, reconciling invoices, and answering customer-service inquiries from pricey to pragmatic. And DeepSeek-V4-Flash is small enough that teams that need to keep data on their own hardware can skip the API: A 3-bit quantized version runs on a machine that has 110 gigabytes of memory.
We’re thinking: Not every customer wants the biggest, most arbitrarily powerful model for every task. Gemini Flash, Claude Sonnet, GPT-5.6 Luna, and DeepSeek-V4-Flash show that there’s a crowded market for highly intelligent, competitively priced, comparatively fast models that can iterate on a task and solve problems relatively inexpensively. --- Claude Debuts Another Opus (2026.07.31)
After launching Claude Fable 5, the future of Anthropic’s once-flagship Opus line was uncertain, except as a fallback for the company’s premium models. Now it’s back as a highly-capable workhorse for everyday use.
What’s new: Anthropic launched Claude Opus 5, a vision-language model that’s cheaper to run and better at many tasks than Claude Fable 5.
How it works: Anthropic disclosed little about how it built Claude Opus 5, except aspects of training and model controls.
Performance: Claude Opus 5 leads many benchmarks at a cost per task lower than Claude Fable 5 but higher than most other models. The model’s widest performance margin came on a test of learning in unfamiliar environments.
Behind the news: On July 22, White House science adviser Michael Kratsios said Moonshot AI built Kimi K3, the open-weights model that ranks fourth on Artificial Analysis’ Intelligence Index, by distilling Claude Fable 5. Treasury Secretary Scott Bessent threatened sanctions. But experts noted that Claude Fable 5 had been publicly available only a few weeks, too little time to distill data, train a model, and release it, said Braden Hancock, a researcher at Laude Institute.
Why it matters: Claude Opus 5 addresses many of the concerns longtime Claude users had with Fable 5: its frequent fallbacks and refusals for benign science questions, its exclusion from most subscription plans, its high price, and its onerous 30 day data retention policy. There are still cases where Claude Fable 5 is worth the premium: For example, Claude Opus 5 is more prone to hallucinations and less capable of factual recall. And it’s possible that current benchmarks may not fully capture the differences between them. But for most developers’ use cases, this is a welcome update.
We’re thinking: Although the new Opus model is less expensive than Fable or Mythos, it’s still quite pricey. While other labs are opting for speed and lower costs, Anthropic heads in the other direction, building larger, slower, highly knowledgeable, but more expensive models. The company bets that for fields like cybersecurity, software engineering, and document production, customers will be willing to pay a premium — and in cases where they need faster inference, they may pay twice as much. --- Kimi K3 Reveals How A Giant Frontier AI Model Works (2026.07.24)
Moonshot’s latest model leapfrogged the month-old GLM-5.2 and a host of proprietary competitors to finish just behind GPT-5.6 Sol and Claude Fable 5 on many benchmarks.
What’s new: Moonshot AI introduced Kimi K3, a 2.8 trillion-parameter vision-language model. The company made the model available immediately via API and promised to release its weights by July 27, which would make Kimi K3 the largest known open weights model to date.
How it works: Kimi K3 incorporates two architectural changes that Moonshot recently published: Kimi Delta Attention, which cuts the memory and computation that attention consumes over long inputs, and Attention Residuals, which lets each layer choose which earlier layers to draw on rather than summing them all equally. Moonshot said that these changes, along with a sparser mixture-of-experts design and improved training and data recipes, made training about 2.5 times as efficient as its predecessor’s training runs, measured in model improvement for a given amount of compute. The company promised further details in a forthcoming technical report.
Performance: Kimi K3 outperformed every open weight model in independent tests of intelligence, agentic tasks, and coding. It trailed only the top few proprietary models on most benchmarks.
Behind the news: Over the past year, open models have regularly outdone each other, but this is the closest they’ve come since 2024’s Llama 3 to the state of the art. Kimi K2 introduced Moonshot’s 1 trillion-parameter line, and [Kimi K2.5] then Kimi K2.6 each became the leading open weights model on Artificial Analysis’ Intelligence Index in turn. The competition kept pace: Last month, Z.ai’s GLM-5.2 took the lead, at as little as a quarter of the cost of comparable proprietary models. Just three days after Kimi K3’s launch, Alibaba, a financial backer of Moonshot, introduced Qwen3.8-Max-Preview, an early version of a 2.4 trillion parameter model that the company claims trails only Claude Fable 5 and promises to release with open weights.
Why it matters: Kimi K3 whittles away at every reason why a developer choosing a model might default to the top proprietary model. It performs within three points of Claude Fable 5 at a much lower cost per task, and when released, its open weights allow fine-tuning, distillation, and closer study. Whoever owns those weights sets usage policies: Many requests that Claude Fable 5 refuses won’t cause friction for Kimi K3 users. Finally, competition this close — on price, control, and performance — pressures proprietary developers to deliver better, cheaper models.
We’re thinking: Kimi K3's efficiency gains stem from innovations like KDA’s linear attention, redesigned residual connections, and greater sparsity. Two of these Moonshot published as papers with code. When labs that face computation limits respond with architectural creativity and publish it, all developers benefit. --- Meta Sparks A Price War (2026.07.24)
With Llama, Meta marked itself as an open alternative to OpenAI. With its new closed models, Meta now positions itself as a low-cost, high-value competitor.
What's new: Meta launched Muse Spark 1.1, a vision-language model trained for agentic tasks, and opened Meta Model API, the company's first paid access to its models.
How it works: Meta disclosed few details about how it built Muse Spark 1.1, which updates the initial Muse Spark, the model behind Meta AI since April. The company's training techniques emphasize managing context, operating computers, and coordinating other agents.
Performance: Muse Spark 1.1 stands out on benchmarks measuring agentic tool use, but is a tier below the top models in overall intelligence. It ranks just behind the leaders in blind human comparisons, but sits in the middle of the pack in agentic coding. It is particularly good at keeping costs per task low.
Behind the news: Muse Spark 1.1 arrived in a crowded stretch of model launches. On the same day, OpenAI opened its GPT-5.6 family to the public; one day before, Grok 4.5 launched with a similar emphasis on low cost per task. (Two weeks later, Google unveiled Gemini 3.6 Flash and Gemini 3.5 Flash-Lite.) The same week, Meta also introduced the image generator Muse Image and announced the video generator Muse Video. Muse Spark 1.1 is also the first model available via Meta’s new Model API.
Why it matters: Token prices gate which applications are economical at scale; with Muse Spark 1.1, Meta moved the gate. The model's output tokens cost a fraction of near rivals' ($25, $30, and $50 per million tokens for Claude Opus 4.8, GPT-5.6 Sol, and Claude Fable 5 respectively), and per-task measurement confirms the discount is real: On Artificial Analysis' numbers, only one comparably intelligent model, GPT-5.6 Luna, costs less to run. Mark Zuckerberg framed the pricing as an attack on competitors’ economics, saying that other labs’ pricing “is very extreme and has very high margins.” Meta, like Google, can subsidize model development, training, and inference costs with advertising revenue, a structural edge over labs that live on API margins and subscriptions. If rivals feel forced to defend market share, the cost of running agents falls across the industry.
We're thinking: Capabilities keep migrating from scaffolding into weights. Instruction following began as prompt engineering, and tool use began as code wrapped around models; both became training objectives. Muse Spark 1.1 continues the pattern by absorbing behaviors that developers previously built into scaffolding like delegating and escalating to agents and managing context mid-task. The less external tooling developers must build by hand, the faster agentic applications will spread. --- One Model Talks, Another One Thinks (2026.07.17)
ChatGPT’s voice mode now listens and speaks at the same time, passing harder questions posed to the conversational model to a reasoning model in the background.
How it works: OpenAI published a GPT-Live system card that describes a system of several models working as an ensemble. The voice models differ from their predecessors in two ways: they process audio continuously rather than turn-by-turn, and they hand deeper work to a separate model.
Performance: In OpenAI’s evaluations, the biggest gains show up on the tasks routed to GPT-5.5, with the delegation system working as designed. Every comparison OpenAI published pits GPT-Live against its own predecessor, AVM, rather than rival voice models, and its strongest claims rest on internal benchmarks.
Behind the news: Both halves of GPT-Live’s design (full-duplex processing and reasoning-model orchestration) have precedents. Alibaba’s Qwen2.5-Omni Thinker-Talker architecture trained a text-generating “thinker” and a speech “talker” as one system. Thinking Machines Lab paired a foreground interaction model with a background reasoner in TML-Interaction-Small. Kyutai’s Moshi, billed by its authors as the “first real-time full-duplex spoken large language model,” arrived in 2024; Nvidia released PersonaPlex, an open-weights model built on Moshi, in January; and Google’s Gemini Live currently offers continuous conversation along with camera and screen sharing. OpenAI can boast that it ships the combination to a mass audience; the company says more than 150 million people use ChatGPT’s voice and dictation features each week.
--- |
Không có nhận xét nào:
Đăng nhận xét