Stripe has reportedly entered exclusive talks to acquire AI model aggregator OpenRouter in a cash-and-stock deal valued at approximately $10 billion. This potential valuation marks a rapid surge from OpenRouter's $1.3 billion valuation in May 2026.
Deal Context and Strategic Fit
Company Roles: OpenRouter acts as an intermediary marketplace, letting developers compare, access, and switch among hundreds of AI models. [1]
Existing Ties: OpenRouter already uses Stripe to handle customer billing and transactions. [1]
Broader Strategy: The acquisition would expand Stripe beyond basic payment processing into core AI infrastructure and token-based economic metering. Other firms, like Databricks, also explored bids before Stripe entered exclusive negotiations. [1, 2, 3, 4]
OpenRouter founders Alex Atallah & Louis Vichy - Nguyễn Hoàng Hải (2023)
Photo Source: Nguyễn Văn Vinh
OpenRouter was co-founded by Alex Atallah (who serves as CEO), along with co-founders including Chris Clark and Louis Vichy (also referred to as Louis Vilgo). [1, 2, 3]
Alex Atallah is widely recognized as the primary co-founder and face of the company, having previously co-founded the NFT marketplace OpenSea. OpenRouter was launched in early 2023 as a unified API gateway and marketplace for large language models (LLMs). [1, 2, 3, 4]
Stripe recently entered exclusive talks to buy OpenRouter in a cash-and-stock deal that would value the startup for close to $10 billion, according to people with knowledge of the discussion.
This step means OpenRouter has effectively taken itself off the market while the companies negotiate price and other details to finalize a transaction. Terms of the deal could still change or it could fall apart. It’s possible that another bidder could still emerge if the exclusivity window expires.
OpenRouter, last valued at $1.3 billion, has been working with an investment bank to evaluate its options, the people said. Other big tech companies had been considering potential deals for OpenRouter, The Information reported.
A deal for OpenRouter could help Stripe, which processes billions of transactions mainly for business customers, direct customers to the cheapest model and the ones best suited for specific tasks. OpenRouter uses Stripe to process payments and for invoicing, tax and other services.
OpenRouter, which helps app developers access hundreds of AI models, was recently generating about $140 million in annualized revenue. If a deal goes through, Stripe would be paying a substantial premium of about 70 times its recent annualized revenue.
A spokesperson for Stripe said the company does not comment on rumors or speculation. The founders of OpenRouter didn’t respond to a request for comment.
```
"Cash-and-stock deal" (thỏa thuận kết hợp tiền mặt và cổ phiếu) là hình thức thanh toán trong mua bán, sáp nhập doanh nghiệp (M&A). Bên mua trả cho cổ đông công ty bị mua lại một phần bằng tiền mặt và một phần bằng cổ phiếu của công ty mua. [1, 2]
Đặc điểm chính
Thanh toán hỗn hợp: Cổ đông nhận tiền mặt ngay lập tức kết hợp quyền sở hữu cổ phần mới.
Chia sẻ rủi ro: Giảm áp lực tiền mặt cho bên mua và giúp bên bán hưởng lợi thế tăng trưởng tương lai từ cổ phiếu. [1, 2]
Ưu và nhược điểm
Lợi ích:
Bên mua không cạn kiệt nguồn tiền mặt dự trữ.
Bên bán có thể hoãn thuế thặng dư vốn đối với phần cổ phiếu. [1]
Rủi ro:
Pha loãng cổ phiếu của công ty mua.
Giá trị phần thanh toán bằng cổ phiếu có thể biến động theo thị trường. [1]
```
---
From The Fintech Society (Facebook group)
Stripe to acquire OpenRouter for $10 billion
Titles are generated by AI from Meta
Stripe’s Ten Billion Dollar Pivot
The Wall Street Journal and The Information reported that Stripe is in advanced discussions to acquire the AI marketplace OpenRouter for 10 billion dollars.
This is not a standard fintech rollup—it is a masterstroke in infrastructure dominance. While the market is distracted by Stripe’s separate 53 billion dollar attempt to buy PayPal, their true target is owning the tollbooth for machine learning.
OpenRouter is a scaling platform enabling developers to route traffic across more than 400 AI models using a single API call. The platform currently processes over 200 trillion tokens per month and serves upwards of 10 million users.
Just two months ago, PitchBook reported that OpenRouter secured a 1.3 billion valuation backed by Menlo Ventures and Alphabet’s CapitalG.
If this acquisition closes, Stripe is paying a massive 7.7x markup to absorb this architecture before competitors can react.
Stripe, currently commanding a 159 billion valuation, already handles OpenRouter's backend billing. This operational relationship provides the fintech giant with the perfect inside track to view the raw velocity of AI developer adoption.
Media consensus indicates an agreement could be announced shortly, though The Wall Street Journal notes that negotiations remain fluid and could still attract competing tech conglomerates.
Market sentiment is split between two major Stripe initiatives. Analysts remain highly skeptical of Stripe’s unsolicited joint bid with Advent International for PayPal, which was quickly rebuffed. Conversely, consensus across technology publications is intensely bullish on Stripe pivoting into the AI distribution layer.
The Infrastructure Shift
Over the next six months, expect a radical evolution in how API consumption is monetized. By acquiring OpenRouter, Stripe is moving beyond basic payment processing to position itself as the default utility provider for software development.
This directly solves the rising enterprise fear of vendor lock-in. Developers are actively seeking ways to diversify their model usage away from deep dependence on dominant providers like OpenAI or Anthropic.
The second-order effect of this deal is profound. By controlling the access point to hundreds of models, Stripe ensures that the AI models themselves become interchangeable commodities.The true profit margin will ultimately be captured by the platform that meters the usage, controls the routing, and processes the final billing.
The race to dominate artificial intelligence has officially moved from raw compute power to payment routing.Are we seeing the beginning of a new era where fintech and AI infrastructure merge into a single utility layer?
I track and break down breaking technology research and market shifts daily. Hit follow so you do not miss the next executive analysis.
---
OpenRouter does not have an official, mainstream page on the main English Wikipedia, but it is documented on alternative community wikis like Miraheze. It functions as a unified API gateway and marketplace for large language models. [1]
Core Features
Single Endpoint: Access hundreds of AI models from providers like OpenAI, Anthropic, Google, and Meta using one API key.
No Vendor Lock-in: Switch between different models easily by changing a single parameter in your code.
Cost Optimization: Route prompts dynamically to find the most affordable or reliable provider. [1, 2, 3, 4, 5]
---
Yahoo Finance
What's behind Stripe's OpenRouter move
Lucinda Shen
2 min read
Payments company Stripe is in talks to acquire OpenRouter for around $10 billion, Wall Street Journal reported Thursday.
Why it matters: It says a lot about Stripe's ambitions as it seeks to represent the "GDP of the internet."
Driving the news: OpenRouter allows companies to switch between different AI models, a high-demand product at a time when companies are looking to control spending.
It's not considered a fintech or payments company — but it is building a network, which is key for payments companies.
OpenRouter also has the potential to represent AI spending, given its role as a gateway to multiple AI models.
Between the lines: In some ways, it's revenue model is similar to a payments firm. It charges a percentage fee for on top of the underlying model's cost. CEO Alex Atallah has notably compared his company to Stripe in the past.
Ramp, the expense management company last valued at $44 billion, is also developing a routing product, as token costs become one of the biggest concerns for companies.
Increasingly, tokens are becoming the new currency.
What they're saying: "As tokens become increasingly fungible with money, streaming payments in real time is an important part of Stripe's economic infrastructure for AI," a Stripe announcement earlier this year read.
State of play: OpenRouter has a wealth of suitors, from what we hear and from media reports.
It was valued at $1.3 billion earlier this year, which could make a deal arapid boon for investors include CapitalG and Menlo Ventures.
The bottom line: If token usage becomes much more spread out, these companies in the middle could be the ultimate (and less volatile) winners.
Cursor launched a routing product, too, earlier this week. Databricks also has such capabilities.
LỜI ĐỀ NGHỊ 10 TỶ USD DÀNH CHO MỘT STARTUP AI KHÔNG MÔ HÌNH, KHÔNG CHIP, KHÔNG DATA CENTER
Stripe đang muốn mua lại OpenRouter với giá khoảng 10 tỷ USD!
OpenRouter chỉ làm duy nhất một việc: Anh chị gửi cho nó một prompt, nó sẽ chọn ra mô hình rẻ nhất nhưng đủ tốt trong số hơn 400 mô hình và gửi yêu cầu đó đi để xử lý. Nó giúp các doanh nghiệp tiết kiệm một khoản ngân sách AI khổng lồ.
Nhưng chỉ mới 2 tháng trước, OpenRouter chỉ được định giá 1,3 tỷ USD. Tại sao giá trị của một bộ định tuyến mô hình AI lại tăng gấp 8 lần chỉ trong 60 ngày?
Đó chính là câu chuyện của tuần này. Và tôi tin rằng đây sẽ là một tin tức tuyệt vời cho Việt Nam.
THỊ TRƯỜNG AI ĐÃ TÁCH LÀM HAI: MÃ NGUỒN MỞ VÀ TIÊN PHONG
Các mô hình tiên phong mới nhất từ OpenAI và Anthropic đảm nhận những việc khó nhằn nhất: nghiên cứu, y tế và những bài toán chưa ai giải quyết được. Nhưng giá thành cao, năng lực lại có hạn.
Các mô hình mã nguồn mở sẽ lo "tất cả những việc còn lại" chính là phần lớn khối lượng công việc tự động hóa thực tế trong một công ty. Chúng được bán như điện năng: chi phí tính toán cộng thêm một biên lợi nhuận nhỏ, nhưng bán với khối lượng cực kỳ khủng.
Cả hai đều đem lại lợi nhuận cao với hai cách hoàn toàn khác biệt. Hãy nghĩ đến Apple và Android: Android có nhiều người dùng hơn và thu lợi nhuận qua số lượng lớn, trong khi Apple lấy lợi nhuận từ phân khúc cao cấp. Cả hai đều chiến thắng theo những cách khác nhau.
Những con số sẽ làm rõ điều này: 1 triệu token đầu ra tốn khoảng 50 USD trên mô hình đóng hàng đầu, nhưng chỉ tốn 0,87 USD trên DeepSeek (thấp hơn tới 98%). Đó chính xác là cách Coinbase vừa cắt giảm một nửa chi phí AI của họ bằng cách chuyển phần lớn công việc sang các mô hình mở.
Thị phần của mã nguồn mở sẽ lớn đến mức nào? Câu hỏi này đã biến thành một cuộc tranh luận công khai trong tuần qua. Jason Calacanis từ All-In Podcast (một kênh mà tôi thực sự khuyên anh chị nên nghe) đã đăng tải rằng các mô hình mở hiện đã xử lý 95% công việc của ông và sự khác biệt so với mô hình tiên phong là không đáng kể, đồng nghĩa với việc hầu như mọi thứ sẽ trở thành mã nguồn mở. Elon Musk đã phản hồi trực tiếp về vấn đề: "Thực ra đó là hai thế giới hoàn toàn khác biệt"-> những công việc khó nhất vẫn cần đến mô hình tiên phong.
Tôi nghĩ cả hai đều đúng. Đó chính là mấu chốt: Không phải là một bên giành chiến thắng. Đó là sự hình thành của hai thị trường riêng biệt.
Và mảng tiên phong vẫn tiếp tục tăng trưởng chóng mặt. Có báo cáo cho thấy doanh thu thường niên của Anthropic đã tăng từ khoảng 10 tỷ USD lên hơn 80 tỷ USD tính đến thời điểm hiện tại trong năm nay. Cùng lúc đó, cả hai phòng lab lớn đều vừa cắt giảm giá token, chính áp lực từ mã nguồn mở đang giữ cho mức giá này ở mức hợp lý.
Mã nguồn mở hoàn toàn không triệt tiêu hay thay thế các mô hình tiên phong. Nó nhiều lắm chỉ làm giảm bớt tốc độ tăng trưởng của mảng này đôi chút.
Một điều nữa tôi dự đoán: Bên cạnh hai phòng lab tiên phong lớn, chúng ta sẽ thấy sự xuất hiện của các phòng lab tiên phong ngách quy mô nhỏ, tập trung giải quyết cực kỳ tốt một bài toán hẹp. Y tế. Vật liệu. Nghiên cứu đặc thù. Không phải mọi thứ thuộc cấp độ tiên phong đều sẽ chỉ nằm trong tay hai công ty.
AI QUẢN LÝ ĐƯỢC LƯU LƯỢNG GIỮA CÁC MÔ HÌNH SẼ LÀ NGƯỜI THẮNG LỚN
Khi có hàng trăm mô hình với các mức giá chênh lệch điên rồ, phải có ai đó đứng ra quyết định cho từng yêu cầu một: Mô hình nào là đủ tốt và mô hình nào là rẻ nhất?
Công việc điều phối đó chính là thứ mà Stripe đang sẵn sàng trả 10 tỷ USD để thâu tóm.
Google cũng đã nhìn thấy bức tranh tương tự. Ngày 4 tháng 8, họ đã ra mắt bộ định tuyến mô hình của riêng mình, một API duy nhất có nhiệm vụ phân luồng từng yêu cầu tới Gemini, Claude hoặc một mô hình mở.
Sự thay đổi nhân sự cấp cao của Google cũng hoàn toàn khớp với định hướng này. Các dòng tít báo chí gọi đây là "chảy máu chất xám": Jeff Dean rời đi sau 27 năm gắn bó, Demis Hassabis chuyển sang vai trò nhà khoa học trưởng. Cách đọc vị của tôi lại khác, Google thực ra đang dịch chuyển dòng tiền từ việc tự xây dựng mô hình đắt đỏ và rủi ro sang việc vận hành hạ tầng để cho tất cả mọi người thuê, cùng với một vài mô hình cực kỳ đặc thù để hỗ trợ hệ sinh thái đó. Google Cloud đã tăng trưởng 82% trong năm ngoái.
Google không hề thua trong cuộc đua này. Họ chỉ đơn giản là chọn một làn đường mang lại nhiều lợi nhuận hơn. Tôi kỳ vọng họ sẽ trở lại vô cùng mạnh mẽ.
TRUNG QUỐC VỪA CHO ĐI MÔ HÌNH MỞ MẠNH MẼ NHẤT LỊCH SỬ. TÔI ĐÃ TẢI NÓ VỀ VÀ MẤT ĐẾN 2 ĐÊM.
Kimi K3: 2,8 triệu tỷ tham số, toàn bộ trọng số được công khai vào ngày 27 tháng 7 và đây cũng là mô hình mở lớn nhất từng được phát hành.
Lưu ý nhỏ: Tôi đã mất tới hai đêm chỉ để tải nó về. Kích thước file thực sự khổng lồ.
Lực cầu lớn đến mức Moonshot đã phải tạm dừng các gói đăng ký cho người dùng mới chỉ trong vòng vài ngày. Đạt gần một triệu lượt tải xuống trong tuần đầu tiên. Đây không còn là một cuộc thử nghiệm nữa.
Điều này có ý nghĩa gì với những người đang xây dựng sản phẩm: Trí thông minh tiệm cận mức độ tiên phong giờ đây đã miễn phí để tải về và chạy trên chính máy chủ của anh chị. Một khi nó đã nằm trên máy của anh chị, không ai có thể ngắt kết nối nó.
Vẫn có hai thứ có thể thay đổi: Các phiên bản phát hành trong tương lai có thể đi kèm với những điều kiện ràng buộc và các quy định luật pháp ở cả hai bờ Thái Bình Dương vẫn đang được đem ra tranh luận. Kết luận của tôi vẫn không thay đổi: Hãy xây dựng sản phẩm của anh chị sao cho có thể hoán đổi lớp mô hình chỉ trong vài ngày. Khi đó, không một sự thay đổi luật chơi nào ở bất kỳ đâu có thể cản bước anh chị.
ĐIỀU NÀY LÀ TIN TỐT CHO VIỆT NAM
Tôi tin rằng sự phân tách này thực sự là một tin tuyệt vời cho Việt Nam.
Hầu hết các công việc tự động hóa hiện nay đều vận hành trên các mô hình tải về miễn phí và chi phí chạy cực rẻ, bao gồm cả việc chạy trên các máy chủ đặt tại Việt Nam.
Phần việc nhỏ giọt thực sự cần đến trí thông minh tiên phong, anh chị cứ đi thuê và nhớ giữ cho chúng luôn có thể hoán đổi linh hoạt.
Và bản thân một bộ định tuyến đã là một ý tưởng startup. Sẽ có người đứng ra xây dựng lớp định tuyến và kết nối (routing and harness layer) dành riêng cho các doanh nghiệp Việt Nam. Công ty đó cũng sẽ không cần sở hữu bất kỳ mô hình nào. Cứ nhìn xem Stripe đang định giá công việc đó là bao nhiêu là biết.
---
Stripe Acquires OpenRouter
Good Morning, AI Enthusiasts!
The frontier is still moving, but the more interesting story may be everything quietly being built underneath it.
M&A
Stripe Acquires OpenRouter
What's happening: Bloomberg reported Saturday that Stripe has agreed to acquire OpenRouter for more than $7 billion, roughly 50 times its annualized revenue of $140 million. OpenRouter routes model calls across more than 400 models for over ten million developers, handles 55 trillion tokens a week, and charges a 5.5% platform fee on every dollar of AI compute that passes through it. It was valued at $1.3 billion just 82 days ago.
How this hits reality: Stripe collects roughly 2.9% on a standard card payment, and its net take rate after interchange and network fees is lower. OpenRouter collects 5.5% on every dollar of AI compute, with no returns, no logistics, no physical goods. OpenRouter currently pays Stripe processing fees out of that 5.5%. Post-acquisition, the fee stays inside the house. Stripe did not overpay for a middleware company. It swapped a payment processor's take rate for a routing platform's take rate, on a transaction base that tripled in months and has no ceiling.
Key takeaway: Stripe paid 50 times revenue to upgrade from collecting 3% of e-commerce to collecting 5.5% of AI compute, and the second number grows faster and costs less to collect.
OpenRouter helps customers to select different AI models to perform different tasks, depending on their specific needs and budget. The company announced in May that it had raised a $113 million Series B, at a reported $1.3 billion valuation. (Investors include Sequoia, Andreessen Horowitz, Menlo Ventures, and Alphabet’s Capital G.)
At the time, OpenRouter CEO Alex Atallah described the company as the equivalent of Stripe for AI, because it provides customers with a single access point for different systems and prevents lock-in. The startup also claimed to have 8 million global users and to provide access to more than 400 models.
The Wall Street Journal reported last month that Stripe and OpenRouter were in acquisition talks. Now, Bloomberg said those discussions have led to a deal price of more than $7 billion.
A Stripe spokesperson told TechCrunch that the company does not comment on rumors or speculation.
---
Cecile G. Tamura's post:
Frontier labs after hearing about Ox Alpha and who the hell is giving out 100T tokens/day for free.
New stealth model: Ox Alpha
Ox Alpha is a frontier model built for efficient coding, sustained agentic work, and real-world production use.
Ox Alpha (`stealth/ox-alpha`) is an anonymous, frontier-tier reasoning model that launched on OpenRouter. It is made temporarily available for free to developers to gather real-world usage data and stress-test serving infrastructure.
Key Technical Specifications
* Context Window: 1,048,576 tokens (1M), making it suitable for feeding entire codebases, long documentation, or multi-hour agent trajectories in a single prompt.
* Output Capacity: Up to 131,072 max output tokens, designed for generating entire multi-file code structures or lengthy analytical responses without hitting completion caps.
* Multimodal Input: Native support for Text, Images, and Video inputs—notably the first stealth release on OpenRouter to explicitly feature video comprehension alongside native multimodal capabilities.
* Target Workloads: Architected as an agentic reasoning model tailored for complex code generation, long-horizon software engineering, and production-level tool call loops.
Industry Context & Speculation
Deploying stealth models—releasing unbranded models to avoid evaluation bias and gather organic production traffic—has become a standard playbook for major AI labs.
* Previous Stealth Precedents:*Previous stealth releases on OpenRouter followed a similar trajectory before their developers formally claimed them (e.g., *Hunter Alpha* became Xiaomi’s MiMo-V2-Pro, while others turned out to be GLM releases from Zhipu AI or LongCat from Meituan).
* Leading Theories: Early community benchmark testing and response formatting habits point heavily toward Ox Alpha being an unreleased, next-generation model iteration from major Chinese research labs—most prominently Zhipu AI (GLM), Xiaomi (MiMo), or MiniMax.
Note on Data Privacy: While the promotional preview lists high rate limits and $0 cost, OpenRouter's stealth terms note that prompts and completions are retained by the underlying provider for operational logging (though not used for training). Avoid routing strictly confidential or sensitive proprietary code through anonymous endpoints.
Ox Alpha appeared anonymously as a “stealth” model and quickly drew attention for its impressive coding, reasoning, multimodal capabilities, and huge 1M-token context window.
Now, independent model-forensics testing has found strong fingerprints linking it to the GLM-5 generation—including an exact tokenizer match across dozens of probes. The identity hasn’t been officially announced by Z.ai, but the evidence is getting hard to ignore.
The mystery may finally be solved.
---
Z.ai formally launched GLM-5.3-Flash, revealing that the previously previewed “Ox Alpha” model is its public identity.
Z.ai announced GLM-5.3-Flash as a natively multimodal model with a 1M-token context window, 320B total parameters / 18B active parameters, released under the MIT License, and available via weights, API, chat, coding plan, and AutoClaw.
The launch also resolved the long-running Ox Alpha mystery: multiple posters explicitly connected Ox Alpha to GLM-5.3-Flash, including SemiAnalysis, rasbt, theo, and Cline.
Artificial Analysis first published an overview with an incorrect 400k context window, then issued a correction to 1M context, aligning with Z.ai’s original announcement.
Community response was unusually strong for an open-weight release, ranging from brief shock reactions like “HOLY” to more substantive claims that the model may now be the best intelligence-per-dollar option, e.g. Artificial Analysis and zainhas.
Independent pushback emerged on at least one modality claim: skalskip92 argued the model looks weak on several vision/object detection tasks despite being “native vision.”
Official claims and launch details
Z.ai’s primary launch tweet is the factual anchor: GLM-5.3-Flash is described as:
A follow-up launch-support post from AutoClaw framed the model as suitable for vision-language understanding, code generation, and long-horizon agentic tasks and paired availability with credits/rebates, but this is mainly rollout information rather than new technical evidence: AutoClaw launch post.
Independent benchmarks and cost/performance positioning
Ties GPT-5.6 Terra and Muse Spark 1.2 at 57, but at much lower cost per task.
$0.09/task vs $0.68/task for GLM-5.3 max.
Claimed ~7.5x lower cost per task than GLM-5.3 max.
Claimed ~5.7x cheaper per task than GPT-5.6 Terra and ~4.4x cheaper than Muse Spark 1.2.
Token-efficiency and reasoning mix
Artificial Analysis notes an interesting tradeoff:
GLM-5.3-Flash used 149M output tokens to run the Intelligence Index
compared with 168M for GLM-5.3
but more than Kimi K3 (133M) and Qwen3.8 2.4T A95B (136M) at similar Intelligence Index score
134M of the 149M tokens (~90%) were reasoning tokens
This is an important nuance: the model’s economics look excellent largely because token pricing is extremely low, not because it is especially token-frugal.
Agentic/work evals from Artificial Analysis
Artificial Analysis also reports that GLM-5.3-Flash is stronger than its raw knowledge metrics might imply on agentic tasks:
GDPval-AA v2 Elo: 1770
tied within margin of error with GLM-5.3 and Grok 4.6
behind only Claude Opus 5 xhigh/max
Terminal-Bench v2.1:84.3% vs 83.9% for GLM-5.3
τ³-Banking:47.2%, trailing GLM-5.3 by 3.1 percentage points
Knowledge/hallucination stats
AA-Omniscience score:+7
Accuracy:28%
Hallucination rate:28%
Compared with GLM-5.3:
GLM-5.3 accuracy 34%
GLM-5.3 hallucination rate 30%
Compared with GPT-5.6 Terra:
Terra accuracy 47%
This suggests a recurring theme in reactions: GLM-5.3-Flash may be much stronger on practical code/agentic workflows than on broad real-world factual knowledge.
Architecture and systems details
Several technically informed reactions tried to reverse engineer or summarize what changed from GLM-5.2 / GLM-5.x.
The most detailed public architecture breakdown in the tweet set came from rasbt, who says GLM-5.3-Flash moves from GLM-5.2’s 744B-A40B backbone to 320B-A18B, and uses:
Kimi Linear-style 3:1 hybrid attention
34 KDA layers (Kimi Delta Attention)
11 MLA/DSA layers
MLA = Multi-head Latent Attention
DSA = DeepSeek Sparse Attention
DeepSeek V4-style mHC residual path
four parallel streams
plus a native vision encoder
The same tweet describes it as “super hybrid” because both major attention components are already “efficient” variants rather than a simple efficient/full-attention hybrid.
Another useful systems-oriented summary from thealexker frames the release as an efficiency story, highlighting:
compared to GLM-5.2:
~1/10 the cost
active params 32B → 18B
layers 92 → 45
hybrid linear + sparse attention
smaller average KV cache per layer
lower attention compute compounding at long contexts
claims that visual intelligence benefited from coding/RL style improvements
says the GLM-5.3 infrastructure agent co-authored parts of the work by helping with kernels, bottlenecks, and serving stack optimization
The broader context post from eliebakouch is opinionated but technically notable because it places GLM in a Chinese open-model trend:
nearly all Chinese frontier models now use linear attention
nearly all use sparse attention / indexer-compression designs
many use fancy residuals like mHC, attention residuals, gated residuals
many use Muon
That post is not a direct GLM paper summary, but it helps explain why the architecture details immediately resonated with model engineers: GLM-5.3-Flash appears to be another data point in a fast-converging efficiency-first Chinese frontier OSS design space.
Không có nhận xét nào:
Đăng nhận xét