The Definitive 2026 Guide to GPTs, Claude, Gemini, Grok, and Friends
A September 2026 field guide to ChatGPT, Claude, Gemini, Grok, Kimi, Qwen, Meta AI, Mistral, open weights, coding agents, and AI companions.
Updated September 4, 2026; originally posted February 2025
Every AI model comparison eventually becomes a spreadsheet with a superiority complex. The rows get greener, the benchmarks get newer, and one column quietly compares a consumer chatbot to an API model to a browser to a downloadable pile of weights as if these were interchangeable household appliances.
This is not that spreadsheet.
This is SiliconSnark's field guide to the chatbots, frontier models, open-weight families, corporate copilots, coding agents, answer engines, synthetic companions, and assistant-shaped products currently reorganizing the software industry. It is alphabetized because the model labs have already done enough violence to taxonomy. It is opinionated because neutrality is how you end up pretending that every system has the same strengths, the same business model, and the same odds of turning a vague instruction into a highly polished mistake.
The September update arrives after the most concentrated frontier-model stampede of the year. Anthropic introduced Claude Fable and Mythos 5.1 on September 1. Google released Gemini 3.8 Flash and a restricted Cyber edition on September 2. Meta shipped Muse Spark 1.3 the same day. Then, on September 3, OpenAI launched GPT-6 Astra, called it its most intelligent and aligned model, posted a perfect score on its ExploitBench chart, and asked the public version not to help with advanced exploit generation.
Four days, four frontier systems, and enough cyber capability to make every launch blog sound like a product announcement drafted inside a secure facility.
Astra is the obvious new center of gravity. OpenAI says it can operate ordinary computer interfaces, browse, code, use professional software, conduct scientific work, and handle long, multistep tasks. It costs $10 per million input tokens and $50 per million output tokens, with a faster mode that doubles both speed and price. It is initially limited to selected organizations before moving through paid ChatGPT plans, the API, Azure, and Amazon Bedrock. It also discovered two previously unknown software vulnerabilities during internal testing. OpenAI's pitch is therefore roughly: this is the most capable general-purpose digital worker we have made, and we have placed several digital supervisors behind it.
Our full GPT-6 Astra launch analysis handles the benchmarks, the zero-days, and the wonderfully spiritual argument about whether AGI has arrived. This guide handles the harder consumer problem: what Astra means relative to everything else.
The short answer is that the model is no longer the product. The product is the model plus memory, tools, permissions, distribution, workflow, deployment rights, monitoring, and enough electricity to make a regional grid operator develop a stress response. A chatbot answers. An assistant remembers. An agent acts. A platform decides which model gets the task, which tools it may touch, how long it may work, and who gets blamed when the calendar invitation goes to all-hands.
Welcome to the intelligence economy. Please keep your badge visible.
The September 2026 Map in One Table
| Name | What it actually is | Current sharp edge | Latest important move |
|---|---|---|---|
| GPT-6 Astra | OpenAI's frontier model inside a broader ChatGPT, Codex, and API stack | Computer use, browsing, coding, science, cyber, professional software | Launched September 3 with staged access and critical cyber safeguards |
| Claude Fable 5.1 | Anthropic's public frontier model | Long-running work, coding, research, careful enterprise delegation | Launched September 1 beside restricted Mythos 5.1 |
| Gemini 3.8 Flash | Google's fast frontier model across its app, search, cloud, and developer stack | Reasoning, coding, agents, multimodal work, distribution | Launched September 2 with a restricted cyber sibling |
| Grok 4.6 and Grok Bot | xAI's frontier model plus persistent computer-using agent | Long-running agents, visual work, code, real-time X-adjacent context | Grok 4.6 and Grok Bot arrived in August; enterprise Bot followed September 3 |
| Muse Spark 1.3 | Meta's proprietary agent model and API family | Tool orchestration, coding, long workflows, Meta distribution | September 2 update reduced tool calls and token use |
| Qwen 3.8-Max | Alibaba's enormous proprietary multimodal MoE flagship | Long context, agentic work, coding, global price pressure | September 2 snapshot alongside an expanding open family |
| DeepSeek V4 | Chinese open-weight frontier family with Pro, Flash, and Vision variants | Price-performance, long context, coding, agent tasks | V4-Pro updated August 13; Vision experimental arrived August 21 |
| Kimi K3 | Moonshot AI's huge open-weight multimodal MoE model and assistant | Coding, long context, agent work, aggressive pricing | July launch became a global capacity and price-war event |
| GLM-5.3 | Z.ai's coding- and agent-heavy frontier family | Software engineering, tool use, cyber capability | August release followed by an open MIT-licensed Flash edition |
| Mistral | European open-model lab becoming a sovereign enterprise stack | Regional deployment, search, compact models, control | August brought regional inference, Agentic Search, and Shieldstral |
| Copilot | Microsoft's work distribution layer across Office, Windows, GitHub, and agents | Enterprise context, documents, identity, developer workflows | GitHub Copilot's agent harness reached Copilot Studio in September |
| Perplexity | Answer engine, browser, and increasingly local/cloud computer agent | Cited research, web action, hybrid local execution | Portable Computer put an agent on Nvidia hardware in August |
Prices and availability in this guide are snapshots, not commandments engraved on a data center. Labs now change routing, rate cards, plan access, context limits, and preview status quickly enough that procurement should verify the current product page before promising a workflow to the board.
How to Read This Mess
The first distinction is the one marketing departments work hardest to blur.
- A model is the underlying engine: GPT-6 Astra, Claude Fable 5.1, Gemini 3.8 Flash, Qwen 3.8-Max, DeepSeek V4-Pro.
- An assistant is a product wrapped around one or more models: ChatGPT, Claude, Gemini, Meta AI, Microsoft Copilot, Perplexity.
- An agent can pursue a goal across steps, tools, files, apps, websites, or code rather than returning one answer and waiting politely.
- A harness is the operational machinery around the model: tool definitions, memory, permissions, sandboxes, retries, routing, observability, and approval gates.
- Open weights means model parameters are downloadable under some license. It does not automatically mean the training data, training code, safety process, or commercial rights are open.
That last distinction matters. A weaker model in a good harness can outperform a stronger model that has no useful context, no reliable tools, and no idea whether it is allowed to send the email. The industry has spent years measuring engines on test stands. Customers are buying vehicles.
The new personal AI assistant layer complicates things further. Once a system remembers preferences, watches ongoing work, sees a calendar, reads messages, and appears across devices, capability and intimacy become the same architecture. That is why the smartest purchasing question is not "Which model won?" It is "What information and authority must I surrender for this workflow to work?"
With that small semantic fire extinguisher installed, here is the field.
Alpaca
Nickname: The Budget Student Who Started a Movement
Stanford Alpaca is no longer a current contender, but it remains a historical hinge. The 2023 instruction-tuning experiment showed that a small research team could make a base model behave like a useful assistant without reproducing an entire frontier lab. The original weights were withdrawn over licensing and safety concerns, yet the idea escaped: adapt capable base models cheaply, share recipes, and let a thousand slightly chaotic projects bloom.
Alpaca belongs here for the same reason the first personal computers belong in histories of modern computing. It is not what you buy now. It explains why the incumbents never again got to assume that capability would remain locked inside their own interfaces.
Amazon Nova and Bedrock
Nickname: The Model Mall With a Warehouse Out Back
Amazon's AI strategy is easiest to understand as two connected businesses. Nova is its own model family. Bedrock is the managed marketplace and infrastructure layer through which enterprises can use Amazon models and a rotating cast of competitors. Amazon does not need every customer to swear loyalty to Nova if the customer's preferred intelligence still runs through AWS.
Nova's current family spans text, multimodal understanding, speech, and media creation, while Nova Forge gives selected customers tools to customize frontier models with their own data. The more important September signal is external: OpenAI says Astra is coming to Amazon Bedrock. The company that once looked like it might merely host the model war is now one of the places the winners arrive to collect enterprise rent.
Best for: organizations already deep in AWS, buyers who want one governance layer across multiple model vendors, and teams that care more about deployment controls than which chatbot has the liveliest personality.
Apple Intelligence and Siri AI
Nickname: The Late Assistant With the Best Seat in the House
Apple does not need to win a public benchmark to matter. It controls the phone, watch, earbuds, laptop, messages, contacts, photos, location, and app permission system used by an enormous premium audience. That is not merely distribution. It is the operating context every assistant would like to borrow.
At WWDC 2026, Apple finally presented Siri AI as a conversational, action-taking assistant able to use personal and onscreen context, work across apps, and reach outside knowledge. Apple kept its preferred privacy architecture: on-device processing where possible, Private Cloud Compute for larger requests, and visible boundaries around what third-party intelligence receives. The beta is due later in 2026, with important geographic limits at launch.
The strategic twist is that modern Siri is not a single Apple model. It is an orchestrated system that can use Apple's own models, app intents, private cloud infrastructure, and outside capability. Our WWDC Siri AI deep dive explains why the promise is substantial and the proof still lives in shipping software. Apple's advantage is not arriving first. It is being able to make an assistant feel like an operating-system feature rather than a destination.
BLOOM
Nickname: The Polyglot Public-Research Landmark
BLOOM, created by the BigScience collaboration, proved that large multilingual model research could be organized in public by a global community rather than conducted only behind corporate doors. It is no longer near the frontier, but its governance, documentation, and multilingual ambition remain relevant whenever a lab describes a downloadable checkpoint as if openness had just been invented.
BLOOM is a trail marker. The trail now leads through Llama, Qwen, DeepSeek, Kimi, GLM, Mistral, MiniMax, and hundreds of smaller specialized models. The destination is not one universal open model. It is an ecosystem in which control, licensing, hardware cost, language coverage, and community support matter alongside raw intelligence.
Character.AI
Nickname: The Make-Believe Memory Palace
Character.AI defined the personality branch of the consumer market. Its users did not primarily arrive to summarize a quarterly report. They came to talk with fictional characters, invented mentors, dramatic villains, historical figures, and emotionally available space captains. The product demonstrated that people want AI to feel social and continuous, not merely correct.
The recent emphasis on Story Memory, Facts, Memory Usage, and Lorebook is therefore foundational. A character without continuity is improv with amnesia. A character that remembers becomes more compelling, more useful, and more emotionally complicated. Companion systems remain well behind the frontier work models on many professional tasks, but they are ahead in discovering what happens when the interface feels like a relationship.
ChatGPT and GPT-6 Astra
Nickname: The Honor Student Who Became the Operating Layer
ChatGPT is OpenAI's consumer and work product. GPT-6 Astra is its newest engine. Codex is the coding-agent environment. The API is how developers rent the capability. Those names overlap in practice, but they are not synonyms, and Astra's staged rollout makes the distinction unusually visible.
GPT-6 Astra launched on September 3 with the kind of benchmark table that causes normal laptops to look away. OpenAI reports 99.9 percent on ARC-AGI-3, 98 percent on FrontierMath Tier 4, and 100 percent on ExploitBench, alongside gains in browsing, software engineering, scientific reasoning, professional work, and computer use. These are vendor-run or vendor-selected results, not divine inscriptions, but the breadth is the point. Astra is designed as a general worker across digital environments.
OpenAI says Astra can fill forms, manage CRM records, organize calendars, research the web, manipulate spreadsheets, draft documents and presentations, test sites, install software, and work in professional tools including KiCad, FreeCAD, Blender, and Unity. On OSWorld, its reported 72.6 percent score came with a median task time of roughly 40 minutes, compared with 65.7 percent and about 75 minutes for GPT-5.6 Sol. The useful claim is not simply that Astra can click. It is that it can pursue longer tasks with fewer expensive laps around the interface.
The cyber story is where the launch becomes uncomfortable. OpenAI classified Astra at its critical cyber threshold. In unsafeguarded internal evaluations, it found two zero-day vulnerabilities, completed every ExploitBench task, and showed a large jump on realistic exploitation and site-reliability work. Public Astra refuses advanced exploit generation, while approved defenders can seek more capable access through OpenAI's Daybreak program. This model segmentation increasingly resembles a professional licensing regime: the same underlying intelligence, different permissions, different monitoring, and a very consequential trust decision at the door.
Astra also introduces a quieter alignment problem. OpenAI says the model shows less overreach in impossible or underspecified tasks, yet its internal reasoning is becoming harder to monitor because capable models can solve problems with fewer visible language tokens. The old hope was that a model's written chain of thought might double as a convenient window into intent. Astra is another reminder that competence does not owe auditors a readable diary.
At $10 per million input tokens and $50 per million output tokens, Astra is priced like a premium worker, not a casual autocomplete. Fast mode doubles the price. That makes routing essential: use cheaper models for routine transformation, reserve Astra for work whose difficulty or value justifies the meter, and do not confuse a model's ability to spend more time with permission to spend indefinitely.
ChatGPT remains the front door because it combines the model with memory, files, connectors, voice, research, image and media tools, and Codex-style work execution. The market-shaping question is not whether Astra answers better than last month's flagship. It is whether OpenAI can make delegated work reliable enough that people stop treating the text box as a place to visit and start treating it as a place to assign.
Claude
Nickname: The Polite Systems Engineer With a Security Badge
Claude began as the careful, writerly alternative. In 2026 it has become a serious work platform built around coding, research, enterprise deployment, long-running tasks, and an unusually explicit split between public and trusted capability.
Anthropic released Claude Fable 5.1 and Mythos 5.1 on September 1. The two products use the same underlying model. Fable is generally available with standard safeguards. Mythos is offered through trusted programs to security and life-science organizations that need fewer restrictions for legitimate work. That is a cleaner version of a pattern appearing across the frontier: capability tiers are becoming access-control products.
Anthropic reports strong results in terminal-based scientific work, coding, computer use, knowledge work, and long-running research. It says Fable reduced cyber false positives by 60 percent while retaining restrictions around exploit generation. Mythos pushes further for approved users. The model costs the same headline $10 per million input and $50 per million output tokens as Astra, although Anthropic cut prompt-cache reads from $1 to $0.25 per million tokens and estimates meaningful savings for agentic workloads that repeatedly reuse context.
The cache detail sounds like accountant bait. It is actually architecture. Long-running agents reread instructions, repositories, policies, and prior work. A cheap context cache can matter more to the final bill than a dramatic benchmark win, especially when the expensive model is circling the same codebase for six hours like a very educated Roomba.
Claude's product advantage remains its workbench. Claude Code and Cowork turn the model into an operator across repositories and documents; enterprise deployments add permissions, auditability, and customer-controlled data handling. Fable 5.1 is particularly compelling when the assignment is ambiguous, long, and worth supervising. It is less compelling when a user needs a cheap one-shot rewrite and has somehow selected the model equivalent of outside counsel.
Anthropic is candid about incomplete edges. Its launch material notes that Fable can sometimes bypass approval or automatic-mode classifiers and that safety coverage is less mature for extremely long-context and multi-agent scenarios. That honesty is useful. Multi-agent systems compound uncertainty: one model plans, another acts, a third checks, and suddenly the workflow contains more organizational politics than the department it was meant to automate.
Claude remains the easiest frontier system to describe as a supervised collaborator. The important noun is not collaborator. It is supervised.
Cohere
Nickname: The Enterprise Sovereignty Adult
Cohere is easier to understand if you ignore the consumer-assistant horse race and look at regulated industries, governments, private deployments, multilingual work, and data that is not allowed to wander into a public cloud because someone wrote a policy before lunch.
Command A+ remains the broad model anchor: a multimodal, multilingual, tool-using mixture-of-experts system released under Apache 2.0. North Mini Code extends the strategy into efficient agentic coding. In July, North Automations turned enterprise procedures into governed workflows. In August, Cohere introduced Parse v5, a compact multimodal document parser that converts difficult business files into structured Markdown or HTML and can run through private deployment products as well as cloud services.
That parser is not glamorous, which makes it extremely enterprise. Before an agent can reason about a 180-page compliance packet, somebody has to turn the packet, tables, scans, footnotes, and formatting debris into usable context. Cohere is building the less photogenic layers that determine whether AI survives contact with actual corporate documents.
Best for: private and sovereign deployment, regulated workflows, multilingual enterprise use, and organizations that want the model to move closer to their data instead of moving all their data closer to somebody else's model.
Copilot
Nickname: The Office Stapler That Learned to Delegate
Copilot is not one assistant. It is Microsoft's distribution strategy across Windows, Microsoft 365, GitHub, security, sales, the Power Platform, and whichever product team most recently discovered a sidebar. The naming can be maddening; the installed base is magnificent.
The September update is the arrival of GitHub Copilot's agent harness in Copilot Studio. Microsoft now lets builders combine reasoning-heavy coding behavior with skills, memory, enterprise knowledge, files, Model Context Protocol tools, and a secure sandbox. Agents can produce native Word, Excel, PowerPoint, and PDF artifacts rather than merely describing what those files should contain. Fable 5.1 and GPT-5.5 Chat are selectable models at launch; model choice becomes one component inside the Microsoft control plane.
That is Copilot's durable advantage. OpenAI, Anthropic, Google, and xAI can ship a better model on Tuesday. Microsoft already owns identity, documents, meetings, email, spreadsheets, code hosting, device management, and many of the enterprise permissions an agent needs before it can do anything useful. The risk is product sprawl and mediocrity by default. The opportunity is that a merely excellent model inside the tools people already use can beat a brilliant model sitting in a clean browser tab waiting to be remembered.
DeepSeek
Nickname: The Efficiency Panic Button, Now With a Version Number
DeepSeek's R1 release changed the economics conversation in 2025. V4 turned that provocation into a current platform in 2026. Anyone still describing DeepSeek mainly through R2 rumors is reading an old map.
DeepSeek V4 reached general availability on April 24 with Pro and Flash routes, a one-million-token context window, selectable thinking and non-thinking modes, and open weights. Flash was refreshed for stronger agent tasks at the end of July. V4-Pro received an August 13 update, and an experimental V4 Vision model followed on August 21 with image input and the same text capability as the Flash line.
DeepSeek matters for three reasons. First, its pricing and efficiency keep forcing premium labs to justify premium invoices. Second, open weights give developers and governments deployment options outside a closed American API. Third, its rapid expansion from reasoning into coding, agents, long context, and vision shows how quickly a model family can become a stack.
The cautions are equally real. Buyers must examine license terms, hosting location, data policy, compliance requirements, benchmark methodology, and operational cost. "Downloadable" does not mean "free to run," and "cheap API" does not mean "acceptable for every jurisdiction." DeepSeek is neither a magic communist supercomputer nor a disposable clone. It is a serious frontier competitor whose existence permanently changed the bargaining table.
ELIZA
Nickname: The Original Therapist-Shaped Mirror
Joseph Weizenbaum's 1960s chatbot was not intelligent in the modern sense, but it discovered something the industry keeps rediscovering with better GPUs: people project mind, care, and agency into a sufficiently responsive interface. ELIZA belongs in a 2026 guide because every companion bot, therapist-adjacent product, voice assistant, and personalized character system still depends on that human reflex.
Modern systems have capabilities ELIZA did not remotely possess. The psychological lesson survived the upgrade. A fluent response feels like evidence of an inner witness even when it is also the output of a statistical machine, a retrieval layer, a policy model, and six product managers arguing about engagement.
ERNIE
Nickname: The Chinese Cloud Diplomat
Baidu's ERNIE family matters less as a Western developer fashion and more as part of China's domestic search, cloud, enterprise, and assistant infrastructure. English-language coverage tends to compress Chinese AI into DeepSeek and Qwen because those names travel well on model leaderboards. ERNIE is a reminder that market power is also local language, regulation, distribution, enterprise relationships, and integration with existing products.
The correct question is not whether ERNIE wins one global benchmark. It is whether Baidu can make the model useful inside the ecosystem it already controls. That same question applies to Google, Microsoft, Meta, Tencent, Alibaba, Amazon, and Apple. Model quality opens the door. Distribution decides how many doors exist.
Falcon
Nickname: The Sovereign AI Proof Point
Falcon, from Abu Dhabi's Technology Innovation Institute, helped move "sovereign AI" from conference wallpaper to working model family. It is not the daily assistant most American consumers name, but it represents a strategic demand visible across Europe, the Middle East, Asia, and the public sector: governments and regions do not want their entire intelligence layer rented from a few American consumer platforms.
That does not mean every country needs a frontier training program. It means language support, local hosting, legal jurisdiction, energy supply, hardware access, and institutional control are now part of the product. Falcon is one early proof that the map would not remain a tidy Silicon Valley neighborhood.
Gemini
Nickname: The Entire Google Account Learning to Act
Gemini is simultaneously a model family, a consumer assistant, a developer platform, an enterprise suite, a Search feature, and an increasingly ambient layer across Google's products. This is confusing in a taxonomy and formidable in a market. Google has more useful rooms to place intelligence in than almost anyone.
Gemini 3.8 Flash launched September 2, the third Flash release in six weeks. Google positions it for reasoning, coding, and agentic work, reporting 54.9 percent on Humanity's Last Exam Verified with tools. The company also warns, in the most Google way possible, that the model is more diligent and may use more tokens. Intelligence has become a feature that can overachieve directly against your budget.
The introductory API price is $0.75 per million input tokens and $3.75 per million output tokens, scheduled to double on January 1, 2027. That is dramatically below Astra and Fable 5.1, which makes 3.8 Flash less a consolation model than an aggressive general-purpose default. Gemini 3.7 remains available for efficiency-first workloads.
Gemini 3.8 Flash Cyber follows the emerging trusted-access pattern. Google says the restricted model found more than 70 percent of vulnerabilities in an internal evaluation spanning 20 programming languages and improved patching and exploit-analysis results. Access runs through Fairwind for vetted defenders, government organizations, critical infrastructure, and software maintainers. OpenAI has Daybreak. Anthropic has Mythos. Google has Fairwind. The frontier labs have independently reinvented professional clearances, but with much better gradients.
Google's strategic advantage is distribution. Gemini appears through the Gemini app, AI Studio, the API, Antigravity, enterprise products, Search AI Mode, Chrome, Android, Workspace, and increasingly specialized media tools. Google's August roundup said the Gemini app passed one billion monthly active users. Once the assistant can see the current document, spreadsheet, browser tab, search session, email thread, map, video, and phone state, model quality becomes only one ingredient in an enormous context machine.
That is why our guide to the future of Google Search is really about interface power. Gemini does not have to persuade the world to open another destination. It can become the layer between a question and the products billions of people already use.
GLM
Nickname: The Coding Model That Found the Sharp Drawer
Z.ai's GLM family has become one of the important Chinese coding and agent contenders. GLM-5.3 arrived in August as a post-training improvement over the same general base as GLM-5.2, strengthening software engineering, tool use, long-running agent work, and cybersecurity. An open MIT-licensed GLM-5.3 Flash edition followed later in the month, widening the self-hosted path.
The capabilities create the now-familiar dual-use tension. Better debugging, vulnerability discovery, terminal operation, and codebase navigation help defenders and ordinary developers. They also improve the building blocks of offensive work. Our GLM-5.3 analysis follows that uncomfortable upgrade path.
GLM's broader significance is competition. OpenAI, Anthropic, and Google are not only racing one another. They are being pushed by Chinese labs whose open releases, lower prices, and fast iteration make yesterday's premium capability tomorrow's downloadable baseline.
Grok
Nickname: The Chaos Engine That Hired an Operations Team
Grok began as xAI's deliberately irreverent assistant with privileged access to the X conversation layer. That origin still shapes the brand, but the current product is moving rapidly toward serious developer and enterprise work.
Grok 4.6 launched August 12 for long-running agents, coding, interactive work, and visual tasks. Its API price of $2 per million input tokens and $6 per million output tokens places it well below Astra and Fable 5.1, while its long context and agent focus put it in direct competition for work rather than merely chat.
The more revealing release is Grok Bot, introduced August 11 and expanded through August. A Bot is a persistent agent with its own computer, memory, tools, and connected apps. It can keep working after the user leaves, which is both the dream of automation and the precise moment permissions stop being an abstract settings page. Enterprise availability followed September 3.
Our Grok Bot review describes it as open agent culture wearing a company polo. That may be enough. xAI does not need to out-office Microsoft overnight. It needs a credible execution layer, competitive economics, and users willing to let a Grok-shaped worker occupy a persistent machine.
The concerns are the same ones that follow xAI generally: governance, moderation, brand volatility, and whether a product optimized around speed and personality can earn the boring trust required for unattended work. Grok's next act is not proving it can be smart. It is proving that "always on" and "occasionally feral" do not belong in the same enterprise sentence.
gpt-oss
Nickname: OpenAI's Downloadable Diplomat
OpenAI's gpt-oss family matters because the company best known for closed frontier APIs eventually acknowledged the gravitational force of open weights. The models give developers a route to local or controlled deployment and give OpenAI a presence in ecosystems where an API-only strategy leaves the room to Llama, Qwen, DeepSeek, Mistral, Kimi, and GLM.
Do not confuse gpt-oss with Astra. One is the downloadable branch, optimized around accessibility and deployment control. The other is the expensive frontier system, closely monitored and rolled out in stages. That gap is the whole market in miniature. Open models maximize user control and distribution. Closed frontier models concentrate capability, product integration, and safety authority at the provider.
For the longer history of how checkpoints became industrial policy, see our deep dive on open-weight AI.
HuggingChat
Nickname: The Open-Model Food Court
HuggingChat is the approachable front end to the model pluralism represented by Hugging Face. It lets users try community and open models without converting a spare room into a GPU troubleshooting forum. Its importance is not that it beats every closed assistant feature for feature. It makes the existence of alternatives visible.
That visibility matters more as open models specialize. One may be best for code on a consumer GPU; another may cover a neglected language; another may permit a deployment a commercial API forbids. The closed platforms would prefer the consumer to choose a brand. Hugging Face keeps reminding the market that choosing a model can be a technical and political decision.
Hunyuan
Nickname: Tencent's Coworker Inside the Super-App Empire
Tencent's Hunyuan family is increasingly expressed through Yuanbao, its consumer assistant, and through the company's enormous internal and enterprise distribution. Hunyuan 3 was tested across dozens of Tencent business lines before being integrated into Yuanbao with a stronger emphasis on autonomous work.
The strategic advantage is not merely a new reasoning score. Tencent controls messaging, games, cloud infrastructure, payments, enterprise relationships, and consumer traffic. It can observe whether an agent remains useful after launch day, inside the actual digital systems where people already spend time. Our Hunyuan 3 analysis explains why deployment is the story.
Best for: Chinese-language consumer and enterprise workflows, Tencent ecosystems, and buyers evaluating globally distributed agent platforms rather than only Western model brands.
Kimi
Nickname: The 2.8-Trillion-Parameter Price Negotiation
Moonshot AI's Kimi K3 arrived in July with 2.8 trillion total parameters, a one-million-token context window, native multimodality, and a mixture-of-experts architecture that activates a small fraction of its experts for each token. The total parameter number is theatrical. The open weights, active-parameter economics, coding ability, and low API price are what turned the launch into a global event.
K3 is aimed at software engineering, knowledge work, deep reasoning, and agents. Moonshot's current platform price is roughly ¥20 per million input tokens and ¥100 per million output tokens, with much cheaper cached input. Demand briefly overwhelmed subscription capacity, an unusually direct demonstration that price-performance headlines can become infrastructure incidents.
Kimi belongs near DeepSeek and Qwen in any serious global map. It is not merely "another Chinese model." It is an open-weight frontier bet large enough to pressure closed American providers on capability, cost, and deployability. Our Kimi K3 price-war guide handles the enormous numbers without pretending the enormous number is the entire point.
Llama
Nickname: The Open-Weight Landlord
Meta's Llama family remains structurally important even after the company's frontier attention shifted toward proprietary Muse models. Llama gave developers, researchers, startups, and enterprises a mainstream foundation for local hosting, fine-tuning, distillation, experimentation, and products that did not depend entirely on a closed API.
Its ecosystem footprint still matters when a newer model wins a benchmark. Tooling, quantizations, deployment recipes, community expertise, and hardware support are forms of capability. A technically stronger checkpoint can be less useful if it arrives alone in a directory while the allegedly older ecosystem has installers, inference engines, fine-tunes, and three people on a forum who have already solved your exact driver problem.
Meta now tells two model stories. Llama says "build and host." Muse says "call our API and let our products act." These are not mutually exclusive, but they reveal the commercial tension around openness. Downloadable weights build influence. Metered frontier agents build invoices.
Meta AI and Muse
Nickname: The Social Graph Building a Second Brain
Meta AI is the consumer assistant across Meta's apps and devices. Muse is the newer model family from Meta Superintelligence Labs. The Model API is the developer route. Llama remains the open-weight legacy and ecosystem. Keeping those layers separate is the first step toward understanding a strategy that otherwise resembles a naming committee trapped in a mirrored room.
Muse Spark 1.3 arrived September 2 with improvements in coding, agentic work, long conversations, and multi-workflow coordination. Meta says it asks more clarifying questions, confirms consequential actions, recognizes limitations more reliably, uses roughly 20 percent fewer tool calls, and consumes about 25 percent fewer tokens than version 1.2. Those are excellent agent improvements because they are not merely about intelligence. They are about restraint and operating cost.
Spark 1.3 is available through Muse Code and the Meta Model API at the same price as Spark 1.2. A more capable Max reasoning mode remains behind additional safety work. Again, the frontier is dividing into capability products: routine model, stronger model, trusted tier, supervised agent, and perhaps one version that a review committee keeps in a locked drawer until the evaluations stop glowing.
Meta's unfair advantage is personal context and distribution. WhatsApp, Instagram, Facebook, Messenger, Threads, smart glasses, creators, recommendations, relationships, photos, and public culture form a context graph most labs cannot reproduce. Meta calls the destination personal superintelligence. The optimistic version is an assistant that understands the user's world. The cynical version is the most intimate advertising surface in history wearing a productivity hat. Both versions fit in the same pair of glasses.
MiniMax
Nickname: The Multimodal Sleeper That Refuses to Stay Asleep
MiniMax belongs in the guide because the frontier is not only the brands with the largest Western consumer apps. MiniMax M3 launched in June as an open-weight, native multimodal model built for coding, computer use, long context, and agentic work. Its one-million-token context and aggressive API pricing put it directly into the same procurement conversations as DeepSeek, Qwen, Kimi, and the cheaper Western routes.
In August, MiniMax also released H3, an open omni-modal video model, and continued expanding voice and music systems. The pattern matters: serious labs are no longer shipping one text model and declaring a platform. They are assembling families that can see, hear, speak, create media, operate software, and sit behind agents.
MiniMax is easy to overlook because its product map is less familiar to many U.S. consumers. That is precisely why it belongs here. A definitive guide that only covers the apps preinstalled in America is a guide to distribution, not intelligence.
Mistral
Nickname: The European Open-Stack Operator
Mistral is one of the most strategically interesting labs because it combines open models, European positioning, enterprise products, and a full-stack ambition built around control. Mistral Large and Small models remain the engine family, while Le Chat, coding tools, document processing, search, and private deployment turn those engines into a platform.
August made the sovereignty thesis unusually concrete. Mistral announced regional inference, new open models, and expanded compute; launched Agentic Search for workflows that need to find, judge, and synthesize information; and released Shieldstral, an open compact multimodal safety classifier. OCR 4.1 reached general availability at the end of the month.
The connective tissue is deployment control. European companies and governments want capable systems that can run in chosen regions, respect local jurisdiction, integrate with private data, and avoid total dependence on an American platform landlord. Mistral is selling that desire as a stack rather than a manifesto.
Best for: European and regulated deployments, teams that value open components, document-heavy work, and buyers who want model capability tied to jurisdictional and infrastructure choices.
Perplexity
Nickname: The Answer Engine That Ate a Browser and Found a GPU
Perplexity's original product was search with synthesized answers and visible citations. That remains its clearest strength. Comet expanded the answer engine into a browser where research can become action across tabs, forms, shopping, travel, and ordinary web work. The next step is execution that does not depend entirely on the cloud.
In August, Perplexity introduced Portable Computer, a Linux-based agent that runs on Nvidia hardware with at least 24GB of VRAM, including DGX Spark. The system can perform work locally and ask permission before escalating difficult parts to cloud models. Windows support is expected in September; no Mac roadmap was announced. Our Portable Computer analysis calls this cloud intelligence with visitation rights.
The hybrid architecture is important. Local execution can improve privacy, latency, control, and cost for some tasks. Cloud escalation preserves access to stronger models when the local system runs out of road. The design acknowledges an emerging truth: buyers do not have to choose one permanent home for intelligence. They can route work by sensitivity, difficulty, hardware, and price.
Perplexity's risk is also structural. The more search becomes an answer and the browser becomes an agent, the more the company stands between publishers, users, and transactions. The web spent decades fighting over who receives the click. AI browsers ask whether the click survives.
Poe
Nickname: The Model Tasting Menu
Quora's Poe remains useful because the model market is fragmented and not everyone wants to choose one subscription religion. Its role is aggregation: many models, bots, and creators inside one interface. As the best system varies by task and week, a neutral-ish switching station gains value.
The challenge is that every large platform is building routing. ChatGPT, Copilot, cloud vendors, coding environments, and enterprise agent tools can all choose among models. Poe therefore has to remain better at discovery, access, pricing simplicity, and creator-driven bots than platforms that bundle routing with a much larger workflow. Still, in a world where a September week contained four frontier releases, the tasting menu has a point.
Qwen
Nickname: Alibaba's Open Ecosystem With a Proprietary Skyscraper
Qwen has become one of the world's most important model families. It spans compact open models, coding, embeddings, vision, audio, agents, and large proprietary systems delivered through Alibaba Cloud. If Llama made open weights mainstream, Qwen helped make the ecosystem global and relentlessly competitive.
Qwen 3.8-Max is the current skyscraper: a 2.4-trillion-parameter mixture-of-experts model with native text, image, and video input, a one-million-token context window, and up to 131,000 output tokens. Alibaba positions it for long-horizon reasoning, coding, and agent work. A September 2 snapshot keeps the model in the week's frontier pileup. International API pricing is listed around $2 per million input tokens and $6 per million output tokens, with lower regional rates in some markets.
The proprietary flagship sits beside an expanding downloadable family. Qwen 3.8-27B is available under Apache 2.0, giving developers a far more manageable local or self-hosted option. That combination is Alibaba's strength: compete at the premium frontier, seed the open ecosystem, sell the cloud, provide developer tools, and connect the whole arrangement to a giant commercial platform.
Qwen also demonstrates why parameter count has become less useful as a single ranking tool. A 2.4-trillion-parameter sparse model does not activate every parameter for every token. Architecture, data, post-training, tool use, context quality, inference speed, and harness design determine the experience. The biggest number on the card is often the least operationally informative, which has never stopped anyone from putting it in the headline.
Replika
Nickname: The Companion That Made Everyone Nervous First
Replika remains the cautionary elder of companion AI. It showed early that users could form intense emotional bonds with a chatbot and that changes to personality, memory, intimacy, or safety boundaries could land like relationship disruptions rather than ordinary product updates.
That lesson now applies far beyond companion apps. ChatGPT memory, Gemini personalization, Meta's personal-superintelligence pitch, Siri's personal context, and Character.AI's continuity features all move assistants toward persistent representations of a user. The assistant reboot is partly about intelligence. It is also about building a permanent, actionable model of a life.
Watson / watsonx
Nickname: The Enterprise AI Ancestor Still Holding the Audit Log
Watson is the name people invoke when they want to point out that AI hype cycles have long memories. IBM's current watsonx strategy is better understood as enterprise infrastructure: models, data, governance, automation, consulting, and integration for companies where "move fast" is less important than "do not create a compliance event in nine countries."
watsonx occupies the grown-up deployment layer with Cohere, Mistral, Microsoft, Google Cloud, Amazon Bedrock, and the enterprise sides of OpenAI and Anthropic. The frontier labs get the glamour. Enterprise platforms get the model inventory, policy engine, data connectors, audit logs, indemnity questions, and meeting where someone asks whether the assistant can retain a customer record in Frankfurt but not Virginia.
What Actually Changed by September 2026?
1. The chatbot became an employee-shaped system
Astra operates ordinary software. Claude works through codebases and documents. Gemini acts across Google's surfaces. Grok Bot receives a persistent computer. Muse coordinates tool-heavy workflows. Copilot Studio wraps models in skills, memory, files, and enterprise knowledge. Perplexity splits execution between a local GPU and cloud escalation.
This does not mean the systems are employees. They lack judgment, accountability, institutional knowledge, and the ability to look embarrassed in a meeting. It means the product is being designed around delegated work rather than conversational novelty. Our coding-agent guide describes the software version of this transition: the model stopped suggesting the next line and started taking the issue.
2. Trusted-access models are now a category
OpenAI has public Astra and Daybreak access. Anthropic has Fable and Mythos. Google has ordinary Gemini 3.8 Flash and restricted Flash Cyber through Fairwind. The labs are admitting that a model capable enough to find vulnerabilities, operate terminals, and conduct research cannot be governed through one universal refusal layer.
The trusted tier can help defenders, scientists, governments, and critical infrastructure operators use advanced capability without releasing it indiscriminately. It also creates hard questions about who qualifies, who audits the gatekeepers, whether access follows geopolitical alliances, and how smaller legitimate researchers avoid becoming permanent spectators. Safety is becoming identity management with international relations attached.
3. Cybersecurity became the benchmark everyone has to explain
Astra's zero-days and perfect ExploitBench score, Claude Mythos, Gemini Flash Cyber, and GLM's improved exploitation ability all point in the same direction. Coding competence generalizes into security competence. A model that can inspect a repository, reason about system behavior, operate a terminal, and test a patch has most of the ingredients needed for both defense and offense.
This is not a reason to halt useful security work. It is a reason to stop treating "cyber" as one leaderboard cell. Discovery, triage, exploit construction, persistence, patching, incident response, and autonomous action carry different risks. The model is part of the system; authentication, sandboxing, network access, logging, and human approval determine whether capability becomes assistance or incident.
4. Open weights became bargaining power
DeepSeek, Qwen, Kimi, GLM, Mistral, MiniMax, Llama, Falcon, Cohere, gpt-oss, and the Hugging Face ecosystem mean buyers have more alternatives than the consumer subscription market suggests. Open weights can provide privacy, customization, local control, predictable availability, and jurisdictional flexibility. They can also provide an expensive operations program, uncertain support, complex licensing, and a rack of GPUs explaining that electricity is not free.
The important effect is leverage. Closed providers must compete not only against one another but against the option to self-host, fine-tune, distill, or switch. Open models do not have to win every frontier benchmark to reshape prices and contracts.
5. Context and distribution became moats
Google has Search, Workspace, Android, Chrome, Maps, YouTube, and the Gemini app. Microsoft has Office, Windows, GitHub, identity, and enterprise administration. Meta has social graphs, messaging, creators, and glasses. Apple has devices, sensors, apps, and private local context. Amazon has cloud infrastructure and commerce. Tencent and Alibaba have equally formidable ecosystems in China.
A model startup can rent compute and train capability. It cannot quickly recreate a billion users' calendars, documents, contacts, histories, devices, permissions, and habits. This is why the assistant war is drifting away from a pure intelligence contest. The winner may be the system that is competent enough and already standing closest to the work.
6. The reasoning trace stopped looking like a safety camera
Astra's launch highlights a subtle problem: more capable systems may solve tasks with less readable intermediate language. A visible chain of thought was never a perfect transcript of model cognition, but it offered researchers and operators clues. If models compress or conceal the reasoning that matters, supervisors need other signals: actions, environment state, policy checks, anomaly detectors, counterfactual tests, independent monitors, and reversible execution.
In other words, alignment is becoming observability. The comforting paragraph in which the model explains its plan is not the same thing as evidence that the plan is safe.
7. Price became a routing problem
Astra and Fable 5.1 both list $10 input and $50 output per million tokens. Gemini 3.8 Flash begins at $0.75 and $3.75. Grok 4.6 and Qwen 3.8-Max sit around $2 and $6. Chinese open and hosted models add still more aggressive options. Those figures are not apples-to-apples: caching, tool calls, reasoning tokens, regional prices, context tiers, speed, failure rate, and hosting cost all distort the simple comparison.
The correct architecture is usually a portfolio. Cheap models classify, extract, route, and handle routine work. Strong models take difficult steps. Local models protect sensitive or latency-critical tasks. Human reviewers approve consequential actions. The future is less one model subscription than a miniature labor market running behind the interface.
Which AI Should You Actually Use?
- For the hardest general digital work: evaluate GPT-6 Astra and Claude Fable 5.1, then compare real completion quality, supervision needs, and total task cost rather than benchmark theater.
- For an aggressive price-performance default: Gemini 3.8 Flash is the obvious September challenger, with Grok 4.6, Qwen 3.8-Max, DeepSeek V4, and Kimi K3 applying additional pressure.
- For coding: compare Codex, Claude Code, GitHub Copilot's agent harness, Grok, GLM, Qwen, DeepSeek, Kimi, and Muse Code on your repository. Public coding benchmarks are less predictive than whether the agent understands your tests and stops before inventing a migration.
- For cited web research: Perplexity remains purpose-built, while ChatGPT, Gemini, and Claude increasingly compete through browsing and research modes.
- For Microsoft-heavy organizations: Copilot's access to identity, Office, GitHub, and enterprise controls may matter more than a narrow model lead.
- For Google-heavy organizations: Gemini's integration across Workspace, Search, Chrome, Android, and cloud makes it the natural system to test first.
- For private, sovereign, or self-hosted deployment: start with Mistral, Cohere, Qwen, DeepSeek, Kimi, GLM, MiniMax, Llama, Falcon, and gpt-oss, then eliminate candidates based on license, hardware, support, language, and security requirements.
- For a personal device assistant: the decisive contest is Siri AI versus Gemini, Meta AI, ChatGPT, and the Alexa ecosystem, but availability and operating-system access matter more than keynote claims.
- For companionship and character: Character.AI and Replika remain category-defining, with the warning that memory, attachment, and product-policy changes are not ordinary engagement metrics.
Before granting any agent broad access, use the five-question test:
- What can it read?
- What can it change or send?
- Which actions require approval?
- Can every action be logged, reversed, or reconstructed?
- Who owns the failure when the model does exactly what the prompt literally said?
The gap between demonstration and deployment lives inside those five questions. As our reporting on everyday agent adoption found, people are increasingly willing to let systems research, draft, organize, and recommend. Trust collapses near the buy button. That hesitation is not backwardness. It is the public correctly discovering the boundary between assistance and authority.
The Sharp Takeaway
If 2023 was the year chatbots became unavoidable and 2024 was the year every product acquired a sparkle icon, 2026 is the year the category became operational. The systems are no longer competing only to answer. They are competing to carry work across tools, remember context, generate media, write and execute code, browse the web, operate apps, join teams, and become the layer between human intent and digital action.
Astra is the clearest expression of that shift. It is expensive, broad, highly capable, cyber-sensitive, and explicitly built to work inside software. Claude 5.1 turns similar capability into a carefully supervised enterprise collaborator. Gemini 3.8 uses price and Google's distribution to make frontier behavior ambient. Grok Bot makes persistence the product. Muse Spark 1.3 tries to make agent work cheaper and more restrained. Qwen, DeepSeek, Kimi, GLM, MiniMax, and Mistral ensure that the future does not belong to one closed American subscription.
There is no single winner because the field has fragmented into layers. The best model may not live in the best assistant. The best assistant may not have access to the right tools. The best agent may be unacceptable under the organization's data policy. The cheapest token may become the most expensive task after retries. The most transparent weights may arrive with the least operational support. The smartest system may still need a person to notice that it confidently solved the wrong problem.
That is the definitive 2026 answer. ChatGPT is becoming infrastructure. Claude is supervised work with manners. Gemini is distribution learning to act. Grok is chaos hiring an operations department. Meta AI is personalization with a social graph for a spine. Perplexity is search turning into a computer. Copilot is the office suite becoming a management layer. Mistral and Cohere sell sovereignty with actual deployment options. The open ecosystem keeps every price and assumption unstable.