A common misjudgment
Meta spent $2 billion on Manus, and Elon Musk gave Cursor a $60 billion acquisition option. When those numbers landed, the most common reactions across the Chinese internet boiled down to two lines: first, aren't both of these just wrappers? They use other people's models underneath — what's so special? Second, Zuckerberg and Musk are impulse buyers — one is "Meta missed AI so it's paying up," the other is "Musk just buys whatever's hot."
The subtext: Manus and Cursor aren't really different from the dozens of other AI agent tools and coding tools out there. They just had good marketing and lucky timing.
This article says that judgment is wrong. Not a little wrong — directionally wrong. Manus and Cursor's level of understanding in their respective fields leads the industry by at least a full step, and that cognitive lead can be verified through specific technical roadmaps and competitor comparisons. Meta and SpaceX/xAI's offers aren't impulse buys. They're pricing that cognitive lead.
Manus has been controversial since its launch in March 2025. The most common criticism is that it's a wrapper: it doesn't train its own models — it uses Claude and Qwen — and just wraps an agent orchestration framework around them. MIT PhD Qin Zengyi's comment represents one camp: it's a great product, but not a technical breakthrough.
To understand what Manus got right, the most effective way is to put it next to its competitors from the same period.
Cognitive difference 1: No role-playing
From 2023 to early 2025, most multi-agent systems were designed by copying human organizational structures. MetaGPT is a textbook example: it splits LLM agents into five roles — product manager, architect, project manager, engineer, QA — each with fixed responsibilities and workflows, executing sequentially like a human software company. This is what's called hat wearing.
The problem starts at the premise. Human societies need specialization because an individual's bandwidth is limited, requiring decades of training to become a senior product manager or senior engineer. Division of labor compensates for human cognitive limitations. LLMs aren't like that. Any off-the-shelf LLM is a generalist — it knows everything about every domain. Telling it in the prompt "you are a senior software engineer" does nothing but constrain its capabilities.
Thinking from first principles, the conclusion is completely different: instead of having multiple agents each play a human role and collaborate sequentially, each agent should retain the full capabilities of a generalist, with task-level division only. Manus's wide research mechanism is the productization of this idea. Its main planner agent breaks the user request into independent sub-tasks, then launches a separate, full-capability Manus instance for each sub-task, with its own context window, executing autonomously in a cloud VM sandbox. No "product manager agent" or "engineer agent" role tags. Each sub-agent can plan, execute, and verify.
This isn't a UI difference or a product strategy difference. It's a different understanding of what LLMs fundamentally are. MetaGPT designs systems from human org structures. Manus designs from LLM capability characteristics. The latter is right; the former is wrong. This judgment was a minority view in March 2025. By 2026, it's become industry consensus: OpenAI's Codex uses Plan/Spec Mode (planner analyzes requests, executor runs in sandbox), Anthropic's Claude Code uses orchestrator-worker (lead agent plans, sub-agents execute in parallel), Cursor uses Planner-Worker-Judge. All the top players have converged on a functional division (plan, execute, evaluate) — none are putting human job hats on agents.
Manus's product-level judgment shows the same depth. In March 2025, while most agent products were still verticalized (research only, generation only), Manus was the first to build an end-to-end pipeline — from autonomous search to code generation to data visualization in one flow. Today that's standard for agent products, but back then it was a minority call. I wrote an analysis that week discussing the compounding effects of Agentic AI across tools, data, and intelligence. Manus was the only product at the time that built all three layers of compounding.
Cognitive difference 2: Creating and distributing User Generated Software
Software has a long-standing supply-demand mismatch: professional software companies serve top-of-the-pyramid needs, leaving long-tail demand unmet. This is similar to media before YouTube — TV networks met head content demand, but long-tail content creation was ignored until User Generated Content platforms arrived.
Manus spotted this and made a product decision that seemed unconventional at the time: letting users deploy and distribute the apps Manus generates. Describe a need, Manus auto-generates frontend, backend, and database, then deploys it all to the cloud with one click and returns a shareable link. That alone already exceeded most agent products of the era. But Manus went a layer further: it provides an API for deployed apps to call Manus's own AI capabilities. In other words, users can not only generate software with AI — the generated software itself can continue to use AI.
This judgment wasn't obvious at the time. In March 2025, most AI agent products positioned themselves as "tools that help you complete a task" — outputting reports, code, or slides, then the job is done. Manus positioned itself as "helping you create a software product that can run and be distributed continuously" — and that product comes with built-in intelligence. These are two completely different product logics. The former treats AI as a one-time productivity tool; the latter treats AI as the infrastructure for User Generated Software.
Market response validated the bet. Manus's waitlist exceeded 2 million after its public demo. What excited users most wasn't just that AI could do research and write code — it was the one-click deployment into a real, usable online product. By the end of 2025, vibe coding and AI app builders had become a $4.7 billion market. Manus was one of the first products to build the complete pipeline of creation, deployment, and intelligence injection.
The cognitive level behind this design choice shows in its completeness of the value chain. Most competitors stopped at generation. Manus thought all the way to distribution and continuous operation. This points back to the same root as the first cognitive difference (no hat wearing): this team thinks from first principles, not by incrementally optimizing existing product forms.
Results and responses
Business results directly reflect these insights: 8 months to $100M ARR, 147 trillion tokens processed, over 80 million virtual computers created. GAIA Level 3 benchmark score of 57.7%, ahead of OpenAI Deep Research's 47.6%.
Two common follow-ups need addressing.
First: "Agent products are everywhere. Manus is a previous-generation form — no direct use for Meta." This is half true. Manus represents the cloud sandbox agent form, while by 2026 the mainstream direction shifted to local terminal agents like Claude Code and OpenClaw, and enterprise-integrated agents like Amazon Q. Product-generation-wise, Manus's form isn't the newest. But the logic of an acquisition has never been about buying the newest-generation product. Meta bought this team's cognitive level, engineering ability, user base, and infrastructure accumulation. Product forms can iterate; the team's understanding and practical experience with agent AI won't expire just because a newer generation appears. Meta integrated Manus's agent capabilities into its Ads Manager workflow in February 2026 — that shows Manus's technical assets found a real landing spot inside Meta's product system.
A more direct piece of evidence is the context engineering blog post Manus's team published in July 2025. That post is information-dense — it reveals that Manus's team's understanding of agentic AI leads the industry by a step. Its three core principles (keep prefix stable, make context append-only, mask tools don't remove them) were later widely referenced and adopted across the entire harness engineering field. More importantly, the post answered a critical technical roadmap question right at the start: should you train an end-to-end agentic model based on open-source models, or should you build agents on top of frontier models' in-context learning capabilities? Manus chose the latter and proved the path's feasibility with product results. This judgment wasn't consensus in mid-2025; by 2026, it had become mainstream industry practice. A technical blog post with that level of foresight and influence is itself proof of the team's cognitive level.
Second: "Manus is just a wrapper from start to finish — no technical substance." In April 2026, China's National Development and Reform Commission (NDRC) used the foreign investment security review for the first time in five years to issue a "prohibition and revocation" order to block the acquisition. If Manus were truly just a wrapper without core technology, regulators would have no reason to deploy the strongest legal tool to protect it. The NDRC's decision was a landmark use of the review mechanism, sending a clear warning to AI startups — a move I explored in detail in my coverage of the block. Regulators determined that this company's core team, R&D capabilities, training data, and IP constitute national security assets worth protecting. That determination carries more weight than any technical evaluation or media debate.
Cursor: The only third-party player training its own model
Cursor faces a similar wrapper question: it uses other people's models underneath and just built an editor. But Cursor made a judgment that no competitor in its space made — and built a complete technical moat around it.
Cognitive difference 1: Deciding that self-trained models are a product necessity, then building them
The core loop of a coding agent is high-frequency tool calls: read files, write code, run commands. Every round has latency, and it accumulates directly into product experience. The Cursor team realized early on that relying on external frontier model APIs for speed and cost couldn't deliver a satisfying interactive experience for developers. Self-training the model was an unavoidable step at the product level. In Cursor's own blog words, their goal was to train "the smartest model that can support interactive use, keeping developers in the flow of programming."
You might wonder: didn't we just say Manus was right to use external model APIs? How is self-training suddenly necessary for Cursor? The difference lies in the core constraints of each domain. For Manus's general agent domain, the key differentiation is at the agent architecture and context engineering layer — underlying model capability differences get absorbed by the agent framework. Programming is different: latency and cost directly determine product usability. The commonality is that both made the right build vs. buy decision based on their domain's actual constraints.
Once committed to this direction, Cursor executed and validated the judgment with product experience. After Composer 1 launched, I replaced Sonnet 4.5 with it in many projects. Subjectively, for about 90% of daily programming tasks (fixing bugs, writing CRUD, refactoring, adding features), Composer 1 and Sonnet 4.5 produce no obvious quality difference. The proportion of tasks that truly need rocket-science-level reasoning is tiny. Most is grunt work where model capability gaps don't show. But the speed advantage is crushing: the same task takes a minute or two with Sonnet 4.5, and a few seconds to ten-plus seconds with Composer 1. Similar quality, multiple times faster — that's a huge experience difference in high-frequency use. This is exactly the judgment Cursor made from the start: the bottleneck for product experience in programming is model speed and cost, not capability ceiling.
On approach, Cursor didn't pretrain a model from scratch. It took an open-source MoE base and ran large-scale RL post-training in an agent harness simulating Cursor's production environment, training the model's tool-calling decisions and response efficiency.
A common question: isn't this just fine-tuning?
The five-month evolution from Composer 1 to 2 answers that. Cursor's training pipeline went through three iterations, each upgrading the training methodology itself, not just tuning parameters. The 1 and 1.5 phases were pure RL: large-scale post-training on an open-source base. For Composer 1.5, RL compute scaled 20x, post-training compute even exceeded the base pretraining itself, and they introduced two new training behaviors: thinking tokens (adaptive reasoning depth) and self-summarization (automatic long-context compression). But they found the marginal returns of the RL-only path diminishing: CursorBench improved only 6.2 points from 1 to 1.5, while compute increased 20x.
For Composer 2, Cursor made a key methodological pivot: adding continued pretraining before RL to improve the starting quality of RL exploration. The base switched to Kimi K2.5 (confirmed by Moonshot), and after continued pretraining plus RL, CursorBench jumped 17.1 points. Composer 2's technical report states clearly: it achieved Pareto optimality with significantly lower inference cost than comparable models. In other words, Cursor's post-training pipeline didn't just layer a fine-tune on top of a base — it compressed cost and latency while maintaining comparable coding capability.
This self-correction in methodology has academic backing. ICML 2025 research (SFT Memorizes, RL Generalizes) and Moonshot's own Kimi K2 technical report point in the same direction: pretraining establishes priors, RL explores efficiently on those priors, and continued pretraining changes the starting quality. The Cursor team independently discovered this before Composer 2 and shipped it into the product.
Look at competitors' choices. The AI coding tools space has many startups: Cline is an open-source VS Code extension that connects to various third-party models; Trae is from ByteDance, using Claude 3.5 Sonnet and GPT-4o; Windsurf (formerly Codeium) is from Cognition. Their product differentiation comes from UI design, workflow orchestration, and pricing — not from model capability itself. Cognition's SWE-1.5 (the model behind Windsurf) does RL on an open-source base but hasn't gone to continued pretraining. Cline and Trae do zero model training. In the AI coding tools space, only LLM providers' first-party products (OpenAI's Codex, Anthropic's Claude Code) and Cursor have completed the full four-link chain: base model selection, continued pretraining, RL post-training, and product integration. Cursor is the only third-party startup to do it. These competitors aren't not trying hard enough. They simply didn't make the judgment that self-training is a product necessity.
Cognitive difference 2: Harness engineering shipped to product
Cursor's cognitive lead also shows in its exploration of harness engineering and agent scaling — and those explorations have shipped directly into the product.
In a self-driving codebases experiment published in February 2026, Cursor used a recursive Planner-Worker architecture to run hundreds of agents in parallel, peaking at roughly 1000 commits/hour, generating over one million lines of Rust code. The experiment tackled a question not of "how to make one agent write good code" but "how to get 10x meaningful throughput from 10x compute input." The mere posing of that question already led the industry.
This experiment later sparked controversy. People examined the public repo and found that none of the recent commits compiled — all GitHub Actions CI failed. Reddit and Hacker News flooded with criticism, calling Cursor a fraud.
The criticism caught a real gap, but "fraud" is inaccurate. Cursor proactively discussed these issues in the original blog. The post stated upfront that the repository wasn't for external use and that code quality would be imperfect. It specifically analyzed why the system design had to tolerate some error rate: when they demanded 100% correctness per commit, the system's effective throughput collapsed — agents overstepped to fix unrelated issues and stepped on each other. The post also discussed agent dependency hallucination (agents pulling in dependencies they shouldn't use) and explained subsequent fixes. Cursor wasn't caught and then made excuses. It published failure patterns alongside the experiment results.
This is more fairly read as a cutting-edge experiment on agent spatial scaling, with Cursor honestly documenting current capability boundaries and failure modes. The output code quality was genuinely poor — the public repo's state was worse than the blog's phrasing implied. But the questions the experiment answered (multi-agent parallel throughput scaling laws, error rate control, task allocation strategies) no other company had explored at the same scale. Linear CEO Karri Saarinen recently noted that AI coding agents add bandwidth rather than replace engineers, a perspective that helps explain Cursor's focus on productivity over agentic replacement.
Meanwhile, Cline was optimizing single-agent permission control and tool-calling workflows. Trae was optimizing its Builder Mode project scaffolding experience. All valuable work — but they answer questions at a different level. Cursor is already thinking about agent spatial scalability. Competitors are still optimizing single-agent interaction quality.
Cursor also introduced background agents and parallel task execution into the product early, surfacing these capabilities in the UI. While most AI coding tools were still discussing single-turn dialogue quality, Cursor was solving the further-out problem of agents continuing to work after a user walks away. Competitors like Cline and Trae didn't start similar exploration until mid-2026.
Responding to the wrapper accusation
LLM providers' first-party products (OpenAI Codex, Anthropic Claude Code, Google Gemini Code Assist) have poured resources far beyond Cursor's, yet Cursor still competes head-on at the product level. Fortune magazine's March 2026 analysis was titled Cursor's crossroads, discussing not whether Cursor can compete with LLM providers, but how it maintains its edge after they go all-in. A real wrapper product wouldn't be placed in that discussion. SpaceX/xAI offering $60 billion isn't buying an editor skin. The option is really about access to Cursor's developer data, as PatronusAI co-founder Anand Kannappan argued.
Common patterns
Put Manus and Cursor together, and the patterns become clear.
First, both teams led their respective fields in cognitive understanding. Manus's understanding of LLM nature (no hat wearing, wide research), its judgment on User Generated Software (complete pipeline of creation, deployment, and intelligence injection), and its grasp of end-to-end agent capability combinations; Cursor's understanding of the post-training pipeline (full four-link chain), its grasp of harness engineering (spatial scalability), and its understanding of agent working patterns (background agents and parallel execution) — all came at least six months to a year earlier than peers. Competitor comparisons confirm this: during the same period, the problems competitors were solving and the problems these two were solving often existed on different levels. This lead isn't achieved by throwing resources at it. It comes from a different fundamental understanding of AI as a medium.
Second, cognitive lead translated into outcome differences. Manus hit $100M ARR and a leading GAIA benchmark score. Cursor achieved the strongest position in AI coding tools outside LLM providers' first-party products. Both outcomes were backed by top-tier buyers with real money: Meta at $2 billion, SpaceX/xAI at $60 billion.
Third, both teams faced the wrapper accusation. That accusation reflects an increasingly outdated technical taxonomy. In 2023, whether a company had a self-trained model was the core criterion for technical substance. By 2025, dimensions like training pipeline design, agent architecture engineering, context engineering methodology, and harness engineering practice have become equally or more important than self-trained models. Judging 2025 products with a 2023 framework systematically underestimates the technical depth of companies like these.
Both companies face very real challenges. Manus's acquisition has been blocked by the NDRC, creating significant path uncertainty. Cursor faces head-on competition from Claude Code and Codex, and core engineering talent is already flowing to xAI. But challenges don't change the fact that these two teams have demonstrated technical judgment that places them in the first tier of the entire industry during the Agent AI era. Zuckerberg and Musk's offers aren't impulse purchases. They're pricing that judgment.