The irony is that community health centers may have the most to gain from AI, precisely because they have the fewest resources to waste.
The answer is to stop chasing large language models and start rethinking how we’re building AI and what we really want to get out of it.
Modern “smart” vehicles may have as much in common with computing as with mechanics in 2026.
In the era of autonomous AI, when an agent reads data from a collaboration tool, updates CRM records or queries a data warehouse, someone has to answer for those actions.
Executives are prioritizing AI as an investment, but accountability rarely makes it into that conversation.
Docker AI Governance treats the developer laptop as production infrastructure, giving CISOs runtime-level policy across sandbox, MCP, network and credentials.
For decades, the IQ test has been one of the most familiar — and most contested — yardsticks for human intelligence. Now, a startup project called AI IQ is applying the same metaphor to artificial intelligence, assigning estimated intelligence quotients to more than 50 of the world’s most powerful language models and plotting them on a standard bell curve.
The result is a set of interactive visualizations at aiiq.org that have ricocheted across social media in the past week, drawing praise from enterprise technologists who say the charts make an impossibly complex market legible — and sharp criticism from researchers and commentators who warn the entire framework is misleading.
“This is super useful,” wrote Thibaut Mélen, a technology commentator, on X. “Much easier to understand model progress when it’s mapped like this instead of another giant leaderboard table.”
Brian Vellmure, a business strategist, offered a similar endorsement: “This is helpful. Anecdotally tracks with personal experience.”
But the backlash arrived just as quickly. “It’s nonsense. AI is far too jagged. The map is not the territory,” posted AI Deeply, an artificial intelligence commentary account, crystallizing a worry shared by many researchers: that reducing a language model’s sprawling, uneven capabilities to a single number creates a dangerous illusion of precision.
AI IQ was created by Ryan Shea, an engineer, entrepreneur, and angel investor best known as a co-founder of the blockchain platform Stacks. Shea also co-founded Voterbase and has invested in the early stages of several unicorns, including OpenSea, Lattice, Anchorage, and Mercury. He holds a Bachelor of Science in Mechanical Engineering from Princeton University.
The site’s methodology rests on a deceptively simple formula. AI IQ groups 12 benchmarks into four reasoning dimensions: abstract, mathematical, programmatic, and academic. The composite IQ is a straight average of those four dimension scores: IQ = ¼ (IQ_Abstract + IQ_Math + IQ_Prog + IQ_Acad).
The abstract reasoning dimension draws from ARC-AGI-1 and ARC-AGI-2, the notoriously difficult pattern-recognition benchmarks designed to test general fluid intelligence. Mathematical reasoning includes FrontierMath (Tiers 1–3 and Tier 4), AIME, and ProofBench. Programmatic reasoning uses Terminal-Bench 2.0, SWE-Bench Verified, and SciCode. Academic reasoning pulls from Humanity’s Last Exam, CritPt, and GPQA Diamond.
Each raw benchmark score gets mapped to an implied IQ through what the site describes as “hand-calibrated difficulty curves.” Crucially, the methodology compresses ceilings for benchmarks considered easier or more susceptible to data contamination, preventing them from inflating scores above 100. Harder, less gameable benchmarks retain higher ceilings. The system also handles missing data conservatively: models need scores on at least two of the four dimensions to receive a derived IQ, and when benchmarks are absent, the pipeline deliberately pulls scores down rather than up. The site states that “every derived IQ averages all four dimensions, so missing coverage cannot make a model look better by omission.”
As of mid-May 2026, the AI IQ charts tell a story of rapid convergence at the top of the frontier — and widening diversity in the tiers below.
According to the Frontier IQ Over Time chart, GPT-5.5 from OpenAI currently sits at the peak of the bell curve, with an estimated IQ near 136 — the highest of any model tracked. It is closely followed by GPT-5.4 (approximately 131), Opus 4.7 from Anthropic (approximately 132), and Opus 4.6 (approximately 129). Google’s Gemini 3.1 Pro lands near 131, making the top cluster extraordinarily tight.
That compression is not unique to AI IQ’s framework. Visual Capitalist, drawing from a separate Mensa-based ranking by TrackingAI, recently observed the same dynamic, noting that “the biggest takeaway is how compressed the top of the leaderboard has become.” On that scale, Grok-4.20 Expert Mode and GPT 5.4 Pro tied at 145, with Gemini 3.1 Pro at 141.
Below the frontier cluster, the AI IQ charts show a crowded midfield. Models from Chinese labs — Kimi K2.6, GLM-5, DeepSeek-V3.2, Qwen3.6, MiniMax-M2.7 — bunch between roughly 112 and 118, making the cost-performance tier increasingly competitive for enterprise buyers who don’t need the absolute best model for every task. One X user, ovsky, noted that the data “confirms experience with sonnet 4.6 being an absolute workhorse as opposed to opus 4.5” — pointing to the way the charts can validate practitioner intuitions that headline rankings often miss.
What distinguishes AI IQ from most other benchmarking efforts is its inclusion of an “EQ” — emotional intelligence — score. The site maps each model’s EQ-Bench 3 Elo score and Arena Elo score to an estimated EQ using calibrated piecewise-linear scales, then takes a 50/50 weighted composite of the two.
The EQ scores produce a meaningfully different ranking than IQ alone. On the IQ vs. EQ scatter plot, Anthropic’s Opus 4.7 leads on EQ with a score near 132, pushing it into the upper-right quadrant — the most desirable position, signaling both high cognitive and high emotional intelligence. OpenAI’s GPT-5.5 and GPT-5.4 cluster in the high-IQ zone but lag slightly on EQ. Google’s Gemini 3.1 Pro sits in a strong middle position on both axes.
One notable methodological choice has drawn attention: EQ-Bench 3 is judged by Claude, an Anthropic model, which the site acknowledges “creates potential scoring bias in favor of Anthropic models.” To correct for this, AI IQ subtracts a 200-point Elo penalty from the EQ-Bench component for all Anthropic models before mapping to implied EQ. The Arena component is unaffected since it uses human judges. That self-correction is unusual in the benchmarking world, and it suggests Shea is aware of the methodological minefield he has entered. Still, the EQ dimension captures something IQ alone cannot: the growing importance of conversational quality, collaboration, and trust in models deployed for user-facing work.
Perhaps the most practically useful chart on the site is not the bell curve but the IQ vs. Effective Cost scatter plot. It maps each model’s estimated IQ against an “effective cost” metric — defined as the token cost for a task using 2 million input tokens and 1 million output tokens, multiplied by a usage efficiency factor.
The chart reveals a familiar pattern in enterprise technology: the best models are not always the best value. GPT-5.5 and Opus 4.7 sit in the upper-left corner — high IQ, high cost, with effective per-task costs north of $30 and $50 respectively. Meanwhile, models like GPT-5.4-mini, DeepSeek-V3.2, and MiniMax-M2.7 occupy a sweet spot in the middle: respectable IQ scores between 112 and 120, at effective costs ranging from roughly $1 to $5 per task. At the cheapest extreme, GPT-oss-20b (an open-source OpenAI model) appears near $0.20 effective cost with an IQ around 107 — potentially the most economical option for bulk classification or extraction workloads.
The site also offers a 3D visualization mapping IQ, EQ, and effective cost simultaneously. A dashed line running through the cube points toward the ideal: higher IQ, higher EQ, and lower cost. Models near the “green end” of that axis are stronger all-around deals; those near the “red end” sacrifice capability, cost efficiency, or both. For CIOs staring at API invoices, the implication is clear: the intelligence gap between a $50 model and a $3 model has narrowed enough that routing — using expensive models for hard problems and cheap ones for everything else — is no longer optional. It is the dominant architecture for serious AI deployments.
The loudest objection to AI IQ is philosophical, and it cuts deep. Critics argue that collapsing a model’s uneven capabilities into a single score obscures more than it reveals.
“IQ as a proxy is fading — we’re seeing reasoning density spikes that don’t map to g-factor,” posted Zaya, a technology commentator, on X. “GPT-5.5 already hit saturation on MMLU-Pro, but still fails ClockBench 50% of the time.”
That observation touches on what AI researchers call the “jaggedness” problem: large language models often exhibit wildly uneven capabilities, excelling at graduate-level physics while failing at tasks a child could do. A composite score can paper over those gaps.
Pressureangle, another X user, posted a more granular critique, calling out “complete lack of transparency” and arguing the site never fully discloses how its calibration curves were created or validated. In fairness, AI IQ does list its 12 benchmarks and shows the shape of each calibration curve in its methodology modal. But the raw data and precise mathematical transformations are not published as open datasets — a gap that matters to researchers accustomed to fully reproducible methods.
Others questioned the premise itself. “As useless as human IQ testing,” wrote haashim on X. Shubham Sharma, an AI and technology writer, offered a constructive alternative: “Why not having the Models take an official (MENSA-Grade) test? Wouldn’t this be the most accurate and most ‘human-comparable’ way to benchmark intelligence?” That approach already exists through TrackingAI, which administers the Mensa Norway IQ test to language models. But Mensa-style tests measure only abstract pattern recognition, while AI IQ attempts a broader composite across coding, mathematics, and academic reasoning. As Visual Capitalist noted, “an IQ-style benchmark captures only one slice of capability.” Each approach has tradeoffs — and neither has won the argument yet.
For all the debate about methodology, the most important signal in AI IQ’s data may not be any single model’s score. It is the shape of the market the charts reveal.
There are now more than 50 frontier-class models available through APIs, from at least 14 major providers spanning the United States, China, and Europe. Each provider publishes its own benchmarks, often cherry-picked to showcase strengths. The result is a Tower of Babel where no two companies measure the same thing in the same way. Academic research has highlighted that “most benchmarks introduce bias by focusing on a particular type of domain,” and the Frontier IQ Over Time chart on AI IQ shows just how fast the targets are moving: in October 2023, GPT-4-turbo sat near an estimated IQ of 75. By early 2026, the top models were brushing 135 — roughly 60 points of improvement in 30 months.
That pace raises a fundamental question about whether any scoring system can keep up. The site compresses ceilings for saturated benchmarks, but as models continue to max out even the hardest tests — ARC-AGI-2, FrontierMath Tier 4, Humanity’s Last Exam — the framework will face the same ceiling effects that have plagued every AI evaluation before it. Connor Forsyth pointed to this dynamic on X: “ARC AGI 3 disagrees,” he wrote, referencing a next-generation benchmark that may already be undermining current scores.
AI IQ is not perfect. Its methodology is partially opaque. Its IQ metaphor can mislead. And its creator acknowledges known biases while likely missing others. But the alternative — wading through dozens of provider-specific benchmark tables, each using different test suites and scoring conventions — is worse. The site offers enterprise buyers something genuinely scarce: a single framework for comparing models across providers, dimensions, and price points, updated regularly, with enough nuance to show that the right answer to “which model is best?” is almost always “it depends on the task.”
As Debdoot Ghosh mused on X after viewing the charts: “Now a human’s role is just to orchestrate?“
Maybe. But if the AI IQ data shows anything clearly, it is that orchestration — knowing which model to deploy, when, and at what price — has become its own form of intelligence. And for that, there is no benchmark yet.
Good news, OpenClaw fans — you can once again use your Claude AI subscription to power the hit, open source, autonomous AI agentic harness! But, there’s a big catch with how it’s being enacted.
A few hours ago, Anthropic announced via its official developer communications account on X, @ClaudeDevs, that it is changing its Claude paid subscription tiers, introducing a new subcategory of “Agent SDK” credits for all paid subscribers, which they can now allocate specifically for “programmatic” uses, including external, third-party agents such as OpenClaw.
The move is a major reversal from the Anthropic’s policy introduced in early April 2026 that expressly prohibited its AI subscriptions from being used to power these kind of non-Anthropic agents and harnesses, after Anthropic said they caused capacity and service issues.
The problem was that some Claude subscribers were paying $20 to $200 per month under Anthropic’s Claude Pro and Max subscriptions, but consuming hundreds, even thousands of dollars of tokens (units of information) above those prices through their OpenClaw (and similar autonomous) agents. This was an unsustainable position for Anthropic’s finances and its limited compute infrastructure for inferencing the models to end users.
To be clear, even when it enacted the old prohibition against OpenClaw and similar agents last month, Anthropic never fully cut off the capability for Claude to be used in OpenClaw. Rather, it redirected users to pay through the company’s application programming interface (API), which is billed by usage (priced per million tokens, rather than a flat monthly rate as the subscriptions offer), or pay for extra usage credits atop their subscriptions.
Now, Anthropic is giving Claude subscribers another way to use their subscription bill to pay for third-party agents.
However, the restoration comes with a significant catch: programmatic usage is no longer subsidized by the general subscription pool but is instead restricted to a fixed, non-rollover monthly credit, also worth $20-$200 depending on your Claude plan, and billed at the API rates.
In other words, if you don’t end up using these new Agent SDK credits, they simply expire at the end of the month. And if you do use them all up, you cannot dip into your general subscription usage limits to cover any additional usage — you’ll need to buy extra usage credits instead.
To understand why this restoration matters, one must look at the technical friction that led to the initial ban on April 4, 2026.
Anthropic’s first-party tools, such as Claude Code and Claude Cowork, are engineered to maximize “prompt cache hit rates”—a method of reusing previously processed text to save on expensive compute cycles.
Third-party tools like OpenClaw, which allow users to run autonomous agents through external services like Discord or Telegram, were often unoptimized for these efficiencies.Boris Cherny, Head of Claude Code, noted that these third-party services were “really hard for us to do sustainably” because they bypassed the caching mechanisms that allow Anthropic to offer flat-rate subscriptions.
The sheer volume of data being re-processed by inefficient agents was threatening the stability of the system for the broader user base. Even with Anthropic’s massive expansion into new hardware—including access to the 300MW Colossus 1 data center and its 220,000+ GPUs—the demand for agentic workflows was outpacing sustainable supply.
The new “Agent SDK credit” system solves this technical bottleneck by shifting the cost of inefficiency back to the user. By providing a dedicated dollar-amount credit, Anthropic no longer has to “eat the difference” on unoptimized third-party code. If an agent is inefficient and burns through tokens, it simply drains the user’s new $20 to $200 Agent SDK credit budget faster, rather than exceeding the value of Anthropic’s fixed monthly subscription tiers.
The restoration of third-party access is segmented across Anthropic’s billing tiers, creating a new hierarchy of “programmatic power.” Here’s how much Anthropic is giving each user in terms of the new Agent SDK credits (in addition to their normal Claude usage through Anthropic Claude products like Claude Code, Claude Cowork etc).
|
Plan |
Monthly, Dedicated Agent SDK Credit (on top of existing subscription plans) |
Usage Context |
|
Pro |
$20 |
Individual scripts and light SDK use. |
|
Max 5x |
$100 |
Moderate agentic automation. |
|
Max 20x |
$200 |
Professional-grade dev environments. |
|
Team (Premium) |
$100 / seat |
Collaborative team automation. |
|
Enterprise (Premium) |
$200 / seat |
Seat-based high-scale enterprise use. |
This system introduces a sharp divide between “interactive” and “programmatic” workflows. If you are chatting with Claude in a browser or using Claude Code in a terminal to write code interactively, you are still drawing from your standard, high-capacity subscription limits.
As Anthropic technical staffer Lydia Hallie wrote in a post on X, “To add some clarity: you don’t pay extra. It’s the same subscription, same price per month.” Hallie also included the following helpful diagram of how the new Agent SDK credits work:
However, the moment you use the claude -p command for non-interactive tasks, run a GitHub Action, or connect a third-party tool like OpenClaw, the system switches to the dedicated Agent SDK credit.
Once the Agent SDK credit limit ($20 for Pro plans, $100 for Max 5X, etc) is exhausted, programmatic usage stops unless the user has enabled “extra usage” billing, which is charged at standard, pay-as-you-go API rates.
Crucially, for those who found the original subscription model to be an infinite resource, this is a hard cap. Credits do not roll over, meaning the “use it or lose it” nature of the system forces a monthly reset of the developer’s budget.
The licensing implications of this move are profound for the “agentic” ecosystem.
By explicitly allowing third-party apps like Conductor and OpenClaw to authenticate via the Agent SDK, Anthropic is legitimizing a workflow it had previously attempted to block.
However, in doing so, it has ended the era of “compute arbitrage”.In the early part of 2026, a $20 Pro subscription could be leveraged via OpenClaw to run agents that would cost hundreds of dollars on a standard API key.
By moving to a metered credit, Anthropic is aligning its subscription model with its Developer Platform (API). While it offers a “free” buffer for subscribers, it ensures that high-volume, production-level automation is moved to predictable, token-based billing.
This protects the company’s margins while still offering a “sandbox” for developers to experiment without the immediate overhead of an API-first account.
While Anthropic executives framed the update as a “simplification”, the developer community has largely branded it as a significant reduction in the value of their subscriptions. The backlash focuses on the sharp disparity between the previous effective usage and the new, metered reality.
Popular AI YouTuber and developer Theo Browne (@theo) of T3.gg warned developers that this change constitutes a massive devaluation for those using external tools. “If you use any of the following with your Claude sub, your usage must got cut by 25x,” Theo stated, listing T3 Code, Conductor, Zed, and Jean as affected platforms. He concluded with a sharp warning: “They’re disguising this as ‘free credits’. Don’t fall for it”.
Kun Chen, a solo builder and former L8 engineer at Meta, Microsoft, and Atlassian, interpreted the move as a full surrender of Anthropic’s market lead. “it’s official. Anthropic pulled the plug on ALL programmatic use of claude subscription,” Chen posted, adding that he had found himself “increasingly bullish about OpenAI” as a result. Chen argued that “Anthropic’s only lead was on coding, and gpt 5.5 has flipped that already,” signaling a potential migration of elite developer talent.
Other builders questioned the practical utility of the credits offered. Ben Hylak, co-founder and chief technology officer at AI agent observability and governance startup Raindrop.ai, voiced concern over the sustainability of Anthropic’s infrastructure. “this is either really silly, or shows how bad of a spot anthropic is in re: gpus,” Hylak noted, before bluntly asking users to “guess how many turns $20 in API credits last”.
The frustration extended to the marketing of the change. EverNever, creator of inkstone.uk, expressed disbelief at the framing of the policy. “Wait what?! You take away more ways to utilize the subscription I am paying for?! And you dare to make it look like a win?”. This sentiment highlights a growing rift between Anthropic and its power-user base, who feel that previously inclusive features are being rescinded under the guise of an “upgrade.”
Anthropic’s “restoration” is a tactical move to retain developers while strictly managing the physical limits of compute. By June 15, the “agentic” era for Claude subscribers will be a metered one.
The company has successfully reclaimed control over its margins, even if it has cost them some of the goodwill of their most vocal power users.
For the individual developer or enterprise AI builder relying on Anthropic models for OpenClaw, however, it’s clearly an improvement over the blanket ban from last month.
For the first time since the AI race began, more American businesses are paying for Anthropic’s Claude than for OpenAI’s ChatGPT.
Adoption of Anthropic rose 3.8% in April to 34.4% of businesses, according to the May 2026 release of the Ramp AI Index. OpenAI’s adoption fell 2.9% to 32.3%. Overall AI adoption among businesses rose 0.2 percentage points to 50.6%.
The crossover — published Tuesday by Ramp, the corporate card and finance automation platform that tracks spending patterns across more than 50,000 U.S. businesses — marks the culmination of a yearlong surge by Anthropic that few in the industry predicted. Anthropic has quadrupled its business adoption over the past year, while OpenAI grew its business adoption by only 0.3%.
But the same report that crowns a new market leader also warns that Anthropic’s position may be more fragile than it appears — threatened by escalating costs, compute constraints, and the very token-based pricing model that has fueled the company’s extraordinary revenue growth.
To appreciate the scale of the shift, consider where the two companies stood a year ago. In April 2025, OpenAI commanded roughly 32% of business AI adoption according to Ramp’s underlying data, while Anthropic stood at under 8%. OpenAI had built an early, commanding lead as the consumer default — ChatGPT was where most people first encountered AI, and that momentum carried into corporate purchasing decisions.
Anthropic’s path was different. The company was popular early on with the earliest adopters — engineers, AI evangelists, the technical vanguard inside organizations. As Ramp lead economist Ara Kharazian noted in the March 2026 edition of the index, Anthropic leveraged that early-adopter base to go mainstream. By February, Anthropic was winning about 70% of head-to-head matchups against OpenAI among businesses purchasing AI services for the first time — a complete reversal of the trends observed in 2025.
The trajectory is visible in Ramp’s underlying data. The company’s adoption figures show Anthropic climbing from 0.03% of businesses in June 2023 to 7.94% by April 2025, then rocketing to 34.44% by April 2026.
OpenAI, meanwhile, peaked near 36.5% in mid-2025 and has been slowly declining since. The engine behind much of this growth is a single product: Claude Code, the company’s agentic AI coding tool, which has become the fastest-growing product in Anthropic’s history. A recent analysis estimated that 4% of all GitHub public commits worldwide were being authored by Claude Code — double the percentage from just one month prior.
Business Insider reported in April that the crossover was imminent. A Ramp spokesperson told the outlet that “at the current pace, Anthropic is on track to surpass OpenAI within the next two months,” noting that it already led “among early adopters, including VC-backed companies, and in key sectors like software, finance, and professional services.” That prediction proved accurate almost to the day.
The Ramp data on business spending finds its complement in a separate workforce survey that underscores just how deeply AI has embedded itself into American economic life. For the first time in Gallup’s measurement, half of employed American adults say they use AI in their role at least a few times a year, up from 46% the previous quarter. Frequent use is also increasing, with 13% of employees now saying they use AI daily and 28% reporting they use it a few times a week or more.
But the Gallup data, based on a February 2026 survey of 23,717 U.S. employees, also suggests that the benefits of AI remain concentrated at the level of individual tasks rather than organizational transformation. Only about one in 10 employees in AI-adopting organizations strongly agree that artificial intelligence has transformed how work gets done. That finding is consistent with firm-level studies across the U.S., U.K., Germany, and Australia showing chief executives reporting minimal broad productivity effects from AI over the past three years — a notable gap between the hype cycle and operational reality.
The Ramp methodology captures a different but complementary signal. Where Gallup asks employees whether they use AI, Ramp measures whether their employer is writing checks for it. The index counts corporate card and invoice-based payments, identifying firms as AI adopters if they have a positive transaction amount for an AI product or service in a given month. As Ramp’s methodology page notes, its results likely underestimate actual adoption because many employees use free AI tools or personal accounts for work tasks. Taken together, the two datasets paint a picture of AI that is ubiquitous in the American workplace but has not yet delivered on its promise to fundamentally transform how organizations operate.
Perhaps the most striking aspect of Ramp’s analysis is its refusal to declare a lasting winner. Kharazian identified three specific risks facing Anthropic even as the company takes the lead — and the most serious one stems from a structural tension baked into the company’s business model.
Anthropic makes more money when businesses purchase more tokens, meaning the company is incentivized to drive users toward more expensive models even when cheaper ones are sufficient. This dynamic is already creating budget crises at major enterprises. Uber’s CTO revealed that the company spent its entire 2026 AI budget in just four months, largely on Claude Code and Cursor, with engineers reporting monthly API costs between $500 and $2,000 per person. Adoption jumped from 32% to 84% of Uber engineers in a matter of months, and about 70% of committed code at Uber now comes from AI. The Uber case is a microcosm of a broader tension: Claude Code works — perhaps too well. When a productivity tool becomes so valuable that an organization’s $3.4 billion R&D operation can’t afford to keep the lights on, the resulting cost scrutiny could push enterprises toward cheaper alternatives.
At the same time, quality and reliability have suffered under the weight of demand. In recent weeks, users have experienced frequent outages, rate limits, and increasing dissatisfaction with Claude’s results. Anthropic has responded by resetting usage limits and by striking a compute deal with SpaceX to access more than 300 megawatts of new capacity at the Colossus 1 data center in Memphis. CEO Dario Amodei said the company saw “80x growth per year in revenue and usage” for Q1 2026, when it had only planned for 10x. And Ramp economist Rafael Hajjar found that Anthropic’s latest model update would triple token costs for any prompt that includes an image — a change that seems at odds with the company’s already-acute cost and compute problems.
The Ramp report points to competitive dynamics that could reshape the market within months. Some of the fastest-growing vendors on Ramp’s platform in April were AI inference platforms that give companies access to cheap, open-source models — offering enterprises a way to get “good enough” AI at a fraction of the cost, particularly for routine tasks that don’t require frontier model capabilities.
OpenAI’s Codex presents an even more direct threat. By most measures, it is a strong product that does many of the same tasks as Claude Code at a lower price point — and the switching cost between models is minimal. Uber itself is already testing Codex as a hedge, a move that could preview a broader pattern across enterprise tech. OpenAI also retains enormous structural advantages. ChatGPT reached 900 million weekly active users by March 2026, dwarfing Claude’s consumer footprint. Enterprise revenue now makes up more than 40% of OpenAI’s total and is on track to reach parity with consumer revenue by the end of 2026. And OpenAI’s $122 billion funding round, closed in March at an $852 billion valuation, gives it vast resources to compete on pricing, capacity, and product development.
Anthropic is not standing still on distribution. AWS recently launched Claude Platform on AWS, giving enterprises direct access to Anthropic’s native platform through existing AWS credentials, billing, and access controls — a move that lowers procurement friction considerably. Anthropic has also announced compute agreements totaling billions of dollars with Amazon, Google, Microsoft, Nvidia, and others, though much of that capacity won’t come online until late 2026 or 2027. Anthropic is reportedly in talks to raise another $50 billion at a valuation approaching $900 billion.
Beneath the spending data and market share charts lies a more intriguing question: Why are businesses choosing Anthropic over a cheaper, comparably performing alternative?
Kharazian explored this in his March analysis. Claude Code and OpenAI’s Codex are roughly comparable products — on certain benchmarks, Codex is arguably better, and it’s also cheaper. Yet Anthropic can’t meet its own demand. Every plan still has usage limits and rate caps. The company is actively turning away revenue because it doesn’t have the compute to serve it. Despite charging more for roughly equivalent performance, Anthropic’s demand is growing.
Kharazian suggested the answer might be cultural. Earlier this year, Anthropic refused to agree to the Pentagon’s terms of use for Claude, resulting in a blacklisting by the Department of Defense. OpenAI stepped in to offer its services in Anthropic’s place. In the wake of that episode, users rallied around Anthropic, and Claude temporarily surpassed ChatGPT on the App Store. The question, Kharazian wrote, is whether choosing an AI model is becoming less like an enterprise procurement decision and “more like the green bubble/blue bubble distinction in iMessage: a signal of identity as much as a choice of technology.”
That observation may sound absurd for an enterprise software category. But Ramp’s data tells a story that pure economics cannot fully explain. In a market where the products perform similarly, where the cheaper option is arguably better on benchmarks, and where switching costs are negligible, something other than spreadsheet logic is driving the biggest shift in AI market share since the industry began. As Kharazian noted in his report: “We have never seen a software industry as dynamic, where newcomers can disrupt market leaders in a matter of months, and where the pace of development overrides the typical forces of vendor stickiness.”
That dynamism cuts both ways. The same forces that propelled a company from 8% to 34% market share in twelve months could just as easily work in reverse. Anthropic’s two-point lead was earned in the most volatile software market in modern history — and in this market, the distance between the throne and the floor has never been shorter.
Browser-based AI tools can streamline workflows, but without clear safeguards, sensitive company information can be exposed through prompts, file uploads or integrations.