Anthropic and OpenAI Stake Concurrent Claims in AI4S, as the Industry Pivots from Model Superiority to Ecosystem Capture
On June 30, Anthropic and OpenAI concurrently entered the AI4S arena.
Anthropic unveiled Claude Science, a research agent workbench, with an explicit declaration that it 'does not rely on new models,' opting instead to bundle existing capabilities into workflows that encompass scientists' routine research.
OpenAI introduced GeneBench-Pro, a benchmark spanning genomics, quantitative biology, and eight additional domains. Per its test results, even the most capable model — GPT-5.6 Sol — posted only a 28.7% end-to-end pass rate across 129 real-world research workflow problems.
The two giants' strategies appear divergent, yet both rest on a shared diagnosis: the AI4S bottleneck is no longer model capability, but the persistent absence of genuine end-to-end functionality.
Operating from this consensus, Anthropic elected to embed existing models into an extensible workbench, offsetting model unreliability through toolchains and workflows. OpenAI moved to preemptively define 'what constitutes task completion in research,' enshrining discourse authority within benchmarks.
Prior to these moves, Google DeepMind had invested years in AI+science via foundational models such as AlphaFold. Its Gemini for Science platform is now bundling proprietary assets with databases, penetrating the same market through platform-level integration.
The AI4S landscape has entered, with little fanfare, a phase of 'ecosystem warfare among giants' — a shift from point comparisons of model prowess to the capture of strategic niches and workflow integration.
What ceiling has AI4S encountered?
Why have the three giants converged on AI4S infrastructure at this precise juncture?
As noted above, OpenAI embedded 129 problems in GeneBench-Pro that fully emulate real research workflows — spanning raw data cleaning, quality control, modeling, diagnosis, and conclusion formation. Scoring follows a strict binary rule: pass is granted only when every decision is correct. An otherwise flawless intermediate analysis yields zero credit if the final conclusion is wrong.
The data reveals that OpenAI's flagship model, GPT-5.6 Sol, posted just 28.7% under Max reasoning settings. Among non-GPT models, the strongest — Claude Opus 4.8 — managed only 16.0%.
This indicates that models can detect data anomalies and identify local diagnostic signals, yet they fail to translate that recognition into downstream methodological adjustments and sound analytical decisions. The problem is noted but action is unchanged — OpenAI's paper terms this shortcoming the 'notice-act gap.'
What gives rise to this chasm between 'recognition' and 'action'? Wu Hao, founder and CEO of Luomi Technology, identified three structural deficiencies of general-purpose large language models in the life sciences:
First, an inherent difficulty in comprehending the distinctive structure of raw biological data.
Second, many biological phenomena resist standard text tokenization — gene expression, for instance, is inherently stochastic.
Third, biological data is pervasively marked by unknown missing values.
Research costs are another non-trivial factor. GeneBench-Pro data indicates that a single problem can cost thousands of dollars in human expert labor. When models are unreliable, institutions must continue to depend on costly human expertise. Moreover, the life sciences impose exceptionally stringent data compliance requirements.
This is why the convergence is happening now. Model capabilities have encountered the 'notice-act gap' ceiling, and the conventional approach of scaling compute no longer applies in research contexts. Engineering integration, ecosystem positioning, and data sovereignty have emerged as the pragmatic frontiers. The three giants' concurrent entry is a structural inevitability born from that ceiling.
One Table, Three Distinct Strategies
Confronted with this ceiling, the three giants have charted markedly divergent AI4S courses. Leiphone observed that each path converges on the same terminus: becoming the indispensable underlying infrastructure for scientific work.
Anthropic's play is the most direct. Claude Science functions as a dedicated workbench: a primary AI assistant decomposes tasks like a project manager, assigns them to sub-agents, and a fact-checker cross-validates results. The platform connects to over 60 scientific databases and ships with pre-built toolkits for genomics, protein structure, chemistry, and related fields.
Wu Hao noted that the technical core lies in invoking external vertical models — such as scGPT for single-cell data and DNABERT for gene sequence analysis — via the MCP protocol for specific computations. Claude itself is confined to natural language understanding, task decomposition, and result interpretation.
This division of labor renders Anthropic genuinely independent of new models, conferring tangible advantages: it circumvents the steep inference cost of applying general-purpose models to biological matrices, and vertical models can iterate independently without waiting for the long update cycles of their general-purpose counterparts. Critically, given the life sciences' stringent data compliance requirements, this architecture permits sensitive data to be processed on local MCP Servers, eliminating the need for cloud uploads.
If Anthropic's strategy amounts to commanding an entire track, OpenAI's logic is to install GeneBench-Pro as the arbiter — defining 'what constitutes good AI4S' — and deploy its specialized model GPT-Rosalind as the athlete chasing high scores.
Prior to GeneBench-Pro, OpenAI had already introduced GPT-Rosalind four months earlier. The model, fine-tuned specifically for biological reasoning, is offered as a research preview to qualified U.S. enterprise clients, contingent on security review.
Google DeepMind wields an unmatched trump card. It possesses foundational science models — AlphaFold, AlphaGenome, among others — all proprietary assets, tightly integrated with Gemini for Science and more than 30 life science databases.
The critical distinction: models that other players can only invoke through tool calls are, for Google, native infrastructure. Competitors may build a superior workbench or define more stringent benchmarks, but the core capacity for protein structure prediction resides with Google.
The three giants' market approaches diverge as well:
Anthropic pursues breadth, betting on subscription-based普及. Pro, Max, Team, and Enterprise tiers all include access to Claude Science. Notably, Anthropic recently introduced a $30,000 credit program for 50 postdoctoral and graduate researchers, with a July 15 deadline — an effort to embed its workbench into the habits of young scientists before they attain independent PI status, thereby capturing the next generation's academic workflow.
OpenAI goes narrow: standards are公开, enabling broader participation, but the model remains closed, erecting barriers through enterprise access control.
Google goes deep, constructing moats through proprietary assets. The model is the platform — usage breeds dependency, and dependency deepens lock-in.
Three strategies, three distinct calculations of approach and risk.
Anthropic wagers the ceiling will hold in the near term, prioritizing engineering-led workflow expansion. The central risk: an earlier-than-expected model breakthrough could relegate Anthropic to a mere combinatorial toolbox.
OpenAI bets the ceiling will eventually break, positioning itself as the standard-setter while awaiting model capabilities to catch up. Yet this self-appointed referee status carries the risk of rejection by the scientific community.
Google bets on a layer beyond the ceiling: the party that controls foundational models at the source will always hold cards. The moat is formidable, but the ecosystem remains comparatively insular.
Each of the three holds distinct chips and distinct blind spots. None possesses a certain winning hand, yet all have committed their chips within the same strategic window.
For now, the outcome resists prediction. At minimum, no marquee client has been captured by a single player: pharmaceutical heavyweight Novo Nordisk appears on both Anthropic's (Claude Science reference customer) and OpenAI's (Rosalind early partner) rosters. The same client is concurrently evaluating multiple solutions, signaling that the market remains in open competition — no vendor's toolchain has yet proven compelling enough to induce scientists to port their entire workflow.
AI4S's endgame will probably not be determined by any single giant. When all three players reached the ceiling on the same day, each chose to enter — yet no consensus on the path forward emerged. The true answer remains with scientists: how they weigh data sovereignty, academic independence, and research efficiency, and to whom they grant their vote of confidence. That decision may prove more consequential than any technical metric.
For further AI4S developments and industry insights, interested parties may contact the author via WeChat at LorraineSummer.