04Deep Dive
Strategic posture, model by model.
Capability scores summarise; strategy explains. Below, each contender's operating thesis, structural advantage, and the single risk a CEO must price into a multi-year contract.
Anthropic has converted "Constitutional AI" from a research talking point into the dominant enterprise distribution strategy. Opus 4.7 is the technical artefact; the business story is 40% of enterprise LLM spend.
Strengths
- 87.6% SWE-bench Verified — current frontier ceiling
- 1M-token context, agentic file-system memory
- Powers Cursor and Windsurf — the dev-tool default
- Highest LMArena Elo (~1503) among major providers
Frictions
- No native image, video, or audio generation
- $5 / $25 per 1M — premium pricing band
- Token compression vs. Opus 4.6 (~1.0–1.35×)
- Compute dependency on AWS / Google infrastructure
Claude becomes the default substrate for agentic and regulated workloads through 2027. The "Mythos" preview signals further reasoning gains. Pricing remains the single largest commercial risk.
No competitor delivers comparable depth across text, image, video, and audio under a single API surface. Veo 3.1 is the most consequential video model in commercial use; Nano Banana 2 is its image counterpart.
Strengths
- 2M-token context — largest commercial offering
- 94.3% GPQA Diamond — frontier reasoning benchmark
- Native Veo 3.1 video with synchronized audio
- Workspace, Android, YouTube distribution
Frictions
- Time-to-first-token at upper end (~27s reasoning)
- Coding ecosystem narrower than Claude / GPT
- Reasoning tokens billed — opaque cost surface
- Brand association still consumer-skewed
The most likely 2026 platform winner. Multimodal saturation, integrated tooling, and Vertex AI distribution compound into a near-monopoly in cross-modal workflows.
The most consequential pricing event in the industry since the launch of GPT-4. DeepSeek's sparse attention architecture, MIT-licensed weights, and gold-medal mathematical reasoning have re-baselined what "expensive AI" means.
Strengths
- $0.14 / $0.28 per 1M — ~96% below closed frontier
- MIT licence — fully redistributable, self-hostable
- Gold-medal IMO/IOI 2025; R1 chain-of-thought
- DeepSeek Sparse Attention (DSA) innovation
Frictions
- Geopolitical exposure (PRC origin, possible restrictions)
- API uptime degradation during demand spikes
- No native image / video generation
- Enterprise governance maturity below US labs
DeepSeek becomes the cost benchmark every CFO measures against. The realistic deployment path is self-hosted V3.2 / V4 for non-sensitive workloads alongside a Western model for client-facing surfaces.
A 100× training-compute jump versus the prior generation. Grok's structural advantage is real-time integration with the X graph — for media, market intelligence, and live narrative tracking, no other model has equivalent ground-truth.
Strengths
- Native, real-time X / Twitter signal
- Grok Imagine: integrated image + video + voice
- 65% hallucination reduction vs. Grok 4.0
- Most aggressive compute roadmap (Memphis)
Frictions
- Brand polarization caps enterprise adoption
- Video quality capped at 720p (vs. Veo / Kling 4K)
- Output speed in 18th percentile of frontier models
- Single-platform data lock-in (X)
Grok 5 (Jan 2026 launch window, 6T-parameter target) is the watch event. If it executes, xAI moves from edge-case to default for media, finance, and intelligence verticals.
Perplexity is not competing on raw model quality — it routes between frontier providers — but on the workflow of grounded research. Comet (free), Deep Research, and citation-native answers redefine how analysts consume the web.
Strengths
- Citations as first-class output — audit-grade
- 45M users, ~1B queries/month
- Comet browser disrupts the default search surface
- Model-agnostic via Council and Sonar API
Frictions
- No native reasoning depth — depends on upstream models
- No image / video generation
- EU AI Act GPAI compliance unclear (deadline Aug 2026)
- M&A target risk in 2026–27
Wins the research vertical, struggles to break into general-purpose AI. Most likely outcome: a $20–50B acquisition or a vertical SaaS pivot toward regulated industries.
Meta has converted open weights into a strategic moat. Llama 4 Scout's 10M-token context window — unmatched by any closed model — paired with Meta's in-house chip programme, makes Llama the de facto open infrastructure standard.
Strengths
- 10M-token context (Scout) — largest open-weight
- Mixture-of-Experts: 17B active / 400B total (Maverick)
- Compatible with every major inference stack
- Zero per-token cost when self-hosted
Frictions
- Community licence (not OSI) — 700M MAU clause
- Frontier benchmark gap vs. Claude / Gemini / GPT-5
- Maverick demands multi-GPU enterprise hardware
- Behemoth 2T release schedule slipping
The default "private deployment" stack for regulated, sovereign, and cost-sensitive workloads. Behemoth GA (H2 2026) closes the frontier gap meaningfully — possibly fully.
Copilot is no longer a model — it is a workflow. Microsoft has integrated GPT-5.4 Thinking, GPT-5.3 Instant, Anthropic Claude (via "Critique" and "Council"), and Microsoft's own MAI / Phi family into a single addressable surface across Word, Excel, Outlook, Teams, and GitHub.
Strengths
- 100M+ MAU; Microsoft 365 distribution
- Multi-model orchestration — vendor-agnostic at the surface
- Sora 2 video generation in Frontier programme
- Agent 365 — autonomous task execution rollout
Frictions
- $30 / seat enterprise — high TCO at scale
- No proprietary frontier model of its own
- Dependency on OpenAI commercial terms
- Adoption ramp slower than headline ROI suggests
Wins enterprise productivity by default. Real strategic story: every Microsoft 365 seat becomes an AI seat. Risk is OpenAI-relationship volatility, not model quality.