Insights
Which AI is right for my company?
There is no single best AI provider, and anyone who tells you otherwise is selling something. The right choice turns on where your work already lives, how regulated your data is, and whether you are buying seats for staff or building on an API. What follows is a cost comparison written for the person who has to defend the spend, then a read on which provider suits which industry, because a manufacturer and a bank should not land in the same place.
Two buying models sit underneath all of it. Per-seat plans put an AI assistant inside the tools your employees already use, for a flat monthly fee. API access lets you build AI into your own products and workflows, billed by the token of text processed. Most companies end up using both.
API pricing: building on the models
If you are embedding AI into your own systems, you pay per million tokens (roughly 750,000 words), charged separately for input and output. These are list prices as of July 2026, in USD.
| Provider | Flagship model | Mid-tier workhorse | Budget tier |
|---|---|---|---|
| Anthropic | Claude Opus 4.8 ($5 / $25) | Claude Sonnet 5 ($3 / $15; intro $2 / $10 to 31 August 2026) | Claude Haiku 4.5 ($1 / $5) |
| OpenAI | GPT-5.6 Sol ($5 / $30) | GPT-5.6 Terra ($2.50 / $15) | GPT-5.6 Luna ($1 / $6) |
| Gemini 3.1 Pro ($2 / $12) | Gemini 3.5 Flash ($1.50 / $9) | Gemini 3.1 Flash-Lite ($0.25 / $1.50) | |
| Microsoft | First-party MAI family (MAI-Thinking-1 for reasoning, a GitHub Copilot coding model, plus image, voice and transcription), launched at Build 2026, alongside OpenAI's models, both via Azure AI Foundry and Azure OpenAI | Priced through Azure AI Foundry | Confirm current per-token rates there |
(Prices are input / output per million tokens.)
A cost owner should take a few things from this table. The flagship tiers cluster close to $5 for input, so the real spread sits at the bottom of the range, where Google's Flash-Lite is far cheaper for high-volume, low-complexity work. Output runs three to six times the input price across every provider, so the models that produce more text cost more than the headline rate suggests. The largest saving, though, does not come from picking a provider. It comes from routing: cheap models for extraction and classification, flagship models held back for real reasoning. Batch processing at around half price and prompt caching at up to 90 percent off on repeated context reduce the bill further, and both work whichever provider you choose.
Per-seat pricing: putting AI in front of staff
If you are rolling AI out to employees inside the apps they already use, you buy seats. List prices, USD, per user per month, July 2026.
| Provider | Business or team seat | Enterprise | The catch |
|---|---|---|---|
| Microsoft 365 Copilot | Copilot Business around $21 (SMB) | $30 add-on | Requires a qualifying M365 base licence (E3 around $39, E5 around $60 after the July 2026 increase), so the true all-in cost is roughly $69 to $90 per seat |
| ChatGPT | Business $25 (annual) or $30 monthly | Enterprise, custom (often around $60) | Enterprise pricing is unpublished and usage-metered |
| Claude | Team $25 (annual), five-seat minimum | Enterprise, custom | Team caps at 150 seats; Enterprise usage is billed at API rates on top |
| Gemini | Bundled into Google Workspace (from around $14 a seat) | Gemini Enterprise add-on around $30 to $36 | Cheapest entry if you already run Workspace; the agentic Enterprise platform is usage-priced |
These four do not price the same way, which is the first thing to hold onto. Copilot layers a flat fee onto a Microsoft subscription you must already hold. Claude and ChatGPT increasingly charge for a seat and then meter usage on top. Gemini comes close to free if you already live in Google Workspace. Which provider counts as cheapest is therefore decided mostly by the platform you already pay for, and the AI line item is often smaller than the platform cost sitting underneath it.
Which provider suits which industry
For a manufacturer, the sensitive asset is intellectual property, the designs, specifications and process know-how, more than regulated personal data. Most manufacturers already run Microsoft 365 across the estate, from shop floor to office, which makes Microsoft 365 Copilot the least disruptive route to staff productivity and Azure OpenAI a natural home for document-heavy work such as spec review, RFQ processing and maintenance manuals, inside an environment already cleared for compliance. Where the job is processing long technical documents in volume, Google's Gemini models are worth a trial, since the large context window and cheaper Flash tier suit high-page-count, lower-complexity extraction. Keep data-residency controls tight so the IP stays in-region.
For a financial services firm, the governance requirements settle the choice before cost enters the conversation. You need contractual training exclusion, a DPA, configurable or zero data retention, audit logging and, often, data residency, which places the decision firmly in the enterprise tiers. ChatGPT Enterprise, Claude Enterprise and Azure OpenAI Service each offer the compliance posture a regulated firm has to evidence, including SOC 2, GDPR alignment and zero-data-retention options. Vertex AI is a strong choice where IP indemnification and tight Google Cloud governance matter. Lead with reasoning quality on the analytical work, benchmarking Claude against the flagship OpenAI and Gemini Pro models, and treat per-token cost as secondary to the audit trail.
For a software company, you will build on APIs, so this becomes a question of benchmarks and cost rather than seats. Claude (Opus 4.8 and Sonnet 5) and OpenAI's GPT-5.6 family lead on coding and agentic work, and they are the two to test head to head on your own codebase. Plan to run a routing strategy, with a cheap model such as Haiku, Gemini Flash-Lite or GPT-5.6 Luna for bulk extraction and a flagship reserved for the hard problems, and lean on caching and batch discounts, which will shape your bill more than the choice of provider. For day-to-day developer productivity, the coding assistants attached to these models are where the early gains show up.
Where this leaves you
Match the provider to the platform you already run and the risk you already carry, rather than to a benchmark leaderboard. A manufacturer on Microsoft, a bank that must evidence every control, and a software firm tuning a token bill are three separate decisions, and each of the three usually ends up using more than one provider, routed by task. The cost tables set the floor. The architecture you build on top is what decides the final bill.
Choosing providers, designing the routing that keeps cost down, and putting the governance in place is the core of Firestarter's six-week accelerator, so the decision rests on your workload and your risk profile, with the numbers to support it.