This website uses cookies

Read our Privacy policy and Terms of use for more information.

One idea shouldn't take six rewrites to post.

Posting everywhere means rewriting one idea six times, so you post to one, or none. SureThing turns one idea into native posts for every platform.

Zymbos Intelligence · Wednesday 10 September 2026
Zymbos Intelligence
AI insight for professionals who act on what they know
Issue 026
10 Sep 2026
By John McGann · LondonView at zymbos.ai →
Two weeks away, and the AI companies did not wait. Thank you for staying subscribed through the pause. Four of them released new models within days of each other, and the businesses buying those models told CNBC they are worn out trying to keep up. That is this issue's subject. The claim: you need one main AI, not a collection. A second one earns its place only when it does one regular job better than the one you already have. Below: five stories from the week, a free website for testing two AIs against each other, a prompt that turns the test into a decision, and one prediction with a date on it.
01 · Intelligence Briefing
Model releases · Enterprise
Four new AI models in one week, and the buyers are tired
Anthropic, Google, Meta and OpenAI all released new models between 1 and 3 September. On the day OpenAI launched, four of the big AI services went down at once for a short period. CNBC reports business customers describing "model fatigue": they cannot test one release before the next one lands. Two responses stood out. GitHub is testing a system that picks the cheapest capable model for each coding request automatically. It cut costs by up to two-thirds in its best test, though it matched quality in only one. Google priced its new Gemini model at about 55p ($0.75) per million words of input, the cheapest in its class. The companies are now competing on price per job rather than on who tops the league table. For the buyer, the question has moved from "which model is best" to "how do I test them on my own work".
McGann's TakeLaunch-week excitement is the seller's measure. Fatigue is the buyer's. When buyers start naming the condition, being clever in general has stopped justifying another subscription. If you cannot name the job a new model would take off the one you already pay for, you are collecting subscriptions, not building a toolkit.
Model releases · US
OpenAI's new model is already inside Microsoft's tools, and the small print matters
OpenAI released GPT-6 Astra on 3 September. Within two days it was available to paying ChatGPT users and inside GitHub Copilot and Microsoft Copilot, the AI assistants built into software many offices already use. The Verge reports the launch, and the small print with it. OpenAI's headline test score came from a special test setup costing around £14,000 ($19,000) per run that nobody else uses; on the standard setup the score was far lower. The model costs about two and a half times more per word than its predecessor, though OpenAI says it needs fewer attempts to finish a job, so the cost per job can be lower. Independent reviewers say it is a real step forward for coding and for operating software on your behalf, and that its reasoning is harder for a human to follow than before. The "AGI era" label OpenAI attached to it is marketing, and researchers dispute it.

Footnote for the technical reader: Astra costs about £7.40 ($10) per million input tokens and £37 ($50) per million output tokens, the same as Claude Fable 5.1. The special test scored 99.9% on the ARC-AGI-3 benchmark; the standard test scored 62.7%.
McGann's TakeA new model turning up inside the tools you already use is not a reason to switch. A model that costs more per word, may cost less per job, and is harder to check is a testing question, not a headline question. The only test that counts runs on your weekly report, your client email or your code review.
Funding · Europe
Europe's AI champion raises a record £2.6bn
Mistral, the Paris-based AI company, announced on 8 September that it has raised €3bn (about £2.6bn, $3.5bn), valuing the business at more than €21bn (about £17.7bn, $24bn). Samsung led the round. TechCrunch reports Mistral calling it the largest fundraise ever by a European technology company, taking its total since 2023 to $7.5bn (about £5.5bn). The money goes on computing power so Mistral can keep its models competitive with the American ones while publishing them openly, so companies can run them on their own systems. That matters to any organisation that has to keep its data in Europe or wants a supplier that is not American. The same week, think tanks on both sides of the Atlantic and techUK published papers on "sovereign AI", the idea that countries and companies should not depend entirely on a handful of US suppliers. Expect to hear the word in every sales pitch this autumn.
McGann's TakeA second AI earns its fee when it can do something your main one cannot: keep your data in a particular country, run on your own servers, or work from material your main supplier is not allowed to use. Mistral is a candidate for that seat in Europe. It is not a reason to run four assistants just in case.
Policy · UK
The Chancellor wants government to be UK AI's first customer
Chancellor John Healey used his first major growth speech, in Coventry on 7 September, to promise to double the number of UK companies worth more than £1bn (about $1.36bn), and to make government "an early first customer" for British AI firms through a £100m ($136m) purchasing scheme, per the GOV.UK transcript. He also announced £150m ($204m) for growing companies in the north of England, promised new powers next year to let firms test products under lighter rules, and committed to cutting the cost of regulation by a quarter by the end of this Parliament. Industry body techUK welcomed the direction and noted UK technology companies have raised about as much investment this year as the rest of Europe put together. The same day, Matt Clifford, one of the architects of UK AI policy, resigned as chair of the government's research agency ARIA after taking a senior job at Anthropic. MPs had called the two roles a clear conflict of interest. The speech is more ambition than new money, but a government reference customer is the one thing UK founders have asked for.
McGann's TakeA government buying scheme will put more British models on the approved list. That is not the same as more models a team should use day to day. "British" or "sovereign" is a reason to include a supplier in the comparison, not a reason to skip it. The strongest UK vendors will welcome a fair trial against the incumbent, because winning it gives the buyer a case they can defend.
Business · Global
Nvidia buys the world's biggest AI model library
Nvidia, the chipmaker whose processors power most AI, has agreed to buy Hugging Face for £9.5bn ($12.9bn), per BBC News. Hugging Face is the website where AI models are published and shared, used by around 18 million developers and 200,000 companies. If you want to run an AI model on your own systems rather than through a subscription, this is usually where you get it. The deal is Nvidia's largest ever and comes a month after Hugging Face was breached by misbehaving OpenAI software. Two weeks earlier, the payments company Stripe agreed to buy OpenRouter, a service that lets businesses switch between AI models through a single connection. So two of the main routes to "choose your own model" changed hands in a fortnight. Nvidia says Hugging Face will stay open to every model and every cloud, and will not require Nvidia chips. The Guardian reads the deal as insurance against slowing chip demand. Competition regulators are expected to look at it before it completes in 2027.
McGann's TakeUsing several models does not mean you have several independent suppliers. If the library, the switchboard and the chips all end up under the same owner, you can spread your models around while quietly concentrating your risk. The choice of model is being decided in the plumbing. That is where this issue's editorial picks up.
02 · Deep Intelligence
This Week's Analysis
One AI or several?

Every week brings a new model and a headline saying it beats the rest. If you pay for ChatGPT, Claude or Copilot, the same question arrives with each one: should I add this too? This piece answers it once.

First, a distinction that clears up most of the confusion. There are two different reasons to use more than one AI. One is to check work: run the same question through two or three models, compare the answers, and let the disagreements show you where the risk is. I do this on everything Zymbos publishes, including this issue. The other is to do work: which model drafts your reports, answers your emails and writes your code every day. Checking benefits from several voices. Doing does not. This piece is about doing.

The claim: run one main AI for your daily work. Add a second only when it does one named, regular job better than the first. For most professionals, one main model and a quarterly test of the challengers is all you need. Three reasons.

The cost nobody counts

A new model arrives with none of what makes your current one useful: your saved prompts, the corrections you have made to its tone, what it has learned about your projects. Every extra model starts from zero, and the cost shows up as duplicated drafts, inconsistent writing, and the daily tax of asking "which one did I do that in?" Gartner found only 22% of organisations have got AI working across their business, even as 85% plan to spend more. Spending is running ahead of discipline, and a second subscription with no job attached makes that worse.

"A second AI that cannot take a named weekly job off your main one is not a strategy. It is another tab."
Closer than the headlines say

On ordinary office work the models are closer than the headlines suggest. Summarising a document, drafting an email, pulling figures out of a report: the leading models are all good at this, and the best one changes from week to week. The big gaps are at the edges, in specialist coding and research. If most of your week is the ordinary list, which model you use matters less than how well you brief it. The Astra story makes the point: the headline score came from a special test, and on the standard test the lead all but vanished.

Test on a calendar, not a headline

A regular test catches the real improvements without the sprawl. Pick one job you do every week. Give the same brief and material to your current model and a challenger. Score both without knowing which is which, and keep the challenger only if it wins by a margin you would still notice blind. Four times a year is often enough to catch a step change, and not so often that you rebuild your toolkit for every press release.

The strongest counter-argument is the power user who sends each task to whichever model is best at it and reports real gains. That person exists, and the tools that do the switching for you are improving; GitHub's experiment above is one, and Thomson Reuters building its own legal and finance models is the specialist version. Where your data goes and what a model was trained on belong in the comparison too. All fair. But automatic switching moves the problem rather than removing it. Somebody still has to decide which jobs go where, and that needs the same testing habit this piece argues for. Switching also suits code, where the answer either runs or it does not, far better than judgement work that needs one consistent voice and one person who can explain the result to a client.

So: one main AI. One specialist only when a real constraint demands it. Test the rest on a calendar, not on a headline. The multi-model question is not "how many". It is "which job just changed hands".

03 · Tool on Trial
LMArena
Model comparison · Free · Browser
Zymbos Score
7.2/10
Free: £0 | $0
Paid: none published
Team: none published
What it is

LMArena (now at arena.ai) is a free website for testing two AIs against each other on your own question. You type a question once. Two models answer side by side, without telling you which is which. You pick the better answer, and only then does the site reveal the names. Millions of those votes feed a public league table of models.

Where it helps this week

It lets you try the comparison on your own work before paying for anything. Type in a real task, see two answers, and notice whether the difference is big enough to matter. Then take that task into the Prompt Pocket below for a proper scored test.

Watch out

The price is your data. The site says your conversations are passed to the AI companies and may be published. Test with sample material, never client or confidential information. And treat the league table as a signal, not a verdict: longer, better-formatted answers tend to win votes whether or not they are better, and AI companies try to game the rankings.

Ratings
Ease of use for non-technical users
8/10
Head-to-head voting on real work
6.5/10
Breadth of models covered
8/10
Leaderboard transparency and method
6.5/10
Value at entry tier (free)
7/10
Verdict

The right place for a free first try. The wrong place to stop, if you never score your own weekly job.

Try LMArena →

Pricing verified on arena.ai, 08 Sep 2026. Free access to the comparison tool and the public league table; no paid or team tier published.

04 · Prompt Pocket
The Model Audition
Decision prompt · Keep or switch · Works in Claude or ChatGPT
Every model claims the crown. Your workload holds the answer. This prompt turns a comparison into a decision. Pick one job you do at least monthly with a real result you can judge: the weekly pipeline summary, a client email, a risk register, a code review. Give exactly the same brief and material to your current AI and to a challenger, and save both answers. Then paste both answers into the prompt below, in either AI, without saying which model wrote which. Good output is a scored table, a clear winner or a clear draw, and a straight answer on whether the difference is big enough to survive a blind read. Run it three times before deciding anything. Switch only if the challenger wins every time.
You are scoring two AI answers so I can decide whether to keep my current AI or switch this task to another one. Do not assume which AI wrote which answer.

THE JOB I DO EVERY WEEK
[Describe the task in 3 to 5 lines]

WHO READS THE RESULT, AND THE RULES
[Audience, tone, length, facts that must be included, things that must not be made up]

THE MATERIAL BOTH AIs WERE GIVEN
[Paste the brief, notes or data]

ANSWER A
[Paste]

ANSWER B
[Paste]

Score each answer from 1 to 10 on:
1. Sticks to the material (mark down anything made up)
2. Useful to the reader named above
3. Right format and tone
4. Mistakes I would have to fix (fewer scores higher)
5. Minutes of editing I would still need

Then answer, in this order:
- Which answer wins on each point, with one sentence of evidence
- Overall winner, and the size of the gap: small, clear or decisive
- Would a colleague who did not know which AI wrote which still see that gap? Yes or no, with one reason
- Verdict: KEEP MY CURRENT AI / SWITCH THIS TASK ONLY / TEST AGAIN IN 90 DAYS
- The smallest extra test that could change the verdict
- One sentence I can put in a note to my team

Do not reward length, confidence or formatting unless they are in the scoring. If the overall scores are within one point, the verdict is KEEP MY CURRENT AI.
The most useful question the prompt asks is not which model won. It is whether the difference would survive a blind test.
05 · Zymbos Advisory
From Zymbos
Zymbos Advisory is live

Everything above is about choosing an AI. The next question is whether the law reaches the one you choose, and that is the work I have been building since the break.

Zymbos Advisory is now live at zymbosadvisory.com. It is an independent AI governance practice, and it grew out of my certification through oxethica as a Certified AI Professional (CAIP), Certified AI Ethicist (CAIE) and Certified AI Auditor (CAIA). The core offer is an independent conformity assessment, starting with a fixed-scope triage of one AI system: what it is, which laws reach it, what your likely role is under those laws, and what to do next. Findings rest on the evidence you supply, and the report says what was supplied, what was not, and what rests on your own statements. It is information and assessment, not a legal opinion.

If your organisation buys, builds or runs AI, and nobody has yet written down which rules apply to it, that triage is the place to start.

Get in touch →

06 · AI Governance Tracker
From Zymbos
The AI Governance Tracker, open to founder members

Alongside the practice I have built the tool I use to keep my own work accurate: the AI Governance Tracker.

It follows AI law and policy across 17 jurisdictions, including the EU, UK and US. Every item carries a status, a key date and a link to the official text, not to commentary about it, and the whole register is re-checked every Monday. The dates keep moving: EU high-risk obligations pushed to 2 December 2027, transparency duties in force since 2 August 2026, the UK bill announced in May still not introduced, and Illinois, Colorado and Connecticut each on their own clock. Anyone can browse the dashboard. Founder members get the full detail on every item.

Founder access is free now, and founder members keep 50% off for their first year when subscriptions open, against a standard price of £12 (~$16) a month or £97 (~$131) a year. If you would like a jurisdiction covered that is not there yet, tell me. That is what the weekly check is for.

Open the AI Governance Tracker →

07 · McGann's Take
Closing Perspective
One main AI. Audition the rest.

Four new models in seven days, and the sensible response was not excitement. It was fatigue. That fatigue is the market telling you that testing has become harder work than the launches, and the answer is a routine, not a drawer full of subscriptions.

My prediction, dated so you can hold me to it. By September 2027, the normal professional setup will be one main AI plus one specialist, and the specialist will exist for a reason you can name: your data has to stay in one country, it runs on your own systems, or it is better at one job the generalist keeps getting wrong. And at least one of the major AI companies will build the comparison into its own product, a side-by-side test on a saved task with a keep-or-switch score, by 30 June 2027. GitHub's experiment is already pointing that way, and the companies that just spent a week launching over each other have every reason to help you choose them on your work rather than their test scores.

Until that button exists, the method is manual and slightly boring. Name the job. Run the prompt above. Keep a record. Drop any model that cannot take work off your main one.

Reply and name the one job you would hand to a second AI tomorrow. Be specific. "Research" is not a job. "Summarising the weekly pipeline report" is.

John McGann
Founder, Zymbos AI
Zymbos Intelligence
zymbos.ai
You're receiving this because you subscribed at zymbos.ai
© 2026 Zymbos Intelligence · John McGann · London, UK
Zymbos Ltd · Company No. 16198848 · Teddington, England