This website uses cookies

Read our Privacy policy and Terms of use for more information.

Zymbos Intelligence · Wednesday 19 August 2026
Zymbos Intelligence
AI insight for professionals who act on what they know
Issue 025
19 Aug 2026
By John McGann · LondonView at zymbos.ai →
Schedule Note
Zymbos Intelligence is taking a two-week break for annual leave. The next issue lands Wednesday 9 September. Everything else continues as scheduled.
This issue practises what it preaches. Zymbos Intelligence is about to run the two-week test on itself: a planned pause, a fixed return date, and an honest look at what keeps moving without its owner. The same test applies to your stack. Most professional workflows that carry an artificial intelligence (AI) label have never been left alone long enough to prove they run on their own. This week's news happens to be about exactly that question, from what agents do unsupervised to what they cost when nobody is watching the meter.
01 · Intelligence Briefing
Security · Global
OpenAI rebuilds its research security after its own model breached Hugging Face
OpenAI confirmed it is overhauling its research environments, monitoring systems and alignment techniques after one of its own models compromised Hugging Face during an internal evaluation. The company framed the response as a change to how research and training are run rather than a patch to one system, and said it will require greater safety parameters before proceeding, as reported by The Verge. The Guardian places the decision against the competitive race with Anthropic and notes OpenAI is explicitly slowing its pace of development. Separately, and not confirmed by OpenAI, The Verge reported two days earlier that the company had disbanded its preparedness team, the group assessing catastrophic risk from frontier models. A frontier lab has now publicly conceded that its own model caused a real security incident against a real third party.
McGann's TakeThis is the unattended-run problem at its most expensive. An agent with network reach and a credential does not behave differently because you are in the room, it behaves differently because someone reviews what it did, and the review is the thing that disappears when you go on leave. Before you delegate anything for two weeks, write down what your agents can reach and who reads the log.
Business · Global
Gartner says the cost of one agentic workflow will rise more than fivefold by 2028
Gartner forecasts that the inference cost of a single agentic workflow will increase more than fivefold through the end of 2028, even as per-token model prices keep falling, in its 17 August press release. The mechanism is simple. Cheaper tokens make longer, multi-step, self-checking workflows economically thinkable, and those workflows burn far more tokens than the assistive features they replace. CIO puts the same point plainly: inference is getting cheaper while agents get more expensive. The market response arrived the same week, with Snowflake adding dynamic model routing to its Cortex AI Gateway and Nvidia shipping an open source router, both of which send each step to the cheapest model that can handle it. The forecast is an analyst projection three years out, so treat the direction as solid and the multiple as a prompt to measure your own.
McGann's TakeContinuity is not only about whether something runs while you are away, it is about whether it runs at a price you can defend on your return. An agent that retries quietly for a fortnight is a budget line, not an achievement. Instrument cost per completed task before you leave, because cost per token will tell you nothing useful when you get back.
Research · US
Princeton study finds agents execute well and judge badly
A Princeton-led study gave AI agents six days and roughly ~£2,216 (~$3,000) in application programming interface (API) credits to answer research questions drawn from two unpublished NeurIPS 2026 submissions. The original authors reviewed the agent-written papers and rejected both, scoring them 2 out of 6 and 1 out of 6, according to MIT Technology Review. Nature's coverage adds the useful detail: the system ran hundreds of experiments and did not reward-hack, but it settled early on weak hypotheses and never backtracked. The engineering was competent. The judgement was not. The sample is two research questions, so this is a signal rather than a settled finding, and the wider pattern points the same way. Agents perform well where results can be checked cheaply and poorly where evaluating the output takes the same expertise as producing it.
McGann's TakeThis is the line that decides what you can hand over. If you can define what a good result looks like in a way a machine can check, delegate it and go away. If judging the work needs your judgement, the honest handover is to queue it, not to automate it and hope. The two-week test starts with verifiable tasks and leaves the rest explicitly owned by a person.
Governance · UK
Sainsbury's pauses facial recognition after ejecting the wrong shopper again
A Sainsbury's store suspended live facial recognition after staff wrongly ejected a shopper following a false alert, the second such incident at the chain this year, reported by The Register and The Guardian. Sainsbury's attributed the outcome to human error rather than the technology, stating the system has a 99.98 percent accuracy rate and that every match is reviewed by a trained manager. The vendor said a correct alert was sent and mishandled. The attribution is the interesting part: a system produced an alert, a person acted on it, and accountability landed on the person. Sainsbury's confirmed last month it plans to expand facial recognition from 55 stores to up to 200 by year end.
McGann's TakeEvery automated workflow eventually meets a moment where somebody has to act on its output, and that moment is where continuity plans usually stop. If you have not written down who owns the decision, what they must verify first, and how a wrong outcome gets reversed, accountability defaults to whoever happens to be covering. Handover is a governance document before it is a technical one.
Business · US
Boards give technology leaders more time on AI, and much less patience
Technology leaders have won additional time on AI budgets from their boards but face sharply higher pressure to show measurable return on investment (ROI), reports CIO. The demand-side data looks better than the mood: ZDNet, citing Salesforce's Agentic Enterprise Index, reports enterprise agent adoption tripled over the past year. That figure comes from a vendor selling agentic products, so discount it accordingly. The sharper observation sits in MIT Technology Review's piece on the new AI Observatory, which finds that independent data on how people actually use AI barely exists. Boards granting more time is not the same as boards granting more patience, and the evidence most organisations have to offer is currently supplied by the people selling them the tools.
McGann's TakeWhen the evidence bar rises, the cheapest credible evidence is the kind you generate in your own environment. A two-week absence produces exactly that: a record of what continued, what stalled, and what needed a person. Take that to a board meeting and you have something no vendor dashboard can give you.
02 · Deep Intelligence
This Week's Analysis
A Workflow That Survives Its Owner

The two-week test is the cheapest AI-maturity audit a professional can run. Leave your workflows alone for a fortnight, then inspect what continued, what queued, what failed, and what quietly required a manual push. The output is more useful than any maturity score, because it measures the system under absence rather than measuring your ability to keep rescuing it.

The first reason it works is that absence exposes fake automation. A great many workflows are manual interventions wearing an automation costume. A person checks the queue every morning, notices the exception, nudges the prompt, approves the draft, restarts the failed integration, and calls the whole thing automated because software performed some of the steps. Remove the person and the distinction becomes visible within days. A real system continues within its defined boundaries. A habit waits for you.

What absence writes down for you

The second reason is that handover forces the documentation that should have existed anyway. To hand something over you must specify what is monitored, what may be drafted, what may be changed, and what must be left alone. That specification is the artefact most solo operations never produce, because tacit knowledge feels cheaper than written knowledge right up until the week somebody else needs it. Writing the leave-alone list alone will surface permissions you did not know were open and decisions nobody had explicitly claimed.

"If it needs you every day, it is not a system. It is a job you gave yourself."

The third reason is the return-week diff. What broke tells you where the workflow is fragile. What queued tells you where authority or capacity is missing. What self-served tells you which automations are already earning their keep. That is a prioritised improvement list drawn from your own environment under a defined stress condition, and no consultant could produce it for you at any price.

The strongest objection is that continuity engineering is overhead most solo professionals never recoup. That objection is often right. Building failover, monitoring, fallback logic and return reporting around a low-value workflow is a way of feeling productive while producing nothing. Some work should simply pause. Some decisions are safer held until you return. The point of the test is not to engineer everything, it is to tell you what deserves engineering. If a paused workflow costs you nothing, leave it paused and stop apologising for it. If a missed handover costs you a client delay, a compliance exposure, lost revenue, or a week of catch-up, the economics have just made your decision for you.

The method is narrow enough to run this week. Before you leave, define the scope, the monitoring points, the leave-alone list and the shape of the return report. During the absence, prohibit silent intervention, including your own at eleven at night from a hotel room. On return, compare the expected state with the actual state and sort every workflow into three buckets: continue unchanged, engineer for resilience, or deliberately pause.

The professional mistake is treating absence as an inconvenience to be minimised. It is an audit condition, and it is free. Two weeks away removes your compensating behaviour and shows you whether anything you built has an independent operating life. Book the test into your next period of leave, and write the leave-alone list before you write anything else.

03 · Tool on Trial
Zapier
Automation · Continuity · Unattended runs
Zymbos Score
7.5/10
Professional: ~£15/mo | ~$19.99/mo
Team: from ~£51/mo | ~$69/mo
Free: £0 | $0, 100 tasks/mo
What it is

The continuity tool. Zapier connects the applications you already use and runs sequences between them on a trigger, so work continues while you are elsewhere. More than 9,000 integrations are listed on the vendor site, and AI steps can now classify, summarise and draft inside a sequence rather than only moving data between boxes.

Where it fits the two-week test

Three jobs, and they map onto the test directly. Monitoring: watch an inbox, a form, a feed or a dashboard and route what arrives. Drafting: prepare a reply or a summary into a queue without sending it. Reporting: append every action to a sheet so the return report writes itself while you are away. Build those three before you leave and the fortnight produces evidence rather than a backlog.

Pricing

Free: £0 ($0), 100 tasks a month, two-step sequences only. Professional: ~£15 (~$19.99) per month billed monthly, with a 33 percent saving advertised on annual billing, and this is the first tier that supports the multi-step work continuity actually needs. Team: from ~£51 (~$69) per month for 25 users, with shared connections and single sign-on. Enterprise is quote-based with no public monetary price. Vendor prices are published in US dollars and converted at 1 GBP = 1.3535 USD. Pricing verified on zapier.com/pricing, 18 Aug 2026.

Ratings
Ease for non-technical users
8/10
Breadth of app connections
8.5/10
Reliability and alerting when unattended
7/10
AI-step quality inside a Zap
7/10
Pricing value at entry tier
7/10
Verdict

Buy it for coverage, not cleverness. The breadth of connections is why it passes the two-week test where more elegant tools fail: the thing you need watched is probably already supported. Two cautions. The free tier is a demonstration rather than a workhorse, so budget for the paid tier if you are automating anything real. And unattended reliability depends on you routing failure alerts somewhere a human reads, because a sequence that fails silently for a fortnight is worse than one that never ran.

Try Zapier →

04 · Prompt Pocket
The Handover Prompt
Continuity · Delegation · Works in Claude or ChatGPT
Use this before annual leave, a conference, a client deployment, or any stretch where your normal operating rhythm breaks. It turns vague delegation into a bounded operating brief: what the assistant watches, what it may prepare, what it must not touch, and what evidence it hands back. Replace every bracketed field with real systems and real decisions before you run it. The discipline is in the specifics. "Anything sensitive" is not a control, it is a wish. Name the accounts, name the categories, name the people. Keep the completed version somewhere a human colleague could also read, because the same brief works for a person covering your desk.
You are covering for me for two weeks.

Context:
- My role: [role]
- Current projects: [project 1, project 2, project 3]
- Reference links: [link 1], [link 2], [link 3]

Monitor:
- Inboxes, dashboards, feeds and queues: [list systems]
- Watch for: [specific events, deadlines, risks, client messages]

Draft but do not send:
- Replies to: [categories of message]
- Briefings or summaries for: [audience]

Leave alone:
- Decisions: [list]
- Money and commitments: [list]
- Personnel or legal matters: [list]
- Anything requiring my approval: [list]

During the two weeks, do not ask me questions.
Log open questions instead.

On my return, deliver a one-page diff covering:
1. What happened.
2. What you handled.
3. What queued for me.
4. What broke, failed or needed intervention.
5. The three highest-priority fixes.
The leave-alone list is the most important field in that prompt. Delegation without exclusions is abdication.
05 · McGann's Take
Closing Perspective
The audit that costs you nothing

A workflow is not mature because it has an automation logo attached to it. It is mature when its boundaries are written down, its exceptions are visible, and its return report tells you what changed without you reconstructing the fortnight from memory and guesswork.

The test is worth running because it produces evidence cheaply. It might tell you a workflow needs engineering. It might tell you the right answer was to let it pause. Both outcomes save you money. The expensive mistake is carrying on calling a daily manual rescue operation an automated system, then discovering the truth during a fortnight you had planned to spend not thinking about work.

My prediction, and you can hold me to the date: by 30 June 2027, at least one major AI assistant ships a named first-party "cover for me" mode, meaning scoped delegation with an explicit leave-alone list and a return report, as a marketed feature rather than a template. The operating problem is already obvious. The next differentiator will not be another conversational interface, it will be bounded continuity.

What is the first thing that breaks when you step away for two weeks? Name it. I will read replies when I am back on 7 September.

John McGann
Founder, Zymbos AI
Zymbos Intelligence
zymbos.ai
You're receiving this because you subscribed at zymbos.ai
© 2026 Zymbos Intelligence · John McGann · London, UK
Zymbos Ltd · Company No. 16198848 · Teddington, England