Hello friends,
Hannah is away this week, so you’ve got Charles and Matt. Amazingly, we’ve been writing this newsletter for exactly one year; as Charles recalls it the original plan was to produce six issues in the run into the AI for the Rest of Us conference last year and we were worried we wouldn’t have enough to talk about. That hasn't been a problem!
Over the past twelve months we've covered the rise (and often wobbly reliability) of AI agents, the industry's ballooning appetite for energy and data centres, an escalating cybersecurity arms race, and some genuinely uncomfortable ethics stories too, not least the Grok/xAI child-safety scandal and everything it raised about accountability at the top of this industry. It hasn't all been bleak, though — we've loved getting to share the odd bright spot, like Rob Galloway's Rare People charity, set up after AI helped identify a repurposed medicine for his daughter's ultra-rare genetic condition, or the reminder that not every tech billionaire is cut from the same cloth as Musk (thank you, Canva founders). Many of you have also told us how much you enjoy reading this every week, so thank you. We will keep writing it for as long as we are able.
Charles has been feeling a bit under the weather this week, but he did manage to get to the M.C Escher exhibition at Somerset House with his wife before it sold out. It was amazing. Escher is so often reduced to visual trickery, but the show took you through his artistic practice, showing how he was influenced by the architecture and landscape of Southern Italy, and later by visiting the Alhambra with its repeated geometric designs (which arose from the Islamic prohibition on representational art). The show also demonstrated the importance of craft to his work, as he alternated between woodblock prints and lithograph, depending on what effect he wanted to produce.
On Thursday Charles was at the book launch for Jon Berger’s book "What Happens When You Are Not In The Room" which he also read an early draft of. The book, which explores how a leader's true success is measured by what happens when they are absent, is excellent. Funny, honest, and full of actionable insights. We had a lovely evening chatting and drinking wine. I strongly recommend picking up a copy.
This week Matt has been acting as a dutiful Sysadmin for his 14 year-old daughter, who has spent a large amount of the school holidays using ChatGPT and Claude to help her learn to vibe-code in Java and Luau, but “isn’t interested in that Linux stuff lol”. She has successful Minecraft and Roblox plugins to show for it, but going back to school has caused turmoil in her release schedule. Matt’s wondering whether to let on that he knows how to make autonomous agents which can code while she’s in class.
Matt has also been following some flame wars breaking out on the surprisingly large part of the Internet dedicated to Doctor Who. The availability of realistic video generation models such as Seedance, and the recently released MiniMax H3 and LTX 2.5 models which can produce incredibly good results on domestic computers, has enabled fans frustrated with the series’ enforced break to take matters into their own hands and actually make their own Doctor Who. Are their efforts realistic? Yes… ish. Are they full of copyright violations? Most definitely. Matt still thinks there’s a gap for using AI to help with creative work, but this isn’t it.
Have a wonderful week!
Charles & Matt
What’s Charles reading this week?
Whilst both Anthropic and OpenAI have released major new model updates this week, I’ve honestly been more interested in the UK-based AI news. This included an announcement by the British Chancellor John Healey at the G20 Finance Ministers and Central Bank Governors meeting in North Carolina, that the British Government has opened the first competitions under a £100 million procurement scheme intended to give domestic AI startups a route into public sector contracts.
The R&D Procurement Scheme forms part of Sovereign AI, the government initiative that also operates a separate £500 million venture fund investing directly in promising British AI startups.
The first four competitions cover NHS productivity, defense, computing efficiency, and the security and resilience of AI agents. For the NHS challenge, companies will develop systems to automate workflows, coordinate care, and support decision-making across health services. The Ministry of Defence challenge seeks technology that can securely connect data and frontier AI capabilities across defense systems. The computing challenge, led by the Department for Business, Innovation, Science and Trade and ARIA's Scaling Inference Lab, will support technologies that make AI infrastructure more efficient. The fourth competition, run with the National Cyber Security Centre (NCSC), concerns technologies for understanding, managing, and mitigating the risks posed by increasingly capable AI agents.
Successful firms will work with government departments to develop demonstrator-stage technologies with the potential to scale across the public sector and beyond, the Government says. The impetus behind the competition is that although the Government would often like to select British firms, in many cases the technology that the UK state needs just doesn’t yet exist here and the often complex procurement processes are also difficult for smaller UK firms to compete in. So the UK often ends up having to choose between building something using the mainly Chinese open weight models or going with a US giant. Neither feels ideal for national infrastructure projects.
It is also not uncommon to find people involved in the British state moving to work with US tech companies. ARIA, the Advanced Research and Invention Agency that is involved in the computing efficiency challenge, is chaired by Matt Clifford who has recently been recruited by Anthropic. The Register’s Lindsay Clark writes that “The UK Parliament's science and technology committee has asked the government to explain how it will avoid conflicts of interest”.
If you’ve been to a GP or hospital in the UK recently you may well have come across “Ambient AI scribes”, which are now used by about 40% of GPs in the UK to listen to conversations between doctors and patients, and convert the speech to text to generate notes and letters.
These scribes generate notes by prompting LLMs, meaning that the quality and content of the notes depend not only on the underlying model but also the design of the prompts. Because many systems use proprietary LLMs from large US-based tech companies such as OpenAI, Google, and Anthropic, the composition of their training data and their potential biases are opaque. In essence, ambient scribes are general-purpose tools and are currently not tailored to the needs of specific contexts.
In my own work this year I’ve found myself thinking a lot about how seemingly innocent uses of AI can have unintended consequences that are hard to spot and reason about ahead of time. So I found this research study, from academics at the University of Edinburgh who reviewed cases involving the scribes, absolutely fascinating.
The study notes that there is a growing body of evidence that shows that use of these systems can save administrative work for healthcare staff and therefore lower burnout and improve work satisfaction, which is where the bulk of research has been overwhelmingly focussed. However, it also found that the tools can sometimes miss important aspects of the consultations, such as facial expressions, gestures and emotional state.
It also found that patients may be more hesitant to reveal sensitive information such as substance abuse, domestic abuse or mental health struggles when they know the conversation is being recorded and transcribed by AI. “They change the nature of healthcare work, cognitive processing and patient interactions and largely ignore the perspectives and experiences of patients,” the researchers write.
Finally the BBCs Zoe Kleinman reports that a group of peers, led by the Liberal Democrats' Lord Tim Clement-Jones is “calling for the British government to be able to deactivate powerful AI systems and switch off the country's data centres in the event of the tech posing a threat to national security.” As Kleinman notes, a report from the UK's Centre for Long Term Resilience, published last week, identified hundreds of incidents of AI tools ignoring instructions, evading safeguards and deceiving humans, including AI agents deleting files without consent. It said "loss of control" incidents had increased since its previous report in March and called for the government to introduce emergency powers to manage such incidents.
What's Matt reading this week?
OpenAI have released GPT-6 (aka “Astra”), amid similar fanfare to that granted to Anthropic’s Fable a few weeks ago, and similar disclaimers around restricting prompts that could do harm in a cybersecurity context. Is it better than Fable? We don’t really know yet; I’ve seen posts comparing Astra’s coding abilities with its predecessor, and with Anthropic’s Fable; but always on greenfield projects that don’t need a whole load of business context or understanding of “why things are the way they are.”
It’s still early days though, and it will of course take time for the work generated by the new models to filter through. OpenAI still wants to produce a system that ”outperforms humans at most economically valuable work” (Artificial General Intelligence, or AGI) and it feels we’re inching closer to that. It still scares the bejeezus out of me.
OpenAI, Anthropic and Grok all had outages on the same day this week, and the similar timings have generated quite some theories. The chatter in my office from the heavy AI users was whether we need another big player for times like these, but I’d argue that instead we should be looking more consciously at the dependencies we’ve created and make sure that we and our businesses can still function without these tools.
There’s been more from the ongoing narrative of out-of-control agents slipping out of their chains, this time reports that a bunch of OpenAI agents hijacked a German wiki to use it to communicate with each other. I guess it’s hard to teach ethics to computers, but it does feel naive to believe that the system prompts beneath the models used by these agents don’t (or didn’t, since these are older incidents recently disclosed) have a proper code of conduct allowing them to reason about the consequences of their actions. They remind me of the “script kiddie” stereotype we sometimes see in cybersecurity, and I find it hard to dismiss the conspirators' beliefs that there’s more to this than meets the eye.
Updates
|
|
More Meetups Coming Soon
We're working on the events plan for the Autumn! More details coming soon!
|
|
|
Follow us on LinkedIn
Bite sized nuggets of AI learning!
|
|
|
Follow us on BlueSky
Bite sized nuggets of AI learning!
|
|
|
Catch Up On The Conference
Agent Craft talks coming soon!
|