Hello friends,
When the sun is shining in London it isn’t uncommon to see pubs and beer gardens start to fill up from the moment they open at 12pm, sometimes earlier. There’s nothing we like more than a cold pint in the sunshine. For event organisers this means we have stiff competition, do you stay inside and listen to a talk about AI or do you join the masses in the beer garden for a sundowner?
This month you don’t need to choose because AI for the rest of us is heading to the beer garden for our summer social on Wednesday July 22nd. This is not a traditional meetup with talks, it’s a gathering for the community to celebrate how much we’ve grown together. You can grab your spot for our summer social here.
Thank you to our sponsor Rezonant for making this happen!
Have a lovely week!
Hannah & Charles
|
|
Teach your agents when to say no
Enforceable boundaries around every AI agent. The autonomy you can't fully predict becomes autonomy you can trust — because it knows where to stop. Trust infrastructure for AI
|
What’s Hannah reading this week?
I’m not sure when I became the sort of person that tracks model release dates more reliably than her friends' birthdays but here we are. Late on Thursday evening my partner and I excitedly took GPT-5.6 Sol out for a spin and it was impressively fast even in light mode.
(Yes, I am slightly ashamed, I will do better at the birthdays.)
This week OpenAI released GPT-5.6, a collection of 3 models, each with a selection of operating modes to give them some extra grr on challenging tasks. Looking at the benchmarks both Sol and Terra are serious challengers to Fable 5.
As the team at Vellum AI have observed, the benchmarks are so impressive for Terra that Sol need only be used for the most challenging work especially the Ultra mode. The cheat sheet:
- Sol: Save for the most challenging work
- Terra: The default choice for coding, cybersecurity, agentic and knowledge work
- Luna: Use it on the simple stuff, repeatable tasks, not high context reasoning
Efficiency and speed were emphasised over and over again in the announcement on Thursday with some impressive numbers. Charles has shared a wonderful explainer about why LLMs are so inefficient at generating tokens so these successes in efficiency should indeed be celebrated.
With Fable 5 moving to a metered “pay as you go” system I was interested to see a price comparison, especially given the speed and efficiency claims from OpenAI.
API pricing per million tokens:
- Sol: $5 input / $30 output
- Terra: $2.50 input / $15 output
- Luna: $1 input / $6 output
- Claude Fable 5: $10 input / $50 output
You could set fire to a lot of money this way if usage of Fable and Sol isn’t carefully controlled… How’s that tokenmaxxing going now? Hmm.
In the world of cybersecurity we’re still very much in the early stages of the Vulnpocalypse. The Athena initiative spearheaded by Chainguard shared their first update, with over 40,000 new vulnerabilities discovered and processed in the first 3 weeks. These vulnerabilities are not yet disclosed, they do not have CVE numbers, your scanners do not know about them yet… but they are real. As John Morello, CTO & Co-Founder of Minimus, shared this week - you are already vulnerable. The best advice for teams today is not only to prepare to patch, it’s that you need to reinforce your defences today, you are vulnerable today.
The Linux Foundation recently announced the Akrites Project to take on a larger role in the triage, deduplication, coordination and fixing of the endless deluge of new vulnerabilities. This group will also take on the role of “Maintainer of Last Resort” should the existing maintainers of a critical open source project be unreachable or unable to act upon all of the fixes. It's really great to see tech giants collaborating like this, in the open, and a vendor neutral home to act as custodian for important fixes.
In an open letter the founding members of Arkrites share their rallying cry:
We all depend on open source, and we will all defend it together
What's Charles reading this week?
One of the joys of writing this newsletter is the fact that we’re lucky enough to have an incredibly wide readership, including people who are experts on AI and people who are just getting to grips with it all. If you are in the latter camp and have found yourself thinking, “What is a token?”, “Why are they getting so expensive?” and “Why are LLMs so incredibly inefficient at generating them?”, or just really want to understand how an LLM works, this video from Dr. Mike Pound on Computerphile, a YouTube channel produced by the University of Nottingham, is the clearest explanation I’ve ever seen, and an excellent way to spend 25 minutes.
The UK Government continues to explore different ways that AI might be helpful to deliver public services. One recent example is Gov Chat, which launched publicly in May of this year as a chatbot that uses the information on the GOV.UK website as its source. “Instead of navigating or searching 80,000 pages of guidance, users can now use the app to ask a question and get a conversational response, grounded in official government information and remaining firmly rooted in the principles of clarity, accuracy and trust,” Deputy Director for GOV.UK AI Shelina Hargrove wrote in May. If it works well, I think this kind of collapsing of information is a potentially great use of LLMs.
Another good use case, which caught the eye of loyal reader Krista this week, is from NHS England, part of the UK’s publicly funded healthcare system, who are rolling out an AI triage tool on the generally excellent NHS app. The AI asks patients appropriate questions and then routes them to a GP, pharmacy, A&E, community service, or self-care advice. It is part of a £10bn tech overhaul, reaching 200,000+ patients within 12 months and all app users by April 2028. The AI won't decide whether patients see a doctor, but is meant to route people faster while clinicians retain judgement. A trial at a Sussex GP practice cut phone queuing by 29%.
Alongside it, NHS trusts are expanding AI notetaking tools that transcribe patient-staff conversations into clinical summaries, starting with four London-area trusts plus existing programmes at Alder Hey and Manchester University trusts. A Great Ormond Street-led trial found staff spent almost 25% more time with patients when using the tool.
But whilst using AI to help doctors and patients seems to me to be a great use of the technology, I’m rather less convinced about the use of LLM by students to “help” them. Admittedly I’m not a teacher, but it seems to me the risk is that the students miss out on both subject learning and acquiring the skills that getting a degree teaches. It's complicated though, as the universities are grappling with the fact that AI tools exist, students will use them (and will likely be required to if they shift from academia to work), and tools for AI detection are unreliable.
Emma Whitford, a journalist for Inside Higher Ed, reports on Brown university economics professor Roberto Serrano, who gave a take-home midterm exam for the first time in nearly two decades after a campus shooting, and got an average score of 96 percent, way above his usual 65-80 percent range. Suspicious, he ran the questions through ChatGPT himself and found his students' answers mirrored the AI's almost exactly, down to an oddly convoluted proof technique. He switched the final to in-person. Eighteen students dropped, nine skipped it entirely, and the average fell to 48.6 percent. In response Serrano voided the midterm, reweighted the final to 80 percent of the grade, and 19 students failed the course.
Brown's own AI-in-teaching committee released a report the same week showing three-quarters of faculty worry about AI cheating, but its recommendation leans away from enforcement. Serrano's take is that "We cannot afford to have a society in which a significant fraction of our best young minds think that cheating is OK”. I’m not sure if we have any teachers or lectures amongst our readership who are grappling with these issues, but FWIW as someone who is neither I thought these guidelines from MIT were a good starting point.
Way back in November we reported on a keynote at Øredev from Simon Wardley of Wardley Mapping fame, in which he suggested that LLMs are a non-kinetic form of warfare designed to embed their makers’ values into the decision making processes of a wider community. Academics from the Oxford Internet Institute and the Hasso Plattner Institute have studied the effect Wardley mentioned, and found that large language models (LLMs) consistently changed the direction of social media posts on contested topics, even when explicitly instructed to preserve the original meaning. The researchers also show, through simulations of real-world social networks, how these small changes could accumulate across millions of interactions and gradually influence broader public opinion.
Updates
|
|
Summer Social
Join us on July 22nd for our annual summer party!
|
|
|
Follow us on LinkedIn
Bite sized nuggets of AI learning!
|
|
|
Follow us on BlueSky
Bite sized nuggets of AI learning!
|
|
|
Catch Up On The Conference
Agent Craft talks coming soon!
|