2026 in LLMs (so far)
On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological exploration of everything that happened in 2026. The video is on YouTube ; here are my annotated slides and notes to accompany the talk.
And as an annotated presentation :
I'm going to give a lightning tour of everything that has happened so far in 2026. The year isn't over yet!
For me, 2026 started a couple of months earlier in November 2025.
November saw the release of two important models: Claude Opus 4.5 and GPT-5.1.
As is usually the case with new models, these were incremental improvements on the models that came before them.
But every now and then when a model improves, it crosses an invisible line where something that didn't really work starts working.
In this case, the thing that started working was their coding agents. Claude Code had been around since February 2025, Codex was a little younger.
These two new models, when paired with their respective coding agent harnesses, improved from "often make mistakes" to "reliable enough to use on a day-to-day basis".
For a couple of years now I've been evaluating new models by asking them to "Generate an SVG of a pelican riding a bicycle". It's probably the world's stupidest benchmark - there's only so much you can learn from it.
But it's still a challenge for models, because drawing pelicans is difficult, drawing bicycles is difficult, and pelicans can't ride bicycles in the first place.
Here's the state of the art for November. Claude still couldn't really draw a bicycle! The GPT-5.1 bicycle frame is pretty crap too.
Also in November, we had the first commit to an obscure GitHub repository called "Warelay". We'll come back to this repository shortly.
An then there were the December holidays, and individual developers took some time off and many started tinkering with these new coding agent model combinations... and it began to dawn on us quite how much they could do that they couldn't do before.
Come January, a lot of us were quite excited to start putting this stuff into action.
Every year I set myself a New Year's resolution, and for as long as I can remember it's been the same thing: stay focused. Take on less new projects. Try to get things done in the projects I already have.
This year I decided that since that had never worked before, I'm going to go the other way.
We've got coding agents now, let's see what they can do. I'm going to take on as many new projects as I like!
(You can ask me at the end of the year if this turned out to be a good idea or not. I have a lot of plates spinning right now.)
"Be more ambitious" has been something of a theme for the year, because the only way to find the limits of this technology is to keep on pushing them until they don't work.
I also went on the Oxide and friends podcast with Bryan Cantrill and Adam Leventhal to share predictions for the next year (and three and six years).
With hindsight, my LLM predictions were pretty unambitious.
I said "it will become undeniable that LLMs write good code" - I think we're there now.
I predicted we would finally solve sandboxing. I counted and around 40 of the 277 sessions at this conference touched on sandboxing or agent security in some way, so we're at least putting a lot of effort into that!
I predicted "a Challenger disaster" for coding agent security. There's certainly been a whole lot of noise around agent security this year, though the exact disaster I predicted (with coding agents being hijacked and causing real-world economic damage) hasn't really played out.
We threw in a joke prediction that the Pope would weigh in on the economic impact of LLMs.
I also predicted that New Zealand's Kākāpō parrots would have an outstanding breeding season this year.
These birds live in New Zealand. They are flightless nocturnal parrots. They're kind of dumpy looking, I think they're beautiful, and there were only 236 of these parrots in the world at the start of the year.
Kākāpō only breed when the Rimu trees have a big fruiting season, and that hasn't happened in four years... but this year the Rimu fruit were looking excellent.
Photo by Kimberley Collins .
Also on that podcast, we coined a term (full credit to Adam) for "that feeling of Al induced ennui where software engineers get listless because the Al can do anything".
We called it Deep Blue .
This has been a major theme throughout the year, and was touched on by several speakers at this conference.
As a software engineer, I've never had a year of my career where everything has changed so quickly and so dramatically.
A lot of what I've been doing this year is trying to come to terms with that and what that means for my own profession.
Also in January, I suffered from what I'm calling AI mania .
This is not the same thing as AI psychosis .
With AI mania, any time your agent isn't building something for you feels like wasted time. You're losing sleep because you could be staying up later getting your agents to do stuff.
My AI mania presented itself in some ridiculously over-ambitious projects.
I built a JavaScript interpreter entirely in Python , vibe-ported from MicroQuickJS by Fabrice Bellard.
Then I built a WebAssembly runtime in Python as well .
These projects were quite useful, in that they sort of cured me of my AI mania... because after I built these things, I got to look at them and ask "does the world need a slow, buggy, half-baked Python JavaScript interpreter?"
I don't think the world does.
n * 2);
console.log('Doubled:", doubled);
var evens = numbers.filter(n => n % 2 === 0);
3 console.log('Evens:', evens);
var sum = numbers.reduce((a, b) => a + b, 0);
console.log('Sum:", sum);
+ [eT
=e p=
output Lom
Doubled: [2, 4, 6, 8, 10, 12, 14, 16, 18, 20]
Evens: [2, 4, 6, 8, 10]
Execution time: 8.00ms
About: micro-javascript is a pure Python JavaScript interpreter with configurable memory and time limits. This playground runs entirely in your browser using
Pyodide (Python compiled to WebAssembly). View on GitHub
" style="max-width: 100%" />
I did get this out of it:
This page runs my JavaScript interpreter built in Python, running in Python using Pyodide , which is Python complied to WebAssembly, running in JavaScript, running in a browser.
It's a beautiful stack of horrors. I've been having a lot of fun with WebAssembly this year.
By the end of January, that repository we saw started in November had renamed itself, first to CLAWDIS, then CLAWDBOT, then Moltbot, and finally to OpenClaw.
At this point OpenClaw had 8,300 commits, less than two months after the project had started. I looked today and it's over 100,000 commits now!
This is the most vibe-coded piece of software in existence.
(Here's how I generated that list of name changes .)
This kicked off the OpenClaw revolution. It effectively defined a new category of software.
There's a generic term for this which I really enjoy. We call software like this a "Claw". There's OpenClaw, NanoClaw , IronClaw , PicoClaw ...
Today they're being rebranded as "personal agents" or "general agents", but I still like to think of them as Claws.
The Apple stores in the Bay Area sold out of Mac Minis because so many people were buying Mac Minis to run OpenClaw!
Drew Breunig said that this is because your OpenClaw is a digital pet, and you buy a Mac mini as an aquarium to keep your claw in, which is kind of delightful.
暂无评论,快来抢沙发~