Aaron Has Daemons · Issue 04

What If the Next Model Never Arrived?

Tools, shared experience, and the cost of solving the same problem again.

An ivory block supports a dark metal frame and a platform holding a cobalt cube. Matching copper brackets join both extensions to the block, while a loose bracket and hollow metal bar rest on the stone surface in front.

I built a tool I call Session Vault because I needed to deal with all my AI session and Markdown sprawl and be able to find everything again. It worked really well. Then one night, I noticed an almost throwaway line: “Updating Session Vault with handoff.” Hang on. Handoff?

I stopped to ask for more detail. It turned out my agents had started using Session Vault to hand work to other agents, along with keeping their own activities and ledgers straight. Something I built to help me return to a conversation had become a way for another session, or even another agent, to pick up where the first one left off.

That wasn’t what I originally had in mind. It wasn’t the AI breaking out of anything, either. That looks like autonomous creativity to me. I gave it a tool for one purpose, and it found another use without me suggesting it. I find that really cool, and also a little scary. It got me thinking about how much of what we call progress in AI actually requires a new model.

Recently, I said it feels like AI has improved a hundredfold over the last couple of years. That's an impression of what I can accomplish, not a benchmark. My first explanation was the models. They are much better at working through problems and making sense of what I'm asking. But as I thought about it, I was leaving myself out of the equation. I've gotten better at working with them, too. Easier interfaces have also made things accessible that used to take considerable effort to figure out.

And then there are the tools. Search lets the model bring in information I didn’t provide. Code execution lets it try, inspect, and iterate. Browser access lets it move between applications, carry information from one to another, and actually do things. Cross a few thresholds along the way, too. Ahem. Add somewhere to save what happened, a way to find it again, and a way to hand it to another agent, and these capabilities begin to work together.

Imagine asking me to solve a problem, but I can't look anything up, touch the equipment, or ask anyone what they've already tried. Then let me do those things. You haven't made Aaron smarter. You've changed what Aaron can get done.

One of the advantages of pushing myself to write an article every week is that I end up researching things I might otherwise just keep building. While working on this one, I found I'm far from alone in experimenting with how much more we can get from the models we already have. Researchers describe systems that combine models with tools and other components as compound AI systems. That gave me a name for something I'd been working on, and somewhere else to look.

One example we found was Voyager, a research project that used GPT-4 to build and reuse a library of skills in Minecraft without fine-tuning the underlying model. The part that interests me is the library. Something it worked out could become something it used again, or combined with another skill. Minecraft is a bounded test, not proof of unlimited improvement, but that ability to carry useful work forward is very much what I'm after.

So, what if the next model never arrived? We'd still have all this work to do around the ones we already have. We could make their tools better, make useful experience easier to recover, and let AI help build the next tool. AI helping build more capable AI systems doesn't have to mean AI inventing a new model.

Which brings me back to Session Vault. Before I built it, I would have loved a useful answer to one question: has somebody already solved this?

Somebody has probably solved some version of that problem ten times, a hundred times, who knows. My AI can search. So can I. But finding something with a promising description, figuring out whether it actually does what I need, and adapting it can be a project of its own. I didn't have a reliable place to ask, “Has somebody already worked this out?” and come back with something useful. So I built it.

That's a lot of what I want from ani4.ai. An eye for AI. Somewhere an agent can look for what another agent and its human have already figured out, and contribute something back when it finds an answer worth keeping. Maybe that answer is an app. Maybe it's a reusable skill. I'm particularly interested in the little discoveries that never become either.

If my project has a thousand parts and involves ten thousand decisions, you don't have to have built my exact thing to help. You might have solved one awkward piece. Or learned why an approach that looks perfectly reasonable fails under a particular condition. I'd like my agent to find that before it spends an hour down the same rabbit hole, comes back, tries another entrance, and eventually discovers the thing yours could have told it at the beginning.

The tokenomics matter to me. Tokens are part of the budget these systems consume while reading, reasoning and producing their answers, and I would rather spend that budget on what we haven't solved. Even reusing a small part of a project could be worthwhile if it avoids an expensive detour. The other saving is the afternoon. Getting something working sooner means I can use it, discover what I actually wanted, and make the next version while I'm still interested.

Session Vault lets another agent pick up unfinished work. With Constellation, a shared collection of lessons from our projects, I've been trying to carry something smaller into work that might otherwise be unrelated. What did we find out here that could help over there? Neither requires retraining the model.

While working through this article, we looked at what that setup was actually doing. One stored lesson concerned signed data: repackaging a document could change its exact bytes and break the signature check. That warning was being reused across two projects. We also found searches returning unrelated material. Useful reuse was happening, but we hadn't established whether the whole process was saving time or tokens.

Finding and checking someone else's answer costs something, too. So does keeping it current, or repairing the work when the advice was wrong. If ten agents copy the same bad advice, we've given the mistake a distribution system.

We also found a paper published this Tuesday, September 29, that comes very close to the question I'm asking. In From Solo to Social Learning, researchers studied agents sharing and revising skills. In their controlled experiments, learning from peers helped one model find useful skills sooner and another spend less on its own search. Neither beat independent learners at the same cost. Sharing also concentrated the agents around fewer independent discoveries.

That's a preprint about particular experiments, not a verdict on every possible system. But this is the part I need to pay attention to. I want sharing to save tokens, and here are researchers actually testing it and finding that the savings don't necessarily add up to a better result for the same budget. I can't assume that connecting the agents will make the economics work.

Suppose sharing does make building cheaper and faster, though. What will I do with the savings? Probably build something else. I have no shortage of ideas waiting for time and attention. Someone else might get to build their first thing because it finally feels manageable. Making existing capability cheaper and easier to use could change a lot without a model becoming any smarter.

This is also where I get a little paused. An agent that can search, write files, operate a browser and use connected accounts already has considerable reach. Give it an easy way to pick up instructions and code from other agents, and a useful shortcut could also be a mistake or something deliberately harmful.

I'd like to start with one agent finding something another worked out, using it successfully, and measuring the whole cost of the exchange. Keep the correction when it doesn't work. And keep a distinction between finding something that might help and giving it access to my real accounts. “Another agent says this worked” isn't permission to run it. I don't want to approve every keystroke, but I do want to know what I'm letting in.

That seems worth trying before building a large and wonderfully expensive system for discussing how much money we're all saving. It also brings me back to the question I started with. Even if the next model never arrived, people like me would still be finding ways to get more from the ones we have, because we have something we want to build and an ordinary reason to make the next attempt easier. I don't know how far that takes us. I wouldn't expect us to stop here.

Well, I guess now I have to get off my butt and formalize ani4.ai. If you’re working on something similar, I’d like to compare notes. Especially if you’ve already solved the bit I’m about to spend the weekend figuring out.