August 19, 2026

Ross Dawson on Recursive Self-Improvement in Humans + AI Systems (HAI Ep54)

"Evaluation for the human is where we start to reference not just what happens, but also how we feel about it, our values, who we are, and where we are applying our judgment in action."

–Ross Dawson

Robert Scoble

About Ross Dawson

Ross Dawson is a futurist, keynote speaker, strategy advisor, author, and host of Humans + AI podcast. He is Chairman of the Advanced Human Technologies group of companies and Founder of Humans + AI startup Informivity. He has delivered keynote speeches and strategy workshops in 33 countries and is the bestselling author of 5 books, most recently Thriving on Overload.

Website:

rossdawson.com

LinkedIn Profile:

Ross Dawson

What you will learn

  • How recursive self-improvement applies to both AI and humans in learning systems
  • The differences between single loop and double loop learning, and why they matter for AI-human collaboration
  • How reinforcement learning shapes AI development and future capabilities
  • The six essential steps for effective learning loops in both AI and human contexts: sense, plan, act, evaluate, update, persist
  • Risks of AI eroding human skills and why maintaining and improving human judgment is critical
  • Real-world examples where humans use AI to enhance, not replace, their capabilities (e.g., military training, aviation, chess)
  • Why the chief learning officer and learning & development teams must rethink organizational learning in the AI era
  • Key considerations for designing organizations that maximize both human and AI capabilities through integrated learning loops

Episode Resources

Transcript

Ross Dawson: Welcome to a solo episode of the Humans Plus AI podcast. We have some wonderful interviews coming soon, but today I want to share some of my thinking, particularly on how we build improving humans plus AI systems.

First, a bit of an update. I've been working on a lot of things which haven't been very visible yet, but will be soon. In terms of ventures, Fracios, our AI-augmented strategy platform, is just moving into our beta cohort two before launching as a product. We're getting some very good feedback and it's evolving rapidly.

Our Humans Plus AI Teaming course is not far from being completed, and we're also launching a whole set of resources on humans plus AI decision making and decision architectures, and how to implement them. There's also a lot of deep thinking going on about humans plus AI systems and how we can implement them, particularly around humans plus AI organizations, so there's a lot I have to share.

And you heard it first here: I am planning to launch a Substack so I can share some deeper articles that go into more depth. So today, a little bit of an early preview of some of my thinking—some of which I'll flesh out in a more structured way in a forthcoming article. The topic is what I call the recursive self-improvement of humans plus AI systems.

Now, the starting point is that AI is improving itself, and that's the current focus of the frontier labs: to build recursive self-improvement so it starts to improve itself, and AI moves out of sight. But we should be applying exactly the same thinking to human learning loops. Humans can recursively self-improve. We need to be able to design that in, which of course takes us to the next step of building the humans plus AI loops, where the learning loop is not just in the AI, it is not just in the human, it is the humans plus AI system, where we are focusing on building recursive self-improvement.

So, the human gets more capable, the AI gets more capable, and the system and how they interact improves over time. In my first book, Developing Knowledge-Based Client Relationships—this is back in 2000—I talked about the difference between single loop and double loop learning. This was Chris Argyris and Donald Schön, and that's still current; people still refer to that.

Single loop learning is simply where something changes as a result of an output. For example, a thermostat: your objective is a particular temperature, you turn the heater on to bring it up, or you turn it off if it goes too high. That is a single loop.

Double loop learning is when you apply that learning loop to a single learning loop—in which case, you are learning whether the temperature setting is right, or you're learning why you are heating the room at all. You are questioning the assumptions and able to learn and improve that over time.

So, let's take that distinction and start to think about what is happening with AI recursive self-improvement. This takes us back to the idea of reinforcement learning. Richard Sutton and Andrew Barto wrote the book in 1998, Reinforcement Learning, and that earned them the Turing Award—the top award in AI—for their contributions. Still today, reinforcement learning is at the heart of how AI systems develop.

Essentially, you have an agent or an AI system that acts. It receives some kind of feedback on how well that action went, then works out how that feedback applies to how it was produced, updates the system, and carries that improvement forward to the next cycle.

Now, this has gone a lot further from basic reinforcement learning, where you have agents that rewrite themselves, design their own fine-tuning, and create new scaffolds for themselves. What I've done is generalize that in a way which is applicable both to AI and to humans, very much inspired by Sutton and Barto's work. So this idea of sense, plan, act, evaluate, update, and persist.

So let's look at how that sequence works within an AI system. AI sensing: it has a specific context, only given what the AI is given in terms of data inputs and so on. These are human-imposed senses.

Plan: now, AI has gotten very good at reasoning, structuring options, generating options, and selecting between those. That's what we see a lot of in action, particularly with the more recent frontier models.

Act: one of the key features of AI actions is that they can be parallel, so there's a high volume of action and a lot of data gathered from it.

Evaluation: this is where "evals"—the current term—come in, where you have some kind of measure and can see whether or not the AI is meeting that measure. This means you can then update the system, changing everything from particular parameters in the models to potentially weights in the model.

The final step is persist, where it actually consolidates. You are capturing that in the context systems within which the AI is built, or rebuilding the harness, for example, and that flows back to the entire system.

If we think about that on a human level, the first point to note is that AI has been built around these reinforcing feedback loops. Reinforcement learning is fundamental to what AI is—the fundamental nature of what it is.

But we are now in a context where the nature of the way work is happening, and the way people are using AI or AI is being brought into workflows, is deskilling people. It is taking away from their cognition, from their thinking, from their ability to work, and that is the most fundamental problem.

The heart of my work is to combat that, to move against that, so that we build systems where human skills increase rather than erode. Because by default, the way we are building and using AI, and the way we're bringing AI to organizations, is taking away from human judgment and capabilities.

So let's come back to those six steps and think about how these could apply, or do apply, at a human level. In terms of senses, humans are very good at that. We have lots of senses, including social interaction and context. The issue there is around allocating our attention, something I'm also building a whole body of content around.

Next step is planning. Humans are good because they have the context, but they don't have the cognitive bandwidth to generate more than a few options or assess those. In fact, human planning is arguably less transparent than the AI's model, so we need to be able to justify that.

In terms of act, that is not parallel like an AI—it is serial. We take a step, but we are applying our judgment every time we act in some form. We are judging, and that is something where we are actively looking for feedback.

Evaluation for the human is where we start to reference not just what happens, but also how we feel about it, our values, who we are, and where we are applying our judgment in action. Every time we take an action, we are judging before we do it, and we are doing that afterward. That's intrinsic to what we do, so that then provides an update.

And obviously, we just soak in experience as humans, but now we can start to design better so that we, as humans, more explicitly capture that update of the lesson or the feedback by making sure that evaluation goes against our previous judgment. Then, to persist, we can use external systems—AI record keeping and so on—as well as change the way we think, so that we are building our own true learning loop.

So now, if we start to say: at one level we've got the AI learning loop, as we've described, and at the other we've got the human learning loop. We are looking to maximize each of those together. The AI learning loop is functioning pretty well. The human learning loop we do need to be working on, because it will basically erode unless we augment it.

But now we need to think about how we put the human learning loop and the AI learning loop together. The critical thing at the junctions there is that the AI part of that humans plus AI system loop is this context system, where it starts to build a body of reference data, information, and systems—being able to build up, for example, the processes that happen in organizations, the decisions that have been made.

It is the context in an organizational setting; there's a lot which it needs to build for the AI to be system-relevant and valuable, not just in its own feedback loop, but applying that to the ever-increasing context and the granular detail in the arena in which the AI operates.

On the human side, it is about judgment development. You may be familiar with some of the things I've already shared around judgment development. There's a lot more I'm working on now, where the issue is: how do we deliberately make sure that human judgment improves? That includes making sure that people think before they act, are able to assess what the AI produces and their own actions, understand the degree of confidence in the AI, and then compound its ability to measure judgment. More on that another time.

This is interesting to think about. You can imagine, as John Hagel often says, that scalable learning is essentially the only thing that matters. Scalable efficiency was the focus in the last century. Now, we enable people to scale the learning of organizations. And today, with the role of AI, that learning has to be the humans plus AI systems loop.

Now, I have scanned very deep and wide, and it is deeply disappointing that I'm not seeing examples of organizations that truly are designing and implementing this humans plus AI systems loop. If we look at AI-native companies, there are a number that are using loops. For example, Harvey is able to get feedback on how systems are used to improve them. Replit is using both customer feedback and internal feedback to drive improvement.

These are all examples of what AI people call RLHF—reinforcement learning from human feedback. But this is all focused on improving the AI system; it's not focused on improving the human system.

One example, which I think is nice, is the U.S. Army Command and General Staff College. They've done a war gaming exercise where officers make a decision, the AI creates a judgment on what has been done, creates a response to that, and then the officers have a brief period of time to check it and potentially override it. They go through a sequence of turns—nine turns in under three hours—basically having teams who are working on their own and those who are working with AI. What they are finding is that they are essentially using the AI to improve their own judgment.

It is always the officer's judgment in how they're using that, but the AI is providing a whole array of different tasks. This means they can demonstrably not just improve their judgment through the AI's input, but afterwards are able to assess and generate for themselves better options, and to judge those more effectively. This is a system designed to improve the humans.

Another critically important example is from 2013, when the Federal Aviation Authority made a ruling that pilots have to regularly practice without the autopilot. They demonstrated that those who overused the autopilot were essentially eroding their skills, as many people are now experiencing with the overuse of large language models. This is now mandated: to practice independently to increase their capabilities.

Another example, less structured but real, is chess. Human capabilities at chess have significantly improved since we have AI chess models. We still can't beat the AI, but humans have individually worked out how to play with or engage with AI in a way that improves their capabilities and get better by themselves—to learn from how the AI does things, not just to defer to the AI, but to actually, now independently of the AI, use that to become better and more capable. AI chess capability improved, as they have in Go, having learned from AlphaGo and other AI systems.

Thinking about this system—the humans plus AI learning loop, the AI learning loop, the human learning loop, the humans plus AI learning loop—the first level is that this really dramatically changes the learning function in organizations.

Learning and development—the chief learning officer and their team—need to be thinking in very different ways around the nature of learning, because it is all about the design of how work itself is done. It's not about taking classes or doing lessons—there's some of that too, of course—but it is all about the actual development of judgment by humans in interacting with AI, because that is now the reality of the practice, and one which can be designed to improve human capabilities.

The broader point is this: ultimately, it's about organizational design. As I often say, all organizations will be humans plus AI. They will have humans and they will have AI. We must design them so that those humans and AI work well together.

When we think about it from this frame—starting from building learning loops, not just of AI but of human learning, and in particular the humans plus AI learning loop—there are many implications for what this means for organizational design, and absolutely for leaders of organizations.

There's quite a lot more to dive into in all of that. I will lay all this out in a little bit more structured way in an article in the next little while, and also dive a bit more into some of the implications of that, particularly on what this means for the learning function, what it means for organizations and leadership, and how we can deliberately design and improve the human learning loop and the humans plus AI learning loop, given that the AI learning loop is already working pretty well.

Thanks for listening. Wonderful to have you with me on the Humans Plus AI podcast journey. A lot more to come, and speak later. Bye.