BI 247 Maxim Raginsky: A Control Theory View on Brains and AI

October 07, 2026 • 01:47:39
BI 247 Maxim Raginsky: A Control Theory View on Brains and AI
Brain Inspired
BI 247 Maxim Raginsky: A Control Theory View on Brains and AI

Oct 07 2026 | 01:47:39

/

Show Notes

Support the show to get full episodes, full archive, and join the Discord community.

The Transmitter is an online publication that aims to deliver useful information, insights and tools to build bridges across neuroscience and advance research. Visit thetransmitter.org to explore the latest neuroscience news and perspectives, written by journalists and scientists.

Read more about our partnership.

Sign up for Brain Inspired email alerts to be notified every time a new Brain Inspired episode is released.

To explore more neuroscience news and perspectives, visit thetransmitter.org.

Maxim Raginsky is a professor at the University of Illinois at Urbana-Champaign. Max describes himself as interested in probability and stochastic processes, deterministic and stochastic control, machine learning, optimization, and information theory. Today we mostly lean on his control theory expertise, although you'll here his knowledge is vast in many other domains, even some neuroscience. I wanted his control theory perspective on neuroscience, AI, and biological autonomy, so we dance around a lot of topics related to those. Max also writes a substack called The Art of the Realizable, from which I drew during parts of our conversation.

0:00 - Intro 3:07 - Low energy lifestyle 4:27 - Engineering and philosophy? 13:52 - Brains vs AI 19:45 - Inferring the inside from behavior 30:57 - Analog vs digital 41:32 - Is the brain a control system? 46:50 - Willems control 1:02:12 - Control vs cybernetics 1:13:26 - A control perspective on AI vs brains 1:17:46 - AGI 1:21:48 - Turing 1950 1:29:07 - Perceptual control theory and active inference 1:40:09 - Passive control in the brain? 1:41:48 - Computation 1:43:36 - Counting spikes

View Full Transcript

Episode Transcript

[00:00:03] Speaker A: Control is about effective interaction, right? Does. Does a system interact with its environment effectively? Is it. I like the term coherence. It's an overall property of a system that has or system of thought, you know, technological system, social system, whatever. It consists of multiple parts. And each of those parts, each of them presupposes the other. But it is a concept that allows you to kind of think about a system in terms of. In terms of its goals and values. Speaking of Turing, you know, the. Again, people read into his thinking, whatever that, whatever they. They think they want to see. But in his paper that introduced the Turing machine, he was very careful to, to argue that certain sources of information should not be modeled computationally. Right? Because if they could be, then you can just model that, interconnect the two together. Now you have a bigger Turing machine to, you know, to work with. Imagine an alternative world which is just like ours, except for this one thing. There's a system. There's this computational arrangement, there's this artificial system that can, let's say, you know, perform any possible economically valuable task or cognitive task at a superhuman level. And the problem with that is that this is not a possible world, which is just like ours, except in this other aspect. These differences are going to amplify so much that, you know, reasoning about that world becomes completely, completely impossible. You're hit. This. [00:01:41] Speaker B: This is Brain Inspired. Powered by the transmitter. Hello, I'm Paul Middlebrooks. Welcome to Brain Inspired. Maxime Raginski is a professor at the University of Illinois at Urbana Champaign. Max describes himself as interested in probability and stochastic process, deterministic and stochastic control, machine learning, optimization, and information theory. Today we mostly lean on his control theory expertise, although you'll hear his knowledge is vast in many other domains, even some neuroscience. I wanted his control theory perspective on neuroscience, AI and biological autonomy. So we dance around a lot of topics related to those. Maybe. Max also writes a substack called the Art of the Realizable, from which I drew during parts of our conversation. Support Brain Inspired on Patreon to hear all full episodes and get the full archive of episodes, go to the Show Notes to learn more information on [email protected] Podcast247. Last thing. Sorry about the occasional echo that you'll hear on my voice. I didn't notice it while recording and I couldn't get rid of it during editing, so sorry about that. And I hope it's not too distracting. What is not distracting is Max. And here is Max. Max, let's talk about the Most important thing first here, what does it mean? So the very, very last thing on your website you claim to lead a low energy lifestyle and there's a link to a poem. What does that mean? [00:03:25] Speaker A: It's sort of a philosophical take on all the demands that are placed on our lives. Not just professional, but personal. Right. I mean, we're constantly asked to plug ourselves into various social media, various environments that quantify rank, gamify all the activities that we're supposed to just enjoy as part of life. And so the low energy lifestyle in the sense is kind of a reference, it's a private joke which no longer has any relevance. But the link takes to one of the chapters in Tao Te Ching that talks about sort of the value of maintaining inner harmony and just letting the natural flow of things take its course. [00:04:27] Speaker B: So I like that very much. And that actually segues nice to my first comment, I guess. And maybe you can comment on. My comment is I don't know. And you can correct me if I'm wrong. I don't imagine that many machine learning geared, control engineering geared folks are constantly referring to philosophical philosophy of science and philosophy otherwise in all of their talks and works. Is that right? I mean, everything you do seems to be imbued with philosophical perspectives. [00:05:07] Speaker A: I mean, in my personal case, it's more of a kind of mid career, midlife reflections. Parts are kind of brought on by a couple of momentous events. One was when my daughter was born and she's now 12, but 12 years ago, that was right before I got in tenure and I realized I don't have to participate in the rat race to the same extent that I've been told to do it. If I have a choice between writing another incremental paper and witnessing how my daughter is growing up and just enjoying these moments, I'll choose the latter. Because if I don't write that paper, either somebody else will if it's worth saying, or it just wasn't worth saying in the first place. And the second momentous event was the COVID pandemic where I thought that this would be an opportunity for everybody to take a collective deep breath and focus on again, sort of these aspects of life that seem to be so elusive and so valuable. But no, I mean, there was a sizable faction of the academic community that chose to actually just up their productivity. Oh, we don't have to go anywhere now. Let's just crank up the paper mill. And so that sort of realization that I'm actually, this is kind of the low energy, this kind of internal Taoist freedom as well. But I went back to my interest in philosophy that I had when I was much younger. And at some point when I was preoccupied with all the usual pursuits of kind of an up and coming faculty member, I lost track of that, and I reconnected with some of these ideas. Partly it was kind of inspired by a lecture series that my wife, Svetlana Lzebnik, who is also an academic, she's a prominent researcher in computer vision. She had given a series of lectures on sort of the past and the Future of Computer vision at Georgia Tech when both of us were there on sabbatical in actually 2019. 2020. And so throughout the process of her preparing these lectures, she was looking at some literature that intersected these key philosophical ideas. There were some questions from cognitive science, there were some questions from philosophy of perception, philosophy of intelligence. So we had discussions about that. She was looking at, for example, the ideas and views of David Marr, J.J. gibson and Jan Kunderink. Maybe the listeners of this podcast would know Mar and Gibson more, but they might not have heard of Jan Kunderink. Jan Kunderink is a psychophysicist, a researcher of vision and perception who, who was working in the Netherlands. And he wrote some of the foundational texts and articles on sort of all aspects of color perception, shape perception, how do we orient ourselves in the world. He had this book called Solid Shape that was on my wife's bookshelf, one of formative texts for her. And so Kunderink was bringing up in his more informal writings questions from, again, philosophy of perception. So he was talking about phenomenology, so people like Edmund Husserl. And he was also talking about the writings of Jakob von Uxkul, who was talking about this notion of Umwelt. And then I, as a control theorist, just kind of noticed, oh, these are interesting parallels to the kinds of things that I do. So I became newly interested in some of these things, started, started kind of going back to this literature. And then again, because of the disruptions brought on by, by the COVID lockdowns, I, you know, I found myself with more time to get back to this. But I think I kind of strayed a bit from your question about whether it's rare for a researcher in control or machine learning to be informed by philosophy of science. Maybe on a larger scale, it's rare, but somehow I saw over and over that various colleagues of mine, whom I never had suspected of having deep philosophical interests also have them. And so maybe it's this phenomenon where or you orient your perception towards certain patterns that you had not noticed before. All of a sudden I started noticing all these references. So I think it's maybe globally, it's rare, but I think I just have been surrounding myself with more reflective kind of set of people. And so they definitely, if not quoting these thinkers in their talks, they're definitely aware of them and are informed by a whole very wide range of thinking. [00:10:20] Speaker B: Well, that's the thing that stands out. It's just. Yeah, sorry, that's the thing that stands out is really every two or three slides you're quoting some philosophical perspective that, that helps undergird whatever message you're delivering about control optimization or, you know, boundedness or the generalizability problem, all of these things. And so you really weave it in nicely. So I just, I really appreciate it. But you said you didn't feel as much pressure to join the rat or to stay in the rat race, but you're still quite prolific publishing, so it's in some sense, maybe, I don't know if you picked up the pace, but I don't know, backing off like that, does it allow for some things to flow better and living that low energy lifestyle? [00:11:08] Speaker A: I mean, I think so. I think on today's scale I would not call myself prolific. If anything, I'm kind of glacial. [00:11:19] Speaker B: But you're polymathic, right? That word polymath does not apply to that many people. And I think you truly are polymathic. Is that. I'm sure you're too humble to accept a descriptor like that. Or maybe you're not, I don't know, because I think it's just accurate. I prefer accuracy over humility. [00:11:35] Speaker A: All right. Well, I don't necessarily like that term because it does have a bit of a self congratulatory aspect to it. But I definitely am interested in all sorts of ways of apprehending and dealing with the world. Because I certainly think that, for example, the technical or scientific view is just a particular projection of it. And then the humanistic view is another, more sort of literary lens on the world is yet another kind of historical perspective, economic ones. I think one has to be kind of interested in this to some extent. You obviously can't be expert at all of these things and you definitely should refrain from confidently making pronouncements about these things. But there's this saying that is, I think it's apocryphal. It's attributed to one of the Athenian, I think, politicians. I don't think he ever said that. But the Saying is something like, you may not be interested in politics, but politics is definitely interested in you. But I think, I don't view it as this kind of external imposition. It's a way for me to kind of channel certain thoughts about, let's say, human embeddedness in their environment, in the world, interacting with other people, with artifacts and culture and things like that. How that filters into my research is I can never explain it, but it definitely does. So I think it was just this broadening of perspectives that is one of the mechanisms that I think helps me maintain sort of the sense of meaning and the worth of what I'm doing, if not necessarily to others. And I hope that that's the case, but at least to me as well. Why do I keep going? [00:13:52] Speaker B: Okay, well, I'm going to ask you some pretty broad questions because I don't know where our conversation is going to go. So we can narrow them if the question is too broad. Too broad, but starting, you know, really broad. I'm, I'm curious, like to a control engineer, if that's part of how you would describe yourself. Right, To a control engineer, do brains and artificial intelligent models, networks, do they look the same? [00:14:26] Speaker A: They can, we have seen already many, many times over that their external attributes, the processes that we can use to sort of monitor their outcomes from the outside perspective, can often look the same. Whether or not that says anything about their internal organization. I prefer to be more skeptical about this because it's just the paths by which brains had developed throughout evolution versus the technological development of artificial intelligence systems are just very different. And so that means that the way that these systems couple to their environment and the way that this mutual interaction shapes them and rejects certain unviable paths and encourages more fruitful directions, they're just very different. So I am wary of saying just because, let's say large language models can produce fluent linguistic streams that actually we perceive as meaningful through our interaction with them, that that says anything necessarily about any sort of isomorphism between how they're internally structured and how our brains are structured. That was actually kind of an interesting perspective that I picked up by watching a four part series by a very renowned mathematician, Mikhail Gromov. He's a geometer, works in France. So he gave a four part series of lectures on these questions underlying generation learning, et cetera, kind of informed by recent developments in LLMs. And there was one interesting thing that he said, and it's that if you think about, let's say the brain as a control system, it's remarkable in that it can take very, very low energy control inputs, let's say words, and generate very, very high energy actions as a result. For example, if somebody says, look out, there's a tiger and you're going to run. So all of a sudden a fairly low energy act, this acoustic signal generates a very high energy metabolic activity. And then you're doing this sort of, you're running away from the tiger. And he said, and I don't know, it could be a very far fetched analogy, but it resonated with some of the ideas I've been thinking about as a control engineer that el cereal Ms. These days are more like the cerebellum, meaning that again, I'm not a neuroscientist, but the analogy made sense to me. The role of the cerebellum is to coordinate very high energy activity based on fairly low energy control inputs, whereas the cortex most of the time generates low energy outputs or control signals to coordinate other activities. And he was basically saying that the reason why he expects to see more sort of archite structural changes in LLMs, that eventually people will realize that they do want to build something more like the cortex as opposed to something like more like the cerebellum. [00:17:56] Speaker B: Okay, yeah. Generally people say that AI looks more like it's supposed to be more like the cortex. That's where we think that's what we're thinking. [00:18:03] Speaker A: Right, right, right. But the fact that, for example, that there's no process by which acquired knowledge is somehow offloaded and stored in a more permanent way, as opposed to, you know, being a sort of, you know, dynamic working, that separation is not yet present in LLMs. It's all distributed through the parameters of the model. But this modularization, which is. We don't want to talk about localization of function, but definitely the brain has some sort of modular structures. It's a control system. It's very complicated. And so perhaps that's the kind of thing that we will build in. But nevertheless, I don't think it'll tell us much about any sort of deep analogy with brains. I think again, there are systems that either have evolved or were constructed to maintain certain relations with their environments. They seem to be pretty good at that. But we have to be mindful of this kind of. I guess philosophers of mind refer to this as multiple realizability. Control theorists or control engineers have this notion as well that given the same desired external behavior, there are multiple ways of realizing that, building that as a system. And so often this process of inferring this internal organization from externally observable behavior, it's obviously a very ill posed problem, right? There's no guarantee unless you open the system up and actually look inside to confirm that that's what's going on. [00:19:46] Speaker B: But as a control engineer, so you do like device modeling, which does this, which tries to essentially infer the goings on internally based only on external observables and you've made. [00:20:00] Speaker A: But there's also the dual problem, which is synthesis. And that's the one where if you want to design a system with particular behavior and you start with a given set of building blocks that come from your domain, if you're an electrical engineer, you're going to use electronic and electrical components. If you're a mechanical engineer, you're going to use mechanical devices. If you're a chemical engineer, you're going to use chemical reactions and things like that. And the commonality exists on, let's say, the level of mathematical equations. There's certain building blocks, say basic feedback routines, let's say systems with linear dynamics for the state and systems where the output is a certain function of the state. You write down the equations mathematically and then of course they translate into different physical instantiations in a particular domain. And then you're asked how to synthesize out of these building blocks a given external behavior. And that's an interesting question. And of course, as an engineer, you're free to, let's say, if the design specs change, you're free to completely throw out existing design and start anew. Unless of course, it's too expensive to do that. Then you're going to figure out the best trade off between sort of modifying a given arrangement and achieving what you want. Somehow it My understanding of, let's say biological evolution is that there's a lot of this sort of a path dependence that's built in. You can't just get rid of whole existing structures in one fell swoop and replace them with something new. There has to be this kind of incremental change. But on the other hand, I also thought this is something that I picked up from a book called the Plausibility of Life by Kirchner and Gerhardt. There's two systems biologists writing about what they call facilitated selection, I think in evolution. But there's this one theme that they talked about, they referred to it as constraints that deconstrain. And that means that certain things are just locked in. But because the architecture of living systems and the architecture of engineering systems is layered, meaning that if you have a certain process, certain set of generative primitives. If you're constrained to use them, that means that there's going to be some incredible diversity of their instantiations at some lower layers. But as long as you're dealing with this particular set of building blocks at this particular layer, that may actually free you from thinking about what's going on underneath and actually implement more complicated systems. We actually see this in, for example, the emergence of digital computers because this kind of a binary logic that digital computers operate means that you take electronic circuits that are analog and you operate them in highly constrained regimes. But as long as you have reliable logical gates, then you can make use of universality of Boolean functions for certain types of tasks and not really worry too much about how the logic is implemented in, let's say, solid state. Right. Sometimes, of course, you may need to do that. If all of a sudden you do not meet certain criteria like speed or accuracy, then you have to go back and say, oh, how am I implementing this digital architecture? Do I use this type of transistor, do I switch to a new generation, et cetera. But for the most part you can just kind build on what's already available at the current layer of abstraction. And I think that's a very useful sort of way of thinking about engineering and control engineering in particular. That comes from the ideas of John Doyle, he's a control theorist at Caltech. Was also very kind about layers and architectures. Very kind of. That idea has been very influential in my thinking as well. [00:24:15] Speaker B: Huh. There's so many directions to go here. But one of the things that you said is you can. Well, what I was thinking when you said, you know, evolution is path dependent and you can't just get rid of stuff and replace it with new stuff. However, I mean, there are brain lesions that people suffer that do change their personality drastically. But you can also get rid of large swaths of the brain and be just fine. Like, right. You don't notice a mental difference and there's no observable, you can not have a cerebellum, for example. And I don't know if you're just fine, but, but you can get by just, just fine for the most part. [00:24:56] Speaker A: But yeah, there's, yeah, this, right. And this notion of graceful degradation, I think that's, that's kind of what motivated, let's say Frank Rosenblatt when he was developing, developing the perceptron. Because the previous, this logic gates view of MC was very much, it was brittle, right? It did not allow for this graceful degradation because if you had arranged a network of McCulloch Pitts neurons to compute a certain function. If you wanted to compute a function that was slightly different, there was quite a distinct possibility you might have to completely rewire the original circuit. [00:25:33] Speaker B: Right. It was von Neumann who noted that in his. What? I don't remember the name of the book. The Brain and the Computer. [00:25:44] Speaker A: The Computer and the Brain, yes, yes. Computer and the Brain, Yes. [00:25:47] Speaker B: Just the high redundancy apparently present in Brains. And that we should design computers after that. [00:25:53] Speaker A: Right. I mean, he pointed that out in earlier paper on sort of his approach to error correction. And so this idea of building redundancy into systems consisting of multiple interacting components, each of the components may be unreliable, but you build in a lot of this redundancy and error correction by replicating, by multiplexing and things like that. So he definitely. I guess he kind of really expanded on that in the Computer and the Brain, which my understanding is it was a set of lectures that he gave [00:26:31] Speaker B: previously, I think, but not put together yet. [00:26:34] Speaker A: Right. [00:26:35] Speaker B: It was posthumously published, I believe. [00:26:37] Speaker A: Posthumously published. But then he did have quite a few very prescient observations. One was this necessity of thinking both about the digital and the analog modes of operation. I think his analogy there was that, let's say, neural activity, if you think about it, let's say. I mean, I don't want to get into all these debates among neuroscientists about the coding in neural spikes, the code. But he was thinking about something digital on, off, all or nothing, versus analog. So the analog system to him was perfectly exemplified by, let's say, something like hormones. We just broadcast some sort of signal to a very large assembly of neurons or any other structures in the brain. And so they act completely differently. And he was saying that this is very important to kind of understand this tension and mutual need for analog and digital systems. We see this with LLMs now. That's actually on this kind of conceptual level of design principles. It's actually very interesting. Again, I really owe this observation to Svetlana, my wife. She said that there's this inversion that took place that this symbolic AI of McCarthy and Minsky, and then to a great extent, also the views of David Marr on vision. The idea was that the world is analog, but you go symbolic almost immediately. Marr's idea was this. The processing of analog external input into symbols starts almost right away. And then you manipulate these representations, these symbolic, discrete representations. This was, again, this kind of classic AI view that you represent the world through propositionally, and then the AI system will Manipulate those propositions using a series of truth preserving transformations, and then at the output, then it gets somehow converted into actions on the analog world. But what we see with LLMs is completely different. In fact, we perceive the world through tokens, but then the internal processing is with vectors in this very high dimensional space in which it's very easy to separate clusters of these activation patterns using almost like linear operations. So that's kind of an interesting thing. But again, it is this mutual interaction between symbolic and analog modes. And the other deep thought in the computer, in the brain, was that [00:29:21] Speaker B: at [00:29:21] Speaker A: some point you need to think about manipulating information which is encoded statistically. So it's not that you take on the usual viewpoint in Shannon's information theory. The use of probability by Shannon was to basically argue that probability distributions for different occurrences or states of the world are models for what's likely and what's not. But what we still manipulate are, let's say we encode data using, once again, bits. Bits are this, what's that phrase, the universal currency in information theory. So we encode the message we wish to send using bits, we apply various transformations to it, error correcting codes, et cetera. And then the decoding then takes place at the receiver. So the bits get decoded into intelligible signals. And the role of probability is just to make sure that, that these bits actually are used economically. You take advantage of various statistical regularities in the world. What von Neumann meant when he talked about statistics is that the signals themselves actually are probability distributions or some sort of statistical rates. And your goal is to manipulate those like they don't necessarily represent anything about the world, they're just instrumentally useful. Manipulating probability vectors is somehow useful. [00:30:56] Speaker B: I want to go back for just a moment to this analog versus digital idea that you were talking about. I've had a couple control engineer like people on the podcast, one of whom was Rodolpho Sepulchre, and he's all about modeling, thinking about like individual neurons as a mixed analog digital feedback system, right, where the, the spike is the positive feedback and the voltage activity is the negative feedback. But you were talking, did I understand you correctly, that the way that you see LLMs is that we feed them, one way that they differ is that we instead, okay, humans take in analog signals and then we're not going to talk about the code because all we do is count spikes, you know, and that's the early ideas. But LLMs take in digital signals and were you saying that their internal processing is more akin to analog processing because of the easily separability of things. [00:31:56] Speaker A: Not necessarily separability, but what I meant was that the internal processing takes place in the Euclidean space. I mean it's some sort of a curve, super high dimensional Euclidean space. You do not necessarily have a one to one mapping between sort of these steps and some sort of truth preserving logical operations. I don't think that that's what's happening in the brain either. Obviously the neural activations that are underlying brain function are definitely also not Boolean processors. Right. I mean, I think that whole. That I think was also Daniel Dennett who had that metaphor that brains are just theorem provers wrapped in an overcoat of sensors and defectors. I don't think that's correct. I mean that aspect of the intentional stance did not age well at all. Right. We're not proving theorems. I think in that sense even somebody like Chomsky had a much more nuanced take where he basically said that that linguistic activity, even something like inner monologue, it's just a manifestation of a whole lot of deeper subconscious processes that we're not aware of. And we don't know what most of them are, how they interact, how they produce this. Even if some people do think in phrases and sentences. And I never really, for example, about myself. Sometimes I do when I'm working on a problem, you know, on a research problem, a lot of it is really just writing. So a lot of my, you know, a lot of my thinking really goes, goes on, you know, externally on paper, but I definitely start thinking about things. And this is when, this is kind of, let's say, when I want to stress test a particular proof, I start thinking about it and I definitely hear myself talk through that. But when you're carrying out, you know, most of your everyday activities, like you just, you know, kind of out there in the world doing, doing your thing, you don't be think about that linguistically or propositionally. [00:34:09] Speaker B: Well, I think that there are varying degrees of that. Right. I think a lot of people do and I'm not one of those people. [00:34:15] Speaker A: Maybe they do, but I think, but I think this is also kind of an interesting philosophical point that a lot of phenomenologists focused on is definitely in Martin Heidegger's thought that most of the time we operate in this reflective mode. And when he called, I think the theoretical attitude only engages when something goes wrong or something is unexpected. Then you do catch yourself thinking propositionally. Right. So that's. [00:34:44] Speaker B: Yeah, otherwise we are in the moment sort of interesting. [00:34:47] Speaker A: Otherwise we're in the moment. Exactly. [00:34:49] Speaker B: Well, honed normal tools of behavior. [00:34:51] Speaker A: I mean, it is kind of an interesting tension between what you would call open loop and feedback control. This is actually something that late Roger Brockett from HARV pointed out. He had this notion of, he was interested in modeling sensory motor control and in particular the process by which you acquire competence in some skill, let's say learning to play a particular musical piece or training in a particular sport or something like that, that initially there is a lot of feedback. You do stop and you do reflect on what you're doing right now. You take stock of your current state and then based on that, figure out what to do next. And you're slow and deliberate. But once it becomes more second nature, it becomes very, kind of, very much an open loop kind of thing. The only signal you pay attention to is the clock, right? Because open loop control or non feedback control, it's not really open loop in the sense that you're not sensing anything. You are sensing, you're sensing time, right? So the time is the input to that controller. And so when you become fluid, let's [00:36:05] Speaker B: say, how are you sensing time? [00:36:08] Speaker A: Well, that depends, right? I mean, if you're, if you're a musician, I mean, first you play with a metronome, right? And then, and then eventually you kind of, you know, you start sort of, you know where you are in a piece. [00:36:20] Speaker B: You're sensing your own behavior. [00:36:22] Speaker A: You're sensing your own behavior, but it's sequentialized, right? And so, and, but the thing is, again, if you and I, you know, I used to play an instrument. I used to play classical guitar. I haven't done that in ages. But it is true, I mean, that I, once you, you have practiced the piece long enough, you are basically operating as an open loop controller. You don't look at the instrument. I mean, you look at the sheet music and the sheet music is sequentialized. That's your kind of sense of time. You don't go back. So that's kind of an interesting thing. And Brockett actually used the term attention to describe this. He used attention in two senses. This one sense was his idea that, that you can think about a general control routine as a mapping from time and your current state to the next action. And the tension to him was a question of sensitivity of that mapping to time and the sensitivity of that mapping to the state. So if that mapping is more sensitive to the state and doesn't really depend much on time, then it's closer to closed loop feedback. And if it's more sensitive to time and does not really pay attention to the state, then it's closer to open loop. And he had this whole idea, he called it minimum attention control. You design control strategies, let's say, in robotic motion, that trade off the two different parts of attention, the attention to time and the attention to the current state. And then if you sort of control the trade off between these two terms, you can sort of see that the sensory motor policies look very different. And then he used attention in this kind of offhand remark in a paper that he wrote in 2001 called I Think something like New issues of Mathematics of Control. It was this collection, I think, put out by. It was a collection of articles from mathematicians and computer scientists and physicists, kind of giving you the shape of mathematics for the 21st century. So his contribution was about mathematics of control. So he did have a paragraph about the use of. He actually used the word tokens in 2001 to indicate this idea that [00:38:48] Speaker B: we [00:38:48] Speaker A: do a lot of compression in the sense that we take this very rich sensory signal from the world and represent it using just a few words. And then on the other end, we can take those few words and expand them into a very rich set of sensory molar feedback loops and primitives that we actually use to accomplish goals. And he was interested in how this could be modeled, how this could be used in robotics, how this could be used in understanding, let's say, neural mechanisms of, let's say, locomotion and things like that. And towards the end of that paper, he also said there's another aspect which psychologists refer to as attention. And perhaps one way to model attention is as a vector that describes the direction in which the more relevant data lie, or something like that. Again, very sort of prescient set of remarks. So that sort of thing has been one of the questions I am really interested in. How do we actually think about these general principles? Right? Because one can either come up with a very fine grained mathematical model for a particular structure or a problem. Let's say, how do LLMs represent certain things? Why are they so effective at particular tasks involving something like planning and sequence prediction? But on the other hand, it is the fact that there are some general principles that underlie it. And even though maybe the details of the architecture may change, but it seems like some set of principles that maintain certain variance is there kind of. There's this great book called, I think it's called Principles of Neural Design by Sterling and Laughlin, where they isolate a list of principles that you can use to understand the Brain on a more sort of qualitative level, you don't necessarily jump from that to sharp quantitative predictions, but at least it kind of gives you a set of organizing principles. And I think that this is actually something I'm interested in in the context of LLMs. And now that there's this fashionable term physical artificial intelligence, that's when we connect. When we connect LLMs to actual devices that can affect and be affected by the world. [00:41:22] Speaker B: Like a robot would be a physical artificial intelligence. [00:41:24] Speaker A: Like a robot would be. Right, exactly, yeah. Using LLMs as modules to control robots. [00:41:32] Speaker B: But do you think of the brain as a control system? Like, I know that's a very simplistic way to put it, but is that how you think of brains? [00:41:41] Speaker A: I think. Well, I think it would be very reductive. I think it's a. It's a collection of control systems that sort of are called to action and, you know, recruited and then maybe de emphasized at certain times, depending on the task. And then there's kind of a meta control that makes sure that for the most part, all of these control systems are operating coherently. Right. This notion of coherence between different senses. I got that idea from a couple of books by this French neuroscientist, Alain Bertoz. He had this book called the Brain Sense of Movement, where he said that one of the biggest challenges in neuroscience is to explain how it is that the brain actually maintains coherence between, let's say, different senses, including proprioception and theoreception, not just, let's say, vision and hearing and taste and things like that. And your sense of balance. I think the first chapter in that book is about the sense of balance and then how that feeds into other senses. And he uses, for example, this phenomenon of motion sickness to explain how, on the one hand, this is actually a very ingenious solution to detecting whether you've been poisoned, that was arrived at that by evolution. There's a lack of coherence between vision and balance. When could that happen? Well, when you've ingested something that's poisonous. But on the other hand, the fact that now we have technology, let's say cars, that allows us to transport ourselves over long distances, this is a setting where this mechanism for checking for coherence between vision and balance now serves as almost like an adversarial example in this kind of machine learning parlance. So that's a very interesting question. I think I strayed a little bit from your original question. I think it's a bit reductive to think about the brain, as I may say that sometimes the brain is a control system, but I think it's a collection of control systems that are operating mostly coherently, globally, so that the organism can maintain certain relations with its environment, which also includes parts of the organism itself. And so from that perspective, it is a particular way of. It's a control architecture. It's a way of assembling different control systems to accomplish these goals. [00:44:15] Speaker B: I mentioned having some previous guests who are control freaks. So another one is Dmitry Tchiokovsky, and he actually thinks of individual neurons as essentially trying to control their inputs. Like, each individual neuron is sending a signal out, and. And as they're getting feedback, like recurrent feedback, they're essentially processing the effects of what that signal must have done and then trying to adjust their behavior so that they actually get the effects that they want. But when you say a collection of control systems, you don't mean every neuron is a control system. You mean going back to that sort of modular architecture of the brain. [00:44:57] Speaker A: I think that depends on how you look at it. Right. Because this also cuts into that whole discussion of levels versus layers that I think John Doyle likes to emphasize, because certainly any system that's embedded in an environment and its goal is to maintain certain relations with that environment as a control system, it's interacting with something. And by means of that interaction, the range of all possible behaviors that the system could in principle, execute is narrowed. I mean, that's the essence of control, right? You have, let's say, two or more systems. Each of them has a particular set of potentialities that describe it in all possible environments. Right. I mean, you can think about it that way. It's a very kind of. Also, maybe kind of an Aristotelian point of view. Right. But once you interconnect two such systems, they now constrain each other. So that means that a large part of their set of potentialities is no longer actualized. A very specific path through their behavioral possibilities has been selected. That's control. [00:46:08] Speaker B: But also, there's a phrase called enabling constraints. You used a phrase earlier called constraints. [00:46:14] Speaker A: That deconstrain. Yeah, enabling constraints. I think that's Alicia Guerrero, if I remember. Correct me, I'm not sure if she [00:46:20] Speaker B: coined it, but she does. Yeah, she does use it, right? [00:46:22] Speaker A: She does use it. [00:46:23] Speaker B: Yeah. Yeah, she probably coined it. She doesn't get enough credit for her. For her. [00:46:26] Speaker A: I agree. I agree. Yes. [00:46:29] Speaker B: But so. So by. So you have these two equally equipotential brain areas. Let's say, as we were talking about and they can both do almost anything. But then if they're connected together now, it constrains what they can do. But by, by mutually constraining each other, then they're then able to do something that neither one of them alone could do. [00:46:47] Speaker A: Exactly. [00:46:48] Speaker B: Which would be enabling. [00:46:49] Speaker A: Yeah, that's right. I mean, this is actually kind of very interesting view. So I imbibed that almost kind of wholesale from the thinking of another control theorist, Jan Willems. He also passed away. [00:47:02] Speaker B: Ah, I was going to bring up this [00:47:06] Speaker A: quite a while ago. So this idea of control as interconnection. He was just mostly interested in broadening this idea of control. If we think about control inspired by cybernetics, it has to be feedback. So feedback involves also processing. So we measure some. Something we process, we maybe compare it with some standard and process that, feed that back. It also has a very kind of reinforcement in this kind of Pavlovian or Skinnerian sense to it. But he also says there are a lot of systems that are passive. They don't really process anything. They do what they do by virtue of being coupled to other systems. [00:47:46] Speaker B: Right. [00:47:47] Speaker A: And they're still controllers. So for example, pressure valves. Pressure valves don't do any computation. They do what they do by virtue of their physical organization. And when they're coupled to another system, they accomplish a given goal. Or heat sinks. Heat sinks are control devices. Once again, restrict. They restrict, let's say, temperature profile of particular electronic circuit. Circuit. So you couple them together, all of a sudden now you have a reduction in the variety of possible behaviors that could be generated. And this is a way of accomplishing control. Feedback and signal process is just a particular way of implementing this. But not every control system is that way. One of the examples from the literature and more theoretical biology, there's this book by Morenoisio on biological autonomy. They're building on the ideas of people like Francisco Varela and Robert Rosen. And they talk about networks of constraints that are operating at a higher, let's say, layer. Right. Because on the physical level, organisms have to be open systems. That's just not possible to not exchange energy, matter and information with our surroundings. But if you look at the structure of that, that has to remain somehow stable and variant. So this is this idea of networks of self maintaining constraints. So they give a nice example of the circulatory system that if you didn't have a network of blood vessels, you had blood that carries, let's say, nutrients and it'd be sloshing all around, but the blood vessels just channel and control it. And so they control this transport. But again, there's no feedback going on. There's no processing. It's just the constraints. [00:49:43] Speaker B: Passive controller. [00:49:44] Speaker A: Yes, it's passive controller. And active controllers, of course, are very important and useful. In particular, I mean, kind of. Yeah. I mean, when Mietosh Klowski talks about neurons as feedback controllers, that is the same sort of thinking as, let's say, coming back to Gibson. So Gibson had this book called the Senses Considered as Perceptual System. And what he meant by that is that the senses are not just these passive receivers, they actually are active control systems. He had this phrase that we have to move in order to perceive, but we have to perceive in order to move. Right. And so this idea of action coupled to perception in this feedback loop. And so thinking about controlling your perception, controlling information is a very useful, useful way of thinking about it. But, you know, of course, you would have various types of control available to, let's say, both biological and artificial systems. Some of it is really of the cybernetic variety, but some of this is of this more kind of, you know, Jan Willems inspired control by interconnection. [00:50:55] Speaker B: Yeah. So that, that's interesting because the. The organizational closure or closure of constraints ideas that you were talking about from Reno and Mosio. [00:51:03] Speaker A: Right. [00:51:04] Speaker B: Is it a perfect analogy to go to the sensor motor coupling, active or sorry, not active action perception coupling? I mean, I've kind of. Because I really love that work with the organizational closure. And I've had multiple people on here to talk about it, but it's unclear to me how to really apply it. Thinking about cognition and mind and brains. [00:51:24] Speaker A: It's unclear to me as well, because again, I mean, I think they are very careful about saying that this. We have to. We have to be very precise in what we mean by closure. Right, right. Because again, organisms have to be open systems. [00:51:42] Speaker B: You know, the closure only has to do with the constraints like that. [00:51:45] Speaker A: Yes. [00:51:47] Speaker B: So the classic example is an individual cell, Right. That it. It produces things that it is dependent on for survival. So in that sense, it's quote, unquote, autopoietic because it creates itself and maintains its itself. [00:52:00] Speaker A: That's right. Right. But I think Moreno and Mosio basically said that what creates and maintains itself is not the physical organization, but sort of the, you know, the formal. Right. The formal description of the cell, in a sense. I mean, others have made that same point, you know, before in multiple contexts. But. Yes, I mean, that's. Because the thing is that this is actually also very important. To realize that control systems actually have three dimensions, we'll mostly talk about one of them, which is when the system is put together, it functions and it functions by being coupled to its environment. Whether it uses a feedback mechanism or just a direct physical coupling is a secondary point. But it behaves, it does what it's supposed to do. But on the other hand, and there's this other dimension of how it maintains that behavior, meaning that what does a system do if something goes wrong? If we're talking about something like a thermostat, a thermostat breaks down, you can fix it yourself, but you do have to go in. The thermostat cannot do that by itself. And then there's the third aspect is evolution, right? If all of a sudden you have a better way of accomplishing that same function, the system must be modified, right? [00:53:35] Speaker B: Are you talking about non living systems? Are you talking about a mechanical device, control? [00:53:41] Speaker A: No, I'm talking. No, not necessarily. I'm talking about control systems understood broadly. I'm using the thermostat as an example because it actually doesn embody all three of these aspects too, right? I mean, there's this notion of somebody has to maintain that system. Either it maintains itself or you have somebody coming in and fixing it. The system does what it does by virtue of the way it's interconnected with this environment. And finally there's this whole notion of how does it respond to really qualitative changes. So in the case of a thermostat, it would be something like a software update. But in case of living organisms, you acquire, let's say acquire new skills, you acquire new ways of sensing your environment. That was actually a very interesting idea idea that Robert Rosen put together, this notion of emergence relative to a model. I don't want to talk about all the thorny issues surrounding emergence because I think there's a lot of confusion in that realm. But he did have an interesting idea, this notion of emergence relative to a model. And he used the immune system as a good example of this. When there are certain phenomena going on in your environment and your current set of, let's say your current repertoire of sensory primitives is not good at modeling or predicting them, you add another one, right? So that's why the immune system is really good at kind of generating antibodies for almost anything, as long as that so of pressure is persistent, right? I mean, it's almost like this evolution going on at the timescale of an organism as opposed to generation. So this actually is an example of all three of These functions, the being, the behaving and the becoming, those are the kind of three aspects of control. And most control engineers mostly focus on the behaving part, right. Like design a control system that does what it's supposed to do, and here's how it's built. But then there's the being, well, what's the life cycle of that system? Again, there are good models for that in the technical realm. There are models for that. The biologists use their neurobiologists. And then there's the becoming part. That's the, you know, how does it evolve, how does it. How does it maintain itself not only to keep doing what it's doing, but what if there's some really pressing need to change the way things are going? So I think this view of. Right. So this idea that you can think about control as a network of constraints, I think that's part of that. And I think that is the part of this system kind of maintaining itself in order to keep doing what it's supposed to do. [00:56:51] Speaker B: Well, so I was interested because I read your biological autonomy piece, right. That talks about those organizational closure ideas from Moreno and Masio and others. But part of your point, I think your point in that was that the idea like the behavioralist, behaviorist approach. [00:57:12] Speaker A: Behavioralist, behavioral approach. [00:57:15] Speaker B: I don't want to say behaviorist. Right. [00:57:16] Speaker A: It's not behaviorist. [00:57:17] Speaker B: Skinnerian. Yeah, but it's kind of behaviorist because. [00:57:20] Speaker A: So kind of. [00:57:21] Speaker B: So I don't, I don't know enough about Willems and the behavioral approach to say this, but was your point that that kind of approach is in line with the idea of constraint closure, or we should think of it more so from the behavioral approach than we should from constraint closure? [00:57:42] Speaker A: No, I think my point there was that constraint closure, disorganizational closure is an important concept and that there is a framework that could be repurposed from control theory to reasonable. [00:58:00] Speaker B: To formalize it. [00:58:00] Speaker A: To formalize it. The reason why I don't want to assimilate, you know, Willems to Skinner, is because Skinner definitely distinguished inputs and outputs, right? I mean, you have stimulus response, but they both. [00:58:22] Speaker B: They both like black boxes. Just. [00:58:24] Speaker A: They're both like black box. Yes, but Will. And first of all, he said that it's a very interesting philosophical kind of thing. He was inspired by cybernetics and general systems theory. He basically said that in several places in citing all these books on. One would call general theory of systems, he said, let's think about systems as sets, meaning that, like I said, a system is Described by all of its potential, all possible situations it can find itself in, because certain potentialities are just completely ruled out. Because if a system could generate absolutely anything, it's not interesting, right? [00:59:07] Speaker B: Is it all the trajectories, for example, [00:59:10] Speaker A: all the trajectories that you could take. And then he said, if you look at these trajectories, you're tempted to say, oh, this is an input and this is an output. And he said that, well, that designation can only come from experience. So, for example, the way that he formalized an input, he said, in principle, it's part of the system's trajectory that's completely unconstrained. You cannot explain it by any law. It could be anything. Right. Whereas an output is something that you can actually see. Like, once you designate an input and then you start thinking about what affects what, then you can sort of argue, okay, well, this part of the system's trajectory seems to depend on this other part, and it does not anticipate it. Therefore, it's legitimate to call it an output. So an output is already constrained. An input is something that's, in principle, [01:00:03] Speaker B: unconstrained, unless the input is constrained by the output. [01:00:07] Speaker A: Unless the input is constrained by the output, in which case, of course, you're going to see something like a closed loop behavior. And in that case, the system would not have an input in the sense of Willems. Right. It would be. So this is what you would call an autonomous system to a control theorist, an autonomous system. This is actually very interesting to a control theorist, an autonomous system. A system does not have an external input. Right. It does its thing. [01:00:35] Speaker B: Let me just read. So I sent that paper to a couple of my friends. Sorry, did you want to say something more? [01:00:40] Speaker A: I did want to say one more thing. Right. That Willems viewpoint is not completely black box oriented. I think a very abstract approach. So it's not mainstream in control at all. There are certain communities within control theory that are really influenced by it, but it's not by any means mainstream for a couple of reasons. One is that it is really good at nailing down basic questions about system identification, going from observed behavior from data to models. [01:01:15] Speaker B: That's the whole. [01:01:15] Speaker A: That's his whole point. [01:01:16] Speaker B: That's his whole point. Pure data to models to models. [01:01:20] Speaker A: However, the other question that the control theorists occupy themselves, and as I mentioned earlier, is synthesis. And synthesis is not something that behavioral viewpoint lends itself easily to. He does talk about these notions of. He talks about tearing and zooming. Tearing just means identifying modules, breaking up a system into Modules that constrain each other and zooming. That's when you zoom in a particular module and you describe that behaviorally and you say, oh, but the way that it plays a part in as bigger system is, well, it shares certain variables with certain other modules. Therefore these variables are constrained. And this is how we explain the whole thing. It's an explanatory framework, but it's not such a good framework for actually thinking about how to build systems. [01:02:08] Speaker B: Right. It's about how to infer what the system is. [01:02:11] Speaker A: Yes, but to me, the big value of that is precisely this idea that you have to expand the scope of what you mean by control. That's not all just this Wiener Ashby style cybernetic feedback. There are tons of controllers out there that accomplish control functions by not being feedback controllers. [01:02:30] Speaker B: Okay, all right. What I was going to say is. So I sent that manuscript to a few of my closure constraint friends because I was trying to better understand what you were trying to say. And I got an email back. I won't say who it's from, but can I just read you that email? I think you actually just cleared up most, if not all of this. But here's the responses. I'm a bit confused by the fact that Raginsky seems to define the traditional cybernetic paradigm, end quote. In a very detached info processing way that's completely opposite to how I think of it. Namely, I've always seen cybernetics as involving a rejection of the idea of inputs and outputs and the reduction of feedback and control to physical coupling. So I typically see it discussed in opposition to the classic cognitivist info processing story. That's probably because they share the same roots and because I focused on Ashby. And those ideas are very much the Ashby and aspect of cybernetics that influenced Varela. And this person says they're not clear on where Willems breaks with Ashby rather than continuing his project. And it goes on. But that's probably enough for you to respond to, which I think you were already touching on some of this. [01:03:45] Speaker A: I think so, yes. So it's actually pretty interesting. It is correct that if we talk about inputs and outputs and output to input feedback and associate that with cybernetics, that would be through Wiener, because Wiener's framework for control is really this very classic. What's called frequency domain control, so developed in particular by people like Bode and Nyquist at Bell Lab. Right. You describe control systems by means of what's called transfer functions. And transfer functions are really transform. They tell you how a system maps an input to an output feedback is then precisely a kind of a thing where you take an output of a system and then you loop it back. So the subsequent input is affected in some part by the output outputs of the system. Right. That's how, for example, you maintain a set point. This is whole notion of, let's say, or this idea of homeostasis, right? I mean, like for example, maintaining a certain level of blood calcium. And it has to be a very, very tightly regulated set point. And so you use different ways of processing that. In a sense it is information processing, right? Because what are you processing? You're processing deviations on the set point. That's informal. So Wiener's viewpoint was really kind of this classical input output view. Whereas Ashby, writing before Wiener, he actually did zero in on state space models before that became mainstream and controlled, that would be 1960s. And he did think about systems. I don't think he really emphasized that much in Introduction to Cybernetics or in Design for a Brain, but in his certain writings say he wanted to model learning by trial and error. And he did have a state space model. It was basically a differential equation where the dynamics. So that's the right hand side that tells you how the time derivative of the state is affected by the current state. Also included parameters that he called and he basically, those are control inputs. And by varying those parameters you can induce different trajectories through the system. And he basically said if you vary those parameters discontinuously, then you can, you know, kind of switch between different types of behavior. But I don't think he kind of fully grasped this idea that. Well, you can also make those parameters actually dependent on the state. That would be the state feedback. Right. [01:06:30] Speaker B: In a sense differ from Willems then. Sorry. [01:06:32] Speaker A: Okay, so no, I mean that's actually. I know I do tend to give kind of long winded explanations. So Willems once again says, and that's actually echoes the point that I was making earlier, that if I have two systems that exhibit similar external behavior, you cannot necessarily say that they are the same internally. So state space model is an internal model. And the idea. So what Willems says is that a state. State is in a sense a property of the. It's not a property of a system, it's a property of a system representation. Because there are multiple ways of thinking about states. So one sort of standard way. What is a state generally or conceptually? It's basically a way to summarize the system's memory for the purpose of predicting and directing future behavior. It Basically means the following. If the system receives some sort of information or stimuli or whatever, or behave that certain way in the past, what do I need to know about it right now in order to predict what it's going to do in the future? Or if I apply future inputs, how do I steer the output towards the desired trajectory? So anything that allows you to decouple past knowledge from this memory description is a state. And then not all states are created equal. For example, one very trivial state is just the entire state past. Because obviously if I know the entire past and if I know the system laws, I know everything's going to do in the future. [01:08:13] Speaker B: Could you substitute model for state here? [01:08:16] Speaker A: In a way, yes, you can, as long as. But again, what I'm saying is that the notion of state has to come with additional sort of descriptors attached to it. So I can talk about something like a minimal state, which basically means that it's something that, that if I process it further, I will lose information. I'll no longer be able to say anything sensible about the system in the future. And so you can observe, let's say, the system's trajectories. And some of those trajectories will be equivalent in the sense that even though they're different from a viewpoint of the system's future, it doesn't matter which of those trajectories was actually instantiated. So I can lump those trajectories into a single label and say, that's going to be my state. So this is the idea of state that came partly from computer science, partly from control theory through these ideas relating to optimal control. And what Willems is saying is that unless you can actually open up the system and look inside that state is just a property of how you choose to measure the state system. I think it's a very important point. It does actually touch on, you know, this, this second order cybernetics. That's, that would be, you know, von Furster and you know, Varela and people like that. [01:09:40] Speaker B: But to Willems, then the, the, the state. Okay, so let's say you could open up the, the system and see its innards. Like let's say a brain or some, yeah. Machine device or whatever, some circuit board. Yeah, but to Willems, those would be equivalent because they can all be described by the same state. From a control perspective, they would be equivalent because that, it's that state that, that gives equivalent outputs, let's say. So it doesn't really matter or what it's made of or how it's shaped [01:10:21] Speaker A: well, okay, there are multiple ways in which the scan matter. This was actually one of the driving forces behind this kind of shift of perspective to state space modeling. Because if you view systems just from input output perspective, it does mask all sorts of internal phenomena going on. For example, example, you could actually have internal state dynamics that are somehow either unstable or uncontrollable from the input output point of view. If you're standing outside the system, you might not see any of that because it doesn't manifest itself in the external behavior. There are also certain state descriptions that are just more useful as explanatory frames. So for example, in certain types of systems there's something called the modality form. And that basically means that it has a high dimensional state. So the state has multiple components, it's a vector. But each of those components evolves independently. They're not coupled, they have their own dynamics. And the reason why you see any sort of correlation is because then all those state coordinates are mixed combined, let's say linearly to produce the output output. Right. Sometimes that kind of framing is very useful. For example, I remember reading an overview on sort of the dynamical systems perspective on general anesthesia by Emory Brown. And his very kind of high level point was that the effect of anesthesia has to with do do with synchronization and desynchronization of different areas of the brain. From that point of view, thinking about this in terms of modes and whether they're coupled or non interacting is a very fruitful perspective. But it might not be, or it could also be fruitful for neurological conditions like epilepsy, for example, that also has to do with synchronization and coupling. But another instance is, is thinking about state description that consists of a bunch of non interacting modes. And the only way to make them interact is to combine them at the output is not useful either for building the system. Right. Because that might mean that you might use some resources less economically than if you had let certain states affect other states. Because that could actually affect something called controllability. Controllability is the system's responsiveness to inputs in order to induce it to, you know, move from one state to another. Right. [01:13:13] Speaker B: And so sensitivity would be another. [01:13:16] Speaker A: It's, it's. Yeah, it's, it's related to sensitivity. [01:13:19] Speaker B: Yeah, sorry I messed that up because I know that's different. Technically it's somewhat different. [01:13:23] Speaker A: But, but yeah. [01:13:26] Speaker B: Okay, then zooming way back out to one of my original questions for you. [01:13:31] Speaker A: Right. [01:13:31] Speaker B: The, the, these differences between biological and artificial intelligence. Does it so Based on how you view what the term intelligence means, do you. Do you see any difference in, like, if a human and an AI system produced the same token sequence? Is the artificial system intelligent in the same way that the human. Human is intelligent? And does it matter? [01:14:02] Speaker A: I don't. First of all, I don't know what intelligence really means. [01:14:06] Speaker B: Glad you said that. I don't either. It's, it's. It's. [01:14:10] Speaker A: It's one of these, you know, really sort of malleable concepts that means, you know, whatever it is that, you know, [01:14:16] Speaker B: there are so many of those in cognitive sciences especially, like, they're all up for grabs kind of. [01:14:22] Speaker A: Yes, exactly. So, you know, there's a very narrow definition of intelligence, I think, due to John McCarthy, and it's basically the ability to solve tasks in varied environments, et cetera, et cetera. But can we reduce everything to solving a bunch of tasks? I mean, I think this is actually kind of detrimental to almost any aspect of life and culture that today's artificial intelligence companies take touch. They reduce it all to this, you know, morass of quantification and, you know, like, list of tasks. [01:15:00] Speaker B: Well, I'm thinking more in terms of the Willems approach that we were just talking about, or from a control engineering kind of approach where that leaves us. Where it should leave me in my thinking. [01:15:15] Speaker A: Okay, so control is about effective interaction. Right. Does a system interact with its environment effectively? Is it. I like the term coherence. It comes from all sorts of domains. I mean, in philosophy, it's an overall property of a system that has or system of thought, technological system, social system, whatever, it consists of multiple parts, and each of those parts, even though you can explain them in isolation, in order to really explain what it does within the context of that system, you kind of have to. Each of them presupposes the other. It's not this indissoluble totality that you kind of see in some of the idealist philosophy, like in Hegel, for example, but it is a concept that allows you to kind of think about a system in terms of its goals and values. Right. For example, the inner Internet, as a system, it's coherent as long as all the users connected to it can accomplish what they want. But from the viewpoint of somebody else, it can become incoherent. For example, if you just see the Internet as a source of social instability or something like that, then, yeah, I mean, that results in loss of coherence in various valuable social, economic arrangements. Right. So just like intelligence, coherence is also this concept that you can sort of know it when you see it. But I'm very skeptical of criteria and tests for intelligence and definitions for intelligence and especially the kind of a thing where you say, oh, an AI system can produce certain behaviors that are often indistinguishable from humans. Well, you know, what is it doing? You know, what is it meant to accomplish? How do we interpret it? Right? I mean, a lot of it is sort of, in a sense, in the eye of the beholder, right? I mean, if the system accomplishes what we want and if it allows us to interact with the world and with us effectively according to some criteria, then that's meaningful. Whether or not we can start adjudicating, you know, like meaning of intelligence. I'm not sure that's a question that control engineering is equipped to answer or should be asked to answer. [01:17:44] Speaker B: Well, I'm the host, so I can ask you anything that. Whether you should or should, of course, but I think I've heard you mention that you don't believe that AGI is a thing or a thing that can be achieved or, or what are your. Am I quoting right? [01:18:05] Speaker A: Okay. It's also a thing that had this kind of a chameleonic definition, right? There's this definition that now, I think, now we don't talk about AGI. Apparently some of the quarters in this, the asi, right, of artificial superintelligence, that basically, basically means. But that goes back to thinking of people like IJ Goode, for example. So he talked about ultra intelligence. And it's basically an artificial system whose capabilities can exceed human capabilities in any economically valuable domain. And that already is where you see that. This is where I see issues because this notion of economically valuable is not stable, right? We don't want to tie ourselves to fixed markers. What's economically valuable now might not be interesting or economically valuable in the future. In the future, something else might also because these are fundamentally open ended arrangements like social economic arrangements. There's that definition. Complex systems. There's that definition. And I think it's incoherent precisely for the fact that. Fact that every time we're asked to imagine a system that's better than humans at something, it's typically framed as a thought experiment where you say, okay, imagine an alternative world which is just like ours, except for this one thing. There's a system, there's this computational arrangement, there's this artificial system that can, let's say, perform any possible economically valuable task or cognitive task at a superhuman level. And the problem with that is that this is not a possible world, which is just like ours, except in this other aspect. These differences are going to amplify so much that reasoning about that world becomes completely, completely impossible. You're going to hit this regress. You're going to start thinking about, okay, how does this affect law? How does this affect economics? How does this affect family relations? How does this affect student teacher relationships? How does this affect the possibility of, I don't know, our notions of our role in the universe, et cetera. So it's not a coherent thought experiment that you can reason about without engaging, let's say, your own predispositions, your own beliefs. And so it becomes this very unstable kind of a foundation. So from that perspective, yeah, I don't think that it's a coherent concept. Talking about specialized intelligence. Intelligence, if you want to use the term intelligence, that's fine. But, for example, take an ordinary house fly. Its ability to evade targets and fly [01:21:00] Speaker B: is confounding, is what it is. [01:21:02] Speaker A: It's confounding, exactly. It's annoying, right? But they're really good at it. Can it solve differential equations? No, but. But that doesn't really matter. [01:21:17] Speaker B: Okay. [01:21:18] Speaker A: And then, of course, now there's. I think I remember there was this interesting thing, that there was a legal definition of OpenAI, I guess, you know, like, encoded in the contract between. I think, legal definition of AGI encoded in the contract. I forgot who it was. Between OpenAI and, you know, whatever it was commercializing its technology, it had to do with. We achieve AGI when the rate of profit from these systems exceeds. [01:21:43] Speaker B: Oh, that's right. That's right. Yeah. Such a bullshit. [01:21:46] Speaker A: Exactly. [01:21:49] Speaker B: Why do you hate Turing so much? [01:21:51] Speaker A: I don't, but I think the. [01:21:58] Speaker B: I said, you've written about how people misread Turing's 1950 famous paper. [01:22:04] Speaker A: Yeah. [01:22:04] Speaker B: Yes. [01:22:05] Speaker A: The computing machinery and intelligence. There are a couple of things. One is that they don't really read the paper. They just kind of project onto it whatever they want it to say. For example, this whole idea that, well, these LLMs, the latest models. How can you say they're not conscious when they do X, Y and Z? And they pass the Turing test. There were no Turing. Turing's paper specifically says that we don't need to bring in anything about consciousness. In fact, he was responding to an objection from consciousness, saying that you need consciousness in order to intelligently interact with your world. Right. So he basically just disavows that. So he's completely agnostic on the question of consciousness. So therefore, falling back on Turing to argue that the question of whether modern AI systems are conscious is contagious. Test it. I think it's disingenuous and part of this kind of originating in a certain kind of vagueness of Turing's language and some of the frankly, very silly points, like he says, well, if ESP is possible, then all bets are off. What can we do? But I think, to me, the main reason why we should not really, really see that paper as anything other than historical curiosity is this fixation on language and text to the exclusion of everything else. Because at some point, I read a book called Cerebral Symphony by William Calvin as a neuroscientist, and he had this interesting idea, and I think he wrote, you know, papers, you know, scholarly papers about this, that language should not be viewed as something that just came out of nowhere. It's more that, you know, our hominid ancestors at some point evolved. They learned to, let's say, like, throw a rock farther and with better aim. And that required orchestrating and planning sequences of actions that can be represented in some convenient chunks. Chunks. And so neural architectures that could do this are so universal. This is, again, this kind of constraint that deconstrains. They're so universal, they can be recruited for all sorts of other tasks where it turns out that sequencing and planning and anticipation are very important. So that's language. And why do we love poetry? Well, because of that. I mean, you read a poem, you want to know what happens next. So this is this kind of idea that you think about breaking up a complicated task or phenomenon into these kind of units that you can reason about sequentially. So this idea that there's this universal architecture that is so flexible that this idea of predicting the next token, which is, of course, very glib, I don't necessarily like that. But this idea that so many other tasks that somehow can be mapped to this can be supported by the same type of basic neural architecture that can then serve to broaden the scope of our activities, create culture, create appreciation for music and poetry and art and beyond just throwing rocks at an animal that you wish to slay, things like sports and music and, you know, and things like that. This, I think, is a more fruitful direction. That, yes, I mean, we do seem to. It does seem like we did stumble upon this idea now constructively, instead of just seeing it in nature, you know, evolution having sort of made use of the same kind of universal type of neural architecture, instead of saying that we just somehow took one particular instantiation of that architecture, which is language and text and Just substituted everything else for it. [01:26:28] Speaker B: Well, there are some special things about language. Right. Like I can write a letter to my great, great, great, great, great, great grandson. [01:26:35] Speaker A: Yes. [01:26:35] Speaker B: And then if he exists, he can read that letter and wow, it's traveled through time. [01:26:40] Speaker A: Of course, yes. But what I'm saying is that this idea that you evaluate whether a system is thinking or intellig. Right. Because for example, these days there are These things called VLAs, vision, language action models. Suppose you don't see the language track, you only see the actions. Right. So you have a system that, you know, this robotic system that's accomplishing complicated tasks in its environment. Suppose you don't even see who's doing that. You're just confronted with some sort of description of some activity being carried out. There's no language involved. It's just pure, let's say, pure vision. Could we then also argue that there's some intelligence underlying it? Sure, but we should be able to argue that as long as it does involve this notion of kind of sequentially orchestrating some actions. Yeah. Language is special, of course. I mean, that is the foundation of the way that we record our knowledge, the way that we externalize it, that way that we can communicate it through time. Yes. I mean it's very special. But to equate that to some sort of an operational criterion for intelligence. I think that's exactly what is so anachronistic about Turing's paper, that we should no longer just operate with that. [01:28:13] Speaker B: I agree, but I think part of it, which you mentioned also is that, that some of the language was maybe antiquated, like there is drift and semantic drift with what terms mean. And I think part of Turing's point actually was almost saying, so maybe during the day people thought like thought was a conscious endeavor. And he's almost redefining it to say, all right, I'm going to operationalize it to mean this and so we can do away without consciousness. So what I'm calling thinking is actually different than what people usually think. [01:28:46] Speaker A: He does say that maybe, you know, he says maybe in the 21st century, let's say, you know, this language of, you know, thinking computers will no longer seem, you know, anything special. We'll just use it. [01:28:58] Speaker B: Right. So yeah, okay, it's just maybe a couple quick hitting questions before I, before I let you go here. So perceptual control theory from William Powers originally and active inference are both control based, you know, ideas, almost grand unified theories of the brain. [01:29:25] Speaker A: Yes. [01:29:25] Speaker B: Do you have just very open ended question thoughts on active inference and, or perceptual control theory. [01:29:31] Speaker A: I'm. Well, I think by now you probably have gathered that I'm very skeptical about grand unified theories. I think active inference has a bit of, a bit of more of that aspect than perceptual control theory in that it's a theory of everything in a sense that when you formulate a particular, let's say, theory, you first formulate kind of general principles and then you can proceed from that to, let's say, approximations, right? I mean, let's go by the analogy with, let's say Newtonian mechanics, right? There are certain laws, laws, and the main law is that the description of mechanical motion, motion and mechanical or physical systems, normally you would need, let's say if these are smooth trajectories, you'd need to know all of the derivatives of that function in order to describe it. And Newton's laws basically say. No, you only need the first two because, for example, the second derivative, that's acceleration, is going to be a function of positions in the velocities. And that means that all the higher order derivatives can be reducible to those. You don't need them anymore. And then based on that, then you can start formulating also you can start postulating what do various forces look like, so gravitational attraction, et cetera, et cetera. And then you realize that you end up with various problems where certain things cannot be done. Exactly. So in order to tractably analyze them, you propose approximately. And approximations are a separate component in that theory. My attempts at reading all the papers on active inference, starting with some of the earlier papers by Friston, is that it's all jumbled together. And if a certain, they're like, oh, but this density is now an approximation of this other density that's not computable directly. And then of course, every time you read a new paper, the equations look very different, different. Some new terms are added or some new explanations or maybe previous explanations that were not necessarily cogent or modified. So I mean, I'm not a puppetarian. I don't think that falsifiability is sort of the definitive criterion in science because [01:31:59] Speaker B: I thought, who is it that you. Is it meal that gives this formal definition of how theory should be updated and stuff. So I thought that you would be perfectly fine with this sort of changing as you go. [01:32:10] Speaker A: Well, I'm. Well, okay. So Mill kind of draws on Imre Lakatosh this whole idea that when you're falsifying, you're not falsifying just a theory. You're falsifying all sorts of other things. They're linked together, right? And so when you make a prediction, a prediction does not pan out. That means that a whole conjunction of various clauses got falsified. You don't know which one of them is false. And then you're going to do this maneuver where you're going to start saying, okay, I can sacrifice this one. Maybe it has to do with the fact that maybe my equipment was not calibrated. Okay, let's check and do that. There's always going to be like a core theory that you're going to hold on to for as long as you can, and you're only going to discard that when absolutely necessary. Sometimes that led to discovery of new planets when discovery of Neptune happened, because there are certain deviations from the predictions of Newtonian mechanics. And instead of saying, oh, Newtonian mechanism mechanics, description of motion of planets must be incorrect. No, there has to be another planet. Right? So the core theory was not modified. You introduced another thing. It's not clear in active inference. What is the core theory, what are the auxiliary theories, what are various assumptions like ceteris paribus clauses, et cetera, et cetera. Right. So from that perspective, and also I think it's a valuable meta theory, meaning that. So Friston calls the free energy principle. Principles are. Or for example, in mechanics, there's the principle of least action. Principles are something that you take for granted. They're devices for generating theories. So from that perspective, I think it's a useful framework. It is a framework that says if what we're manipulating are these representations of our world as probabilistic objects, and there's this notion of kind of causal separation, this whole notion of a Markov blanket, et cetera, these are all very useful concepts. But somehow I think that there is a bit of. There's a bit of a. I don't want to use that comparison to epicycles, but there's a little bit of that [01:34:25] Speaker B: in active inference to manipulate to make it fit. [01:34:28] Speaker A: Right. With perceptual control theory, if I understand correctly. The main tenet of that is that organisms don't control their world, they control their perception of it. Right? In a sense. In one sense it's definitely true, because we never have pure, unmediated access to the world. I mean, we can only perceive it. And when we take an action, the only way we can assess that that action succeeded is again by perception. But it is also this whole idea, for example, in contrast to, again, somebody like Gibson who would say who would talk about direct perception, right? There's no processing going on. We just direct. Directly. He talks about information pickup, right? [01:35:26] Speaker B: You know, information is in the energy [01:35:28] Speaker A: arrays in the environment, and we directly pick those up. The thing is that, you know, like, as somebody who's, you know, all of whose degrees are in electrical engineering, I always wanted to ask, well, has he ever put together a radio? All right. [01:35:42] Speaker B: I mean, that's. [01:35:44] Speaker A: He did make fundamental contributions. For example, there's this very fundamental concept in computer vision called the optics flow. And that's basically that. That represents, you know, what your retina sees when you move about in the world and when the world moves about you. That was Gibson. But some. [01:36:02] Speaker B: The information to perceive what is actually out in the world is in that optic flow. And you don't have to do any representation of the signal and transformation because it's. [01:36:11] Speaker A: I mean, the right. Representation. Right. Representation is also very loaded term, right? Because, for example, right now there's this fashionable talk of world models, right? [01:36:22] Speaker B: And I still don't know what a [01:36:24] Speaker A: world model really is. Nobody knows. Nobody knows what that is. But, for example, in control, there's something called the internal model principle. The internal model principle. I mean, it's kind of similar to this good regulator theorem from cybernetics, Conant and Ashby, of course, their result is very vague and mathematically somewhat shaky. But the rigorous version of that is the following. If I have a system that in some sense functions effectively in its environment, it adapts to it somehow. Let's say there's a useful signal in that environment and there are disturbances, and the system is good at rejecting the disturbances and honing in on the useful signal, then the system is in some sense equivalent to one that has a subsystem that is capable of simulating that environment. Basically means, if one understands how this adaptation takes place, we can always reason internally about the system that has some module in it that has enough knowledge about the environment in order to interact effectively with it. But on the other hand, if you take, for example, going back to our sense of balance, it doesn't really have an internal model of gravity. It's more that it's the physical coupling between some neurobiological circuits, but also just your hardware, like the semicircular canals that your system of balance uses to sense. [01:37:50] Speaker B: The hardware is almost the model itself. [01:37:53] Speaker A: Is the model itself. Exactly. And so that means that this idea that modeling is necessarily some sort of representation that you can manipulate in this kind of sense of information processing, I think that's a good Metaphor. But you should always be careful about not carrying it too far. So perceptual control theory, I think, I think it's a good frame, right? I mean, in a sense, again, this sort of senses as perceptual systems is very useful. But the thing is that I can say, if I pick up my empty glass, yeah, I'm perceiving its weight, I'm perceiving its shape, et cetera. But that's the only thing I can control. But the way I control it is, let's say, by anticipating. If I, if I let it go right now, what am I going to perceive? Well, I'm going to perceive the remainder of my sparkling water spilling onto my pants. I know that that's some information about a perception. So I'm controlling that by not dropping the glass. So I think in some sense it's a very sort of sensible way of looking at things. But once you start also doing the thing where you define like 20 different variables and you say, well, this is responsible for that and this is responsible for this other thing, you start making sharp quantitative predictions. I think the organizing principles and metaphors are more useful than precise equations in both, I think, PCT and active inference. [01:39:27] Speaker B: Now, this really makes me want to ask you about the role of anticipation in these kinds of control systems. But, but that's kind of a large topic and we don't have that much time. But it's discussing yay or nay. [01:39:40] Speaker A: Yeah, of course. I mean, it's even in the principles of neural design, when you talk about feed forward. [01:39:47] Speaker B: Feed forward is anticipation, which is a precise thing. But anticipation is a little bit fuzzier of a notion. [01:39:53] Speaker A: Sure. Yes, it's a fuzzier notion. I mean, that's for example, again, a book by Rosen on anticipatory systems. A lot of them are a lot. So the interesting thing, a lot of these systems are open loop. They're not feedback systems. Right. [01:40:07] Speaker B: Okay. Yeah. Going way to about the middle of our conversation. Just a couple more minutes, I promise. You said you used the. A pressure valve as an example of an inactive or a passive control system. [01:40:25] Speaker A: Yes. [01:40:26] Speaker B: It made me think like, what in the brain would be a passive control system? Is it the architecture itself, even though there is plasticity in the architecture? Or does that even make sense to think about what in our brains would be serving? Because our brains are so active, right. All the neurons are spiking astrocytes or calcium signaling, and there's all sorts of stuff always going on. There's all this spontaneous activity generated within. But is There an example of a passive control controller in brains. Something to think about, perhaps. [01:41:01] Speaker A: For me, I don't know. I mean, [01:41:06] Speaker B: because they're useful. It's like no energy consumption. I mean, maybe it is the architecture. [01:41:13] Speaker A: It could be just the way that. Right, because the way that the brain is put together, its organization is, in a way, a control system. Right. Because it enforces. It basically imposes constraints on how different subsystems in the brain can interact with each other. I mean, I'm almost inclined to think about digestive system as a passive control system. But then, of course, we have the gut brain axis, which kind of complicates that whole description. But I do like digestion as an example of. Well, you can think about everything as computation. Okay, what is my digestive system computing? [01:41:48] Speaker B: Oh, gosh. Do you think the brain is computing? Do you think of brain activity in computational terms or control terms? [01:41:55] Speaker A: I think sometimes it's a useful lens, but that's not an exclusive. It's not an exclusive sort of a description. Right. Because even so, speaking of Turing, again, people read into his thinking whatever they think they want to see. But in his paper that introduced the Turing machine, he was very careful to argue that certain sources of information should not be modeled computationally, because if they could be, then you can just model that, interconnect the two together. Now you have a bigger Turing machine to work with. So certain things should not be modeled as computation. Our models of it are very computational useful. For example, we certainly don't think that. One wouldn't think that, let's say the planets in the solar system and the sun are solving differential equations. They're moving by virtue of the physical coupling and laws of gravity. But the way that we can make predictions about them is that we can, let's say, speed up our predictions. We can do the predictions about the evolution of the solar system in, you know, on a computational timescale without waiting for what's actually going to happen. Whether or not those predictions are reliable is a different story. But we don't have to wait for that system to actually accomplish its trajectory in order to, you know, and talk about it. Right. But to say that, what was it Aristotle who said that, you know, nature does not move by counting? [01:43:32] Speaker B: Oh, I don't know. [01:43:34] Speaker A: It's one of the Greeks. [01:43:37] Speaker B: All right, so last question, I promise. You know, the history of neuroscience science has been almost exclusively about counting spikes. And, you know, you mentioned like, well, I don't want to talk about what the code. Is it inner spike intervals, is it timing Is it rate codes, et cetera? Do you think in a hundred years time that scientists will look back and think, man, that was really silly for neuroscience to be counting spikes. [01:44:05] Speaker A: Possibly, right? I mean it is, it could be one of those, you know, looking for your keys under the street light because, you know, because you can see better. [01:44:12] Speaker B: It's the cool thing, it's like it's the almost digital readout of an analog signal, right. And propagates. I mean it's awesome, but we put all of our eggs in that basket. [01:44:23] Speaker A: Right. But for example, there was a paper in the information literature, information theory literature called Bits through Queues. So the idea was that you have a queuing system and basically just so this is one of the ways in which you can model spike trains, right? I mean the random arrivals of spike trains are these events and then you can think about a processor that can process those. So in queuing theory literature you talk about a server and you talk about jobs that are being processed and you can talk about various trade offs in that. And so this paper actually said that let's say you have these events, an event happens, it doesn't carry any information other than that an event happen, happened. But what you're going to do is you're going to encode bits in. Can you send those events in such a way that it's the inter arrival timings that actually convey useful information, not the events themselves? And they showed that yes, again, under certain conditions on the arrival, the process of these arrivals and the way that these jobs are serviced, you can actually communicate useful information. So it's probably a little bit of both. Sometimes it could be that inter arrival timing is useful and robust in a particular scenario, but it would not be in some other scenario. So why not both? Why not also something else? Why not take multiple spike trains in a large neural population? And so it could be that if you were to aggregate them, the information about timings is more reliably preserved than information about how many spikes occur in a given second. So yeah, I think it's an interesting, it's more of an manifestation of some social aspects or sociological aspects of science than anything else. [01:46:17] Speaker B: Sure, yeah, yeah. Plus evolution is like use what you have, you know, that's what evolution does, right? [01:46:23] Speaker A: That's right. [01:46:24] Speaker B: Okay, Max. From one world model intelligence to another. I knew that this would be fun and interesting, so I appreciate your time and explanation expertise. [01:46:35] Speaker A: Well, thank you for having me. [01:46:43] Speaker B: Brain Inspired is powered by the Transmitter, an online publication that aims to deliver useful information, insights and tools to build bridges across neuroscience and advance research. Visit thetransmitter.org to explore the latest neuroscience news and perspectives written by the journalists and scientists. If you value Brain Inspired, support it through Patreon to access full length episodes, join our Discord community and even influence who I invite to the podcast. Go to Brain Inspired Co to learn more. The music you hear is a little slow jazzy blues performed by my friend Kyle Donovan. Thank you for your support. See you next time. [01:47:25] Speaker A: Sam.

Other Episodes

Episode 0

October 25, 2024 • 01:29:31
Episode Cover

BI 197 Karen Adolph: How Babies Learn to Move and Think

Support the show to get full episodes and join the Discord community. The Transmitter is an online publication that aims to deliver useful information,...

Listen

Episode 0

December 15, 2020 • 01:42:12
Episode Cover

BI 092 Russ Poldrack: Cognitive Ontologies

Russ and I discuss cognitive ontologies - the "parts" of the mind and their relations - as an ongoing dilemma of how to map...

Listen

Episode 0

November 29, 2022 • 01:22:27
Episode Cover

BI 154 Anne Collins: Learning with Working Memory

Check out my free video series about what's missing in AI and Neuroscience Support the show to get full episodes and join the Discord...

Listen