I don't like the term artificial intelligence and superintelligence I think it can mean way too many things so instead I would say because they are so many definitions of intelligence I would just drop the word intelligence and I would say that in general we are talking about information processing systems so that somehow exist in physics so you can have for example carbon-based information processing systems which are biological systems and silicon based information processing systems which are most of the current AI systems and on top of that I would say you can have superhuman silicon based information processing systems and that are general or very narrow so they are equivalent or super human in some narrow task or more generally for more than one task and the more tasks it can do the more general it is but different architectures can be general in different ways so like xlstms for instance still work pretty well in time series analysis while Transformers are steamrolling a lot of other domains while xlstms are still general because neural networks are generally pretty general since they can be applied to lot of tasks as they're universal function approximators and a lot of narrow systems are better in various ways than more general ones and have different advantages and disadvantages and sometimes hybrids are the best approach and it's constant fight between bitter lesson and human crafted systems. And because there are different definitions of intelligence I would just take them apart and prepend what's in their definition, so for instance there is a definition of intelligence by Francois Chollet which says that it's the ability to generalize so that it's already included in the general information processing part and you can also distinguish between trained in general set of tasks vs ability to generalize to unseen tasks on the fly and then there is the definition of intelligence by Pei Wang which is about adaptation under the limited resources so you could prelend adaptive information processing system or you could say resources efficient information processing system and that also depends like which resources it's connected to then like there's the definition of intelligence by Joscha Bach that's as it's like ability to make models and you could maybe say it's like modeling information processing system or ability to make predictive models so it's predictive modeling information processing system and then some people also put in agency and autonomita and I don't see agency come often in the definitions of intelligence but sure you could say agentic information processing system and autonomous information processing system and I think that's what you actually require in order for all those doomy AI killing everyone scenarios since if you are if you are superhuman information processing system that doesn't have agency and autonomity then it will not cause harm because it's just doesn't have access to the actions you know if it doesn't have general agency, you could say one type of agency is using the Python or lean compiler as a tool and another type of agency that's more general is like connecting it to the whole internet so technically the AI system could hire humans there or buy robotics to make itself also physical if it's just a software system, you could also say that there's alignment to human values, either more broadly to some social or political values or just not harm a human or just not kill human or all of humanity so you could say that would be notkilleveryoneist information processing system or aligned information processing system. And if you buy the instrumental convergence thesis then which I don't really buy fully I think since if you look at mathematicians and scientists, the more smart a human is then on average the less power seeking he is and less wants to kill everyone he is since if you compare mathematicians and scientists to politicians and other people actually in power like pure mathematicians mostly love to scribble equations if I oversimplify it and don't really want to kill humans I think the same can be for superhuman mathematical information processing systems, and self-preservation instinct isn't given in all information processing systems and same for going for any goals at all autonomously, so what I think the AI risk people are afraid of in all this terminology is autonomous agentic killeveryoneist superhuman adaptive generally capable resource acquiring power seeking silicon based information processing system that also has access to a lot of our infrastructure and somehow can gather so many resources to kill everyone which I'm not sure is like practically easy but I don't think it's impossible so I think it would be nice if they would specify much more what kind of information processing system they're afraid of because if they just put all superhuman information processing systems that are more general into just this one superhuman AI bin then you will basically ban nonagentic superhuman silicon-based mathematicians and I think that's dumb but yeah otherwise makes sense even though I suspect it's actually relatively improbable like I'm missing more empirical evidence in order to see how it's the more probable path and I wanna see less guessing and because if you chain a lot of relatively improbable events because there are lot of other events that can happen then the resulting probability can be very low but it depends what is your probability of the everything leading to it. Actually, when I think about it, my argument against instrumental convergence is pretty weak here. " I don't like how vaguely are the terms "artificial intelligence", "AGI", "superintelligence" and "intelligence" used. I think it can mean way too many things. So instead, I would say, because there are so many definitions of intelligence, I would just drop the word "intelligence." I would say that in general we are talking about information processing systems that somehow exist in physics. So you can have, for example, carbon-based information processing systems, which are biological systems. And silicon-based information processing systems, which are most of the current AI systems. And on top of that, you can have superhuman silicon-based information processing systems that are general or very narrow. So they are equivalent or superhuman in some narrow task, or more generally for more than one task. And the more tasks it can do, the more general it is. But different architectures can be general in different ways. Like xLSTMs for instance still work pretty well in time series analysis, while Transformers are steamrolling a lot of other domains. While xLSTMs are still general, because neural networks are generally pretty general, since they can be applied to a lot of tasks as they're universal function approximators. And a lot of narrow systems are better in various ways than more general ones, and have different advantages and disadvantages. And sometimes hybrids are the best approach. And it's a constant fight between the bitter lesson and human-crafted systems. And because there are different definitions of intelligence, I would just take them apart and prepend what's in their definition. So for instance, there is a definition of intelligence by François Chollet, which says that it's the ability to generalize. So that's already included in the "general information processing" part. And you can also distinguish between "trained on a general set of tasks" vs "ability to generalize to unseen tasks on the fly." All of that lives on a spectrum. Then there is the definition of intelligence by Pei Wang, which is about adaptation under limited resources. So you could prepend "adaptive information processing system." Or you could say "resource-efficient information processing system." And that also matters to which resources it has access to. Then there's the definition of intelligence by Joscha Bach, that it's the ability to make models. And you could maybe say it's a "modeling information processing system." Or "ability to make predictive models," so it's a "predictive modeling information processing system." Then some people also put in agency and autonomy. And I don't see agency come up often in the definitions of intelligence, but sure, you could say "agentic information processing system" and "autonomous information processing system." And I think agency and autonomity is what you actually require for all those doomy "AI killing everyone" scenarios. Since if you are a superhuman information processing system that doesn't have agency and autonomy, then it will not cause harm, because it just doesn't have access to the actions that would lead to that. You could say one type of agency is using the Python or Lean compiler as a tool. And another type of agency that's more general is connecting it to the whole computer or the whole internet. So technically the AI system could hire humans there, or buy robotics to make itself also physical if it's just a software system. But there can be a mathematics solving superhuman general information processing system that just perfectly finds equations in data it gets that doesn't have agency. You could also say that there's alignment to human values. Either more broadly to some social or political values, or just "not harm a human," or just "not kill a human," or "not kill all of humanity." So you could say that would be a "notkilleveryoneist information processing system." Or an "some moral values aligned information processing system." And if you buy the instrumental convergence thesis (which I don't really buy fully), I think, since if you look at mathematicians and scientists, the more smart a human is, then on average the less power-seeking he is, and the less "wants to kill everyone" he is. Since if you compare mathematicians and scientists to politicians and other people actually in power, you can see that pure mathematicians mostly love to scribble equations, if I oversimplify it, and don't really want to kill humans. I think the same can be true for superhuman mathematical information processing systems. And self-preservation instinct isn't a given in all information processing systems. And same for going for any goals at all autonomously. So what I think the AI risk people are afraid of, in all this terminology, is an autonomous, agentic, killeveryoneist, superhuman, adaptive, generally capable, resource-acquiring, power-seeking, silicon-based information processing system that also has access to a lot of our infrastructure and somehow can gather so many resources to kill everyone. Which I'm not sure is practically easy. But I don't think it's impossible. So I think it would be nice if these people would specify much more what kind of information processing system they're afraid of. Because if they just put all superhuman information processing systems that are more general into just this one "superhuman AI" bin when wanting to ban systems, then you can accidently ban non-agentic superhuman silicon-based mathematicians, which would be unfortunate. But yeah, otherwise makes sense. Even though I suspect it's actually relatively improbable. I'm missing more empirical evidence in order to see how it's the more probable path. And I wanna see less guessing. Because if you chain a lot of relatively improbable events (because there are a lot of other events that can happen), then the resulting probability can be very low. But it depends what is your probability of everything leading to it. Actually, when I think about it, my argument against instrumental convergence is pretty weak here. " You have to do reinforcement learning of reinforcement learning tasks to actually properly deeply understand reinforcement learning "AGI is when it can do what I can do in idealistic form" - many people " Anything older than 6 months is stone age in AI Older than 1 year is dinosaur age Older than 10 years is prebigbang /s " Scaling law of scaling laws LLMs aren't bitter lesson pilled enough. Too many humans biases. Architectures, learning algorithms, objectives, etc. are created by humans. Data in pretraining is mostly created and curated by humans. Reinforcement learning from human feedback signals are created by humans. Reinforcement learning with verifiable rewards environments are mostly created by humans. This will change. General meta reinforcement learning (and beyond) agents learning from experience from scratch will emerge. https://x.com/yacineMTB/status/2039130969077133781 AI is the optimal field for ADHDers because there's something novel daily The best kind of AI will be fully automated scientific method Schmithuber's history of AI taky záleží kde přesně dáš tu hranici a kam přesně koukáš, protože různý součástky moderních systému se nacházeli inkrementálně, a bylo k nim hodně prekurzorů "The first non-learning RNN architecture (the Ising model or Lenz-Ising model) was introduced and analyzed by physicists Ernst Ising and Wilhelm Lenz in the 1920s." "In 1972, Shun-Ichi Amari made the Lenz-Ising recurrent architecture adaptive such that it could learn to associate input patterns with output patterns by changing its connection weights." "Remarkably, already in 1948, Alan Turing wrote up ideas related to artificial evolution and learning RNNs. This, however, was first published many decades later." "In 1676, Gottfried Wilhelm Leibniz published the chain rule of differential calculus in a memoir (albeit with a sign error of all things!); Guillaume de l'Hospital described it in his 1696 textbook on Leibniz' differential calculus.[LEI07-10][L84] Today, this rule is central for credit assignment in deep neural networks (NNs)." "This answer is used by the technique of gradient descent (GD), apparently first proposed by Augustin-Louis Cauchy in 1847" "In 1805, Adrien-Marie Legendre published what's now called a 2-layer linear neural network (NN)." "In 1958, Frank Rosenblatt not only combined linear NNs and threshold functions (see the section on shallow learning since 1800), he also had more interesting, deeper multilayer perceptrons (MLPs). " "Successful learning in deep feedforward network architectures started in 1965 in Ukraine (back then the USSR) when Alexey Ivakhnenko & Valentin Lapa introduced the first general, working learning algorithms for deep multi-layer perceptrons (MLPs) or feedforward NNs (FNNs) with many hidden layers (already containing the now popular multiplicative gates). " "In 1967, however, Shun-Ichi Amari suggested to train MLPs with many layers in non-incremental end-to-end fashion from scratch by stochastic gradient descent (SGD),[GD1] a method proposed in 1951 by Robbins & Monro. " "In 1970, Seppo Linnainmaa was the first to publish what's now known as backpropagation, the famous algorithm for credit assignment in networks of differentiable nodes. In 1960, Henry J. Kelley already had a precursor of backpropagation in the field of control theory" "Backpropagation is essentially an efficient way of implementing Leibniz's chain rule[LEI07-10] (1676) (see above) for deep networks " https://people.idsia.ch/~juergen/deep-learning-history.html Schmidhuber exists before the Big Bang's singularity where he invented all that will ever be created in this universe The Cambrian explosion of the space of possible intelligences Copernican view of intelligence https://x.com/PhilosophyOfPhy/status/2043098014634643539 The platonic representation hypothesis is true in weaker form for AI systems, and more weaker form for AI x biological systems https://arxiv.org/abs/2405.07987 In my old days (assuming aging won't be reversed) I will be growing neural networks in my garden and be in endless awe, just like now. Eval awareness I wonder, to what degree is this trained and to what degree is this emergent from RL or other methods? https://x.com/MariusHobbhahn/status/2044390695507501146 I wish there were more people trying to twist and apply these image generating visual models to visual scientific reasoning You image people aren't really using mixture of experts much like the language people? TIL https://x.com/i/status/2044728293178442149 Too many people are conflating objective with the mechanism when trying to reason about how LLMs work Looped transformers start the Neuraleseian age of AI Interesting to see more and more research on looped transformers. So this is/will be the next trend? https://arxiv.org/abs/2604.07822 https://fxtwitter.com/i/status/2044229171627639004 https://fxtwitter.com/i/status/2043953033428541853 https://arxiv.org/abs/2604.09168 https://x.com/burny_tech/status/2044663935815635252 I love this analogy "Terence Tao gave the analogy of mathematicians trying to climb “a big mountain range with lots of tall mountains and lots of foothills.” Humans can only climb one step at a time, but they can plan a route to the top of a mountain like Everest. Meanwhile, Tao said, current AIs are like jumping robots. They can sometimes “parkour their way to the top of a 6-foot wall” that a human couldn’t climb. But they can’t do long-term strategic planning. Those 6 feet might become 10 feet, or 100, Tao imagines, but “the little jumping robots are nowhere near the Mount Everests of math.”" https://www.quantamagazine.org/the-ai-revolution-in-math-has-arrived-20260413/ and the jumping robots can also accidentally jump into the opposite direction, or fly into outer space The laws of physics are the only limit of AI capabilities