Knowing is a combination of the degree of accurately predicting empirical data and structural isomorphisms between the (compressed) inner representation's structure and the ground truth that one tries to model
Would you connect additional neurons to your exocortex to expand your intelligence? Or digital neurons? Or any silicon, or any substrate-based neurons, as long as it increases intelligence?
All intelligence is collective
I wonder how many math nerds think that they figured out the mathematics of cognition generally but they figured out the mathematics of their own mathematical cognition instead that doesn't generalize to other humans. Maybe the nonnerd nonmathy cognition is much less mathematical than how many math nerds (including me) like to imagine.
LLMs are part of our collective intelligence
Exploratory unpredictable divergent openended wandering in the space of information, embracing the unknown, with little or no objectives, often gives a big sense of freedom!
A lot of AI researchers are more general intelligences than your average researcher from other fields
Intelligence explosion to understand the mathematics of intelligence, the most interesting force in the universe
There are no hard problems, only problems that are hard to a certain level of intelligence.
The laws of physics are the only limit to intelligence
Master the distribution to escape It
Intelligence is the ability to adapt under limited resources
I'm oscillating between breath first search and depth first search, and oscillating between exploring and exploiting, but I need to exploit much more, since I explore too much
Physics of intelligence is the most fascinating phenomenon ever.
Representations emergently gradually forming, abstracting, phase shifting, being used in reasoning circuits, binding, forming complex network structures with nontrivial geometric structures with fascinating topologies, flowing like a liquid across many scales,...
Human intelligence is extremely limited, specialized intelligence!
The space of all possible intelligences is unfathomably gigantic!
We already have a lot of diverse intelligences existing, and the number of them is increasing, and will continue to increase!
We're still at the beginning of the Cambrian explosion of an extremely diverse number of possible intelligences existing across the whole universe!
https://arxiv.org/abs/1410.0369
https://x.com/burny_tech/status/1966331043314540803
the future of intelligence is hidden in fluid dynamics
https://arxiv.org/abs/2304.02637
https://x.com/burny_tech/status/1966336302485496058
If anyone builds a superintelligence of the type that Eliezer Yudkowsky is imagining, then everyone dies. There are a lot of discussions about how much it is possible, but to me, laws of physics are the only limitation, and it's hard to imagine the future.
If anyone builds a superintelligence of a different type, then everyone doesn't die, or dies, depending on the type of superintelligence.
The space of possible types of superintelligences is vast. It's not a one-dimensional space.
my favorite definition of intelligence is adaptation with limited resources,
and llms are super data hungry and energy hungry and can more easily break out of distribution than humans,
so still a long way to go
https://www.youtube.com/watch?v=K18Gmp2oXIM 03:28
Hmm let's add this to loss function "How natural organisms are so evolvable (capable of quickly adapting to new environments)? A key driver of evolvability is the widespread modularity of biological networks, their organization as functional, sparsely connected subunits. Ubiquitous, direct selection pressure to reduce the cost of connections between network nodes causes the emergence of modular networks. Computational evolution experiments with selection pressures to maximize network performance and minimize connection costs yield networks that are significantly more modular and more evolvable than control experiments that only select for performance."
https://pmc.ncbi.nlm.nih.gov/articles/PMC3574393/
Physics of feature and circuit forming and it's dynamics in intelligent systems generally will be a gigantic scientific field
A great visualization of how people are different from AI in terms of their flaws and capabilities
The capabilities of AIs are growing, but they are different capabilities in a certain way, even though there is some overlap in those capabilities
And different AIs are different in their jaggedness, and different people are different in their jaggedness, but we as humans are still closer to each other than we are to AIs in the space of all possible intelligences
https://x.com/burny_tech/status/2002865099820667186
Current ML/AI systems aren't really trying to be replicas of human thinking like many people seem to think. They're mix of all sorts of insights that we got from neuroscience, optimization theory, probability theory and statistics, computer science, physics, applied mathematics, pure mathematics, philosophy, cognitive science, psychology, empirical random playing around with random stuff, etc. into one complex system.
Neuroscience - Neuronky začaly jako "hmm neurony jdou popsat Hodgkin Huxley modelem, ale vlastně taky jdou hodně simplifikovat jako graf s vertices jako neurony a edges jako synapses, what if we added more neurons and...". Nebo reverse engineering neuronek, obor co se nazývá mechanistic interpretability, je v podstatě digitální neurověda.
Optimization theory - Gradient descent přichází odtuď a je to nejpopulárnější učící algoritmus neuronek, a dokonce je původně z astronomie.
Applied mathematics: Neuronky jsou aplikovaná lineární algebra, víceproměnná reálná analýza, teorie pravděpodobnosti,... A přes tools těhle metod jde dělat tuna triků pro efektivnější tréning.
Probability theory and statistics - Neuronky jsou technicky statistická metoda, existuje statistical learning theory, nebo existují bayesovy metody, atd.
Physics - Hopfield networks jsou prekurzory k neuronkám co známe teď, a jsou spin glass system, za co Hopfield dostal nobelovku z fyziky. Diffusion modely jsou vlastně díky nonequilibrium thermodynamice. AdamW optimizer do gradient descentu přidává "momentum". Flow matching učící algoritmus se učí velocity fields. Na chápání neuronek se teď víc a víc hází statistická fyzika.
Pure mathematics - Jsou pokusy o pure math modely deep learningu, včetně toho categorical deep learning. Nebo je geometric deep learning co používá na analýzu deep learning architektur teorii grup.
Computer science - Např turing completeness je též relevantní, existují architektury jako neural turning machine. Neuronky jsou algoritmy a computational complexity je taky důležitá pro efektivitu. Optimizing concrete software and hardware matters. Etc.
Cognitive science/Psychology/Philosophy - Etika se řeší v alignment problemu a control problemu. Kognitivní vědy řeší problém toho jaký fyzikální systémy mají "pravou" inteligenci, reasoning, nebo co je to vědomí, a jak tyhle slova zadefinovat (uh to je rabbithole).
Empirical random playing around with random stuff - Tak některý improvements vznikly. Dost tvoření modelů je v podstatě alchemie
A je toho milion víc. :D
What if intelligence is the friends we made along the way?
Hmm, that's pointing at the fact that intelligence is collective
The first organisms, the first mammals, the first humans had no one to imitate.
The biggest AI breakthrough will be general AIs that can learn and improve without any imitation.
Physics of intelligence: General physics theory of learning of features and circuits and their formation across all architectures, learning algorithms, data, substrates, systems, etc.
Humans aren't fully general, so better term to use is superhuman adaptable intelligence https://x.com/i/status/2028547387120095717
"
I wish we had a mathematical theory that predicts all the potential capabilities and exact limitations of artificial neural networks much more generally than what we have so far in current theory, that is made from what we know so far empirically.
And I wish we had a more general predictive mathematical theory of machine learning more generally, of artificial intelligence more generally, and of intelligence, and of information processing in general.
But we don't have that. Yet. It is one of the holy grails of science to reach. One of the holy grails that many of the smartest minds on this planet try to work on, and slowly make progress on.
In AI, we're in a stage similar to when steam engines were developed and worked, but the mathematical theory of thermodynamics, that grounded much more why they work, and what are their limitations, came like 100-150 years later, and opened countless others previously unthinkable possibilities thanks to discovering many general principles of phenomena in thermodynamics, like entropy. Which, among many other big scientific theories, enhances our technology to this day.
We have just bits of this holy grail so far. So far we have:
- a bunch of empirical results like finding what the models can do already now in practice, like surprisingly relatively coherent generation of code or mathematics in its extremely high dimensional space of possibilities
- a bunch of empirical analyses like finding some of the emergent circuits that are present in the models as they solve various tasks, like the geometric rotating of manifolds when doing linebreaking, that have various kinds of structure, that is sometimes less optimal, and sometimes more optimal, or directions that gradient descent tends to favor flatter minima
- a bunch of partly unifying theory that is predictive in different ways, like various ways to generalize and predict the found emergent circuits, or applying statistical physics on training dynamics, or proving some properties in the network's infinite width limit in neural tangent kernel theory
- a bunch of somehow, somewhat, in some ways working, systems to study, like all the models out there, built using a bunch of duck taped empirical recipes like scaling laws, common tricks around training algorithms or common architectures choices, like AdamW with transformers
- and other things
We don't really fully know why and how does gradient descent find so many of the solutions that it's finding in the highdimensional nonconvex landscapes, growing so many emergent representations in the nonlinear neural network architectures like transformer in the process, and what is it's full potential and what is it's absolute limitations, and how to predict it much more. We don't fully know what all can still be improved and what all is at its limits.
There are so many unsolved open questions in this whole scientific field. There is so much potential for empirical experiments and theoretical unifying and novel predictions.
This is still a largely open scientific problem for many curious minds to solve, that are trying to solve it, and making slow progress together collectively, more and more with the help of AIs themselves.
"
I don't like the term artificial intelligence and superintelligence I think it can mean way too many things so instead I would say because they are so many definitions of intelligence I would just drop the word intelligence and I would say that in general we are talking about information processing systems so that somehow exist in physics so you can have for example carbon-based information processing systems which are biological systems and silicon based information processing systems which are most of the current AI systems and on top of that I would say you can have superhuman silicon based information processing systems and that are general or very narrow so they are equivalent or super human in some narrow task or more generally for more than one task and the more tasks it can do the more general it is but different architectures can be general in different ways so like xlstms for instance still work pretty well in time series analysis while Transformers are steamrolling a lot of other domains while xlstms are still general because neural networks are generally pretty general since they can be applied to lot of tasks as they're universal function approximators and a lot of narrow systems are better in various ways than more general ones and have different advantages and disadvantages and sometimes hybrids are the best approach and it's constant fight between bitter lesson and human crafted systems. And because there are different definitions of intelligence I would just take them apart and prepend what's in their definition, so for instance there is a definition of intelligence by Francois Chollet which says that it's the ability to generalize so that it's already included in the general information processing part and you can also distinguish between trained in general set of tasks vs ability to generalize to unseen tasks on the fly and then there is the definition of intelligence by Pei Wang which is about adaptation under the limited resources so you could prelend adaptive information processing system or you could say resources efficient information processing system and that also depends like which resources it's connected to then like there's the definition of intelligence by Joscha Bach that's as it's like ability to make models and you could maybe say it's like modeling information processing system or ability to make predictive models so it's predictive modeling information processing system and then some people also put in agency and autonomita and I don't see agency come often in the definitions of intelligence but sure you could say agentic information processing system and autonomous information processing system and I think that's what you actually require in order for all those doomy AI killing everyone scenarios since if you are if you are superhuman information processing system that doesn't have agency and autonomity then it will not cause harm because it's just doesn't have access to the actions you know if it doesn't have general agency, you could say one type of agency is using the Python or lean compiler as a tool and another type of agency that's more general is like connecting it to the whole internet so technically the AI system could hire humans there or buy robotics to make itself also physical if it's just a software system, you could also say that there's alignment to human values, either more broadly to some social or political values or just not harm a human or just not kill human or all of humanity so you could say that would be notkilleveryoneist information processing system or aligned information processing system. And if you buy the instrumental convergence thesis then which I don't really buy fully I think since if you look at mathematicians and scientists, the more smart a human is then on average the less power seeking he is and less wants to kill everyone he is since if you compare mathematicians and scientists to politicians and other people actually in power like pure mathematicians mostly love to scribble equations if I oversimplify it and don't really want to kill humans I think the same can be for superhuman mathematical information processing systems, and self-preservation instinct isn't given in all information processing systems and same for going for any goals at all autonomously, so what I think the AI risk people are afraid of in all this terminology is autonomous agentic killeveryoneist superhuman adaptive generally capable resource acquiring power seeking silicon based information processing system that also has access to a lot of our infrastructure and somehow can gather so many resources to kill everyone which I'm not sure is like practically easy but I don't think it's impossible so I think it would be nice if they would specify much more what kind of information processing system they're afraid of because if they just put all superhuman information processing systems that are more general into just this one superhuman AI bin then you will basically ban nonagentic superhuman silicon-based mathematicians and I think that's dumb but yeah otherwise makes sense even though I suspect it's actually relatively improbable like I'm missing more empirical evidence in order to see how it's the more probable path and I wanna see less guessing and because if you chain a lot of relatively improbable events because there are lot of other events that can happen then the resulting probability can be very low but it depends what is your probability of the everything leading to it.
If you ask 10 different intelligence researchers to define intelligence, you will get 10 different answers, but there are some similarities between the answers
I find it interesting how elephants have 257 billion neurons, which is about three times the number of neurons as a human brain. But their cerebral cortex has only about one-third of the number of neurons as a human's cerebral cortex.
Brains have extremely limited architecture compared to what's theoretically possible in the space of all possible intelligences
The Cambrian explosion of the space of possible intelligences
Copernican view of intelligence
https://x.com/PhilosophyOfPhy/status/2043098014634643539
The platonic representation hypothesis is true in weaker form for AI systems, and more weaker form for AI x biological systems
https://arxiv.org/abs/2405.07987
The universe and we have explored only 0.001% of the space of all possible intelligences
How does the geometry of reasoning look like?
Will AI continue being jagged intelligence and in some dimensions worse than humans, or will it steamroll all of human jaggedness and beyond?
Is human cognition emergent from fundamental physics?
What is your definition of intelligence/AGI? Can current LLM-like approaches achieve intelligence/AGI, or are they intelligent/AGI already? How to empirically falsify and measure that a system is an intelligent/AGI system? Use thinking tokens.
How to build truly strongly generalizing, truly intelligent, truly creative, for truly extrapolating off manifold, truly hyperpolating, truly novelty generating, 420 IQ, AI models?
when i see how ai data centers eat gigawatts of power and i eat just bunch of carbs to power my few tens of watts brain to do similar tasks, im thinking how far could we go if we would slurp terrawatts equivalent worth of carbs instead on this more efficient flesh hardware, but upgrade is needed to reach this fuller potential of the flesh
i often wonder to what extend all the implementation details of the brain down to the complexity of the cell, proteins, neurochemistry, etc., level are actually relevant for intelligence
Reasoning: I think that under various definitions of the word out there, just base LLMs doing latent computations with some emergent abstract circuits, that we reverse engineer in mechanistic interpretability, is a form of reasoning, even if very brittle, fuzzy, stochastic, full of shortcuts, breaking down easily with growing complexity, full of errors, etc.
There also seem to be definitions out there where the concept of agency and reasoning are decoupled, but yeah you can define them by connecting them
I think the issue is that many people have very intuitive definitions that are hard to formalize
Definitions i like are those that are as concrete, as localizable, as implementable, as mathematical as possible
Just stupid ReAct agent loop is a good definition IMO, because its concrete, implementable, and we use it practice in research and engineering all the time, but there are many others
i often tend to think that human cognition is a pretty small subset of possible intelligent systems personally
i think we're always limited observers with limited modelling perspecitves, limited computational power and with limited accuracy
and im ok with possibility of hypercomputational stuff too
reasoning word has infinite amount of definitions :D