"
"LLMs are only matrix multiplications!"
There is much more to the math underlying LLMs than that.
There are multiple layers and levels of analysis.
You have architectures, training algorithms, objectives, data.
You have development, and analysis. You have science, and engineering. You have theory, and empirical practice.
The core that is used the most is linear algebra, multivariate calculus, probability, information theory, convex and non-convex optimization theory.
To expand more, and if I simplify a bit, most (but not all) large language models:
- Use transformer deep learning architectures, that use a lot of matrix multiplications and in order to make those matrix multiplications work, you use various mathematical machinery on top of them (linear algebra, numerical analysis), the neural networks are are directed acyclic graphs (graph theory). You also need tokenization (discrete mathematics, compression theory, coding theory). Low-rank approximations are useful (linear algebra).
- Are trained using backpropagation with stochastic gradient descent variations such as AdamW (multivariate real analysis, convex and non-convex optimization theory, stochastic approximation theory) with cross entropy loss (mathematical information theory) over the next token given previous token objective in language models (probability theory) or diffusion/flow matching objective in image/video models (vector calculus, stochastic/ordinary differential equations, optimal transport theory), and with reinforcement learning from verifiable rewards and from human feedback (stochastic processes, markov decision processes).
- You can look at the trained weights, or at the evolution of the weights in the process of training, or at the process of training itself (dynamical systems), the loss landscape (differential geometry, information geometry, fractal theory) using all sorts of math from statistical mechanics (statistics, probability theory), mean field theory. Mechanistic interpretability reverse engineers all sorts of circuits (sparse coding, harmonic analysis, topology). Spline theory is also a useful theoretical lens.
- For more analysis, (stochastic) approximation theory is good, or statistical learning theory (statistics, functional analysis) is where you begin with learning theory, but deep learning theory goes beyond that, such as with singular learning theory (algebraic geometry, bayesian statistics, statistical physics). Neural Tangent Kernel theory studies neural networks in infinite width limit (functional analysis, stochastic analysis, random matrix theory, differential equations). Geometric Deep Learning uses the lens of symmetry to classify architectures (group theory, representation theory). You can study what transformers can recognize and express (theoretical computer science: formal language theory and automata, computational complexity). Scaling-law modeling is fun. Categorical Deep Learning uses the lens of categories (category theory). Multiagent systems are a thing (game theory).
And so on.
And we still don't know a lot in the science of AI. For example, one of the central questions in the theory of deep learning is: As you train these giant neural networks in realistic practical empirical settings, there are many nontrivial emergent properties we struggle to analyze and reverse engineer, where we don't really fully know why and how does gradient descent in the process find so many of the local minima and saddle point solutions that it's finding in the billion/trillion parameters regimes, in the high-dimensional non-convex landscapes, finding so many emergent representations with features and circuits, arranged in all sorts of geometries, and why does it pick those solutions over others, and why do they generalize to the degree that they generalize, all using the nonlinear architectures. We don't fully know what all its full potential is, and what all it's absolute limitations are, and how to predict all of that. We don't fully know what all can still be improved and what all is at its limits.
"
How does the geometry of reasoning look like?
Example celebrated in the deep learning theory community is muP (Maximal Update Parametrization) from Tensor Programs, which leads to a technique called muTransfer, for taking certain hyperparameters, like the learning rate or initialization variance for example, of a smaller model, and transfering it to a scaled larger model, by stabilizing activations and doing maximal feature learning as you take an infinite limit of network width, which was used for LLMs like GPT-4 or Cerebras-GPT. https://blog.eleuther.ai/mutransfer/
In AI, just like in physics and biology, experiments inform theory and theory informs experiments, in bidirectional interaction. But compared to physics, in AI, experiments currently lead theory much more than theory leads experiments.
AI agents have gotten meaningfully better overall over the last 6 months (still far from perfect) thanks to scaling and algorithmic advances in long horizon reinforcement learning with both verifiable and nonverifiable rewards (and thanks to other things too), and this is being improved and scaled up more.
"
Flow matching is a fascinating generative AI technique when you look more into it.
It's cool that the probability distributions, at the start (input), and at the end (output), are always normalized to integrate/sum up to 1 by definition to actually be proper probability distributions, and at each intermediate step as well, since each transformation is a diffeomorphism, and thanks to continuity equation, since its rearranging the probability mass of the source distribution (e.g., a Gaussian) into the shape of the data distribution, without creating or destroying mass.
And in practice, you sample from the source distribution and data distribution to train velocity fields for paths, using neural networks with mean squared error loss and backprop, and you use numerical integration at inference to numerically solve the learned ordinary differential equation.
https://arxiv.org/abs/2210.02747
"
Do LLMs have some reverse-engineerable emergent more complex reasoning circuits when doing more complex mathematics?
Grok is less sloppy than the average X user
"
Tbh I dont like much of the focus on the next token prediction objective by the mainstream. Many people seem to think that it explains almost everything, but technically thats an objective, which is important part that tells you a lot, but less about how it predicted the next token, about which architecture is used, which learning algorithm is used, what representations were learned, etc. There are many levels of analysis to AI systems.
Like, you can also give a human an objective to predict a next token/word, and it will do that using a different architecture, learning algorithm, etc., which matters a lot.
And if you switch next token prediction for for example diffusion denoising (diffusion LLMs), you can still use same architectures, learning algorithms, get a lot of shared other structure in learned representations with some differences, etc., that matters a lot.
Most of the time someone tells me "its did this failure mode because its only predicting the next token", I feel like that person mostly has little knowledge about everything else that I just mentioned. Sometimes tThe actual biggest causal reason for that failure mode is the objective, but often its at a different part of the whole system than the objective.
"
Will AI continue being jagged intelligence and in some dimensions worse than humans, or will it steamroll all of human jaggedness and beyond?
I think LLMs are absolutely great for search of scientific articles. They're finding out that the "novel" research ideas I am getting were actually implemented by some Chinese version of Jürgen Schmidhuber from the times before the Big Bang, at an accelerated rate, that traditional search engines cannot compete with.
Is current AI conscious? If current AI isn't conscious, can future AI be conscious? How? What's missing? What's your reasoning?
"
https://x.com/lucasmeijer/status/2050890323920859601
Not saying that AI is definitely conscious, but just knowing code and math to implement transformers doesn't tell you everything about the philosophy and science of AI.
When we go into consciousness, there is no academic consensus on the best philosophy of mind positions, no consensus on defining consciousness, no consensus on the model of consciousness, and no consensus on empirically measuring consciousness. There is a lot of unknown and uncertainty.
Other than that, one of the central questions in the theory of deep learning is:
As you train these giant neural networks in realistic practical empirical settings, there are many nontrivial emergent properties we struggle to analyze and reverse engineer, where we don't really fully know why and how does gradient descent in the process find so many of the local minima and saddle point solutions that it's finding in the billion/trillion parameters regimes, in the high-dimensional non-convex landscapes, finding so many emergent representations with features and circuits, arranged in all sorts of geometries, and why does it pick those solutions over others, and why do they generalize to the degree that they generalize, all using the nonlinear architectures.
We don't fully know what all its full potential is, and what all it's absolute limitations are, and how to predict all of that. We don't fully know what all can still be improved and what all is at its limits.
And we still didn't reverse engineer a lot.
What we don't know is that if consciousness is a physical functional empirically measurable pattern defined in some tractable way, which assumes certain position in philosophy of mind that might be not true but it might be true:
Can consciousness emerge as a circuit using gradient descent on top of transformers? Still unanswered.
"
being wrong is the best way to get peer reviewed by the entire AI research twitter
the viral paper estimating parameter count of frontier llms by extrapolating scores on custom knowledge benchmark in open source llms got destroyed pretty hard
https://x.com/burny_tech/status/2052138358780940799
https://www.lesswrong.com/posts/veFMEzDDyWaer2Sms/sanity-checking-incompressible-knowledge-probes https://x.com/itsclivetime/status/2049531517899284792 https://x.com/kalomaze/status/2049556747178873188?s=46 https://x.com/justanotherlaw/status/2050399317782155726?s=20 https://x.com/MLStreetTalk/status/2049750044014956646
Will jailbreaks ever be solved?
Does reinforcement learning with verifiable rewards on LLMs go beyond the human distribution?
Will AI be able to solve the Riemann hypothesis in the future? What's missing for it to be able to do that?
Any problem that can be expressed in language can be casted as next token prediction problem (in math, computer science, etc.)
What do you think about sparse autoencoders, crosscoders, and similar methods? Do they work in a satisfying way in mechanistic interpretability or not?
Physics laws with all it's emergent laws is the only limit to AI capabilities
Are physics laws with all it's emergent laws the only limit of AI capabilities?
its interesting how Lecun is the inverse of Hinton and Bengio in many aspects
I wonder how did that happen
Recursive Multi-Agent Systems
https://arxiv.org/abs/2604.25917
https://x.com/burny_tech/status/2052505240952291683
https://x.com/askalphaxiv/status/2050238223876567129
neuralese attempts are progressing
this will be insane to do mechanistic interpretability on
Yoshua Bengio is lately less "pause ai" and more "international coordination, safety testing, alternative architectures closer to how scientists work" https://www.youtube.com/watch?v=G1ARvwQntAU
https://www.youtube.com/watch?v=PZqDFs2sbiY
"this is how you look when you have multiple agents running btw" - @yacinelearning
I would add:
and sometimes the juggling balls suddenly fly into outer space
https://x.com/yacinelearning/status/2052426977483624931
Error bars are the nightmares of ML researchers
What is your definition of intelligence/AGI? Can current LLM-like approaches achieve intelligence/AGI, or are they intelligent/AGI already? How to empirically falsify and measure that a system is an intelligent/AGI system? Use thinking tokens.
I'm experimenting with the whole spectrum from manually doing everything to vibing everything. How good it is depends on the context. For research I definitely want to manually 100% understand the mathematical equations, the architecture code, the eval code (since they love to reward hack!), etc. Vibing some graphs is ok when you know the plotted data is manually checked. You can direct a LLM to write down an equation you already have in mind so there its syntax work. And writing it as a human is often more succinct and actually what you have in mind. But sometimes letting the LLM go more open ended can help you brain storm a not so bad idea or direction that you might not have thought before, but it struggles to think more outside of the box by default for example as its more biased towards median on average, but you can do some scaffolding to bias it more towards diversity/plurality/rarity somewhat a bit, but it can get into slop territory too quick.
"
capable expert human steering of slop models gives better results
capable expert human steering of capable models gives even better results
human slop steering of capable models gives worse results
human slop steering of slop models give even worse results
in the future, human expert capable steering of "supercapable" models will give worse results as the model can potentially steer itself more effectively than human could steer it
going beyond the centaur phase
which will likely happen in some specialized domains first, similar to how it already happened in chess
if it will happen in general is debatable
what exact architecture, learning algorithm, hardware, etc. will be needed is another aspect
the unit distance erods problem solution by openai was done with relatively minimal steering, and I wonder, would it be possible also with human steering? maybe, but it would have to be a human that didnt assume what most humans seemed to assume about the problem, since contrary to most human approaches, the model went mostly to disprove the conjecture instead of refuting it, and it used tools from algebraic number theory on the discrete combinatorial geometry problem, which wasnt expected
"
some general objective function in my favorite benchmark would be the whole human and machine collective intelligence having better predictive and explanatory grip on the universe across all sciences