End to end pipeline with LLM agents grounded in web and docs generating RL environments, training LLMs using RL and grading them, in a loop. Does that work on scale? Hmmm. https://x.com/dair_ai/status/2023748787949498804?s=20 https://arxiv.org/abs/2602.10090 We need better terminology for the degree of delegation you do when using LLMs to solve a problem. Vibe coding implies a lot of delegation. Agentic engineering implies a bit less delegation with more human control, steering, supervising, double-checking, fixing, etc., so that's better. "The invention of general relativity from newtonian physics is just interpolation at some sufficiently grandiose level of abstraction." - Adam Brown https://www.youtube.com/watch?v=LjY0i2B-Avc True AI is defined as everything that doesn't work yet frontier labs primarily RL their frontier models in their own harnesses, so it doesnt really make sense to use any other one, unless you do some other heavy specialization that cannot be easily adapted in their harnesses, or want open source harness When CEOs can be automated by AI at a human level, then you will have meta human CEOs managing AI CEOs, but those will be able to be partially automated as well The horror of politization of AI is people making technical claims about it, with zero understanding, based on ideology The first organisms, the first mammals, the first humans had no one to imitate. The biggest AI breakthrough will be general AIs that can learn and improve without any imitation. "AIs are still next token predictors and will never be anything other than that!!" have you ever heard about diffusion LLMs next token prediction is like one possible objective out of many and there's also the learning algorithm, the architecture, the emergent structure, all the different mathy theory and empirical tools you can use to analyze the system's components, learning dynamics, learned representations, etc. Humans aren't fully general, so better term to use is superhuman adaptable intelligence https://x.com/i/status/2028547387120095717 Causal-JEPA https://arxiv.org/abs/2602.11389 This is very important in the limitations section, that I had in mind while reading the paper: "Performance depends on the quality of the object-centric encoder, which can limit the performance ceiling." And object-centricness is implemented by slot attention in their examples, and I'm not sure if that will work as well outside of toy tasks in even mildly realistic scenarios, I'm not sure if it will scale. " https://x.com/heynavtoor/status/2030010676157239600 https://arxiv.org/abs/2509.04664 the tweet it wrong and misleading in many ways? how does this slop tweet have 3 million views? it looks like it was engineered either for engagement, or you didnt read the paper, or the tweet was hallucinated by an LLM, or all three? the paper did not mathematically prove that it will always hallucinate? the paper's theoretical result establishes a lower bound on error rates for base models under its assumptions? in the paper literally they sketch path toward minimizing them by changing evaluation and posttraining incentives? literally quote from the paper... "Hallucinations are inevitable only for base models." "However, a non-hallucinating model could be easily created, using a question-answer database and a calculator, which answers a fixed set of questions such as “What is the chemical symbol for gold?” and well-formed mathematical calculations such as “3 + 8”, and otherwise outputs IDK." and they dont have empirical experiments to test/validate/falsify more of their theoretical model's predictions beyond reviewing existing models and benchmarks scoring rules, and few anecdotal examples? and the tweet talks about old models on subset of benchmarks, while the newer models do empirically hallucinate less on benchmarks? and this is one model of part of the hallucination mechanism out of many, from one perspective on the whole complex system? the tweet generally misinterprets a lot of the paper's actual contents and makes stuff up? but also the paper's confidence that this is almost theory of everything of hallucinations ("non-hallucinating model could be easily created by x") is IMO also likely wrong, because it's more complex issue than that, in definitions, architecture, training, data, etc., and their math isnt sufficient, and will IMO require more modifications than what they propose, some of which they kind of partly mention in limitations section " " I wish we had a mathematical theory that predicts all the potential capabilities and exact limitations of artificial neural networks much more generally than what we have so far in current theory, that is made from what we know so far empirically. And I wish we had a more general predictive mathematical theory of machine learning more generally, of artificial intelligence more generally, and of intelligence, and of information processing in general. But we don't have that. Yet. It is one of the holy grails of science to reach. One of the holy grails that many of the smartest minds on this planet try to work on, and slowly make progress on. In AI, we're in a stage similar to when steam engines were developed and worked, but the mathematical theory of thermodynamics, that grounded much more why they work, and what are their limitations, came like 100-150 years later, and opened countless others previously unthinkable possibilities thanks to discovering many general principles of phenomena in thermodynamics, like entropy. Which, among many other big scientific theories, enhances our technology to this day. We have just bits of this holy grail so far. So far we have: - a bunch of empirical results like finding what the models can do already now in practice, like surprisingly relatively coherent generation of code or mathematics in its extremely high dimensional space of possibilities - a bunch of empirical analyses like finding some of the emergent circuits that are present in the models as they solve various tasks, like the geometric rotating of manifolds when doing linebreaking, that have various kinds of structure, that is sometimes less optimal, and sometimes more optimal, or directions that gradient descent tends to favor flatter minima - a bunch of partly unifying theory that is predictive in different ways, like various ways to generalize and predict the found emergent circuits, or applying statistical physics on training dynamics, or proving some properties in the network's infinite width limit in neural tangent kernel theory - a bunch of somehow, somewhat, in some ways working, systems to study, like all the models out there, built using a bunch of duck taped empirical recipes like scaling laws, common tricks around training algorithms or common architectures choices, like AdamW with transformers - and other things We don't really fully know why and how does gradient descent find so many of the solutions that it's finding in the highdimensional nonconvex landscapes, growing so many emergent representations in the nonlinear neural network architectures like transformer in the process, and what is it's full potential and what is it's absolute limitations, and how to predict it much more. We don't fully know what all can still be improved and what all is at its limits. There are so many unsolved open questions in this whole scientific field. There is so much potential for empirical experiments and theoretical unifying and novel predictions. This is still a largely open scientific problem for many curious minds to solve, that are trying to solve it, and making slow progress together collectively, more and more with the help of AIs themselves. " " Co se neví není definice basic gradient descentu! It is not this easy! Ta definice samotná ti neřekne spoustu dalších properties celýho toho optimizačního procesu u highdimensional nonconvex problems s různýma architekturama v praxi! Tam se neví věci! Neví se pořádně např to co jsem napsal v tý stejný větě co quotuješ. Neexistuje pořádná matematická teorie co ti reliably obecně přepoví proč a jak ty řešení, co gradient descent a *varianty a extensions* gradient descentu používaný v praxi nachází, jsou tak relativně good, a dávájí takový smysl at all, a tak relativně reliably, a proč favoruje ty řešení co favoruje a ne jiný, v tom celým highdimensional nonconvex landscapu u realistických bambilion parameterových neurálních architektur a v plný obecnosti, kde se pořádně obecně plně neví u praktických realistických modelů jaký má bias a: - proč se nezasekává v bad minimas, saddle points, atd., a jak, - proč ty najitý řešení generalizují tak jak empiricky relativně generalizují a jak - proč důsledek jsou ty emergentní reprezentace a obvody co nacházíme co reverse engineerujeme a jak - jak pořádně vypadá geometrie loss landscapu - jak se ta celá dynamika mění s různýma architekturama jako transformer co jsou různě nastavený s různými layers, nelinearitama, initial podmínkama, hyperparamterama, noisem, datama, atd. - jak se to mění u různých variant a extensions gradient descentu jako adaptive moment estimation s decoupled weight decay regularizací s různýma hyperparametrama, proč mají různý jiný properties - atd. atd. Intuice classical learning theory a optimization theory dostatečně nevysvětluje proč to funguje tak relativně dobře a člověk by nejdřív čekal že by to nemělo fungovat. Je hodně pokusů o matematický a empirický vyřešení těhle problémů, který jsou ale většinou nekompletní, partial, fragmented, a často ve special regimes (u convex landscapes, nebo neural tangent kernel theory, nebo u toy settings, stochastic gradient descent favoring flat minima, spline theory, atd.), a tam kde je to v praxi nejdůležitější je to často v plenkách a zatím dostatečný. If you will find a predictive mathematical theory that solves these problems, then you will solve tons of open scientific problems. " I see relationship between memorization and overfitting as memorization being one of the primary mechanisms through which overfitting occurs. But not all memorization leads to overfitting (it's useful to memorize some rate facts), and not all overfitting happens though memorization (instead it can be through spurious correlations). And for example regularization helps to prevent this. " Are the current new results in math using LLMs interpolation or extrapolation? Like solutions to Erdos problems, Donald Knuth's Claude's Cycles, AlphaEvolve's new algorithm for multiplying 4x4 complex-valued matrices, etc. What is your favourite definition of interpolation and extrapolation? The definition where interpolation is within the convex hull of training points and extrapolation is outside it is not sufficient, as in high dimensions, where big deep learning models operate, basically everything is extrapolation by this definition. So, let's define it other way: If the training data is supported on some low-d manifold, then let's define interpolation in the low-d latent space as basically the same thing as filling in the rest of that manifold, and extrapolation as if it works off manifold. Would you say that the current new recent results in math using LLMs are on the training data's low-d manifold and the LLMs filled some part of the rest of that low-d manifold with those results? I think... Maybe? It feels like it's on a boundary of the definition? I'm uncertain. Maybe it's more of a fuzzy continuum, as the training data manifold can have more densely and more sparsely sampled regions? And its more extrapolative in more sparse regions, and more interpolative in more dense regions? So the more common structures and ways of structuring that structure (the more common proof techniques, common algebraic manipulations, and common ways of composing them, the more common mathematical structures, etc.) are present, and the more similar the generated overall structure and its substructures are to the training data's structures, the more interpolative it is? And the more OOD it is, the more search on avarage you usally need to do? But even Grothendieck et al. creating modern algebraic geometry didnt create it in vacuum with no connection to structures and ways of structuring those structures known before him, and same for Von Neumann, Einstein, Shrodinger, etc., so it never feels truly completely off manifold. There always seems to be at least some connection, some bridge, some shared language of patterns, otherwise it would probably be incomprehensible to us? Sometimes it's more bridges, sometimes it's fewer bridges. It's never completely off the manifold as the output always has patterns that are somehow shared with the training data? The Erdos problems seem more interpolative, while Donald Knuth's Claude's Cycles and AlphaEvolve's matrix multiplication algorithm seem more extrapolative relative to the Erdos problems? But no equivalent of Grothendieck et al. creating modern algebraic geometry with much more off manifold extrapolative jump into very sparse region happened yet. And how do you see Move 37 in AlphaGo? That seems more off manifold relative human play manifold, but maybe more on manifold relative to the self-play manifold? Or is everything interpolation on a sufficiently grandiose level of abstraction? Or do you need to build new bridges via extrapolation/hyperpolation to create new dimensions and build in them, like Grothendieck's modern algebraic geometry? What was happening in his mind when he invented it? It would be nice if we could measure this empirically. https://www.lesswrong.com/posts/DY8PKdPsCoWonA7nQ/hyperpolation https://arxiv.org/abs/2409.05513 https://www.youtube.com/watch?v=86ib0sfdFtw https://mathstodon.xyz/@tao/115639985263560286 Unsolved mathematical problems follow a "long tail distribution", and current AI automation "harvesting" is concentrated at lowest-hanging fruit the end of the long tail. " A lot of things in technology and engineering evolve quickly and much of the knowledge about it becomes outdated quickly as a result. But the knowledge that will forever remain relevant is the fundamental mathematics that governs that technology with the mathematics of the physical reality across all scales in which it all lives. One day there will be a meta model that does prompt to optimal model weights "Slop is lack of constraints." - Tim Scarfe Terrence Tao is one of the best geniuses of our age I love his nuanced perspectives on everything, including AI, and how it relates to math and science So much intelligence juice in him https://www.youtube.com/watch?v=Q8Fkpi18QXU It's quite surreal to look at my 2022 notes about AGI when almost no one in the mainstream knew about it and now it's very much present in the mainstream " i feel like most people saying LLMs are stochastic parrots are basically stochastically parroting it lol, which is pretty ironic almost anytime i try to understand their model, i notice there's basically no technical knowledge from current research trying to understand LLMs beyond stuff like the architecture using probabilities (and they often dont even know where exactly) which says basically nothing for what they're often wanting to believe, assuming they even understand the technical problems, which they often dont tbh and there are still tons of unsolved problems in reverse engineering LLMs and in proper mathematical theory, we're far from fully understanding them but we have some not really unified parts in empirical exploration and theory with a lot of parts still missing "