How to build truly strongly generalizing, truly intelligent, truly creative, for truly extrapolating off manifold, truly hyperpolating, truly novelty generating, 420 IQ, AI models? " And I would generalize this to attempts to build theoretical (ideally mathematical) models and experiments of information processing systems more broadly, not just deep learning, but also include deep learning! Something in the spirit of using tools from [information theory](<https://en.wikipedia.org/wiki/Information_theory>), [computation theory](<https://en.wikipedia.org/wiki/Theory_of_computation>), [statistical mechanics](<https://en.wikipedia.org/wiki/Statistical_mechanics>), [signal processing](<https://en.wikipedia.org/wiki/Signal_processing>), [probability theory](<https://en.wikipedia.org/wiki/Probability_theory>), [information geometry](<https://en.wikipedia.org/wiki/Information_geometry>), [dynamical systems](<https://en.wikipedia.org/wiki/Dynamical_system>), [algorithmic information theory](<https://en.wikipedia.org/wiki/Algorithmic_information_theory>), [game theory](<https://plato.stanford.edu/entries/game-theory/>), etc. more broadly for more broad classes of systems. Make theories and experiments for other forms of [AI systems](<https://en.wikipedia.org/wiki/Artificial_intelligence>) or information processing systems more broadly, and connect them, even if deep learning is the most hot right now. And in the human and non-human biological brains. Chris fields also has this [physics as information processing](<https://www.youtube.com/watch?v=RpOrRw4EhTo>) framework based on free energy principle. There is for example [free energy principle](<https://en.wikipedia.org/wiki/Free_energy_principle>) (active inference) using some of these, but the general formalism itself is so general that author Friston himself said its not falsifiable, even if concrete models in this formalism seem to be. [Wolfram's supertheory](<https://x.com/i/status/2060896771107353071>) might or might not be also be in this camp, but idk as much details. While [deep learning](<https://en.wikipedia.org/wiki/Deep_learning>) and ML has some [specific theories and empirical results](<https://arxiv.org/abs/2604.21691>) in stuff like [deep learning theory](<https://arxiv.org/abs/2106.10165>) with [neural tangent kernel](<https://en.wikipedia.org/wiki/Neural_tangent_kernel>) or [muP hyperparameter transfer](<https://arxiv.org/abs/2203.03466>), [statistical learning theory](<https://en.wikipedia.org/wiki/Statistical_learning_theory>) with [VC dimension](<https://en.wikipedia.org/wiki/Vapnik%E2%80%93Chervonenkis_dimension>), or [mechanistic interpretability](<https://en.wikipedia.org/wiki/Mechanistic_interpretability>) with feature [circuits](<https://transformer-circuits.pub/2025/attribution-graphs/biology.html>) and [geometries](<https://arxiv.org/abs/2311.03658>), etc. Or [spline theory of deep learning](<https://www.youtube.com/watch?v=l3O2J3LMxqI>) connects neural networks with splines. [Neuroscience](<https://en.wikipedia.org/wiki/Neuroscience>) also loves [manifolds](<https://arxiv.org/abs/2203.11874>) like [mechinterp](<https://en.wikipedia.org/wiki/Mechanistic_interpretability>) recently but call it [neural manifolds](<https://arxiv.org/abs/2203.11874>). [Neuroscience](<https://en.wikipedia.org/wiki/Neuroscience>) also likes some forms of [reinforcement learning](<https://arxiv.org/abs/2311.07315>). Or both fields also love [attractor dynamics](<https://en.wikipedia.org/wiki/Attractor_network>), or [criticality](<https://en.wikipedia.org/wiki/Critical_brain_hypothesis>). There are biologically many inspired approaches thanks to some high level approximate bridges existing, like [hebbian learning](<https://en.wikipedia.org/wiki/Hebbian_theory>), [predictive coding algorithms](<https://arxiv.org/abs/2202.09467>) or [forward forward algorithm](<https://arxiv.org/abs/2212.13345>). In neuroscience there is also [predictive coding](<https://en.wikipedia.org/wiki/Predictive_coding>) and [bayesian brain](<https://en.wikipedia.org/wiki/Bayesian_approaches_to_brain_function>) more broadly, etc. [Bayesians](<https://plato.stanford.edu/entries/epistemology-bayesian/>) also love to cast many [ML algorithms](<https://en.wikipedia.org/wiki/Machine_learning>) as [bayesian](<https://en.wikipedia.org/wiki/Bayesian_inference>), like the [transformer being bayesian here](<https://arxiv.org/abs/2112.10510>). On top of which you can use various [definitions of intelligence](<https://arxiv.org/abs/0712.3329>) and various ways to (often approximately) [measure](<https://arxiv.org/abs/1911.01547>) those different definitions of [intelligence](<https://en.wikipedia.org/wiki/Intelligence>) or [AGI](<https://en.wikipedia.org/wiki/Artificial_general_intelligence>) or similar words. Like [Chollet's algorithmic information theoretic definition of intelligence](<https://arxiv.org/abs/1911.01547>), [Hutter's definitions of intelligence intelligence](<https://arxiv.org/abs/0712.3329>) (and [AIXI](<https://en.wikipedia.org/wiki/AIXI>) and its [approximations](<https://arxiv.org/abs/0909.0801>)), [epiplexity](<https://arxiv.org/abs/2601.03220>), etc. Or you can just look at pure [capabilities](<https://en.wikipedia.org/wiki/Capability>) and [benchmarks](<https://en.wikipedia.org/wiki/Benchmark_%28computing%29>) without caring about this in your theory, but that will limit it to some degree. I want good [predictive](<https://en.wikipedia.org/wiki/Predictive_modelling>) and [explanatory models](<https://en.wikipedia.org/wiki/Explanation#Scientific_explanation>) with good [mathematical](<https://en.wikipedia.org/wiki/Mathematics>) [theoretical grounding](<https://en.wikipedia.org/wiki/Theory>). I want [formalisms](<https://en.wikipedia.org/wiki/Formal_system>)/[frameworks](<https://en.wikipedia.org/wiki/Conceptual_framework>)/[models](<https://en.wikipedia.org/wiki/Mathematical_model>) that support specific and general [proofs](<https://en.wikipedia.org/wiki/Mathematical_proof>). I want concrete [theories](<https://en.wikipedia.org/wiki/Theory>), general [theories](<https://en.wikipedia.org/wiki/Theory>), and unifying [bridges](<https://en.wikipedia.org/wiki/Interdisciplinarity>) between them. Gotta collect and unify them all, to the extent that its realistically possible, without losing predictive and explanatory bits, and maximizing those instead. I wanna maximize those. It feels like there is still so much unknown in this [science](<https://en.wikipedia.org/wiki/Science>) of [information processing systems](<https://en.wikipedia.org/wiki/Information_processing>). I'm also still ignorant of many [subfields](<https://en.wikipedia.org/wiki/Discipline_%28academia%29>) since I havent looked at everything there is. There are also Duggar et al's turing completeness arguments using theory of computation. and Ken Stanley's fractured entangled representation, and singular learning theory " The smarter/more knowledgeable/more intelligent you are, the better you can make use of AI The cosmic horror of AI/ML is not that things don't work. It's that most things work to some degree. Shocking! How well or bad various AI systems perform on various diverse tasks is relatively nonlinearly nuanced in different contexts and not a single linear story of absolutely bad or good in general. Nonlinear multidimensional jagged intelligence. Welcome to the Era of Experience ︀︀storage.googleapis.com/deepmind-media/Era-of-Experience%20/The%20Era%20of%20Experience%20Paper.pdf Duggar doesn't believe that RL reward is enough need like the Suttonian reinforcement learning people in big part believe Unsolved problem is how you do this flexible adaptation of rewards which iirc they don't discuss in that paper And since Sutton also has a paper titled "reward is enough", I suspect his idea would something like using some meta reward to guide the generation of rewards, all using reinforcement learning Duggar is afaik fan of adding in evolution, Kenneth Stanley style (correct me if I'm wrong) At least I think the same, I'm a lot for adding in more Kenneth Stanley ideas like guided novelty search I should revive my project where I try to do that Hmm, or could you do some guided novelty search at inference time And then finetune on the most distant ones (distant and correct in verifiable domains) https://x.com/slow_developer/status/1931497651926892673 Sutton preaching the bitter lesson! If i get it correctly, he's thinking that even LLMs are too full of human biases, and future AIs wont be, and they will scale better because of that. "well yeah, Sutton is a universalist - he basically thinks that there is a god function and all you need is lots of compute. I think there are slightly more sane variations of this idea i.e. Friston's FEP not sure about us having a "reward function" btw, it seems extremely complex" Haha this describes Sutton's view so well god (meta?) reward function scaled with humongous compute " Is there an video or article or book where a lot of real world datasets are used to train industry level LLM with all the code? Everything I can find is toy models trained with toy datasets, that I played with tons of times already. I know GPT3 or Llama papers gives some information about what datasets were used, but I wanna see insights from an expert on how he trains with the data realtime to prevent all sorts failure modes, to make the model have good diverse outputs, to make it have a lot of stable knowledge, to make it do many different tasks when prompted, to not overfit, etc. I guess "Build a Large Language Model (From Scratch)" by Sebastian Raschka is the closest to this ideal that exists, even if it's not exactly what I want. He has chapters on Pretraining on Unlabeled Data, Finetuning for Text Classification, Finetuning to Follow Instructions. https://youtu.be/Zar2TJv-sE0 In that video he has simple datasets, like just pretraining with one book. I wanna see full training pipeline with mixed diverse quality datasets that are cleaned, balanced, blended or/and maybe with ordering for curriculum learning. And I wanna methods for stabilizing training, preventing catastrophic forgetting and mode collapse, etc. in a better model. And making the model behave like assistant, make summaries that make sense, etc. At least there's this RedPajama open reproduction of the LLaMA training dataset. <https://www.together.ai/blog/redpajama-data-v2> Now I wanna see someone train a model using this dataset or a similar dataset. I suspect it should be more than just running this training pipeline for as long as you want, when it comes to bigger frontier models. I just found this GitHub repo to set it for single training run. <https://github.com/techconative/llm-finetune/blob/main/tutorials/pretrain_redpajama.md> <https://github.com/techconative/llm-finetune/blob/main/pretrain/redpajama.py> There's this video on it too but they don't show training in detail. https://www.youtube.com/live/_HFxuQUg51k?si=aOzrC85OkE68MeNa There's also SlimPajama. Then there's also The Pile dataset, which is also very diverse dataset. <https://arxiv.org/abs/2101.00027> which is used in single training run here. <https://github.com/FareedKhan-dev/train-llm-from-scratch> And more insights into creating or extending these datasets than just what's in their papers could also be nice. I wanna see the full complexity of training a full better model in all it's glory with as many implementation details as possible. It's so hard to find such resources. Do you know any resource(s) closer to this ideal? There's also OLMo 2 LLMs, that has open source everything: models, architecture, data, pretraining/posttraining/eval code etc. https://arxiv.org/abs/2501.00656 I might try to run GRPO on OLMo. I want a completely fully open from scratch replicable reasoning model this way. Ah I see they already did RLVR I see there's also Meta's OPT-175B logbook documenting problems in the process https://github.com/facebookresearch/metaseq/blob/main/projects/OPT/chronicles/OPT175B_Logbook.pdf https://arxiv.org/abs/2205.01068 Looks like there's less tricks in the pretraining domain than I thought In terms of how data is poured into the transformers I think I found the closest thing to what I wanted! Let's pretrain a 3B LLM from scratch: on 16+ H100 GPUs https://www.youtube.com/watch?v=aPzbR1s1O_8 " Reasoning: I think that under various definitions of the word out there, just base LLMs doing latent computations with some emergent abstract circuits, that we reverse engineer in mechanistic interpretability, is a form of reasoning, even if very brittle, fuzzy, stochastic, full of shortcuts, breaking down easily with growing complexity, full of errors, etc. There also seem to be definitions out there where the concept of agency and reasoning are decoupled, but yeah you can define them by connecting them chatgptnese language evolves according to the new data coming in! i wonder if AI labs will figure out something like reinforcement learning for general writing to optimize some some objective text metrics like maybe coherence, so that the imitation learning part becomes less relevant for text in general Co OLMo od AllenAI? Ti mají open source everything: models, architecture, data, pretraining/posttraining/eval code etc., i když jejich modely nemaji trilliony parametrů a jenom billiony https://arxiv.org/abs/2501.00656 Blbý je že testovat tolik nových věcí co pořád vychází by stálo bambiliony a HuggingFace často vydá open source varianty closed source věcí Ale bral bych nějakou větší entitu co by tohle systematicky dělala pro co nejvíc researche Inb4 to dopadne takhle https://xkcd.com/927/ Další firma, co se snaží vytvořit ten nejlepší model, co si vybere svoji podmnožinu approaches, jako milion dalších Alespoň je to hezky diverzní decentralizovaný ekosystém, ale bral bych ještě větší diverzitu a decentralizaci Ale bylo by fajn mít třeba jednu stranku co má na jednom místě nejvíc performing a trendy benchmarky, architektury, learning approaches atd., a ke všemu papers, a incentivizivat aby každý paper měl replikace a k těm linkovat HuggingFace je k tomu v některých aspektech blízko, hlavně když se snaží dělat open source varianty closed source big tech researche Nad dáváním všech těchto informací na jedno místo přemýšlím často přemýšlel jsem že začnu benchmarkama už teď si dělám menší vlastní wikinu, ale tam sbírám hodně resources a myšlenek z hodně oborů bez restrikcí a různě strukturovaný, takže to není úplně tohle Dost věcí by šlo myslím automatizovat, např updatovat benchmarky co mají realtime updatovani Dát na jednu stránku nejlepší existující modely s architekturama a relevantním researcherchem, benchmarky, jaký research skupiny existují, jaký R&D firmy existují,... alespoň AllenAI papers jdou 1:1 replikovat protože dávají open úplně všechno On the other hand one good thing about automating certain aspects of healthcare is that it might also mean more cheaper healthcare for everyone, as there's not enough doctors having not enough time for everyone (at least here in Europe i imagine that) So you get bigger combined total power and capabilities of human-machine collaboration in healthcare Radiology is probably the most developed then it comes to AI Recently I've seen this paper: https://pmc.ncbi.nlm.nih.gov/articles/PMC10301708/ "A survey among employees at the Department of Radiology showed significantly lower acceptable error rates for AI (6.8 %) than humans (11.3 %)." I'm trying to watch the AI x healthcare space a lot. If i didnt want to get into AI x physics intersection more, I would want to go into this intersection. it seems like more generally when it comes to technology adoption at least here in Czechia, healthcare is terribly behind in terms of technology adoption When I was at a meetup of people doing AI and other technologies in healthcare, the number 1 topic was that we should reduce some of the regulations that don't keep up with technological advances. 😄 Like some hospitals here still use CDs to transfer data between doctors....... that was an experience i think its fair to call it AI for variety of reasons, for example since deep learning emerged from connectionist paradigm in cognitive science trying to understand the brain i like to define intelligence very generally as that allows for easier creation of various taxonomies for all the various systems etc. or because we find pretty complex emergent feature/circuit learning in deep learning systems, which is one of the most fascinating things ever to me :D or im mindboggled daily why do these systems even generalize in the first place etc. i think more people should be fascinated why all these systems even work we're still so early in the science of reverse engineering them after we grow them this is currently probably my favorite paper that reverse engineers various emergent circuits: https://transformer-circuits.pub/2025/attribution-graphs/biology.html but this research is about the old paradigm models, and particularly a low performing version one :D the science can't keep up with the progress at all im still waiting for the reverse engineering of the bigger SoTA foundational models and of the new reinforcement learning paradigm models that just from just mostly reinforcement learning seem to master lots of benchmarks i think we will find more sophisticated circuits in the new reasoning models trained via reinforcement learning but there are still so many issues like unreliability But I think in the future the models will be more neurosymbolic, where you will get symbolic composing of abstract concepts like in DreamCoder neurosymbolic architecture: Growing generalizable, interpretable knowledge with wake-sleep Bayesian program learning https://arxiv.org/abs/2006.08381 Current systems have mostly fuzzy neural circuits living in gazilion dimensional vector space that easily break, so in this neurosymbolic AI paradigm you could have the fuzzy neural circuits living combined with symbolic programs that you could search for neurally or symbolically And instead of just neural circuits, you could have an ability to switch between purely neural or purely symbolic processing and reasoning, because e.g. for visual recognition, neural networks are best now, but e.g. for sorting a list, a symbolic algorithm is best. And then the ability to do different hybrid processing. But for example the new o3 model from OpenAI is starting to go in that direction. So even though fundamentally in the architecture afaik it's still neural autoregression using deep learning and doesn't learn a symbolic library of reusable strongly generalizing abstract symbolic concepts like DreamCoder does, and even though it has internal circuits, at least when inferencing in its reasoning tokens in chain of thought it is able to execute arbitrary python code, or use arbitrary tool use like web search or do multimodality with images, where they did another specific reinforcement learning training phase on all of this while learning (not just reinforcement learning for alignment and for math/code using verifiers), where before answering it is able to do tool use several hundred times and the result can make sense. reasoning word has infinite amount of definitions :D it depends to which camp of people in the AI field you talk to, for example for some deep learners just emergent feature composition with circuits at some level of complexity is ok, or then the new reinforcement learning paradigm is ok for some different subset of people because there's more emergent chain of thought chains when the model generates answer, or for some symbolists think reasoning must be inherently fundamentally strongly symbolic so they completely don't accept anything else, or some neurosymbolists fundamentally need combination of these approaches, or some only look at raw generalization power, or deduction accuracy, etc.