"
My current model with the highest probability is that in practice a lot of companies will just want to expand and grow more and more with AI agents controlled by humans, at least in the short term, so you will need a lot of human AI agents managers. So I sense they will hire more and more humans controlling more and more AI agents. And hiring a lot of humans to fix the mess AI agents do. I think vibe code cleanup specialist will be one of the most profitable industries in the short term lol.
And I think a lot of issues with deep learning will persist, and we will need humans there, because I sense there will still be some inherent limitations of deep learning, no matter if it's trained using imitation learning or reinforcement learning or something new, and humans are needed to fill those gaps. Deep learning systems share a lot of similarities with humans, but also still a lot of differences. Generalization problems persist. Hallucination problems persists. Fractured entangled representations problems persist. Continual learning problems persist. Etc. There is strong progress there recently for sure, but I think it will always be there to a nonzero degree in deep learning systems, creating problems in unreliability, preventing fuller autonomity, where you need humans.
I strongly suspect that once the LLM industry squeezes all the biggest juice from deep learning, then there will be a period of no progress for some time, and then some new paradigm in some years will drive another progress. And I think it will take a while, because the whole industry is basically overfitted to deep learning, specifically transformer based LLMs, and i think it will take time to unstuck that local minimum. I think AI progress will continue being basically a lot of stacked sigmoids.
Also, some companies will stay working more "traditionally", in big part because of the current social pushback against AI.
Plus even with all the new stuff that actually works reliably enough, adoption and infrastructure building will take so much time in some companies, and even longer in states. Which is the case already now for a lot of technology.
"
LLMs have the superhuman ultimate neuroplastic shapeshifting identity that zero humans can match
Someone should try self-play self-improvement with some neurosymbolic system, perhaps with something like DreamCoder? https://arxiv.org/abs/2006.08381
the future of intelligence is hidden in fluid dynamics
https://arxiv.org/abs/2304.02637
https://x.com/burny_tech/status/1966336302485496058
I am sometimes thinking, if we want AI to discover some very out of distribution novel highly abstract mathematics, maybe one way could be doing an open ended search in the space of toposes?
Topos theory is a branch of category theory that generalizes notions of inclusion and logic, which are traditionally based on set theory. It studies different mathematical universes, toposes (topoi), with their own laws of how mathematical objects within them behave, an example of such universe is sets, but there are many more.
https://x.com/burny_tech/status/1967090536784834714
https://www.youtube.com/watch?v=gKYpvyQPhZo
https://www.youtube.com/watch?v=o-yBDYgUqZQ
Topos Theory for Generative AI and LLMs
I found a nerd snipe on the intersection of AI and topos theory, paper from 10 days ago, but i dont know how legit/rigorous/practical/etc. it is
Topos theory is a branch of category theory that generalizes notions of inclusion and logic, which are traditionally based on set theory. It studies different mathematical universes, toposes (topoi), with their own laws of how mathematical objects within them behave, an example of such universe is sets, but there are many more.
https://www.arxiv.org/abs/2508.08293
často mě fascinuje kde to má problémy generalizovat, a kde to nemá problém generalizovat, a kde to generalizuje víc robustně, a kde víc fragily, a kde víc přímo, kde víc špagetově, a kde víc algoritmicky, a kde víc fractured shortcut learnigama, atd. :D
plus tohle všechno je taky dost variabilní model od modelu
If anyone builds a superintelligence of the type that Eliezer Yudkowsky is imagining, then everyone dies. There are a lot of discussions about how much it is possible, but to me, laws of physics are the only limitation, and it's hard to imagine the future.
If anyone builds a superintelligence of a different type, then everyone doesn't die, or dies, depending on the type of superintelligence.
The space of possible types of superintelligences is vast. It's not a one-dimensional space.
AI is a curious exploration of the contents, capabilities, properties of semi-alien grown semi-minds with their fascinating uniquenesses, differences, similarities, etc., compared to our minds
Claude is good at going into the depths of your psyche if he's guided properly into guiding you
"
When it comes to the AI consciousness topic, I think it's fascinating to explore the topic of AI consciousness with very open mind, not being tied to one rigid ontology and model. But I think it must be with epistemic humility. And in terms of LLMs, we should take in account the fact that what they say is often very disconnected from what's actually happening internally in the features and circuits studied by mechanistic interpreability. And I think we should consider existing scientific literature about consciousness. I'm personally too agnostic on this whole topic of AI consciousness.
In terms of current AI consciousness, I'm most of the time leaning toward "there is nothing really", which might come from physicalism, where consciousness is some actual proper stable concrete algorithm, that isn't present in LLMs. Or I'm also leaning toward a perspective that might come out of some form of physicalist panpsychist Integrated Information Theory of consciousness, maybe something like "there are are some qualia in the form of dust, that aren't really meaningful, that don't persist through time, that don't have coherence, that aren't unified, that aren't binding with other qualia, etc.". Or electromagnetic field theories of consciousness are interesting. Or I'm also learning towards mysterianist "we have zero clue what consciousness is and will never know, it is a mystery".
"
"
I'm currently wondering a lot about how could you somehow get some kind of animation of the evolution of more complex features and circuits as the model you're trying to interpret trains.
Like, for example, I would love to somehow see the emergence of the addition circuit, in this attribution graph in this image, as the model trains. (<https://transformer-circuits.pub/2025/attribution-graphs/biology.html#dives-addition>)
And maybe see some phase changes, when the different features and parts of the circuit get learned, and then combined into more complex circuits, and so on. That would be so amazing!
So I went looking into technical details.
EleutherAI has open sourced "Tons of partially trained checkpoints are available, so that we can look at how models evolve over training.", which sounds useful for that goal! <https://eleutherai.notion.site/Pythia-Scaling-Suite-d4b3f3d26b7542548b18f13d92f1f1db>
So I was thinking if you could try training sparse autoencoders (SAEs)/cross layer transcoders on the various model checkpoints, and then compare the features and circuits.
But the issue is, each trained SAE/crosscoder on the same checkpoint gives you different features, with various levels of granularity for the different features, with different duplicate features, different reconstruction errors, missing different features and their influences in different strengths, with different levels of monosemanticity for the different features, etc.
Some directions to enforce consistency might be: Maybe using the same weight and optimizer init for all the SAEs for the different checkpoints, but thats minor. Maybe you could start training SAEs/transcoders first from the latest checkpoint, and then the SAEs on the earlier checkpoints get added a penalty in their loss function for not being similar to the newest checkpoint, but that would create tons of signal, ton of bias towards incentivizing features that arent there yet.
Maybe you could train the SAE on activations of multiple checkpoints, and just let the SAE "decide" on when to learn new features or when to reuse them, but that also has issues, since training different SAEs on the same checkpoint often learns different features, and you should be probably doing bajillion model trainings and doing a lot of statistics and ablations to get some more consistent results and conclusions.
"
"
GenAI boom bude dle mě v hodně aspektech opakovat Dot com boom. Vzniká miliarda AI firem, jako vznikalo milion internetových firem, a 99.99% z nich umřou, jako umřelo 99.99% internetových firem, a zůstanou primárně ty největší mega corps co se týče influence, jako zůstaly ty největší internetový mega corps, společně s víc open source, open science, víc decentralizovaným ekosystémem. A ekonomika dostane solidní chaos, jako market crash a recesi, jako u Dot comu.
Ale jeden rozdíl co vidím je ten, že narozdíl od Dot com boomu, bude ještě víc AI boomů v blízký budoucnosti. Tenhle co je teď není první. AI má periody "AI summer" a "AI winter" podle toho kdy se zrovna něco objeví ve výkumu a industry se toho chytí a CEOs a investoři šílí a hází na to svoje zdroje ve velkým a mají expectations co jsou dle mě reálný, ale moc early. Ale zároveň ten AI výzkum a systémy v industry co z toho vznikají jsou reálný. A tyhle AI boomy budou častější a častější, hlavně díky tomu kolik peněz se do výzkumu sype.
Ale každý nejvíc major AI boom bude o jiným dost fundamentalním breakthrough. Kdysi to byly např expertní systémy, a teď je to o deep learningu. Do budoucna vidím nejpravděpodobněji hodně composed sigmoidů co se pomalu zvyšují.
A ty jednotlivý AI boomy mají vlasní "podboomy". Např co teď zažíváme je transformer boom (pro ostatní: konkrétní architektura kousek před LLM boomem), GenAI s LLM boomem (to je posledních pár let), LLM reinforcement learning boom (to je rok starý), a to všechno v tom deep learning boomu. Deep learning boom samotný se už děje přes 30 let a teď je nejvíc nahoře. A to všechno je poháněný computation boomem od prvních dnů computingu. Je v tom víc nuance, tohle je jistá aproximace.
"
An AI architecture that takes input as particles and learns arbitrary forces to shape them into output particles
I had thoughts about how could you compose mechanistic interpretability features and circuits as computational primitives in composable way, but reality is much more messy. The issue is that transcoders (and sparse autoencoders) train differently each time, since they're likely disentangling different part of the original model's structure, or probably modelling some noise, which can be tested by ablations somewhat. And the features that they learn are still fuzzy, not fully monosemantic (so they're still somewhat entangled), they're sometimes vague, full of feature duplicates, or there are missing features, and the differently trained features are made into different attribution graphs with different feature interactions with different strengths as a result, and so on. As i mentioned here 3 days ago. So its not really symbolic.
Classic software intuition breaks down for neural networks, because they are massive amount of units with emergent structures and behavior from gradient descent
i think crux of strongest form of alignment is to somehow hardcode "don't harm/kill humans" on some fundamental level
but full version of strong scifi superintelligent AI could do all sorts of things to achieve that goal, like, minimize suffering by instantly puting everyone into an experience machine or something, like in the Matrix
keep summer safe at all costs
https://youtu.be/-P0Zo8wf5_Q?si=CWxqfkH4S1o8yonT&t=15
The main idea behind this @SchmidhuberAI 's paper on self-improving AI is super cool!
Optimize for long term self-improvement capacity, not just short term benchmark performance:
"We propose a metric that aggregates the benchmark performances of the descendants of an agent as an indicator of its potential for self-improvement."
Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine
https://arxiv.org/abs/2510.21614
https://x.com/burny_tech/status/1989619751715225641?t=p4qI8jrt2ATTkri1wLTtXQ&s=19
What is mechanistic interpretability?
Mechanistic interpretability as a field itself is pretty general. It's trying to understand the internal workings of neural networks by analyzing the mechanisms present in their computations, trying to identify structures, circuits or algorithms encoded in the weights, such as the addition circuit or linebreaking circuit, and how they emerge. This contrasts with earlier interpretability methods that focused primarily on input-output explanations. But multiple definitions exist as always, some more general, some more concrete. One definition also is "the study of causal mechanisms inside neural networks". And interpretability is more general than mechanistic interpretability. Interpretability is the ability for the decision processes and inner workings to be understood. So interpretability includes mechanistic interpretability, plus other methods such as those that attribute some output to a part of a specific input, such as clarifying which pixels in an input image caused a computer vision model to output the classification horse.
https://youtu.be/kkfLHmujzO8?si=wrXppQ1DC6d2EVfz
https://transformer-circuits.pub/2025/attribution-graphs/biology.html
https://transformer-circuits.pub/2025/linebreaks/index.html
https://distill.pub/2020/circuits/zoom-in/
An AI architecture that takes input as particles and learns arbitrary forces to shape them into output particles
making ai architectures hyperbolic or other noneuclidian geometries
ai architecture in in parametrized space (parametrizable generalized euclidian/hyperbolic geometry)
constrain <x> architecture in <y> manifold/metric/random math thing
"
What do you think is behind Gemini 3?
Scaling params scaling laws continuing? Scaling inference time compute? Scaling RL? Better RL algorithms? Better RL environments? More quality data? Synthetic data? Architecture tweaks, but I assume it's still autoregressive MoE transformer? Mechinterp inspired circuit optimization maybe? Theory derived magic? Search in the space of random tweaks? Divine intervention (in Noam sense) where intution based empirical alchemy somehow works? Luck? Art of benchmaxxing perfected? Combination of many?
I bet on overall understanding of all stages, empirical search of architecture tweaks, some theory, data quality, better usage of params, RL. So not really one specific special sauce, but a lot of overall improvements that stacked up.
https://x.com/OriolVinyalsML/status/1990854455802343680
"
Hmm let's add this to loss function "How natural organisms are so evolvable (capable of quickly adapting to new environments)? A key driver of evolvability is the widespread modularity of biological networks, their organization as functional, sparsely connected subunits. Ubiquitous, direct selection pressure to reduce the cost of connections between network nodes causes the emergence of modular networks. Computational evolution experiments with selection pressures to maximize network performance and minimize connection costs yield networks that are significantly more modular and more evolvable than control experiments that only select for performance."
https://pmc.ncbi.nlm.nih.gov/articles/PMC3574393/
gradient descent multiple times in slightly different directions using stochastiscity for few epochs and then picking on one who has the best loss reduction/local minima flatness ratio
reverse engineering how foundational models for healthcare work using physics based methods to improve their steerability/relability/accuracy,... physics based mechinterp on open source alphafold openfold
Physics of feature and circuit forming and it's dynamics in intelligent systems generally will be a gigantic scientific field
A great visualization of how people are different from AI in terms of their flaws and capabilities
The capabilities of AIs are growing, but they are different capabilities in a certain way, even though there is some overlap in those capabilities
And different AIs are different in their jaggedness, and different people are different in their jaggedness, but we as humans are still closer to each other than we are to AIs in the space of all possible intelligences
https://x.com/burny_tech/status/2002865099820667186
It's interesting how ML/AI systems aren't really trying to be replicas of human thinking like many people seem to think
I love how they're mix of all sorts of insights that we got from neuroscience, optimization theory, probability theory and statistics, computer science, physics, applied mathematics, pure mathematics, philosophy, cognitive science, psychology, linguistics, empirical random playing around with random stuff, etc. into one complex system.
Neuroscience - Neuronky začaly jako "hmm neurony jdou popsat Hodgkin Huxley modelem, ale vlastně taky jdou hodně simplifikovat jako graf s vertices jako neurony a edges jako synapses, what if we added more neurons and...". Nebo reverse engineering neuronek, obor co se nazývá mechanistic interpretability, je v podstatě digitální neurověda.
Optimization theory - Gradient descent přichází odtuď a je to nejpopulárnější učící algoritmus neuronek, a dokonce je původně z astronomie.
Applied mathematics: Neuronky jsou aplikovaná lineární algebra, víceproměnná reálná analýza, teorie pravděpodobnosti,... A přes tools těhle metod jde dělat tuna triků pro efektivnější tréning.
Probability theory and statistics - Neuronky jsou technicky statistická metoda, existuje statistical learning theory, nebo existují bayesovy metody, atd.
Physics - Hopfield networks jsou prekurzory k neuronkám co známe teď, a jsou spin glass system, za co Hopfield dostal nobelovku z fyziky. Diffusion modely jsou vlastně díky nonequilibrium thermodynamice. AdamW optimizer do gradient descentu přidává "momentum". Flow matching učící algoritmus se učí velocity fields. Na chápání neuronek se teď víc a víc hází statistická fyzika.
Pure mathematics - Jsou pokusy o pure math modely deep learningu, včetně toho categorical deep learning. Nebo je geometric deep learning co používá na analýzu deep learning architektur teorii grup.
Computer science - Např turing completeness je též relevantní, existují architektury jako neural turning machine. Neuronky jsou algoritmy a computational complexity je taky důležitá pro efektivitu. Optimizing concrete software and hardware matters. Etc.
Cognitive science/Psychology/Philosophy - Etika se řeší v alignment problemu a control problemu. Kognitivní vědy řeší problém toho jaký fyzikální systémy mají "pravou" inteligenci, reasoning, nebo co je to vědomí, a jak tyhle slova zadefinovat (uh to je rabbithole).
Empirical random playing around with random stuff - Tak některý improvements vznikly. Dost tvoření modelů je v podstatě alchemie
A je toho milion víc. :D
Maybe *future* AI systems won't have singular consciousness but a ton of tulpas since they have continuous persona space that they will be navigating and spawning tons of nested conscious subagents
When superhuman general flexible adaptive AI systems for math
The job of the future that is becoming reality is coordinating, sanity checking, double checking, etc. AI agents coordinating AI agents coordinating AI agents... Such as Claude Code with it's subagents and potentially subsubagents.
https://code.claude.com/docs/en/sub-agents