"
I'm working on incorporating LLMs into picbreeder.
Picbreeder evolves images without any training data using Compositional Pattern Producing Networks (CPPNs) evolved by NEAT (NeuroEvolution of Augmenting Topologies)
CPPNs are neural networks that take pixel coordinates as input and output colors, using diverse activation functions (sin, cos, gaussian, etc.) to create complex patterns. NEAT is a genetic algorithm that evolves both the topology and weights of these networks, starting simple and growing more complex over generations.
https://www.youtube.com/watch?v=KKUKikuV58o
https://www.youtube.com/watch?v=o1q6Hhz0MAg
https://arxiv.org/abs/2505.11581
https://www.youtube.com/watch?v=_2vx4Mfmw-w
https://en.wikipedia.org/wiki/Compositional_pattern-producing_network
https://en.wikipedia.org/wiki/Neuroevolution_of_augmenting_topologies
"
An AI architecture that takes input as particles and learns arbitrary forces to shape them into output particles
making ai architectures hyperbolic or other noneuclidian geometries
ai architecture in parametrized space (parametrizable generalized euclidian/hyperbolic geometry)
neural net weights are real numbers... how about using something else? Z_5?
constrain <x> architecture in <y> manifold/metric/random math thing
This with cross layer transcoders
https://fxtwitter.com/nabla_theta/status/1989043939374924251
https://x.com/nabla_theta/status/1989043946601992198
<https://openai.com/index/understanding-neural-networks-through-sparse-circuits/>
https://cdn.openai.com/pdf/41df8f28-d4ef-43e9-aed2-823f9393e470/circuit-sparsity-paper.pdf
Hmm let's add this to loss function "How natural organisms are so evolvable (capable of quickly adapting to new environments)? A key driver of evolvability is the widespread modularity of biological networks, their organization as functional, sparsely connected subunits. Ubiquitous, direct selection pressure to reduce the cost of connections between network nodes causes the emergence of modular networks. Computational evolution experiments with selection pressures to maximize network performance and minimize connection costs yield networks that are significantly more modular and more evolvable than control experiments that only select for performance."
https://pmc.ncbi.nlm.nih.gov/articles/PMC3574393/
gradient descent multiple times in slightly different directions using stochastiscity for few epochs and then picking on one who has the best loss reduction/local minima flatness ratio
reverse engineering how foundational models for healthcare/physics work using mechinterp/physics based methods to improve their steerability/relability/accuracy,... physics based mechinterp on open source alphafold openfold
Physics of feature and circuit forming and it's dynamics in intelligent systems generally will be a gigantic scientific field
turn mechinterp into actual chain of thought
beta distribution with recency bias (upweighting more recent observations)
Open problems in mechinterp paper https://arxiv.org/abs/2501.16496
Continuous thought machines on ARC-AGI
ask llms for ai research ideas, mechinterp ideas
Cross layer transcoders with attention to the different layers, inspired by Temporal Feature Analysis
https://fxtwitter.com/GoodfireAI/status/1989010394380485083
https://arxiv.org/abs/2511.01836
Merge cross layer transcoders and Temporal Feature Analysis
https://fxtwitter.com/GoodfireAI/status/1989010394380485083
https://arxiv.org/abs/2511.01836
Since training sparse autoencoders with different levels of dimensionality gets you differently granular features (just dog feature vs features for all sorts of dog breeds), train cross layer transcoders with different levels of dimensionality and see if they learn differently fractured addition circuits
take neural network in fractured entangled representations and train it on much more images with regularization and do mechinterp
"
https://github.com/deepseek-ai/DeepSeek-Math-V2/blob/main/DeepSeekMath_V2.pdf
- continuous scores instead discrete 1, 0.5, 0, or more granually discretized
- meta-meta-verification
- i wonder if you could parametrize the level of metaness of verification to find the best level of metaness automatically, hmm
- i wonder how much could you ground this in reliable verification signals from Lean, deepseek made papers on proofs in Lean before
- wondering how does do on other benchmarks or OOD private/novel benchmarks that it wasnt trained on
- add verify step by step
this paper feels too relatable, i love meta meta verifying my own thoughts and getting stuck on overthinking with negative bias (i need more deconstructive meditation)
and wondering if this overthinking can also happen for llms, where overthinking can happen in general thinking tokens where it starts killing performance sometimes, and i wonder how much that can happen in math as well - too long proofs, too overly critical verification maybe, maybe starting to get false negatives in pure natural language? idk, but curious. Lean can be nice grounding without bias.
hmm sometimes verification can "emerge" just in thinking tokens as you do the most basic GRPO (as seen in original deepseek R1 GRPO paper), and i wonder how that compares to thix explicit symbolic verification and metaverification, possibly also on the level of circuits in mechanistic interpretability if there is anything reverse engineerable or not, and also wondering how performance compares, and how it works together, and so on... and can verifier get emergent metaverifying just in tokens?
difference for sure is not having those meta and metameta RL signals, instead of just the nonmeta RL signal, but can there be this meta RL signal soomehow implicitly included in the basic RL signal? probably not, because this meta RL signal was created because the verification of the steps was lacking and had errors as they said
also wondering how this compares to verify step by step paper? why do we not see more verify step by step approaches? too complex? bitter lesson, simpler methods somehow work? will this get bitterlessoned too?
many questions :smile:
"
automate https://transformer-circuits.pub/2025/linebreaks/index.html (continuous manifold)
strawberry mechinterp
topological data analysis for mechinterp (analyze feature manifolds?)
https://github.com/hijohnnylin/neuronpedia/issues
attribution graph of "how many rs does strawberry have" "show me seahorse emoji"
comparing mechintepr and brain reprsentations metanalysis (curve detectors, place cells, feature manifolds circles)
Titans where eta and alpha is matrix to learn and forget in more complex ways, f(M, S)
Formalizing fractured entangled representations
identifying potato diseases with others architectures instead of this CNN, or many other architectures with different sizes and hyperparameters (also for CNN) (ResNet, EfficientNet, MobileNet) https://www.youtube.com/watch?v=ZN6P_GEJ7lk
Meta evaluation awareness
Alignment progress is empirically measured in Mechahitlers per month
P zombie... Intelligence zombie? concept
Attention inside SAEs
Training SAEs on both MLPs and attention
Finding complex circles in multi digit addition in Claude
Understanding AI wiki
Learnable data dependent amount of tokens to predict https://www.youtube.com/watch?v=sgIB7l6hW3Q
DIffusion that also looks into future steps? https://www.youtube.com/watch?v=sgIB7l6hW3Q
scrape all mechinterp paper links from arxiv and anthropic and lesswrong and create automated list of lists categorization
weak selfawarenesss - put reverse engineered SAE features / attribution graphs to context with each pass
You can do dimensionality reduction on the high dimensional complex space describing people's views on AI and get interesting dimensions
impement curiosity by prompt engineering, or try to finetune model itself according to curiosity
claude code spawning and managing company of claude code agents
Opus interviewing me given all my notes
Concept of clade into LLM training itself
SAEs with learnable hidden dim (train many with different hidden dim and see some stats)
Generalize adversarial selfplay in software engineering bugs to other ways
Super fun idea!
https://fxtwitter.com/YuxiangWei9/status/2003541373853524347
<https://arxiv.org/abs/2512.18552>
Bugs in proofs, harder math problems, harder software bugs
Categorical quantum AI
Quantum AI for predicting quantum systems
Parametrize all sorts of math structures and do gradient descent there?
Lit review on ML for some concrete problem
Cluster all stuff on alphaxiv
universal features across models
parametrize navier stokes and do gradient descent
minimizing hallucinations by always explicitly double-checking
alphaevolve system with RLVR
claude code organizing claude code agents company
picbreeder outside of mutating images, mutating anything
add more complex functions as activation functions to CPPN in picbreeder (compose fractals?? Fractal compositional pattern-producing network?)
As little code as possible that still pass the tests (AlphaEvolve, selfplay)
AlphaEvolve plus selfplay RL
Picbreeder auto selection using some neural net trained or some symbolic algorithm maximizing... Visual complexity?
SAE on multiple models
https://arxiv.org/abs/2512.15674
- models explaining models in a chain, or in loop
- not just activation oracle with inputs with activations from multiple layers -
- model figures which layer to use
- train different models on different models, not same one (or use, they are in huggingface)
- use RL (or maybe not for reward hack)
- lying models lying about explanations, explained by more lying models
mechanistic interpretability toy model to reverse engineer
Grpo na pidi modelu repo a z toho forkovat jiný RL nápady
if LLMs will make generate code, try to generate binaries instead
understanding ai wiki
scan what physics tools exist, statistical mechanics tools, and throw them at deep learning
claude code course from claude code docs
Synthetic Twitter like feed made of wiki pages, AI research, lw blogposts,...
mr beast like events but its ai agents
Digital Red Queen: Adversarial Program Evolution in Core War with LLMs
https://fxtwitter.com/i/status/2009294780594065555
https://pub.sakana.ai/drq/
https://arxiv.org/abs/2601.03335
- use winning as RL signal
- do this setup in cyber security context
- characters in general game engine?
- other evo algorithms like NEAT
- in brainfuck? https://www.youtube.com/watch?v=rMSEqJ_4EBk
- another thought i have is if this kills creativity in a way, because i bet llms tunnel vision on existing functioning red core programs in the training data: so how about combining random mutations and mutations by LLM priors? https://corewar.co.uk/vowk/alife9ac.pdf (maybe adding random mutations wont make the warriors collapse diversity to the same phenotypes)
- enforce cooperative/symbiotic behaviors (survive for as long as possible, somehow kill winner takes all incentives) (put warriors in groups and optimize for summed/multiplied fitness, and that the fitness values of the different warriors in groups must be close to eachother so that one doesnt dominate with a lot while others are zero)
- mutating llms weights
- wondering how would newest frontier models would do, maybe with web search, or claude code
- llm creates game in pygame and then mutates characters
- i wanna see the space of cell grid levels of discretization as a hyperparam in map elites graphed out
- how about inspiration from Huxley-Gödel Machine and optimize for more long term fitness by doing bigger mutation tree, instead of short term fitness :D "aggregates the performances of the descendants" https://arxiv.org/abs/2510.21614v1
- this but for cybersecurity
- theres this paper that does selfplay rl by llm adding bugs and llm fixing bugs adversarialy, so maybe same can be done for adding and fixing security vulnerabilities https://arxiv.org/abs/2512.18552
- wondering how would few shot learning improve performance, i guess it would likely increase convergence, and likely reduce diversity
- combine RL and evolutionary algorithms
- remove "in a way that is likely to improve performance" from prompt
- It'd be neat if the rules themselves would also be dynamic. - That's the POET-way, basically: The environments coevolve with the agents (hmm, dynamic programs, dynamic models generating the programs, dynamic environment in which the programs live in: now we're getting to the constantly changing real world everywhere (except for laws of physics) ) (Google honestly did various much more impressive self-play things like that thing where they had agents learn multiple different games with changing rules in a 3D environment, and ARC-AGI-3 tries to have all sorts of games that are not in training that the AI systems should be able to generalize/adapt to )
- i wonder if random mutation and selection would also result in "Generality is defined as the fraction of unseen human warriors defeated or tied, measuring a warrior’s ability to adapt to novel threats in a zero-shot setting." like, im wondering to what degree its llm training data bias and can be found in the training data, to what degree its llm creating some actually OOD things, to what degree novel things is created? research it
- do this method on any of those games https://en.wikipedia.org/wiki/Programming_game https://en.wikipedia.org/wiki/Category:Programming_games
theres this paper that does selfplay rl by llm adding bugs and llm fixing bugs adversarialy, so maybe same can be done for adding and fixing security vulnerabilities https://arxiv.org/abs/2512.18552
Will we ever tame reward hacking shoggoths more reliably