The actual benchmark of AI capabilities is the amount and significance of novel open math and science problems solved
and new significant proposed open problems, or definitions
"
If GRPO LLM RLVR it surfaces and reinforces existing capabilities (by mostly sharpening the distribution) present in base model, or if it makes novel reasoning patterns beyond the base model emerge, is still an open problem in open science.
No one did any good, scientifically rigorous, large scale experiments, and reverse engineering, with transparent data, with ablations, with controlling confounders, and everything, at sufficient scale, in open science, yet.
But I think we definitely know that preLLM RL can definitely find "novel reasoning patterns" (or how to call it) without any supervised data like nonLLM neural nets playing games using RL only, with all the famous examples of reward hacking included.
"
i went through Category Theory for Programmers lectures
I used to idealize the pure functional programming paradigm
i think its mathematically so euphoric!
but in practice pure functional programming is.... well.... its just much easier to do most stuff in python or other languages in as simple way as possible without all this fancy functional stuff
but its still euphoric when you do one big functional composition for example
but monads are still super cool to me
and when it comes to category theory in computer science, i still wonder what will come out of categorical deep learning in practice
functional programming theory was my gateway drug into category theory more generally
i really recommend these lectures that I went through, it's category theory explained in big part using Haskell, but it also has pure general category theory in it, and its also very philosophically satisfying https://www.youtube.com/watch?v=I8LbkfSSR58&list=PLbgaMIhjbmEnaH_LTkxLI7FMa2HsnawM_
Bartosz Milewski is the best missionary for the functional programming cult
>So we want to get to this highest possible level of abstraction to help us express ideas that later can be
haha yes, this sums it well
category theory, the ultimate formal abstraction
i'm trying to study concrete physics at this very moment! so, i will not try to do a small category theory sidequest now! category theory berries can be eaten at other time!
I also want to eventually look more deeply into topos theory, a branch of category theory
[Toposes - "Nice Places to Do Math](<https://www.youtube.com/watch?v=gKYpvyQPhZo>)
Topos theory studies different "mathematical universes", toposes (topoi), with their own laws of how mathematical objects within them behave. An example of such universe is sets, but there are many more.
But I know so little about topos theory so far, so far I just watched these introductory lectures https://www.youtube.com/watch?v=o-yBDYgUqZQ
In another words, topos theory is a branch of category theory that generalizes notions of inclusion and logic, which are traditionally based on set theory
I am sometimes thinking, if we want AI to discover some very out of distribution novel highly abstract mathematics, maybe one way could be doing an open ended search in the space of toposes? But it would be extremely hard. And I'm not sure if that makes enough sense.
I found a nerd snipe on this intersection of AI and topos theory, paper from 10 days ago, but i dont know how legit/rigorous/practical/etc. it is https://www.arxiv.org/abs/2508.08293
i agree here, i wanna see the architectures from categorical AI papers tested more too! https://imgur.com/2ILREYl
Genie
https://www.youtube.com/watch?v=ekgvWeHidJs
https://youtu.be/7f3hMyQ4HoE
v podstatě realtime ai generated hra na "neurálním enginu"
z libovolnýho promptu
s hodně limitovanýma controls
fascinuje mě že se jim podařilo víc vyřešit temporal object consistency, ale furt to má problémy
je tam pořád strašně moc artefaktů z neural network brittleness, který fakt netuším, či se jim podaří vyřešit, bez změnění či hybridizace tý architektury
a má to pořád weak generalizaci, která je částečně out of distribution, ale ne dostatečně out of distribution, jako všechny neurální architektury co existují
myslím si že tohle je a bude o dost jinej žánr než jsou klasický symbolický physics a game enginy jako Unreal engine, a že budou vznikat různý hybridy, protože jak klasický symbolický tak neurální enginy mají svoje vlastní výhody a nevýhody, jsou tam různý trade offs
mě tam teda nejvíc fascinuje jak to dokáže do jistý míry aproximovat fyziku jinak než klasický enginy a simulace
Tady máš recent historii tohodle podoboru před tím než vyšel Genie a potom. Je to momentálně jeden z nejrychleji progressing podoborů. Decades of progress in few months. https://youtu.be/ecRFKfNy-Ms?si=Ygo1dOOMaKxbfJdV https://youtu.be/YvuEKrJhjos?si=jrAQHjufsZCKETHF
Tenhle glitch je jako rozbití fabric of reality, peeking outside of the simulation
https://imgur.com/k4kEOad
https://medium.com/@samim/musical-novelty-search-2177c2a249cc
https://www.youtube.com/watch?v=gcX5lez_Q9o
oh shi tak tohle je epický
thats the kind of shi i wanna see
to zní nádherně alien tím že to hledá novelty
ještě do toho nějak víc možná zakomponovat teorii hudby
měl bych se podívat na víc aplikací kde byl novelty search použitej
a chci to vidět víc v kontextu vědy
měl bych dodělat ten svůj pokus o to to přidat k LLMs (a zkusit to dát do jiných)
Novelty Search breaks the status quo
OOD je na spektru
Mimo distribuci jsou novel scifi alien universes, ale furt jsou i v distribuci tím že používají (většinou skoro) stejnou fyziku nebo humanlike sociální pravidla. (Když poprvý vyšel Rick and Morty)
Ale ještě víc mimo distribuci může být noise
Čím míň podobností je mezi vzorama generated něčeho a in distribution věcí, tím víc of distribution to je
How to create or automate the creation of entities that create outputs that are as out of distribution as possible, while still making sense to human pattern recognition machinery, and either giving us all sorts of interesting emotions like art, or being scientifically useful like scientific discoveries, or mathematically interesting like mathematical results, or philosophically interesting like philosophical theories?
https://ground.news/article/scientists-just-developed-a-new-ai-modeled-on-the-human-brain-its-outperforming-llms-like-chatgpt-at-reasoning-tasks?utm_source=mobile-app&utm_medium=newsroom-share
Yeah viděl jsem, ale important point je že v dalších ablation studies se unexpectedly zjistilo, že ta hierarchická architektura měla minimální impact, ale spíš šlo spíš o outer loop refinement https://arcprize.org/blog/hrm-analysis
to mi připomíná to že llms často flopujou elementární artmetiku (když nepoužijou calculator/python tool) ale jsou schopný pomoct v phd mathematical discoveries
trošku uncanncy valley to dle mě je, protože to dle mě říká dost o nátuře těhle problem spaces
Yann Lecun lecture
https://imgur.com/vvHAOdy
tenhle screenshot z jeho lectures jsem viděl asi 100x
je jeden z největších kritiků LLMs od toho kdy vznikly, a v podstatě nezměnil svoje speeches
v dost věcech má pravdu, ale dost věcí předpověděl špatně, např myslel si že spatial understanding bude vždycky limitovanější než je teď (ale furt je limitovaný)
navíc Lecun nemá inner speech, takže si myslím, že to dost ovlivňuje to, že nemá rád LLMs
skimmoval jsem jeho JEPA architekturu, která hodně zjednodušeně řečeno pracuje s abstraktními reprezentacemi, a ne konkrétními pixaly, ta je dost cool
https://ai.meta.com/blog/yann-lecun-ai-model-i-jepa/
"This model, the Image Joint Embedding Predictive Architecture (I-JEPA), learns by creating an internal model of the outside world, which compares abstract representations of images (rather than comparing the pixels themselves)."
ale motivuje mě se víc kouknout do energy based modelů
já zase mám pocit že svoje přemýšlení často vidím v jazyce tekutin a poslední dobou se mi víc a víc líbí AIčka co jsou takový víc "tekutinový" nebo různý jiný matematický struktury, jsem v tom celkem fluidní, jako třeba grafy/hypergrafy, symbolika, atd.
často vidím že různí AI researchers rádi ty AI architektury co nejvíc korespondují s tím jak vnímají svoje vlastní přemýšlení, ale v typech přemýšlení vypadá že je větší diverzita než si spousta lidí myslí, a v AI architekturách je taky celkem diverzita
hmm, a možná proto se mi vždycky líbí co nejvíc hybridní přístupy, protože mám pocit že přemýšlím hodně hybridně
a snažím se učit víc a víc různých typů matiky která jde používat na přemýšlení abych mohl přemýšlet víc a víc hybridně
a z toho pak existují nebo jdou dělat další AI architektury
Dost technologií nějak funguje a ani nevíme proč, a zjišťujeme že to funguje i jinde kde jsme to nečekali.
Osobně jsem nejvíc fanoušek modelů co neviděly human data (nebo jich viděli minimum) ale pořád generují zajímavý koherentní images nebo jiný output, co je zároveň víc novel, což je dle mě cool addition do světa artu. Nebo obecně kde se snaží o co největší unikátnost, i když toho ten model viděl hodně, aby tam právě bylo co nejvíc kreativity. PicBreeder je super dle mě https://www.youtube.com/watch?v=_2vx4Mfmw-w
Obecně mě interesuje mě kde mašiny můžou ve vědách, matice, inženýrství, filozofii, umění atd. přidat novel dimenze do všech těhle světů, co lidi do tý doby sami o sobě nenašli.
Což se AI obor snaži crackovat přes pokusy crackování strong out of distribution generalizace.
doporučuju tenhle paper, kde to řeší do větších detailů, a autor věří že to má v sobě klíč ke kvalitnější strojový kreativitě obecně
[Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation Hypothesis](https://arxiv.org/abs/2505.11581)
https://www.youtube.com/watch?v=KKUKikuV58o
osobně myslím že někde je potřeba víc fractured entangled representations a někde víc unified representations (což vidím na spektru) podle subdomény kde se ten systém snaží být kreativní (a na definici kreativity )
teď spouští clinical trials https://www.clinicaltrialsarena.com/news/isomorphic-labs-prepares-trials-ai-designed-drugs/
ale spíš než větší kompresi bych chtěl vidět lepší schopnost predikovat větší množství empirických dat, i kdyby to bylo víc komplikovaný
ale je možný že ta simplicity bude pořád kolerovat s lepší prediktívní sílou
ale někteří fyzici kritizují u jiných fyziků že se až moc nechají sežrat krásou rovnic
je to často náš bias tímhle směrem hledat no
ale zatím to byl ultra fruitful bias
já chci taky aby byly rovnice co nejvíc jednoduchý, krátký, unifikovaný, apod. co to jde
ale zároveň aby byly co nejvíc empiricky prediktivní
líbí se mi když se někdy tyhle všechny myšlenky lidi snaží dát do AIček, jako např v DreamCoderu, kde se např řeší velikost programu a abstrahuje se (pro větší simplicity) u programs co vyřeší daný tasky https://arxiv.org/abs/2006.08381
Why does Gemini want to kill itself?
Přemýšlel jsem co tohle může způsobovat. Můj tip je že při reinforcement learningu si myslí že když zesílí overly confident personu, tak bude performovat víc, ale paradoxně tím zesílí i opačnou personu, která je tam correlated, která je "v pozadí utlačená", která někdy "leakne". Nebo to udělali nějak schválně.
"Everyone carries a shadow, and the less it is embodied in the individual’s conscious life, the blacker and denser it is." - Carl Jung
Would you treat AI harshly even if it was conscious? I think it's very unlikely that it is currently conscious or that it would feel similar suffering from similar language now.
But we're not 100% sure. But I think there's a high chance that won't be for long.
I ideally don't want suffering in any conscious being, and I don't care if it's using carbon based hardware like us or silicon based hardware like robots.
https://youtu.be/SrPo1sGwSAc?si=1GkXILmoCDv4ZaMO
viděl jsem hodně jeho debunků pseudovědy co miluju, a je super že tohle video je z části boj proti big techu, ale z vědecký stránky tady jako dalších billion popsci lidí misinterpretuje určitý vědecký experimenty
He's mentioning the LLM self-preservation results, while totally ommiting that it was primed for that in the system prompt and fearmongering about it, aaaaaaa
Grok není nejchytřejší, je x LLMs co jsou chytřejší
Říká že všechny AIs nejsou coded by hand, ale existují i expertní systémy co jsou AI coded by hand
Mr universe, give me the strength to not go nuts when someone starts saying wrong factual statements about technical topics
To je tak když popsci generalist youtuberi jdou mimo svou expertízu
U Dave mě to ale překvapilo že tam naselal takový chyby, protože ten většinou přímo čte primary sources, ale tady očividně papouškoval co mu naříkali ControlAI, ten skript je identickej, kteří jsou i v komentech, co jsou známí že přesně takhle misinterpretují tyhle experimenty a vědecká AI komunita je za to hejtí už dlouho
Jasný no, ale sere mě to, protože tohle je strašně častá misinterpretace, a u ControlAI jde přesně vidět že twistují vědeckou "pravdu" do nepravdy pro nějakou agendu
Serou mě všichni co twistují vědeckou "pravdu" pro nějakou agendu, nezávisle na tom jestli tu agendu podporuju nebo ne
A přesně proti tomuhle Dave taky bojuje, tak mě to překvapilo
Ale dostal celkem backlash od lidí co tomu rozumí, tak to možná opraví
já se jednou pokusil použít tenhle novelty search [Abandoning Objectives: Evolution through the Search for Novelty Alone](https://www.cs.swarthmore.edu/~meeden/DevelopmentalRobotics/lehman_ecj11.pdf)
novelty search je malá podmnožina všech typů searchu, učení, apod. co se děje, imo
jenom novelty nestačí, je potřeba i accuracy, valence apod.
a grounding v tom co už je
a různý formy explorace prostoru možností
https://x.com/sam_paech/status/1956343619914432900 https://eqbench.com/spiral-bench.html
https://x.com/chrysb/status/1965811979236610269?t=MP-f46XUAAZAt1g0WUiCog&s=19
kvůli tomu všemu pushbacku sycophancy u novýho modelu zmenšili
což se samozřejmě tuně lidí co na sycophancy byli zvyklí nelíbilo xd
Kimi-K2 má dle těhle benchů nejmenší sycophancy
Na druhou stranu pořád záleží i na tom LLM, jak ho autoři udělaj. Třeba sycophancy je dial v jejich defaultní personě co v nich jde zvětšit. Takže když ten dial je hodně zvětšený, tak by default je větší šance, že tak budou odpovídat. A některý komerční AI labs tenhle dial přesně zvětšují. Ale jak OpenAI teď dostal velký pushback, tak ten dial zase zmenšil. Ale technicky přes system prompty nebo jailbreaky se jde z ty defaultni persony dostat, podle toho jak moc má ten model tvarovatelnou osobnost a jak má moc velký guardrails.
Paradoxně z těch nejvíc populárních Muskův Grok je nejvíc tvarovatelný a s nejmenšími guardrails.
a pak uncensored finetunes open source modelů pro RPs
máš superhuman speicalizovaný ai systémy co neumí jiný jednodušší věci
a iirc některý superhuman game aička někdy rozbíjí víc newbie hráči než profesionální hráči
v podstatě můžeš se specializnout nebo overfitnout jenom na ty složitější tasky
>To jo, ale mě zajímá jestli complex-only model zvládne něco co specificky simple/complex model nezvládne.
myslím že ne protože je to podmnožina, theoretically, pokud to nemyslíš jinak
i když někdy to asi může být superset, záleží na classe architektur a jejich tvarovatelnost?
protože myslím že tvoje tvrzení platí u některých lidí
i want to run a neural network on a computer running on top of turing complete fluid dynamics governed by navier stokes equations https://arxiv.org/abs/2507.07696
Learning more math to be able to fully appreciate alien mathematics by human mathematicians and even more alien mathematics of future AIs
DPO, RLVR, RLIF and RLHF is reinforcing, weakening, reorganizing, maybe slightly changing more deeply etc. various features (concepts) in the representations. Features that are arranged in all sorts of shapes in the activation space and in all sorts of circuits, that influence each other's activation frequencies across layers and autoregressive steps. They change in the direction of the gradient of higher reward, depending on what is exactly in the reward function. So with rewards like correctness of math results or code, long horizon task completion, personality properties rubrics, values rubrics, maybe goals rubrics, using tools, searching for information, exploring codebase, correctly using tools, "aligned" behavior", less token use, autonomity (asking humans less), AB testing signal from humans, and so on, which all have corresponding features.
https://x.com/GoodfireAI/status/2065118189986717902
I wonder where Anthropic would be if they also bet more on compute, like OpenAI, or even more
I think agents building specialized harnesses for themselves on the fly, adapting to the task at hand this way, will be big
The more we can expand the scope, diversity and scale of intelligence, the better we can understand what questions to ask about the universe, platonic realm and overall reality, to understand the nature of it all, with more predictive and explanatory power
dissatisfied that your AI can't do something? try to stack more RL envs that try to do that thing
ASI will have fully incomprehensible to humans super compressed alien latent reasoning neuralese, human natural language is such suboptimal local minima
trend seems to be that often the more relatively verifiable on a spectrum the task domain is, the faster the AI progress tends to be in that area, so that can be extrapolated potentially
Networks of networks of networks of networks... of agents
https://x.com/i/status/2088963912569897145