https://x.com/polynoamial/status/1946478249187377206 "t @OpenAI achieved a milestone that many considered years away: gold medal-level performance on the 2025 IMO with a general reasoning LLM—under the same time limits as humans, without tools."
No Lean
Just pure general LLM with no tools
Bitter lesson strikes once again 🙈
this came from pouring compute into a scaling a general RL method according to current public information
We should benchmark creating whole papers etc.! Problem is that it's hard to verify that so it's hard to scale that in automated benchmarks, or as verifier signals for RL
but i think the best benchmark for AI is novel scientific discovery, especially in physics, like discovering quantum mechanics or general relativity wit hthe data that scientists have before they discovered it
i keep coming back to the bitter lesson, because i always start believing that we need more complex neurosymbolic systems, and then some general simple neural method surprises me, and repeat :D
i guess depends which subset of physics do you mean, it is already helping to some degree, and i believe it can help more
but for CERN related stuff, data is probably a bigger bottleneck
I'm currently also obsessing over LLMs combined with Lean
for example this line of research is interesting to me Automated discovery of fundamental variables hidden in experimental data https://www.nature.com/articles/s43588-022-00281-6
it found number of variables very close to the ground truth for those dynamical systems where we know it, but it also found some number of variables for those where we dont know ground truth! but we struggle to extract the equations https://imgur.com/a/bET3vO4
video on it
https://youtu.be/XRL56YCfKtA?si=cEepwmpkRH9hYAfZ
but there are tons of other different AI x physics intersections im also interested in, this is one of them
but i also have this super ambitious hope that some form of future AI will one day help us get unstuck when it comes to quantum gravity and i wanna work towards that
either by better theories, or by designing more feasible experiments, as we are missing empirical data there to test even current theories that exist
beyond current ML analysis in CERN
beyond current LLMs producing stuff that is not mathematically coherent
beyond current initial attempts at neurosymbolic systems (connecting LLMs with Lean or maybe systems like DreamCoder)
beyond current physics inspired ML architectures for predicting dynamical systems
beyond current AI assisted design of experiments
beyond automated discovery of fundamental variables hidden in experimental data
etc.
so i'm currently collecting and learning everything about the intersection of AI and physics that is relevant to this
maybe one day we'll make more progress there
Yea AI designed experiments is also a subfield I see future in, like https://youtu.be/T_2ZoMNzqHQ?si=Lg-aIVz_XbK7MgWt
https://x.com/PhysRevX/status/1910788071701741760?t=X83RD0PdMTHjpIPFnMwCqA&s=19
My current model is: GPT-5 release is interesting. Super over hyped release, but the model is solid when you look at all the benchmarks, but it's not groundbreaking at all, so everyone is disappointed. They basically caught up to Google and Anthropic in some aspects, a bit overtook them in other aspects, but failed to catch up to them in other aspects. OpenAI has been imploding for a while now, and is being gutted from all sides. A lot of the best researchers recently escaped to Safe Superintelligence lab, Thinking Machines lab, or got poached by Zuck. I do wonder if some other AI lab will soon overtake them in user count as they continue imploding. Their primary moat currently is user base capture. Most normies have no idea about all the alternatives. OpenAI should have called o1 or o3 models GPT-5, those were actual breakthroughs. Calling this model GPT-5, when expectations were on the moon, was total mistake lol.
i still can't grasp how we made computers speak human language and pass turing test
i will always be mindblown by it, and never bored, and i wanna understand, HOW
what the f@ is the nature of language??????????? why can we do that???????????????
why are llms so coherent??????????
it breaks all my intuition about combinatorial explosion of possibilities in the structure of language
> its just predicting the next token
i know thats the objective, but why can deep neural nets reduce that combinatorial explosion of possibilities in the structure of language so effectively????????
imagine the space of possible linguistic structures, thats unimaginable number
and even some couple of millions parameter models have coherent language, and those relatively small numbers are god tier compression relative to that
images are maybe even more mindblowing to me from combinatorial explosion perspective, how can couple of gigabytes include all the autistic visual universes in it, all the trains, all the art, all the cats
its not infinitely granular, etc., but still
like sure, i know the architectures, and my current model that all these generative models include compressed version of their training data with weakly generalized brittle features and circuits etc.
but still, HOW? why does gradient descent, why does the diffusion etc. know how to make all this internal structure we study in mechanistic interpretability????
thats not intuitive at all IMO and still a big mystery
all the theories, like flat minima generalizing, etc., are still not enough for me
i think we need better math of emergence to explain all the physics of how circuits form in mechanistic interpretability
i want more mathematical models that will predict why all this deep learning empirical alchemy works! https://imgur.com/0AgDyMa
https://www.youtube.com/watch?v=HR-_U0Pzl1Y
Beff Jezos
The comments here are crazy tbh. I understand some of the scepticism behind his company, wanting concrete benchmarks, concrete hardware in reality, etc., which I want too, and I'm both hyped and skeptical.
Or I understand him probably hyping it too much, as the probabilistic paradigm is technically an accelerator for certain probabilistic algorithms. I mean, there are bazillions of papers describing probabilistic computing with probabilistic bits, and they describe a lot of this? I don't get why it's so controversial.
Or I understand there's the whole memetic tribal warfare between the e/acc ideological AI tribe and other tribes.
But saying that everything he says doesn't make sense is just completely wrong.
It makes sense that probabilistic hardware can act as an accelerator for probabilistic algorithms.
Or he talked about how he contributed to Google's quantum machine learning software, Tensorflow Quantum, which is a very real thing that exists and works, and you can play with it now.
Or he described various of parts of physics well imo, and its frontiers.
Or I think his philosophy inspired by physics is very interesting.
I'm so confused about so many of these comments.
I definitely have a bias, because like me, in big part he wants to use AI to understand the physics of the universe, so i resonate with that deeply
yeah there are countless different types of AI accelerators
Groq is probably the most famous and widely used one rn (if we don't count NPUs everywhere)
But Beff seems to be in the probabilistic computing / thermodynamic AI accelerator camp, and he isn't alone there
even though he showed some physical chips, or he showed some graphs, i also want more stuff
he apparently now has a paper in peer review
i think i also heard him they already give some api to some small amount of people
i'm both hyped and sceptical, and also want more demos, papers, etc., but still as I say, probabilistic/thermodynamic computing/AI has bazilions of papers, so his company probably isnt exactly that unique
and he has background in physics based AI already, he contributed to google's quantum machine learning stuff
and he might be overhyping his acceleator company to get that investor money
it feels like people are way too polarized on Beff overall
but im definitely biased, because im also in love with physics based AI like him
and i also love seeing biology from the lens of physics like him
and i also love theories of everything in physics like him
and e/acc was also an influence on me that weakened the monopoly of doomerism in my set of perspectives, but i dont deeply subscribe to it, as i just like some of its ideas, just how i like some of the ideas from doomers
many takes about AI by general public that seem so extremely uninformed, and without technical nuance, to me, but it's probably not attempts at factual claims but just emotivism
i get angry about some of these claims periodically, and I then write a walls of text of technical details that dissects their factually incorrect claims, but because they have zero knowledge in actual technical details, and they don't even wanna hear it, because its emotivism, as you say, then it's like talking to a wall
on both sides, not just the negative tribe
I need to find some better technical models for this social phenomenon to tattoo it into my brain in my langage more, because I think I kind of keep coming back to it in different forms
https://imgur.com/fveamMk
I'll try to save this into my brain more. But it's hard, lol.
My default assumption is mostly that social contexts are mostly epistemic contexts, not mostly attunement contexts
And its hard for me to not assume it by default, I feel like I have to do conscious effort for that, i have to periodically remind myself that to assume it less by default
But IMO this still doesn't excuse the people spreading the factually incorrect claims. They should just say directly that they distrust the politics around it, and that they know little about the actual technical details of the technology and science. Some other non-technical or technical people then read their claims as factual technical claims and then believe in factually incorrect claims.
I debunked so much bullshit about AI that my dad heard and almost believed because of those people, before I explained it to him.
Sorry for maybe being harsh, it just makes me angry, especially when my dad started believing some of those incorrect claims as technical facts, so it's also misinformation/disinformation.
>One thing I have found about AI skeptics is that they are often distrustful of the technical details being a smokescreen. They often equate it with the way that NFTs etc. were obfuscated in technicality.
Yeah, and people who equate all of AI/ML with cryptocurrency, while knowing so little about it on technical level, drive me completely nuts, don't get me even started lmao
Yeah there is some overlap. But like, those people also use computers generally, and nobody mocks all computers as a whole, because some grifters also use computers, while many normal people also use computers?
Many are IMO being uninformed/misinformed about everything happening in the AI field, doing incorrect overgeneralizing, and spreading misinformation about the whole AI field, which is harmful IMO.
I rant often about non-technical uninformed/misinformed media spreading misinformation/disinformation to the masses
Because of this you can't have a single rational technical discussion in certain heavily left wing circles. Their mind was completely eaten by all the polarizing politics about AI without zero rationalist thinking about the technical details.
For me it's the case mostly in some less technical spaces that I value for different reasons
Technical lefty spaces are a bit better, as some of those people can actually think more technically/rationally about it
But still not all of them
Some of the popular beliefs about how it works technically by non-technical people in those spaces are completely wrong
For example the claim that the models ONLY memorize, bit by bit, no abstracting, no generalizing, nothing else at all, completely denying all mechanistic interpretability research, completely denying fundamentals of ML with bias and variance trade off, certain empirical findings, theoretical research into DL generalization, etc. No amount of evidence can change their mind.
Even as a critic that says that current models do not generalize enough, like Francois Chollet and Ken Stanley for example describe it, this drives me nuts (which some super AI optimists, who are on the exact opposite extreme, hate when I point that out lol)
Then there's also DreamCoder, who's art was technically from a different form of generalization by neurosymbolic methods, when they made it draw shapes, which is super cool
Or PicBreeder using evolutionary methods
I came back to DeepDream few weeks ago, being nostalgic about older times with older AI art, when it used to be niche, before the whole culture war around it
Or live diffusion filter is cool
And they still have some weaker form of primarily in distribution combinatorial creativity IMO
Or another thing I dislike about some lefty internet spaces is that they sometimes allow for posting anything negative related to AI, but ban or create (often muted) containment for posting anything positive about AI
And some other spaces on the other hand want to censor all negative stuff instead and only allow for positive things about AI, which is an equally bad extreme
But I feel like censorship of positive things about AI is stronger overall
And that then results in echo chambers without nuance
in big part why I'm a lot more for open source recently is to give the power to the hands of everyone as much as possible
líbí se mi jak strašně duální technologie to je :smile:
vymýšlí to nový lepší algoritmy ale zároveň to maže produkční databáze
dělá to absolutně stupid chyby v matice a mezitím objevuje nový lepší výsledky v matice a poráží olympiády
atd. :D
průser je prostě větší míra false positives než člověk, a větší brittleness, a easy jailbreaky, apod.
existují alespoň nějaký metody jak to % minimalizovat, ale rush to deploy je větší, tudíž lidi serou na security
řešili jsme to teď tady, na tomhle kanálu jsem pomohl s interview ohledně LLM security https://www.youtube.com/watch?v=c_hmxRVDXBE
The rush to deploy AI without safeguards, echoing past tech blunders like SQL hacks."
ale zároveň záleží na systému, ale to "jen" zmenšuje určitý šance určitých průserů, tím že většina systémů co firmy, co závodí, deployují mají fakt blbou security, protože být na trhu první je priorita, a spousta security issues není vůbec solved
na druhou stranu je dost memů o tom jak junioři taky mažou prod databáze ale většinou jinačíma způsobama no and we have more power over how we design our machines, but limitations exist, but we'll see what will the future bring
https://www.reddit.com/r/MachineLearning/comments/t4axj8/n_using_analog_computers_for_ai/
neviděl jsem tohle konkrétní video ale třeba dost neuromorfickýho hardwaru je analogový
a teď ještě vzniká hardware co využívá stochastickou thermodynamiku ve statistický fyzice, takže nejsou deterministický, ale pravděpodobnostní, což je lepší pro např energy based algorithmy a jiný výpočty ze statistický fyziky
např tady je to hezky popsaný https://arxiv.org/abs/2108.09836 Probabilistic computing with p-bits
Autogenerating titles in email:
Tohle LLM pravděpodobně je, nějaký mini pravděpodobně, protože old school natural language processing metody pro tento usecase, v podstatě typ sumarizace, byly z velký části steamrollnutý transformerama
LLM je většinou transformer a ty jsou všude kde potřebuješ komplexnější natural language processing
I když použiješ Google search tak máš v pozadí jazykovej transformer. Nemyslím AI overview, myslím to sémantický hledání.
Ale je ještě možný že tam budou nějaký víc old school metody, jako možná nějaký podobný TextRanku, nebo co používají TextRank jako část. Záleží či to používá abstractive nebo extractive metody. Tady je to víc popsaný: https://en.wikipedia.org/wiki/Automatic_summarization
TextRank je celkem cool grafivý algoritmus, je to podobný PageRank, ale místo rankování websites rankuješ sentences/části sentences https://en.wikipedia.org/wiki/PageRank
Ale nebude to vanilla TextRank, protože to summary úplně netvoří dobrou sentence, to ti vyrankuje jaký části textu jsou asi nejdůležitější
Proto abstractive metody teď vedou, a ty jsou teď primárně LLMkový
Cool mapa metod https://aws.amazon.com/blogs/machine-learning/techniques-for-automatic-summarization-of-documents-using-language-models/
Poslední dobou sem tam koukám jak by se mohly zlepšit extractive metody, nebo zkombinovat s abstractive, pro minimalizaci halucinací. Možná nějak extrahovat nejdůležitý části, a generovat jenom syntaktickou omáčku kolem nich, aby to byly gramaticky validní věty se 100% existujícím extrahovaným textem, nebo summarizovat po kousíčkách jednotlivý nejdůležitější věty pomalu a s co největším double chekingem.
DeepMind AI math proofs are more readable than OpenAI's? https://x.com/burny_tech/status/1947380891178963238
homogenní bullet point format začínající s emojis je první giveaway že je něco AI generated
ale nedivil bych se kdyby teď už nějaký realtime personalizační finetuning samotných vah začali dělat u těch slabších modelů, protože vím že to plánovali, a některý jiný firmy tohle už dělaj
u těch silnějších modelů by se to spíš asi finančně nevyplatilo s tím co teď jde, ale možná už vyplatilo
https://arxiv.org/abs/2204.00188
tak ten titulek je absolutní nerd snipe pro mě, to musím prozkoumat
z abstraktu dostávám intelektuální orgasmus
zajímá mě jak definují novelty
"mean dissimilarity metric of the k-nearest neighbors"
yes, to je přesně definice novelty co jsem se pokoušel dát i do svý architektury
Někdy přemýšlím že přehnaná aggreableness určitých LLMs je trochu takovej overalignment problem
můžeš jenom používat hotový AI agent systémy
nebo je můžeš adaptovat nebo tvořit na high level úrovni pro tu firmu, což je software engineering a AI engineering nebo někdy research
nebo je můžeš adaptovat nebo tvořit víc low level úrovni, což je AI engineering nebo někdy research
já jen že spousta lidí a manažerů si teď představuje že integrace LLMs do firmy je jenom vymýšlení promptů, ale to je často daleko od toho aby to nějak stačilo pro určitý usecases, a realita je často mnohem komplexnější, kde je potřeba víc domain specific adaptace přes různý engineering, aby to fungovalo blíž k tomu jak si to manažeři představují, ale i tak to na dost věcí co si představují ještě nemá
jako člověk, co je kolem toho všeho jeden z nejvíc optimistických lidí, a zkouším to všechno hned jak něco výjde, a zkouším všechno možný kolem toho tvořit, jsem taky zažil, že mi někdo implikoval, že na něco dosavadní LLM samotný bez úprav stačí, ale realita byla jinde, někteří jsou někdy fakt moc naivní
hmm, a když tohle všechno vidím, tak se pak někdy míň divím, když někteří lidí zvolí pohled, co je přesně opačný extrém, ultra pesimistický pohled na celý tohle téma, kdy si myslí, že to neumí skoro nic, což je podobně mimo no
hmm, teď jsem přes noc nechal automaticky testovat jeden LLM coding systém pro benchmark, a spálil pro ty výsledky dost tisíc Kč
a manažeři teď beztak viděli jak teď novej model drtí coding competitions intelligence , a představují si, že je to to stejný jako dosavadní public modely co teď jdou použít bez alespoň nějaký adaptace (ne není), a to stejný jako veškerej software engineering (ne není), a dělají špatný závěry, podle nějakých jejich postů co jsem viděl a často neví o metodách v engineeringu co jsou často potřeba na tu adaptaci, jako různý context engineering, napojování na best practices firem a na data přes programming, AI agent architektury, někdy i finetuning, atd.
konvoluční sítě jsou používaný na vizuální klasifikaci už snad i 25 let, ale teď poslední roky je začínají nahrazovat vision transformery
transformery začínají nahrazovat až moc všechno xd