Most influential LLM papers and the ideas they introduced (post 2017)
https://fxtwitter.com/goyal__pramod/status/1921419933231038820?t=7APGsTpIj_z6DSGeaWIG2w&s=19
LLM agentic RL survey
https://arxiv.org/abs/2509.02547
Stanford CS329A Self-Improving AI Agents https://cs329a.stanford.edu/
o1 like open source model https://github.com/NovaSky-AI/SkyThought
large reasoning models blueprint https://arxiv.org/abs/2501.11223
https://www.latent.space/p/2025-papers
Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective https://arxiv.org/abs/2412.14135
https://github.com/henrythe9th/AI-Crash-Course
https://github.com/mlabonne/llm-course?tab=readme-ov-file
https://cloud.google.com/blog/products/ai-machine-learning/guide-to-jax-for-pytorch-developers
https://www.zenml.io/llmops-database
https://arxiv.org/abs/2501.07391 RAG SoTA
list of autonomous ai agents https://github.com/e2b-dev/awesome-ai-agents
https://arxiv.org/abs/2501.09223
https://github.com/PatWalters/resources_2025
they put r1 in a loop for 15minutes and it generated: "better than the optimized kernels developed by skilled engineers in some cases"
https://fxtwitter.com/abacaj/status/1889847093046702180?t=aRllQoTGbyzV_E0j4EMAfQ&s=19
https://www.promptingguide.ai/
https://x.com/arankomatsuzaki/status/1889522977185865833
Competed live at IOI 2024
o3 achieved gold
General-purpose o3 surpasses o1 w/ hand-crafted pipelines specialized for coding resultss
good dissecting of the limitations of Deep Research in practice here:
hallucinations are still big problem,
but it sometimes generates cool useful insights and puts together information nicely and shows you rabbithole paths for you to explore!
https://www.youtube.com/watch?v=Kqjd_RzhSSY
Benchmarks like this showing how R1 actually seems to generalize more poorly compared to OpenAI models makes me think there's more secret sauce behind o1/o3
https://x.com/gm8xx8/status/1888831941161451536?t=UZnBkyEIyiD3TZXUDluzyg&s=19
https://x.com/WenhuChen/status/1888691381054435690
Incredibly exciting work! https://t.co/aJEGtpF9eL. Cooperative, safe, driving can arise at scale from self-play training without human data, just as my group saw in Overcooked a few years ago. Caveat: in simulation. Bravo @EugeneVinitsky and coauthors!
https://x.com/edwardfhughes/status/1887492625453793471?t=bXqO9PkNQOO_rgcbhCjGFw&s=19
Types of memory in ai agents
https://x.com/Aurimas_Gr/status/1892196166973977034?t=4ioqtVKxIWb7gqSVgmvuWw&s=19
https://magazine.sebastianraschka.com/p/the-state-of-llm-reasoning-model-training
llm selfplay without human data
https://fxtwitter.com/AndrewZ45732491/status/1919920459748909288
https://arxiv.org/abs/2505.03335
New Sutton interview about his new paper about superiority of RL https://www.youtube.com/watch?v=dhfJfQ5NueM
https://storage.googleapis.com/deepmind-media/Era-of-Experience%20/The%20Era%20of%20Experience%20Paper.pdf
https://fxtwitter.com/stenichele/status/1924816238363942966
ARC-NCA shows how Neural Cellular Automata —incl. memory-rich EngramNCA—crack tasks from the ARC-AGI benchmark, hitting GPT-4.5-level accuracy at a fraction of the cost. 🌱🤖
OpenEvolve: An Open Source Implementation of Google DeepMind's AlphaEvolve https://huggingface.co/blog/codelion/openevolve
Bottlenecked Transformers: Periodic KV Cache Abstraction for Generalised Reasoning https://arxiv.org/abs/2505.16950
no verifiers LLM RL
https://arxiv.org/abs/2505.21493
https://arxiv.org/abs/2505.03335
https://arxiv.org/abs/2505.19590
https://x.com/natolambert/status/1927404027735617541
Introducing Continuous Thought Machines https://fxtwitter.com/SakanaAILabs/status/1921749814829871522
sakana.ai/ctm/
fully decentralized and open source https://www.primeintellect.ai/blog/intellect-2-release
Planetary-Scale Inference: Previewing our Peer-To-Peer Decentralized Inference Stack https://www.primeintellect.ai/blog/inference
Diffusion of Thoughts: Chain-of-Thought Reasoning in Diffusion Language Models https://arxiv.org/abs/2402.07754
https://fxtwitter.com/stenichele/status/1924816238363942966
ARC-NCA shows how Neural Cellular Automata —incl. memory-rich EngramNCA—crack tasks from the ARC-AGI benchmark, hitting GPT-4.5-level accuracy at a fraction of the cost.
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models https://arxiv.org/abs/2505.24864
https://www.verses.ai/blog/whitepaper-mastering-gameworld-10k-in-minutes-with-the-axiom-digital-brain
Atlas (A powerful Titan): a new architecture with long-term in-context memory https://fxtwitter.com/behrouz_ali/status/1928522388100010383
https://arxiv.org/abs/2505.23735
https://fxtwitter.com/LingYang_PU/status/1925385712670830753 https://arxiv.org/abs/2505.15809
We present MMaDA, first diffusion that unifies text reasoning, multimodal understanding, and image generation through Mixed Long-CoT, and unified RL - UniGRPO.
How DeepSeek Built The Current "Best" Math Prover AI DeepSeek-Prover-V2 https://www.youtube.com/watch?v=vhXDKif9mPU https://arxiv.org/abs/2504.21801
Decentralized AI training models https://x.com/0xPrismatic/status/1930243900712857786
https://fxtwitter.com/jyo_pari/status/1933350025284702697 https://arxiv.org/abs/2506.10943
https://fxtwitter.com/SakanaAILabs/status/1932972420522230214?t=hLzpZaQSzHbF2rm5p4Sg0w&s=19
arxiv.org/abs/2506.06105
The LLM's RL Revelation We Didn't See Coming https://www.youtube.com/watch?v=z3awgfU4yno
Papers discussed: Understanding R1-Zero-Like Training: A Critical Perspective https://arxiv.org/abs/2503.20783 Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model https://arxiv.org/abs/2504.13837 Reinforcement Learning Finetunes Small Subnetworks in Large Language Models https://arxiv.org/abs/2505.11711 Spurious Rewards: Rethinking Training Signals in RLVR https://arxiv.org/abs/2506.10947
https://arxiv.org/abs/2507.21046
https://x.com/jiqizhixin/status/1950458686843081200?t=gvkkqmAsgcaS45KB5QkJpQ&s=19
https://arxiv.org/abs/2506.21734
https://x.com/VictorTaelin/status/1950512015899840768?t=hApcuk48rawLlVuL48ATog&s=19
https://arxiv.org/abs/2507.18074?s=09
Open-Ended AI resources https://github.com/jennyzzt/awesome-open-ended
https://arxiv.org/abs/2509.03646
https://x.com/omarsar0/status/1965429326264107070
LLM agents resources https://github.com/kaushikb11/awesome-llm-agents
https://magazine.sebastianraschka.com/p/the-big-llm-architecture-comparison
https://youtu.be/rNlULI-zGcw?si=gpJPPK3B98v_TULl
deepseek parallel communicating residual streams
https://www.youtube.com/watch?v=jYn_1PpRzxI
https://arxiv.org/abs/2512.24880
https://github.com/khangich/machine-learning-interview
https://udlbook.github.io/udlbook/
https://fxtwitter.com/i/status/2014449503546605998
https://arxiv.org/abs/2601.14525