Most influential LLM papers and the ideas they introduced (post 2017) https://fxtwitter.com/goyal__pramod/status/1921419933231038820?t=7APGsTpIj_z6DSGeaWIG2w&s=19 LLM agentic RL survey https://arxiv.org/abs/2509.02547 Stanford CS329A Self-Improving AI Agents https://cs329a.stanford.edu/ o1 like open source model https://github.com/NovaSky-AI/SkyThought large reasoning models blueprint https://arxiv.org/abs/2501.11223 https://www.latent.space/p/2025-papers Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective https://arxiv.org/abs/2412.14135 https://github.com/henrythe9th/AI-Crash-Course https://github.com/mlabonne/llm-course?tab=readme-ov-file https://cloud.google.com/blog/products/ai-machine-learning/guide-to-jax-for-pytorch-developers https://www.zenml.io/llmops-database https://arxiv.org/abs/2501.07391 RAG SoTA list of autonomous ai agents https://github.com/e2b-dev/awesome-ai-agents https://arxiv.org/abs/2501.09223 https://github.com/PatWalters/resources_2025 they put r1 in a loop for 15minutes and it generated: "better than the optimized kernels developed by skilled engineers in some cases" https://fxtwitter.com/abacaj/status/1889847093046702180?t=aRllQoTGbyzV_E0j4EMAfQ&s=19 https://www.promptingguide.ai/ https://x.com/arankomatsuzaki/status/1889522977185865833 Competed live at IOI 2024 o3 achieved gold General-purpose o3 surpasses o1 w/ hand-crafted pipelines specialized for coding resultss good dissecting of the limitations of Deep Research in practice here: hallucinations are still big problem, but it sometimes generates cool useful insights and puts together information nicely and shows you rabbithole paths for you to explore! https://www.youtube.com/watch?v=Kqjd_RzhSSY Benchmarks like this showing how R1 actually seems to generalize more poorly compared to OpenAI models makes me think there's more secret sauce behind o1/o3 https://x.com/gm8xx8/status/1888831941161451536?t=UZnBkyEIyiD3TZXUDluzyg&s=19 https://x.com/WenhuChen/status/1888691381054435690 Incredibly exciting work! https://t.co/aJEGtpF9eL. Cooperative, safe, driving can arise at scale from self-play training without human data, just as my group saw in Overcooked a few years ago. Caveat: in simulation. Bravo @EugeneVinitsky and coauthors! https://x.com/edwardfhughes/status/1887492625453793471?t=bXqO9PkNQOO_rgcbhCjGFw&s=19 Types of memory in ai agents https://x.com/Aurimas_Gr/status/1892196166973977034?t=4ioqtVKxIWb7gqSVgmvuWw&s=19 https://magazine.sebastianraschka.com/p/the-state-of-llm-reasoning-model-training llm selfplay without human data https://fxtwitter.com/AndrewZ45732491/status/1919920459748909288 https://arxiv.org/abs/2505.03335 New Sutton interview about his new paper about superiority of RL https://www.youtube.com/watch?v=dhfJfQ5NueM https://storage.googleapis.com/deepmind-media/Era-of-Experience%20/The%20Era%20of%20Experience%20Paper.pdf https://fxtwitter.com/stenichele/status/1924816238363942966 ARC-NCA shows how Neural Cellular Automata —incl. memory-rich EngramNCA—crack tasks from the ARC-AGI benchmark, hitting GPT-4.5-level accuracy at a fraction of the cost. 🌱🤖 OpenEvolve: An Open Source Implementation of Google DeepMind's AlphaEvolve https://huggingface.co/blog/codelion/openevolve Bottlenecked Transformers: Periodic KV Cache Abstraction for Generalised Reasoning https://arxiv.org/abs/2505.16950 no verifiers LLM RL https://arxiv.org/abs/2505.21493 https://arxiv.org/abs/2505.03335 https://arxiv.org/abs/2505.19590 https://x.com/natolambert/status/1927404027735617541 Introducing Continuous Thought Machines https://fxtwitter.com/SakanaAILabs/status/1921749814829871522 sakana.ai/ctm/ fully decentralized and open source https://www.primeintellect.ai/blog/intellect-2-release Planetary-Scale Inference: Previewing our Peer-To-Peer Decentralized Inference Stack https://www.primeintellect.ai/blog/inference Diffusion of Thoughts: Chain-of-Thought Reasoning in Diffusion Language Models https://arxiv.org/abs/2402.07754 https://fxtwitter.com/stenichele/status/1924816238363942966 ARC-NCA shows how Neural Cellular Automata —incl. memory-rich EngramNCA—crack tasks from the ARC-AGI benchmark, hitting GPT-4.5-level accuracy at a fraction of the cost. ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models https://arxiv.org/abs/2505.24864 https://www.verses.ai/blog/whitepaper-mastering-gameworld-10k-in-minutes-with-the-axiom-digital-brain Atlas (A powerful Titan): a new architecture with long-term in-context memory https://fxtwitter.com/behrouz_ali/status/1928522388100010383 https://arxiv.org/abs/2505.23735 https://fxtwitter.com/LingYang_PU/status/1925385712670830753 https://arxiv.org/abs/2505.15809 We present MMaDA, first diffusion that unifies text reasoning, multimodal understanding, and image generation through Mixed Long-CoT, and unified RL - UniGRPO. How DeepSeek Built The Current "Best" Math Prover AI DeepSeek-Prover-V2 https://www.youtube.com/watch?v=vhXDKif9mPU https://arxiv.org/abs/2504.21801 Decentralized AI training models https://x.com/0xPrismatic/status/1930243900712857786 https://fxtwitter.com/jyo_pari/status/1933350025284702697 https://arxiv.org/abs/2506.10943 https://fxtwitter.com/SakanaAILabs/status/1932972420522230214?t=hLzpZaQSzHbF2rm5p4Sg0w&s=19 arxiv.org/abs/2506.06105 The LLM's RL Revelation We Didn't See Coming https://www.youtube.com/watch?v=z3awgfU4yno Papers discussed: Understanding R1-Zero-Like Training: A Critical Perspective https://arxiv.org/abs/2503.20783 Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model https://arxiv.org/abs/2504.13837 Reinforcement Learning Finetunes Small Subnetworks in Large Language Models https://arxiv.org/abs/2505.11711 Spurious Rewards: Rethinking Training Signals in RLVR https://arxiv.org/abs/2506.10947 https://arxiv.org/abs/2507.21046 https://x.com/jiqizhixin/status/1950458686843081200?t=gvkkqmAsgcaS45KB5QkJpQ&s=19 https://arxiv.org/abs/2506.21734 https://x.com/VictorTaelin/status/1950512015899840768?t=hApcuk48rawLlVuL48ATog&s=19 https://arxiv.org/abs/2507.18074?s=09 Open-Ended AI resources https://github.com/jennyzzt/awesome-open-ended https://arxiv.org/abs/2509.03646 https://x.com/omarsar0/status/1965429326264107070 LLM agents resources https://github.com/kaushikb11/awesome-llm-agents https://magazine.sebastianraschka.com/p/the-big-llm-architecture-comparison https://youtu.be/rNlULI-zGcw?si=gpJPPK3B98v_TULl deepseek parallel communicating residual streams https://www.youtube.com/watch?v=jYn_1PpRzxI https://arxiv.org/abs/2512.24880 https://github.com/khangich/machine-learning-interview https://udlbook.github.io/udlbook/ https://fxtwitter.com/i/status/2014449503546605998 https://arxiv.org/abs/2601.14525