If I had the budget/compute power and time to research anything at the moment, this is what I would try researching. It’ll make sense in a bit, I swear.
The Problem Training robotic policies to perform complex tasks is incredibly …
tldr; A paper caught my eye, proposing a way to treat individual agent calls in multi-turn agent training as individual steps. The key contribution - an easy, intuitive method of assigning credit to each step, for quicker training.
[Paper] …
Recently, I gave a talk on several of DeepSeek’s innovations, which were as extensive as they were complicated. The particular clever discovery that best captured my imagination was their development of GRM and SPCT.
Most AI models — …
tldr I recently gave a talk at the SDx paper club which I run on the paper DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. I wanted to take a moment to blog out some of the talk, specifically how their …
tldr I created a custom robotic arm environment where the arm is rewarded if it sorts shapes a certain way, and then punished if it knocks them off a table. I used OpenAI’s PPO method to be somewhat successful. If you want to skip to …