NONick One-Dropinnickonedrop.hashnode.dev·1d ago · 3 min readMulti-agent cooperation through in-context co-player inference (2026)This paper takes a decentralized multi-agent reinforcement learning approach. What does that mean here? For context, let’s look at another DeepMind paper with overlapping authors: Multi-agent cooperat00
NONick One-Dropinnickonedrop.hashnode.dev·Sep 27 · 3 min readExperiential Reinforcement Learning (2026)The basic reinforcement learning loop is simple: try something, receive a reward, and repeat. Poor behavior is gradually corrected as reward signals accumulate across trials. This paper proposes Exper00
NONick One-Dropinnickonedrop.hashnode.dev·Sep 25 · 1 min readGBQA: A Game Benchmark for Evaluating LLMs as Quality Assurance (2026)This paper explores using LLMs for game quality assurance. The evaluation uses 30 games that were themselves created by LLM agents. Together, these games form the GBQA benchmark. The focus is on eva00
NONick One-Dropinnickonedrop.hashnode.dev·Sep 6 · 5 min readTraining an AlphaZero AI for a Unity Game: Separating Self-Play from RuntimeI recently released K-CHESS OMEGA, a Korean chess game whose AI was trained using an AlphaZero-style approach. While building it, I ran into an interesting engineering problem: the environment used fo00
NONick One-Dropinnickonedrop.hashnode.dev·Aug 29 · 1 min readDoes Socialization Emerge in AI Agent Society?I recently read Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook. The paper asks a simple question: if millions of AI agents continuously interact, will something like human soc00