NONick One-Dropinnickonedrop.hashnode.dev·2d ago · 3 min readMulti-agent cooperation through in-context co-player inference (2026)This paper takes a decentralized multi-agent reinforcement learning approach. What does that mean here? For context, let’s look at another DeepMind paper with overlapping authors: Multi-agent cooperat00
DSDipen Shahintechecho.hashnode.dev·5d ago · 12 min readSimulating Natural Selection: Why Unconditional Kindness Dies, and How Clusters of Trust Conquer the WorldOpening question In Episodes 1 through 3, our computational tournaments operated under a fixed social contract: The population was static. The experimenter handpicked the participants, and every arche00
DSDipen Shahintechecho.hashnode.dev·Sep 27 · 12 min readSimulating Noise and Forgiveness: Why Pure Reciprocity Collapses When Communication BreaksOpening question In Episodes 1 and 2, we proved that reciprocal trust can survive against raw exploitation and build an interconnected society. Under round-robin tournament conditions, Robert Axelrod'00
DSDipen Shahintechecho.hashnode.dev·Sep 20 · 5 min readBuilding a Multi-Agent Tournament Engine: Simulating Emergent Cooperation in PythonIn Episode 1: Building the Laboratory, we constructed a deterministic, isolated foundation for repeated two-agent Prisoner's Dilemma interactions. We verified that when interactions repeat, reciprocal00
DSDipen Shahintechecho.hashnode.dev·Aug 25 · 7 min readI Built a Simulation to Explore Why Selfish Agents CooperateBuilding the domain foundation for an open-source simulation of cooperation, defection, and emergent behaviour. What happens when self-interested agents interact with each other again and again? Do t10
SSangharshainnoob6t5.hashnode.dev·Aug 24 · 3 min readDeceptive AI Doesn’t Break the Rules. It Optimizes Around Them: with game theoryEveryone assumes deception in AI will look obvious — a system glitch, a sudden compute spike, or a glaring anomaly. That assumption is fundamentally wrong. The most effective deception doesn’t violate00
OOmnithiuminomnithium.hashnode.dev·May 31 · 9 min readMulti-Agent Negotiation Protocols: How AI Agents Should Bargain for ResourcesStatic quotas are the death of agentic autonomy. If you're still using Kubernetes-style resource limits to manage your AI swarms, you're likely leaving 30% to 40% of your compute capacity on the table00
MMacaulay001inforgottentheorieshashnodedev.hashnode.dev·May 11 · 7 min readI reran Axelrod's 1981 prisoner's dilemma tournament. Tit-for-tat still earns its reputation.Part 49 of Forgotten Theories, a series re-testing old scientific claims with modern tools. Find the rest under the #forgotten-theories tag. In 1980, the political scientist Robert Axelrod sent letters to game theorists, economists, psychologists, a...00
BSBerkan Seseninsesenai.hashnode.dev·May 11 · 16 min readQ-Learning for Games: Teaching an Agent Tic-Tac-Toe Through Self-PlayTic-tac-toe is a solved game. Any competent adult can force a draw every time. But can an agent figure that out with zero human knowledge? Give two agents a blank board, a few simple rules about wins 00
Ssuboptimal.aiinsuboptimal-ai.hashnode.dev·Mar 15 · 2 min readThe Safety Eval That Must Work Is the One That Can'tThe AI safety community has built an evaluation apparatus specifically to detect dangerous capabilities before deployment. The premise is clear: test the model before you ship it. If it can help synthesize pathogens or scheme against its operators, c...00