Multi-Agent RL Market-Making Engine (MARL)

View Repository Source

Trained dual-swarms of AI agents to compete and cooperate in a simulated stock exchange, written from scratch with PyTorch and formal mathematical specifications.

What happens if you drop a handful of AI agents into a simulated stock exchange — some trying to profit at each other’s expense, some trying to work together — and let them practice tens of thousands of times? This project builds a multi-agent reinforcement learning (MARL) engine with a centralized-training / decentralized-execution paradigm over a custom limit-order-book simulator and an offline tape replayer.

Built From Scratch

I wrote MADDPG and MAPPO from scratch with per-agent actor–critic networks, consolidated into a unified dual-swarm engine (adv_coop). This runs atop a shared abstraction layer (marl_core) providing a swarm registry, a divergence monitor, dataset specs with walk-forward splits, and a common environment and transaction-cost contract.

Formal Math Specifications

A formal math spec defines the state, observation, action, transition, reward, and learning rule as one consistent object. Training and live paths execute against that exact same definition. The framework handles five agent configurations (adversarial, cooperative, hierarchical, bull/bear, and asset-level), each with its own strict mathematical specification.

The Robots Didn’t Always Behave

  • The Adversary: The adversarial trader’s reward declined during its own training phase, losing 99.2% of episodes. Probing the trained actors revealed a zero-cost adversary parking its entire L1 budget on one channel — a corner solution offering no minimax gradient.
  • The Cooperative Swarm: The cooperative swarm flatlined at −71 bps/episode, which the fee structure correctly attributed to churn under a zeroed turnover penalty.
  • The Uninterpretable Win: One hierarchical manager improved by 56% across 35,000 episodes, but worker rewards ran 2–4 orders of magnitude outside bounded specs due to an unfloored normalizer and uncapped participation ratio. The fix was to mandate a held-out evaluation harness rather than continuing to train blindly.

Transparent Reporting

A negative result you understand is worth more than a positive one you can’t explain. I published a repo-level gap analysis explicitly comparing the conceptual objectives against the built reality — simulator fidelity, data tier, and action-space discretization — rather than describing the initial intent as the final implementation.