Can adaptive reward allocation prevent cumulative advantage from creating long-run sequencer concentration without materially reducing system efficiency?
Behavior module is open: participants may later be RL agents, LLM agents, or humans. This first study isolates the reward-allocation rule.
Each participant starts with a different cumulative reward and sequencing advantage.
Selected sequencer: —
Redistribution α: —
| Participant | Cumulative Reward | Advantage | Selection Probability | Wins |
|---|
Reward Gini and top reward share show whether the system becomes more concentrated over time.
Baseline: selected sequencer gets the full reward pool.
Adaptive: when cumulative reward concentration rises, a larger fraction of the current pool is redistributed equally to the other four participants.
Efficiency proxy: total transaction value minus a small redistribution friction.
Static browser-only implementation. No backend, database, API key, secret, paid service, or Hugging Face CPU/GPU runtime.