Sequencer Reward Allocation Lab

Can adaptive reward allocation prevent cumulative advantage from creating long-run sequencer concentration without materially reducing system efficiency?

Behavior module is open: participants may later be RL agents, LLM agents, or humans. This first study isolates the reward-allocation rule.

1. Five Participants

Each participant starts with a different cumulative reward and sequencing advantage.

2. Mechanism

Selection: Pᵢ(t)=Aᵢ(t)/ΣAⱼ(t)
Reward pool: R(t)=Transaction Fees+MEV
Advantage update: Aᵢ(t+1)=Aᵢ(t)+β·rewardᵢ(t)
Adaptive share: α(t)=min(αmax, κ·Gini(cumulative rewards))

3. Run

Current Round

Round0
Fees—
MEV—
Reward Pool—

Selected sequencer: —

Redistribution α: —

Observed Outputs

Reward Gini—
Top Reward Share—
Sequencer HHI—
Efficiency—

4. Participant State

ParticipantCumulative RewardAdvantageSelection ProbabilityWins

5. Concentration Over Time

Reward Gini and top reward share show whether the system becomes more concentrated over time.

6. Interpretation

Baseline: selected sequencer gets the full reward pool.

Adaptive: when cumulative reward concentration rises, a larger fraction of the current pool is redistributed equally to the other four participants.

Efficiency proxy: total transaction value minus a small redistribution friction.

Static browser-only implementation. No backend, database, API key, secret, paid service, or Hugging Face CPU/GPU runtime.