I play a lot of Arena limited and I'm bad at reading what's open mid-draft, so I trained a model to help. You paste in the pack you're looking at and it ranks the cards against what you've already picked.
I trained it on the public 17lands data for Marvel Super Heroes (MSH) Premier Draft: 68,633 drafts, 2.88M individual picks, the first month of the format. Every number below is from that set, but none of the code is tied to it, so you can point it at any 17lands draft export and retrain.
The full write-up with charts is in exploration.ipynb.
Accuracy here means "did the model take the same card the human took". The test set is the last 5 days of the format, and only picks from drafters with a 60%+ lifetime win rate, so it's being graded on whether it agrees with good players on a metagame it never trained on.
| model | top-1 |
|---|---|
| card ratings only, ignores your pool | 50.0% |
| transformer (cross-attention) | 67.2% |
| MLP | 67.6% |
| both averaged | 68.2% |
Top-3 is 94.9%. assist.py uses the average of the two.
The transformer being worse than the MLP surprised me. It was still improving when I ran out of patience, so it probably wins on a GPU, but it cost about ten times as much to train and never caught up. I kept it because the two models get different picks wrong, so averaging them beats either one.
Rank isn't skill. My first plan was to train on diamond and mythic drafters.
Then I actually checked, and median lifetime win rate is 0.58 for bronze against
0.54 for mythic. Rank measures how much you play, not how well, and season resets
drop good players back down to bronze. I ended up selecting on
user_game_win_rate_bucket >= 0.60 with at least 100 games behind it.
Knowing your pool is worth 18 points. A card rating list that ignores what you've drafted gets 50% on the same test set. So roughly half of drafting is knowing which cards are good, and the other half is knowing what your deck needs. The second half is the part a static tier list can't do.
68% is closer to the ceiling than it sounds. I assumed that number was bad until I measured how often two good drafters facing the same pack take the same card. On packs down to 2 cards they only agree 84% of the time, and nothing that looks at just the pack can beat 88%. At 3 cards the cap is already 70%. Most picks don't have one correct answer. The model gets 90% on those same small packs, which is above the pack-only cap, because it can see the pool and a card rating can't.
It knows when it's guessing. When it reports 95%+ confidence it's right 98% of the time, and when it reports 40% it's right 43%. So a low number isn't the model failing, it's telling me the pick is genuinely close and I should think about it.
Every pick is a ranking problem over the 14 or fewer cards actually in front of you. The model scores each card in the pack, softmaxes over that pack only, and trains with cross-entropy against the card the human took:
scores = model(pool, pack, pack_no, pick_no) # one score per card in the pack
loss = F.cross_entropy(scores, chosen_index)
Keeping the softmax on the pack is the part everything else depends on. If you score all 339 cards and mask afterwards, the model burns capacity learning which cards weren't in the pack, which is not something it ever needs to know.
PickMLP treats the pool as a bag of counts. PickCross lets the pack cards
attend to each other and then to the pool, so a card gets scored against the
specific cards you already own.
Drafts are stored as sequences instead of pools:
picks[D, 42] the card taken at each step
packs[D, 42, 14] the cards offered at each step, -1 padded
so the pool at step t is just picks[d, :t], built at batch time. I checked
that against the dataset's own pool_* columns before relying on it. It turns
4.3 GB of CSV into 34 MB.
pip install -r requirements.txt
set MTG_DRAFT_CSV=C:/path/to/draft_data_public.MSH.PremierDraft.csv.gz
python prep_data.py # ~4 min, one time
python train.py --arch mlp --epochs 8 --out mlp_all.pt
python train.py --arch mlp --good --init mlp_all.pt --epochs 2 --lr 4e-4 --out mlp_good.pt
python train.py --arch cross --epochs 2 --bs 2048 --limit 1200000 --out tf_all.pt
python train.py --arch cross --good --init tf_all.pt --epochs 2 --bs 2048 --lr 3e-4 --out tf_good2.pt
python assist.py
Training is pretrain on everyone, then fine-tune on the good-drafter cohort. The bulk data teaches it the cards, the fine-tune teaches it whose judgement to copy.
In assist.py, paste the pack one card per line or ;-separated, then a blank
line. Names are fuzzy-matched and the pool updates as you pick.
Download any draft export from 17lands.com/public_datasets and run the same commands. The set code comes from the filename, and the card list, number of packs, picks per pack and pack size are all read off the file, so nothing needs editing:
set MTG_DRAFT_CSV=C:/path/to/draft_data_public.FIN.PremierDraft.csv.gz
python prep_data.py # writes data/FIN/
python train.py --set FIN --arch mlp --epochs 8 --out fin_all.pt
python assist.py --set FIN --model fin_all.pt --arch mlp
Prepared sets live in data/<SET>/, and if you have more than one you pass
--set or set MTG_SET. Checkpoints are tied to the set they were trained on;
loading one against the wrong set fails with a message rather than silently
producing nonsense.
Other environment variables, all optional: MTG_SET, MTG_DATA_DIR,
MTG_GOOD_WR and MTG_GOOD_GAMES (what counts as a good drafter, default 0.60
over 100+ games), MTG_TEST_DAYS (default 5).
| file | what |
|---|---|
prep_data.py | 17lands CSV to packed arrays |
common.py | loading, splits, batch construction |
models.py | the models |
train.py | training |
eval.py | score a checkpoint |
assist.py | the draft-time tool |
exploration.ipynb | data, findings, results |
The model copies good drafters, it doesn't optimise for winning, because it never
sees a game result during training. The 17lands game data file has per-game
outcomes and would let me train against win rate directly, which is the next thing
I want to do. After that, reading Player.log so I don't have to type the pack in
every pick.
Jupyter Notebook
88.4%
Python
11.6%
I play a lot of Arena limited and I'm bad at reading what's open mid-draft, so I trained a model to help. You paste in the pack you're looking at and it ranks the cards against what you've already picked.
I trained it on the public 17lands data for Marvel Super Heroes (MSH) Premier Draft: 68,633 drafts, 2.88M individual picks, the first month of the format. Every number below is from that set, but none of the code is tied to it, so you can point it at any 17lands draft export and retrain.
The full write-up with charts is in exploration.ipynb.
Accuracy here means "did the model take the same card the human took". The test set is the last 5 days of the format, and only picks from drafters with a 60%+ lifetime win rate, so it's being graded on whether it agrees with good players on a metagame it never trained on.
| model | top-1 |
|---|---|
| card ratings only, ignores your pool | 50.0% |
| transformer (cross-attention) | 67.2% |
| MLP | 67.6% |
| both averaged | 68.2% |
Top-3 is 94.9%. assist.py uses the average of the two.
The transformer being worse than the MLP surprised me. It was still improving when I ran out of patience, so it probably wins on a GPU, but it cost about ten times as much to train and never caught up. I kept it because the two models get different picks wrong, so averaging them beats either one.
Rank isn't skill. My first plan was to train on diamond and mythic drafters.
Then I actually checked, and median lifetime win rate is 0.58 for bronze against
0.54 for mythic. Rank measures how much you play, not how well, and season resets
drop good players back down to bronze. I ended up selecting on
user_game_win_rate_bucket >= 0.60 with at least 100 games behind it.
Knowing your pool is worth 18 points. A card rating list that ignores what you've drafted gets 50% on the same test set. So roughly half of drafting is knowing which cards are good, and the other half is knowing what your deck needs. The second half is the part a static tier list can't do.
68% is closer to the ceiling than it sounds. I assumed that number was bad until I measured how often two good drafters facing the same pack take the same card. On packs down to 2 cards they only agree 84% of the time, and nothing that looks at just the pack can beat 88%. At 3 cards the cap is already 70%. Most picks don't have one correct answer. The model gets 90% on those same small packs, which is above the pack-only cap, because it can see the pool and a card rating can't.
It knows when it's guessing. When it reports 95%+ confidence it's right 98% of the time, and when it reports 40% it's right 43%. So a low number isn't the model failing, it's telling me the pick is genuinely close and I should think about it.
Every pick is a ranking problem over the 14 or fewer cards actually in front of you. The model scores each card in the pack, softmaxes over that pack only, and trains with cross-entropy against the card the human took:
scores = model(pool, pack, pack_no, pick_no) # one score per card in the pack
loss = F.cross_entropy(scores, chosen_index)
Keeping the softmax on the pack is the part everything else depends on. If you score all 339 cards and mask afterwards, the model burns capacity learning which cards weren't in the pack, which is not something it ever needs to know.
PickMLP treats the pool as a bag of counts. PickCross lets the pack cards
attend to each other and then to the pool, so a card gets scored against the
specific cards you already own.
Drafts are stored as sequences instead of pools:
picks[D, 42] the card taken at each step
packs[D, 42, 14] the cards offered at each step, -1 padded
so the pool at step t is just picks[d, :t], built at batch time. I checked
that against the dataset's own pool_* columns before relying on it. It turns
4.3 GB of CSV into 34 MB.
pip install -r requirements.txt
set MTG_DRAFT_CSV=C:/path/to/draft_data_public.MSH.PremierDraft.csv.gz
python prep_data.py # ~4 min, one time
python train.py --arch mlp --epochs 8 --out mlp_all.pt
python train.py --arch mlp --good --init mlp_all.pt --epochs 2 --lr 4e-4 --out mlp_good.pt
python train.py --arch cross --epochs 2 --bs 2048 --limit 1200000 --out tf_all.pt
python train.py --arch cross --good --init tf_all.pt --epochs 2 --bs 2048 --lr 3e-4 --out tf_good2.pt
python assist.py
Training is pretrain on everyone, then fine-tune on the good-drafter cohort. The bulk data teaches it the cards, the fine-tune teaches it whose judgement to copy.
In assist.py, paste the pack one card per line or ;-separated, then a blank
line. Names are fuzzy-matched and the pool updates as you pick.
Download any draft export from 17lands.com/public_datasets and run the same commands. The set code comes from the filename, and the card list, number of packs, picks per pack and pack size are all read off the file, so nothing needs editing:
set MTG_DRAFT_CSV=C:/path/to/draft_data_public.FIN.PremierDraft.csv.gz
python prep_data.py # writes data/FIN/
python train.py --set FIN --arch mlp --epochs 8 --out fin_all.pt
python assist.py --set FIN --model fin_all.pt --arch mlp
Prepared sets live in data/<SET>/, and if you have more than one you pass
--set or set MTG_SET. Checkpoints are tied to the set they were trained on;
loading one against the wrong set fails with a message rather than silently
producing nonsense.
Other environment variables, all optional: MTG_SET, MTG_DATA_DIR,
MTG_GOOD_WR and MTG_GOOD_GAMES (what counts as a good drafter, default 0.60
over 100+ games), MTG_TEST_DAYS (default 5).
| file | what |
|---|---|
prep_data.py | 17lands CSV to packed arrays |
common.py | loading, splits, batch construction |
models.py | the models |
train.py | training |
eval.py | score a checkpoint |
assist.py | the draft-time tool |
exploration.ipynb | data, findings, results |
The model copies good drafters, it doesn't optimise for winning, because it never
sees a game result during training. The 17lands game data file has per-game
outcomes and would let me train against win rate directly, which is the next thing
I want to do. After that, reading Player.log so I don't have to type the pack in
every pick.
Jupyter Notebook
88.4%
Python
11.6%