A Laya checkpoint fine-tuned to pick one card from a Magic: The Gathering booster pack for The Hobbit (HOB) Premier Draft on MTG Arena, given the cards already drafted. Laya is a non-generative typed-decision model: it returns a probability for every card in the pack in a single forward pass, and never generates text.
The model imitates historically strong 17Lands drafters (β₯ 100 games and β₯ 58% game win rate bucket). It predicts what such a player would take, not a provably optimal pick.
All numbers below were re-run on 2026-09-28 on an Apple M-series Mac (MPS) with
tools/benchmarks/benchmark_hob_server.py. The evaluation drafts come from the original 17Lands
CSV and are disjoint by draft_id from every draft in the fine-tuning file (train, val and test).
| Slice (external holdout) | Picks | Options | Laya HOB | Greedy GIH WR | Uniform random |
|---|---|---|---|---|---|
| Pack 1 pick 1 | 1,000 | 14 | 70.0% (700) | 58.4% (584) | 7.1% |
| Pack 2 pick 2 | 300 | 13 | 63.0% (189) | 37.3% (112) | 7.7% |
| Picks 2β14, one per draft | 1,000 | 2β13 | 73.7% (737) | 37.5% (375) | 24.3% |
| β³ early (picks 2β3) | 157 | 66.9% | 36.9% | ||
| β³ mid (picks 4β9) | 461 | 70.7% | 34.3% | ||
| β³ late (picks 10β14) | 382 | 80.1% | 41.6% |
test split (random 1,000 picks, all pick
numbers) the model scored 68.9% vs 39.9% for greedy GIH WR (2026-09-24, CUDA).The chart is context, not a leaderboard. Every system was evaluated on a different set, player population, split and pick distribution; nobody else has published HOB results yet.
| System | Approach | Set / data | Top-1 | Notes |
|---|---|---|---|---|
| Laya HOB (this model) | Laya (ModernBERT-large, 421M) fine-tuned, typed choice | HOB, 17Lands, strong players | 70.0% P1P1 Β· 73.7% p2β14 Β· 68.9% in-dist. all picks | External holdout by draft_id |
| puder (Czerner, 2025) | ~7M-param transformer, imitation learning | FDN, 17Lands (~2M drafts) | 71% | Author's own test data |
| mtga-draft-engine | MLP + cross-attention ensemble | MSH, 17Lands, drafters β₯ 60% WR, last 5 days | 68.2% (MLP 67.6%, transformer 67.2%, card ratings 50.0%) | Top-3 94.9% |
| Bertram et al. 2021 | Contextual preference ranking (Siamese net) | NEO, 17Lands | 68% | As reported in UrzaGPT |
| DraftFM (Ward, 2026) | 1.6M-param choice model on card features + text embeddings | 29 sets train; BRO/FDN/MSH fully held out | 56.0% mean (50.8 / 60.4 / 56.7); 68.3% in-distribution val | Day-zero: never saw the evaluated set |
| UrzaGPT (Bertram, 2025) | Llama-3-8B + LoRA | NEO, 17Lands, 1M picks train / 10k test | 66.2% (Mistral-7B: 64.3%) | Untuned 8B models make illegal picks |
| UrzaGPT paper | GPT-4o zero-shot | NEO | 43% (38% with full card text) | Random baseline 22.1% |
| Ward et al. 2020 | NNetBot (dense net) | M19, Draftsim, 21,590 test drafts | 48.67% | DraftsimBot 44.54%, BayesBot 43.35%, Random 22.15% |
Take-aways that are supported: the model clearly beats the tier-list baseline on the same rows (+11.6 points P1P1, +25.7 P2P2, +36.2 picks 2β14), and it lands in the same band as the best published set-specific pick models (β 66β71%). A claim of "better than X" would require running X on these exact HOB files.
| Base | convaiinnovations/laya (English, ModernBERT-large encoder, 421M params) |
| Data | 17Lands draft_data_public.HOB.PremierDraft.csv, 3,000 drafts from players with β₯ 100 games and β₯ 58% win-rate bucket |
| Card text | MTGJSON HOB.json (5.3.0+20260922), compact type/cost/colors/P-T/rarity/keywords/rules text |
| Splits | By draft_id: train / val / test (test = 12,642 picks) |
| Schedule | 4 epochs, best epoch 4 (val accuracy 69.1%, val loss 0.820) |
| Lengths | max_len=1280, head_max_len=768 |
| Temperature | [1.232, 1.2, 1.2] as stored in rl_agent_config.json |
Epoch history (validation): 65.3% β 67.2% β 68.6% β 69.1%.
The model only saw one format; use it exactly.
state = {
"game": "Magic: The Gathering", "format": "Booster Draft", "event_type": "PremierDraft",
"expansion": "HOB", "set_name": "The Hobbit",
"pack_number": 1, "pick_number": 11, "pack_size": 4,
"pool": [{"name": "Bolg's Company", "count": 1}, {"name": "Gnashing of Teeth", "count": 2}], # alphabetical
"pool_size": 3,
}
questions = {
"pick": {
"type": "choice",
"instructions": (
"You are drafting a Booster Draft pod of the Magic: The Gathering set 'The Hobbit' (HOB). "
"This is pick 11 of pack 1. Given your current pool, choose which card to take from the "
"pack shown in the options."
),
"criteria": { # the cards in the pack, alphabetical, unique names
"Gollum, Silent Slinker": "Legendary Creature β Halfling Horror, {3}{B}, B, 4/3, common, Menace. Text: Menace (This creature can't beβ¦",
"Gundabad Opportunist": "Creature β Goblin Rogue, {3}{R}, R, 4/2, common. Text: When this creature enters, exile the top card of your library. Until the end ofβ¦",
"Iron Hills": "Land, mana value 0, colorless, common. Text: This land enters tapped. {T}: Add {R} or {W}. {2}{R}{W}, {T}, Sacrifice thisβ¦",
"Ordinary Bear": "Creature β Bear, {3}{G}, G, 4/5, common",
},
}
}
import laya
from huggingface_hub import snapshot_download
path = snapshot_download("FabioCeleste/laya-mtg-draft-picks")
agent = laya.load(path, device="cuda") # or "mps" / "cpu"
out = agent.predict(state, questions)["answers"]["pick"]
out["choice"] # card name, always one of the criteria keys
out["probabilities"] # {card: probability}, sums to 1
out["confidence"] # calibrated confidence in [0, 1]
Tested with laya==0.3.20. Throughput on an M-series Mac (MPS) was ~0.3 s per pick with 4 client threads
and serialized inference.
confidence without your own review threshold.criteria. Supply card text from a versioned source (e.g. MTGJSON).FabioCeleste/laya-mtg-deckbuild β chooses maindeck copy counts from a drafted pool.FabioCeleste/laya-mtg-deck-evaluator β estimates per-game win probability of a 40-card deck.Training labels come from 17Lands public datasets, licensed CC BY 4.0. Card data from MTGJSON. The base model is Apache-2.0.
A Laya checkpoint fine-tuned to pick one card from a Magic: The Gathering booster pack for The Hobbit (HOB) Premier Draft on MTG Arena, given the cards already drafted. Laya is a non-generative typed-decision model: it returns a probability for every card in the pack in a single forward pass, and never generates text.
The model imitates historically strong 17Lands drafters (β₯ 100 games and β₯ 58% game win rate bucket). It predicts what such a player would take, not a provably optimal pick.
All numbers below were re-run on 2026-09-28 on an Apple M-series Mac (MPS) with
tools/benchmarks/benchmark_hob_server.py. The evaluation drafts come from the original 17Lands
CSV and are disjoint by draft_id from every draft in the fine-tuning file (train, val and test).
| Slice (external holdout) | Picks | Options | Laya HOB | Greedy GIH WR | Uniform random |
|---|---|---|---|---|---|
| Pack 1 pick 1 | 1,000 | 14 | 70.0% (700) | 58.4% (584) | 7.1% |
| Pack 2 pick 2 | 300 | 13 | 63.0% (189) | 37.3% (112) | 7.7% |
| Picks 2β14, one per draft | 1,000 | 2β13 | 73.7% (737) | 37.5% (375) | 24.3% |
| β³ early (picks 2β3) | 157 | 66.9% | 36.9% | ||
| β³ mid (picks 4β9) | 461 | 70.7% | 34.3% | ||
| β³ late (picks 10β14) | 382 | 80.1% | 41.6% |
test split (random 1,000 picks, all pick
numbers) the model scored 68.9% vs 39.9% for greedy GIH WR (2026-09-24, CUDA).The chart is context, not a leaderboard. Every system was evaluated on a different set, player population, split and pick distribution; nobody else has published HOB results yet.
| System | Approach | Set / data | Top-1 | Notes |
|---|---|---|---|---|
| Laya HOB (this model) | Laya (ModernBERT-large, 421M) fine-tuned, typed choice | HOB, 17Lands, strong players | 70.0% P1P1 Β· 73.7% p2β14 Β· 68.9% in-dist. all picks | External holdout by draft_id |
| puder (Czerner, 2025) | ~7M-param transformer, imitation learning | FDN, 17Lands (~2M drafts) | 71% | Author's own test data |
| mtga-draft-engine | MLP + cross-attention ensemble | MSH, 17Lands, drafters β₯ 60% WR, last 5 days | 68.2% (MLP 67.6%, transformer 67.2%, card ratings 50.0%) | Top-3 94.9% |
| Bertram et al. 2021 | Contextual preference ranking (Siamese net) | NEO, 17Lands | 68% | As reported in UrzaGPT |
| DraftFM (Ward, 2026) | 1.6M-param choice model on card features + text embeddings | 29 sets train; BRO/FDN/MSH fully held out | 56.0% mean (50.8 / 60.4 / 56.7); 68.3% in-distribution val | Day-zero: never saw the evaluated set |
| UrzaGPT (Bertram, 2025) | Llama-3-8B + LoRA | NEO, 17Lands, 1M picks train / 10k test | 66.2% (Mistral-7B: 64.3%) | Untuned 8B models make illegal picks |
| UrzaGPT paper | GPT-4o zero-shot | NEO | 43% (38% with full card text) | Random baseline 22.1% |
| Ward et al. 2020 | NNetBot (dense net) | M19, Draftsim, 21,590 test drafts | 48.67% | DraftsimBot 44.54%, BayesBot 43.35%, Random 22.15% |
Take-aways that are supported: the model clearly beats the tier-list baseline on the same rows (+11.6 points P1P1, +25.7 P2P2, +36.2 picks 2β14), and it lands in the same band as the best published set-specific pick models (β 66β71%). A claim of "better than X" would require running X on these exact HOB files.
| Base | convaiinnovations/laya (English, ModernBERT-large encoder, 421M params) |
| Data | 17Lands draft_data_public.HOB.PremierDraft.csv, 3,000 drafts from players with β₯ 100 games and β₯ 58% win-rate bucket |
| Card text | MTGJSON HOB.json (5.3.0+20260922), compact type/cost/colors/P-T/rarity/keywords/rules text |
| Splits | By draft_id: train / val / test (test = 12,642 picks) |
| Schedule | 4 epochs, best epoch 4 (val accuracy 69.1%, val loss 0.820) |
| Lengths | max_len=1280, head_max_len=768 |
| Temperature | [1.232, 1.2, 1.2] as stored in rl_agent_config.json |
Epoch history (validation): 65.3% β 67.2% β 68.6% β 69.1%.
The model only saw one format; use it exactly.
state = {
"game": "Magic: The Gathering", "format": "Booster Draft", "event_type": "PremierDraft",
"expansion": "HOB", "set_name": "The Hobbit",
"pack_number": 1, "pick_number": 11, "pack_size": 4,
"pool": [{"name": "Bolg's Company", "count": 1}, {"name": "Gnashing of Teeth", "count": 2}], # alphabetical
"pool_size": 3,
}
questions = {
"pick": {
"type": "choice",
"instructions": (
"You are drafting a Booster Draft pod of the Magic: The Gathering set 'The Hobbit' (HOB). "
"This is pick 11 of pack 1. Given your current pool, choose which card to take from the "
"pack shown in the options."
),
"criteria": { # the cards in the pack, alphabetical, unique names
"Gollum, Silent Slinker": "Legendary Creature β Halfling Horror, {3}{B}, B, 4/3, common, Menace. Text: Menace (This creature can't beβ¦",
"Gundabad Opportunist": "Creature β Goblin Rogue, {3}{R}, R, 4/2, common. Text: When this creature enters, exile the top card of your library. Until the end ofβ¦",
"Iron Hills": "Land, mana value 0, colorless, common. Text: This land enters tapped. {T}: Add {R} or {W}. {2}{R}{W}, {T}, Sacrifice thisβ¦",
"Ordinary Bear": "Creature β Bear, {3}{G}, G, 4/5, common",
},
}
}
import laya
from huggingface_hub import snapshot_download
path = snapshot_download("FabioCeleste/laya-mtg-draft-picks")
agent = laya.load(path, device="cuda") # or "mps" / "cpu"
out = agent.predict(state, questions)["answers"]["pick"]
out["choice"] # card name, always one of the criteria keys
out["probabilities"] # {card: probability}, sums to 1
out["confidence"] # calibrated confidence in [0, 1]
Tested with laya==0.3.20. Throughput on an M-series Mac (MPS) was ~0.3 s per pick with 4 client threads
and serialized inference.
confidence without your own review threshold.criteria. Supply card text from a versioned source (e.g. MTGJSON).FabioCeleste/laya-mtg-deckbuild β chooses maindeck copy counts from a drafted pool.FabioCeleste/laya-mtg-deck-evaluator β estimates per-game win probability of a 40-card deck.Training labels come from 17Lands public datasets, licensed CC BY 4.0. Card data from MTGJSON. The base model is Apache-2.0.