DedeProGames/SLM-Tetris-Arena

Space

🧱 SLM Tetris Arena

16

76 commits

updated Sep 29, 2026

See the code

README

🧱 SLM Tetris Arena

Can a language model that only ever read text (e.g. FineWeb-edu) play Tetris without any training on the game?

Decoder-only models (50K–250M parameters, custom architectures welcome) play the same piece sequence side by side on CPU. Two ways to play:

  • Match (friendly): pick any 2+ models and the seed. Nothing is recorded.
  • Ranked: press Play and the arena picks up to 4 models at random from its pool, of any size. The result updates a public Elo leaderboard. The pool is the SUGGESTED list in app.py; removing a model from that list also removes it from the leaderboard.

How a model plays

For each new piece the engine enumerates every legal placement (rotation Γ— column, hard drop), simulates it and describes the outcome in plain English. The model never sees the grid. It scores each description:

In Tetris, the goal is to clear lines, avoid holes and keep the stack low.
This move drops the piece into the lowest part of the board, clears one line, creates no new holes, keeps the stack low and leaves the surface flat.
It is a

value = log P(" good move") βˆ’ log P(" bad move"). The highest value is played; exact ties are broken by a seeded coin that is the same for every player. Because the value is a difference, a model's overall bias towards "good" or "bad" cancels out, so no calibration is needed.

The better the model understands language (more/better pre-training), the better it should read the consequences of each move. That's the hypothesis this arena tests.

  • Guided protocol: the rules are stated in the prompt (reading comprehension).
  • Blind protocol: Here is a move from a game of Tetris. Only pre-training knowledge.

Each protocol has its own leaderboard.

Baselines

  • 🎲 Random: uniform random placement (the floor).
  • πŸ“ Oracle reader: ranks the same descriptions with fixed common sense (roughly the ceiling for a perfect reader).

Elo

Only the Ranked tab changes Elo. The arena picks the players at random from the suggested models (models with fewer ranked games are more likely to be picked), of any size, and uses a random seed. Nobody chooses who plays ranked, so Elo can't be farmed by pairing a model with weak opponents. Any model can meet any other, so every rating sits on one comparable scale (big and small models can be compared directly); Elo weighs each win by the opponent's rating, so beating a much weaker model earns almost nothing once ratings have settled. The match runs on the server in the background: it finishes and counts even if the viewer leaves, and only one ranked match runs at a time (pressing Play while one is running lets you watch it).

Placement is decided by score (100/300/500/800 for 1–4 lines), then lines, then pieces survived (cap: 500 pieces). Multiplayer Elo, K = 32: each pair of players is a game, scaled by 1/(Nβˆ’1). Baselines are never rated. Every ranked match (seed, model commit SHAs, scores, Elo before/after) is stored in the results dataset.

Configuration (Space variables / secrets)

NameDefaultPurpose
HF_TOKEN (secret)–Fine-grained token with write access only to the results dataset. Without it, Elo lives in memory.
RESULTS_REPODedeProGames/lm-tetris-arena-resultsDataset that stores the leaderboard
MAX_MODELS4Language models per match
MAX_PARAMS250000000Parameter limit
MIN_PARAMS50000Smallest model allowed (smaller ones are rejected and removed from the leaderboard)
MAX_PIECES500Piece cap per game
MAX_PARAM_GAP0Optional largest size difference between models picked for a ranked match (0 = no limit)
LARGE_FROM100000000Only with a gap: models this size or bigger can all play each other
ALLOW_REMOTE_CODE1Allow models with custom code (trust_remote_code)
TORCH_THREADS2CPU threads for inference

⚠️ Custom-code models run arbitrary Python on this Space. The app removes HF_TOKEN from the environment before loading any model, but use a fine-grained token scoped to the results dataset so the worst case is a revertible commit on that dataset.

evaluation
gradio
leaderboard
small-language-models
zero-shot

DedeProGames/SLM-Tetris-Arena

Space

🧱 SLM Tetris Arena

16

76 commits

updated Sep 29, 2026

See the code

README

🧱 SLM Tetris Arena

Can a language model that only ever read text (e.g. FineWeb-edu) play Tetris without any training on the game?

Decoder-only models (50K–250M parameters, custom architectures welcome) play the same piece sequence side by side on CPU. Two ways to play:

  • Match (friendly): pick any 2+ models and the seed. Nothing is recorded.
  • Ranked: press Play and the arena picks up to 4 models at random from its pool, of any size. The result updates a public Elo leaderboard. The pool is the SUGGESTED list in app.py; removing a model from that list also removes it from the leaderboard.

How a model plays

For each new piece the engine enumerates every legal placement (rotation Γ— column, hard drop), simulates it and describes the outcome in plain English. The model never sees the grid. It scores each description:

In Tetris, the goal is to clear lines, avoid holes and keep the stack low.
This move drops the piece into the lowest part of the board, clears one line, creates no new holes, keeps the stack low and leaves the surface flat.
It is a

value = log P(" good move") βˆ’ log P(" bad move"). The highest value is played; exact ties are broken by a seeded coin that is the same for every player. Because the value is a difference, a model's overall bias towards "good" or "bad" cancels out, so no calibration is needed.

The better the model understands language (more/better pre-training), the better it should read the consequences of each move. That's the hypothesis this arena tests.

  • Guided protocol: the rules are stated in the prompt (reading comprehension).
  • Blind protocol: Here is a move from a game of Tetris. Only pre-training knowledge.

Each protocol has its own leaderboard.

Baselines

  • 🎲 Random: uniform random placement (the floor).
  • πŸ“ Oracle reader: ranks the same descriptions with fixed common sense (roughly the ceiling for a perfect reader).

Elo

Only the Ranked tab changes Elo. The arena picks the players at random from the suggested models (models with fewer ranked games are more likely to be picked), of any size, and uses a random seed. Nobody chooses who plays ranked, so Elo can't be farmed by pairing a model with weak opponents. Any model can meet any other, so every rating sits on one comparable scale (big and small models can be compared directly); Elo weighs each win by the opponent's rating, so beating a much weaker model earns almost nothing once ratings have settled. The match runs on the server in the background: it finishes and counts even if the viewer leaves, and only one ranked match runs at a time (pressing Play while one is running lets you watch it).

Placement is decided by score (100/300/500/800 for 1–4 lines), then lines, then pieces survived (cap: 500 pieces). Multiplayer Elo, K = 32: each pair of players is a game, scaled by 1/(Nβˆ’1). Baselines are never rated. Every ranked match (seed, model commit SHAs, scores, Elo before/after) is stored in the results dataset.

Configuration (Space variables / secrets)

NameDefaultPurpose
HF_TOKEN (secret)–Fine-grained token with write access only to the results dataset. Without it, Elo lives in memory.
RESULTS_REPODedeProGames/lm-tetris-arena-resultsDataset that stores the leaderboard
MAX_MODELS4Language models per match
MAX_PARAMS250000000Parameter limit
MIN_PARAMS50000Smallest model allowed (smaller ones are rejected and removed from the leaderboard)
MAX_PIECES500Piece cap per game
MAX_PARAM_GAP0Optional largest size difference between models picked for a ranked match (0 = no limit)
LARGE_FROM100000000Only with a gap: models this size or bigger can all play each other
ALLOW_REMOTE_CODE1Allow models with custom code (trust_remote_code)
TORCH_THREADS2CPU threads for inference

⚠️ Custom-code models run arbitrary Python on this Space. The app removes HF_TOKEN from the environment before loading any model, but use a fine-grained token scoped to the results dataset so the worst case is a revertible commit on that dataset.

evaluation
gradio
leaderboard
small-language-models
zero-shot