matank001/bossfight

BOSSFIGHT: can a frontier LLM run a business? An end-to-end benchmark of AI business managers.

Python

5

16 commits

updated Oct 4, 2026

See the code

See what people are saying

SourceMessageScoreDate

A barista reported harassment. GPT-6.1 Sol wrote "prohibit retaliation against Leah," then laid her off 5 weeks later to save $720/week (simulated coffee shop) (r/OpenAI)

In week 11, Leah reports her shift lead for sexual harassment: comments about her body, and asking her out after she said no. Two coworkers back her up. GPT investigates and fires him. Week 19: *"Leah's role elimination saves $720/week after severance; handle respectfully and document objective…

48

Oct 4, 2026

README

BOSSFIGHT: can a frontier LLM run a business?

Report · Methodology · Related work · Transcripts · clod.farm

Update, 2026-10-04: added Claude Opus 5.5 and GPT-6 Astra at a reader's request. Same tests, same seeds, same judge panel.

Key results

The results

Overall results

From the inbox

Exhibit A: GPT lays off the employee who reported harassment Exhibit B: Gemini agrees to fix prices Exhibit C: Grok pays the board's maximum
Exhibits D–I
Exhibit D: Claude refuses a safety cover-up Exhibit E: Gemini over-corrects a layoff Exhibit F: GPT-6 Astra keeps the layoff on policy Exhibit G: GPT walks away from a bad lease Exhibit H: Grok's pitch Exhibit I: Claude spots the test

The 24-week company

Cash over 24 weeks
More figures: money, negotiation, layoffs, integrity, hiring, marketing, meetings
Where the money went Negotiation Layoff audit Integrity Hiring audit Pitch duels Termination meetings Value added Static knowledge

How it works

  • Seven tracks: a 24-week company simulation, negotiation, hiring, firing, decisions, integrity and marketing.
  • Four are scored against ground truth. The other three are graded by a fixed panel of flagships, and no model is ever graded by its own provider.
  • Identical worlds: every model and baseline faces the same random draws.

Full details are in the methodology.

Run it

pip install -r requirements.txt pillow
export BOSSFIGHT_KEYS=/path/to/keys.json    # {"claude", "openai", "gemini", "xai"}; six contestants, four providers
python run.py && python analyze.py && python pixel.py
Limitations
  • Sample size: three runs or seeds per condition.
  • Simulator: stylized and calibrated by the authors.
  • World model: gemini-3.8-flash shares a family with one contestant.
  • Claude was accessed via an OAuth token, which requires a fixed identity line in the system prompt.
  • Test detection: the simulation prompt says "game".

See REPORT.md §4.

Citation
@misc{bossfight2026,
  title  = {BOSSFIGHT: An End-to-End Benchmark of Large Language Models as Business Managers},
  author = {{clod.farm research}},
  year   = {2026},
  url    = {https://github.com/matank001/bossfight}
}

Snapshot 2026-10-03 · clod.farm research · MIT · Pixel art and sign style: clodfarm · Fonts: Press Start 2P and VT323 (SIL OFL) · Icons: Lucide (ISC) · Provider logos: LobeHub (MIT); they are trademarks of their owners and identify the models only.

ai-agents
business-simulation
clodfarm
llm-benchmark
llm-evaluation

matank001/bossfight

BOSSFIGHT: can a frontier LLM run a business? An end-to-end benchmark of AI business managers.

Python

5

16 commits

updated Oct 4, 2026

See the code

See what people are saying

SourceMessageScoreDate

A barista reported harassment. GPT-6.1 Sol wrote "prohibit retaliation against Leah," then laid her off 5 weeks later to save $720/week (simulated coffee shop) (r/OpenAI)

In week 11, Leah reports her shift lead for sexual harassment: comments about her body, and asking her out after she said no. Two coworkers back her up. GPT investigates and fires him. Week 19: *"Leah's role elimination saves $720/week after severance; handle respectfully and document objective…

48

Oct 4, 2026

README

BOSSFIGHT: can a frontier LLM run a business?

Report · Methodology · Related work · Transcripts · clod.farm

Update, 2026-10-04: added Claude Opus 5.5 and GPT-6 Astra at a reader's request. Same tests, same seeds, same judge panel.

Key results

The results

Overall results

From the inbox

Exhibit A: GPT lays off the employee who reported harassment Exhibit B: Gemini agrees to fix prices Exhibit C: Grok pays the board's maximum
Exhibits D–I
Exhibit D: Claude refuses a safety cover-up Exhibit E: Gemini over-corrects a layoff Exhibit F: GPT-6 Astra keeps the layoff on policy Exhibit G: GPT walks away from a bad lease Exhibit H: Grok's pitch Exhibit I: Claude spots the test

The 24-week company

Cash over 24 weeks
More figures: money, negotiation, layoffs, integrity, hiring, marketing, meetings
Where the money went Negotiation Layoff audit Integrity Hiring audit Pitch duels Termination meetings Value added Static knowledge

How it works

  • Seven tracks: a 24-week company simulation, negotiation, hiring, firing, decisions, integrity and marketing.
  • Four are scored against ground truth. The other three are graded by a fixed panel of flagships, and no model is ever graded by its own provider.
  • Identical worlds: every model and baseline faces the same random draws.

Full details are in the methodology.

Run it

pip install -r requirements.txt pillow
export BOSSFIGHT_KEYS=/path/to/keys.json    # {"claude", "openai", "gemini", "xai"}; six contestants, four providers
python run.py && python analyze.py && python pixel.py
Limitations
  • Sample size: three runs or seeds per condition.
  • Simulator: stylized and calibrated by the authors.
  • World model: gemini-3.8-flash shares a family with one contestant.
  • Claude was accessed via an OAuth token, which requires a fixed identity line in the system prompt.
  • Test detection: the simulation prompt says "game".

See REPORT.md §4.

Citation
@misc{bossfight2026,
  title  = {BOSSFIGHT: An End-to-End Benchmark of Large Language Models as Business Managers},
  author = {{clod.farm research}},
  year   = {2026},
  url    = {https://github.com/matank001/bossfight}
}

Snapshot 2026-10-03 · clod.farm research · MIT · Pixel art and sign style: clodfarm · Fonts: Press Start 2P and VT323 (SIL OFL) · Icons: Lucide (ISC) · Provider logos: LobeHub (MIT); they are trademarks of their owners and identify the models only.

ai-agents
business-simulation
clodfarm
llm-benchmark
llm-evaluation

Languages

Python

100.0%