The benchmarks used are:
See also the original site to solve 5-minute mysteries as humans.
Inspired by LMs for Rationality,
Inpsired by DetectBench:
git submodule initgit submodule updatepip install git+https://github.com/huggingface/transformers.gitpip install "outlines @ git+https://github.com/outlines-dev/outlines.git@main"pip install z3-solverpip install python-satpip install datasetspip install ollamapip install openaipip install outlinespip install wandbollama pull qwen2.5ollama pull deepseek-r1Python
94.2%
HTML
5.5%
The benchmarks used are:
See also the original site to solve 5-minute mysteries as humans.
Inspired by LMs for Rationality,
Inpsired by DetectBench:
git submodule initgit submodule updatepip install git+https://github.com/huggingface/transformers.gitpip install "outlines @ git+https://github.com/outlines-dev/outlines.git@main"pip install z3-solverpip install python-satpip install datasetspip install ollamapip install openaipip install outlinespip install wandbollama pull qwen2.5ollama pull deepseek-r1Python
94.2%
HTML
5.5%