Cleaned project layout focused on the main workflow.
gen_data/.test/run_tests.sh.test/result/.analysis/.gen_data/: dataset generation and sampling scripts.test/: evaluation runner and test logic.test/result/: generated evaluation outputs.analysis/: post-processing and statistics scripts.test/run_tests.sh: primary test runner.test/test.py: evaluation logic used by the runner.test/questions.py: question definitions and evaluation helpers.unified_mllm.py: unified model wrapper.recheck/ or only keep latest runs).1 commits
Python
95.5%
Shell
4.5%
Cleaned project layout focused on the main workflow.
gen_data/.test/run_tests.sh.test/result/.analysis/.gen_data/: dataset generation and sampling scripts.test/: evaluation runner and test logic.test/result/: generated evaluation outputs.analysis/: post-processing and statistics scripts.test/run_tests.sh: primary test runner.test/test.py: evaluation logic used by the runner.test/questions.py: question definitions and evaluation helpers.unified_mllm.py: unified model wrapper.recheck/ or only keep latest runs).1 commits
Python
95.5%
Shell
4.5%