Supported datasets:
Edit and run:
bash scripts/eval.sh
Run the template:
bash api_eval/run_batch_exp7.sh
The main pipeline supports use_llm_judge.
There are also standalone judge scripts for PMC-VQA and OmniMedVQA:
Common output files:
results.jsonmetrics.jsontotal_results.json8 commits
Python
96.8%
JavaScript
1.4%
HTML
1.1%
Supported datasets:
Edit and run:
bash scripts/eval.sh
Run the template:
bash api_eval/run_batch_exp7.sh
The main pipeline supports use_llm_judge.
There are also standalone judge scripts for PMC-VQA and OmniMedVQA:
Common output files:
results.jsonmetrics.jsontotal_results.json8 commits
Python
96.8%
JavaScript
1.4%
HTML
1.1%