driver-scene / prompt eval 8-way VLM prompt comparison · BDD100K
no data

no run loaded

Generate a comparison run, then load its output files here (multi-select the whole run directory contents, or drag them anywhere onto this page):

python -m src.pipeline compare --limit 10

comparison_table.json  ·  significance.json (optional — adds CIs + paired tests)  ·  evaluation_results_<variant>.json (optional — per-image drill-down)