no run loaded
Generate a comparison run, then load its output files here (multi-select the whole run directory contents, or drag them anywhere onto this page):
python -m src.pipeline compare --limit 10
comparison_table.json · significance.json (optional — adds CIs + paired tests) · evaluation_results_<variant>.json (optional — per-image drill-down)
variant comparison
ranked by hallucination rate — click a column to re-sort, click a row
for per-image detail
means & uncertainty
per-image
sorted by hallucination, worst first — click the variant row again to
close
analyzer notes
from analysis_<variant>.json