Visible failure modes
Evaluation should reveal when systems become unreliable rather than treating aggregate accuracy as sufficient evidence of trustworthiness.
This area of work examines the reliability, validation, provenance and accountability of AI-assisted investigative technology. Detailed project-level methods, datasets, metrics and experimental identifiers are intentionally not published here while related scholarly work is under anonymous peer review.
Evaluation should reveal when systems become unreliable rather than treating aggregate accuracy as sufficient evidence of trustworthiness.
Automated output should remain subordinate to qualified human review where evidentiary or investigative consequences are material.
Inputs, model/tool conditions, uncertainty, provenance and human decisions should be documented in ways that support review and challenge.