Governance
How AI use stays accountable: policy, audit trails, model approvals, risk frameworks, and the human checks that keep autonomous systems within bounds.
No practices in this stage match the current filters.
-
WhenYou are commissioning an interpretability audit of a production AI model, or scoping the deliverables from an interpretability research engagement.
UseRequire deliverables to include both a proposed intervention and a validation result, not just an explanation. Ask auditors to specify what to change in the model and to provide evidence that the proposed change produces the expected behavioral outcome without degrading other capabilities. Findings that name a pattern without proposing a specific intervention are exploratory outputs; include them in a separate exploratory annex rather than the actionable findings section.
EvidenceThe two-dimension rubric proposed by Orgad, Barez, Haklay et al. (2026) separates concreteness (how specific is the proposed intervention?) from validation (has the intervention been tested?). An audit that delivers only explanations without proposed interventions, or proposed interventions without validation evidence, does not clear the actionability bar the paper defines. Scoping audits explicitly against both criteria ensures deliverables are usable for production decisions rather than only for research awareness.