Abstract
Forensic verification often uses a binary “real vs. fake” label that groups fully synthetic, tampered, and AI-retouched images despite their different consequences. We study these modifications through two complementary channels: a camera channel sensitive to capture and processing traces, and a semantic channel capturing scene content. The channels provide continuous evidence rather than deterministic signatures of manipulation history. We instantiate this perspective in 2CAP (2-Channel Authenticity Protocol), pairing a contrastively trained camera encoder with a frozen semantic encoder for (i) reference-free four-class classification through reliability-weighted cross-attention fusion and (ii) reference-based evidence generation. For the latter, query-reference channel similarities and a patch-level saliency map guide a frozen Vision Language Model (VLM) through an Observe–Generate–Refine loop, without forensic instruction tuning of the VLM. On the evaluated benchmark, 2CAP achieves overall AUC .963 and F1 .844; retouching F1 improves from .868 for the strongest compared baseline to .961. Across five VLM configurations, shared-parser paired evaluation shows model-dependent effects: evidence improves change-type accuracy over evidence-free refinement. These results support manipulation-type discrimination while delimiting the benefits of frozen-VLM refinement.