Reimagining service delivery in the agentic era with Google Public Sector
cloud.google.com·1d ago
TL;DR
Public medical vision-language models (VLMs) may be contaminated by pretraining data, affecting their reported accuracy. A controlled audit was conducted using techniques like image-side near-neighbour overlap and canonical-order exchangeability.
✦ Why It Matters
Engineers should critically assess VLMs for pretraining contamination to ensure accurate model evaluations.
Key Takeaways
How It Works
The authors employed four detection methods to identify overlaps between benchmark images and pretraining datasets. These included image-side near-neighbour overlap and canonical-order exchangeability, which helped quantify the extent of contamination.
Related