Frozen vision-language safety score tracks scene template, not hazards directly

Controlled experiments show that CLIP-based safety signals in reinforcement learning reflect caption-bank scene similarity rather than actual hazard detection.

Academic
Samuel Tetteh · Cody Fleming

Iowa State University

Research Digest··3 min read
Samuel Tetteh and Cody Fleming from Iowa State University audit a frozen CLIP prompt-margin safety score in a driving RL setup.

The authors used three VLM-free FormulaOne driving policies (which never received the score) to generate 180 episodes with 130 isolated contact onsets.

Why this paper

From Iowa State University

In one line

A frozen CLIP safety score measures scene resemblance to its caption bank, not hazard detection.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.

How we workSubscribe