Deepfake Detection Based on Ratio Images

Detecting manipulated videos through interpretable temporal differences between adjacent frames.

I investigated temporal differences between adjacent frames as an interpretable signal for video-level deepfake detection. This complements methods that inspect only the appearance of individual frames.

  1. Adjacent video frames
  2. Face and background regions
  3. Ratio-image construction
  4. Engineered temporal features
  5. Random Forest video classification

Problem

A manipulated frame may appear realistic when inspected on its own. Artifacts can instead emerge in how facial and background regions change across time, so static inspection may discard useful evidence.

Why it matters

Manipulation artifacts can emerge over time rather than within a single frame. Representing temporal behavior provides evidence that a frame-based detector may discard.

Approach

I construct ratio images from adjacent frames and extract features from temporal changes in face and background regions. A Random Forest combines those engineered signals into a video-level classification.

Evaluation and evidence

The repository contains the implementation needed to reproduce the research workflow. The publication reports the experimental evaluation; this summary does not add metrics that are not available in the repository.

Core hypothesis: temporal behavior can expose manipulation artifacts that static appearance may miss.

Technical implementation

Python and OpenCV handle video and region processing; scikit-learn supports feature-based classification. The explicit feature pipeline keeps the relationship between temporal artifacts and the classifier input inspectable.

What I learned

When a failure develops over time, the evaluation must preserve that temporal context. Video-level evidence can reveal behavior that frame-level appearance obscures.