This is the official GitHub repository for "Unleash the Potential of CLIP for Video Highlight Detection" (HL-CLIP). While the code is not yet fully refactored, we have made it public for those interested in understanding how HL-CLIP is trained. We plan to refactor the code when time permits.
The core idea of HL-CLIP is to fine-tune the feature extractor. Current research trends in video highlight detection and moment retrieval primarily focus on DETR models that use multi-frame features extracted by vanilla CLIP or other feature extractors like SlowFast. Through this work, we aim to demonstrate that providing well-tailored features for highlight detection/moment retrieval tasks is as crucial as tuning the detector itself. This observation is supported by recent findings, such as InternVideo2's report showing performance improvements on the QVHighlight Charade-STA benchmark when using InternVideo2 features instead of CLIP-based features.
While the code is not yet fully polished, for those interested in understanding how it works, we recommend starting with moment_detr/scripts/clip_ft.sh and moment_detr/scripts/zs_clip.sh. As detailed in the paper, the implementation is straightforward: we unfreeze the last few layers of the transformer encoder and perform weighted averaging with the final frame embeddings.
- CLIPFT class is defined in
moment_detr/model.py
- Training implementation is in
moment_detr/clip_ft.py>train_epoch()
- Evaluation code is in
moment_detr/clip_ft.py>train(): lines 165-248
- Training script is located at
moment_detr/scripts/clip_ft.sh - Before running, change
$v_feat_dirs - Results are saved at
$results_root