Extract scene-change frames from YouTube educational videos, select the ones you want, and export them as a PDF for annotation in Notability.
- conda env
ytframes— Python backend runs inside this env - ffmpeg — must be installed and on PATH (verified at startup)
- yt-dlp — must be installed and on PATH (verified at startup)
- Node.js (v18+) — for the frontend
conda activate ytframes
pip install imagehash Pillow
cd frontend
npm install./start.shThis starts the backend and frontend, waits until both are ready, and opens http://localhost:5173 in your browser. Press Ctrl+C to stop both servers.
Optional — add a shell alias so you can launch from anywhere:
echo 'alias ytframes="~/Documents/Projects/yt-frame-extractor/start.sh"' >> ~/.zshrc
source ~/.zshrcThen just type ytframes.
# Terminal 1 — backend
conda activate ytframes
cd backend && uvicorn main:app --reload --port 8000
# Terminal 2 — frontend
cd frontend && npm run dev- Paste a YouTube URL and click Process (or press Enter).
- Wait for download + extraction — the status bar shows progress.
- Click thumbnails to select/deselect frames. A blue border + checkmark indicates selection.
- Use the Threshold slider to adjust how many frames are extracted, then click Rescan — the already-downloaded video is reused, no re-download needed.
- Click Export N frames as PDF to download
frames.pdf. - Import into Notability on iPad.
Open backend/extract.py and adjust the constants at the top:
| Constant | Default | Effect |
|---|---|---|
SCENE_THRESHOLD |
3.0 |
scdet threshold (0–100 scale). Lower = more frames. AV1/YouTube videos: try 2–5. h264 videos: try 5–20. |
DEDUP_HASH_THRESHOLD |
5 |
Hamming distance cutoff for duplicate removal. Lower = stricter dedup. |
DOWNLOAD_RESOLUTION |
720 |
Max video height in pixels. 720p keeps slide text readable. |
| Method | Path | Description |
|---|---|---|
POST |
/process |
Download + extract frames; streams ndjson progress |
POST |
/rescan |
Re-extract from cached video with a new threshold |
GET |
/frames/{id} |
Serve a full-resolution frame image |
POST |
/export |
Build and return the PDF |
- "No scene changes detected" — lower
SCENE_THRESHOLDinextract.pyor drag the slider left. AV1-encoded YouTube videos need values around 2–4. - Too many near-identical frames — raise
SCENE_THRESHOLDor lowerDEDUP_HASH_THRESHOLD. - "Could not reach backend" — ensure uvicorn is running on port 8000 with the
ytframesenv active. - Download errors — yt-dlp can fail on age-restricted or members-only videos.