Source-grounded study notes from videos, transcripts, and subtitle files.
This skill is for agents that need to help users learn from video content instead of only summarizing it. It prioritizes timestamped evidence, watch/skip guidance, concept explanations, active recall, and clear source boundaries.
- Normalizes
.srt,.vtt, and timestamped.txttranscripts. - Extracts source context from video URLs or
yt-dlpinfo JSON. - Runs a one-command preparation pipeline for source context plus transcript artifacts.
- Checks local dependencies and cache settings with a doctor script.
- Keeps metadata separate from transcript-backed claims.
- Produces Chinese study notes by default, preserving important source terms.
- Supports compact, default, and deep study outputs.
Recommended: install with Codex's skill installer from the GitHub repo path:
python ~/.codex/skills/.system/skill-installer/scripts/install-skill-from-github.py \
--repo Captain-Pam/video-study-digest-skill \
--path skills/video-study-digestOr ask Codex:
Install the skill from https://github.com/Captain-Pam/video-study-digest-skill/tree/main/skills/video-study-digest
After installation, restart Codex to pick up the new skill.
Manual install:
Clone this repository and symlink or copy the skill folder into your agent's skills directory:
mkdir -p ~/.codex/skills
ln -s "$(pwd)/skills/video-study-digest" ~/.codex/skills/video-study-digestOn Windows PowerShell:
New-Item -ItemType Directory -Force $env:USERPROFILE\.codex\skills
Copy-Item -Recurse .\skills\video-study-digest $env:USERPROFILE\.codex\skills\video-study-digestUse $video-study-digest to turn this video transcript into concise study notes.
Quick environment check:
python skills/video-study-digest/scripts/doctor.pyRecommended one-command source preparation:
python skills/video-study-digest/scripts/video_digest_pipeline.py "https://www.youtube.com/watch?v=VIDEO_ID" --output-dir outputs/videoIf captions are missing and local transcription is acceptable:
python skills/video-study-digest/scripts/video_digest_pipeline.py "https://www.youtube.com/watch?v=VIDEO_ID" --output-dir outputs/video --transcribe-if-neededThe one-command pipeline writes results to the directory passed with --output-dir.
| File | When it appears | Purpose |
|---|---|---|
source_context.json |
URL or yt-dlp info JSON input when metadata extraction works |
Machine-readable title, uploader, duration, chapters, description, tags, thumbnails, and available captions. Use for triage, not transcript evidence. |
source_context.md |
Same as source_context.json |
Human-readable source metadata for quick review. |
transcript.md |
Local transcript/subtitle input, or URL captions found by yt-dlp |
Normalized timestamped transcript for reading, summarizing, and citing. |
transcript.json |
Same as transcript.md |
Machine-readable transcript segments for downstream processing. |
transcript_whisper.vtt |
URL transcription fallback with --transcribe-if-needed |
Timestamped transcript generated from audio by faster-whisper. |
transcript_whisper.md |
Same transcription fallback | Human-readable transcription output. |
transcript_whisper.json |
Same transcription fallback | Machine-readable transcription segments and metadata. |
run_report.json |
Every pipeline run | Machine-readable status, transcript method, cache root, output paths, warnings, and errors. |
run_report.md |
Every pipeline run | Human-readable run summary and output index. |
The pipeline does not write final study notes by itself. After preparing source context and transcript files, use the skill to produce the learning output, commonly saved as study_notes.md when you want a file artifact.
For URL metadata context:
python skills/video-study-digest/scripts/extract_video_context.py "https://www.youtube.com/watch?v=VIDEO_ID" --output source_context.jsonFor captions or transcript normalization:
python skills/video-study-digest/scripts/prepare_transcript.py transcript.vtt --output transcript.mdFor videos without public captions, audio-only transcription can use the default non-C-drive cache:
python skills/video-study-digest/scripts/transcribe_audio.py "https://www.youtube.com/watch?v=VIDEO_ID" --cache-root F:\cc_project\CodexMediaCache --model-size baseThe cache stores audio, transcripts, temporary files, and faster-whisper model files under the configured cache root.
For cross-platform use, set VIDEO_STUDY_CACHE_ROOT or pass --cache-root:
export VIDEO_STUDY_CACHE_ROOT="$HOME/.cache/video-study-digest"On Windows PowerShell:
$env:VIDEO_STUDY_CACHE_ROOT = "F:\cc_project\CodexMediaCache"To persist it for future Codex sessions on Windows:
[Environment]::SetEnvironmentVariable("VIDEO_STUDY_CACHE_ROOT", "F:\cc_project\CodexMediaCache", "User")On Windows, if F:\cc_project\CodexMediaCache already exists, the scripts use it as the default non-C-drive cache. Otherwise they fall back to the platform user cache directory.
- Python 3.10+ recommended.
yt-dlpis optional but recommended for URL metadata and caption extraction.faster-whisperis optional for audio transcription when captions are missing.
The tests use only the Python standard library.
Metadata such as title, description, tags, chapters, thumbnails, and available captions can guide triage and learning goals. They are not transcript evidence. Claims about what the video says should come from timestamped transcript content or be explicitly labeled as inference.
python -m unittest discover -s testsMIT