Skip to content

Repository files navigation

Video Study Digest Skill

Source-grounded study notes from videos, transcripts, and subtitle files.

This skill is for agents that need to help users learn from video content instead of only summarizing it. It prioritizes timestamped evidence, watch/skip guidance, concept explanations, active recall, and clear source boundaries.

What It Does

  • Normalizes .srt, .vtt, and timestamped .txt transcripts.
  • Extracts source context from video URLs or yt-dlp info JSON.
  • Runs a one-command preparation pipeline for source context plus transcript artifacts.
  • Checks local dependencies and cache settings with a doctor script.
  • Keeps metadata separate from transcript-backed claims.
  • Produces Chinese study notes by default, preserving important source terms.
  • Supports compact, default, and deep study outputs.

Install

Recommended: install with Codex's skill installer from the GitHub repo path:

python ~/.codex/skills/.system/skill-installer/scripts/install-skill-from-github.py \
  --repo Captain-Pam/video-study-digest-skill \
  --path skills/video-study-digest

Or ask Codex:

Install the skill from https://github.com/Captain-Pam/video-study-digest-skill/tree/main/skills/video-study-digest

After installation, restart Codex to pick up the new skill.

Manual install:

Clone this repository and symlink or copy the skill folder into your agent's skills directory:

mkdir -p ~/.codex/skills
ln -s "$(pwd)/skills/video-study-digest" ~/.codex/skills/video-study-digest

On Windows PowerShell:

New-Item -ItemType Directory -Force $env:USERPROFILE\.codex\skills
Copy-Item -Recurse .\skills\video-study-digest $env:USERPROFILE\.codex\skills\video-study-digest

Usage

Use $video-study-digest to turn this video transcript into concise study notes.

Quick environment check:

python skills/video-study-digest/scripts/doctor.py

Recommended one-command source preparation:

python skills/video-study-digest/scripts/video_digest_pipeline.py "https://www.youtube.com/watch?v=VIDEO_ID" --output-dir outputs/video

If captions are missing and local transcription is acceptable:

python skills/video-study-digest/scripts/video_digest_pipeline.py "https://www.youtube.com/watch?v=VIDEO_ID" --output-dir outputs/video --transcribe-if-needed

Output Files

The one-command pipeline writes results to the directory passed with --output-dir.

File When it appears Purpose
source_context.json URL or yt-dlp info JSON input when metadata extraction works Machine-readable title, uploader, duration, chapters, description, tags, thumbnails, and available captions. Use for triage, not transcript evidence.
source_context.md Same as source_context.json Human-readable source metadata for quick review.
transcript.md Local transcript/subtitle input, or URL captions found by yt-dlp Normalized timestamped transcript for reading, summarizing, and citing.
transcript.json Same as transcript.md Machine-readable transcript segments for downstream processing.
transcript_whisper.vtt URL transcription fallback with --transcribe-if-needed Timestamped transcript generated from audio by faster-whisper.
transcript_whisper.md Same transcription fallback Human-readable transcription output.
transcript_whisper.json Same transcription fallback Machine-readable transcription segments and metadata.
run_report.json Every pipeline run Machine-readable status, transcript method, cache root, output paths, warnings, and errors.
run_report.md Every pipeline run Human-readable run summary and output index.

The pipeline does not write final study notes by itself. After preparing source context and transcript files, use the skill to produce the learning output, commonly saved as study_notes.md when you want a file artifact.

For URL metadata context:

python skills/video-study-digest/scripts/extract_video_context.py "https://www.youtube.com/watch?v=VIDEO_ID" --output source_context.json

For captions or transcript normalization:

python skills/video-study-digest/scripts/prepare_transcript.py transcript.vtt --output transcript.md

For videos without public captions, audio-only transcription can use the default non-C-drive cache:

python skills/video-study-digest/scripts/transcribe_audio.py "https://www.youtube.com/watch?v=VIDEO_ID" --cache-root F:\cc_project\CodexMediaCache --model-size base

The cache stores audio, transcripts, temporary files, and faster-whisper model files under the configured cache root.

For cross-platform use, set VIDEO_STUDY_CACHE_ROOT or pass --cache-root:

export VIDEO_STUDY_CACHE_ROOT="$HOME/.cache/video-study-digest"

On Windows PowerShell:

$env:VIDEO_STUDY_CACHE_ROOT = "F:\cc_project\CodexMediaCache"

To persist it for future Codex sessions on Windows:

[Environment]::SetEnvironmentVariable("VIDEO_STUDY_CACHE_ROOT", "F:\cc_project\CodexMediaCache", "User")

On Windows, if F:\cc_project\CodexMediaCache already exists, the scripts use it as the default non-C-drive cache. Otherwise they fall back to the platform user cache directory.

Dependencies

  • Python 3.10+ recommended.
  • yt-dlp is optional but recommended for URL metadata and caption extraction.
  • faster-whisper is optional for audio transcription when captions are missing.

The tests use only the Python standard library.

Source Discipline

Metadata such as title, description, tags, chapters, thumbnails, and available captions can guide triage and learning goals. They are not transcript evidence. Claims about what the video says should come from timestamped transcript content or be explicitly labeled as inference.

Test

python -m unittest discover -s tests

License

MIT

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages