A lightweight, self-hosted Telegram media downloader bot built with Python.
This bot downloads media from 1000+ platforms using yt-dlp and uploads the files back to Telegram. It's designed for homelab usage with minimal resource consumption and no external dependencies beyond yt-dlp and ffmpeg.
Key Characteristics:
- Pure utility bot - no AI, no LLM calls
- Async architecture using aiogram 3.x
- Access control via an allowlist of Telegram user IDs
- Per-user rate limiting
- Automatic temporary file cleanup
- Live download progress bar that updates in place (works the same in DMs and groups)
- Instant re-sends: a previously downloaded URL is resent from Telegram's cache (by
file_id) without re-downloading - Friendly, actionable error messages (e.g. "age-restricted — set a COOKIES_FILE")
- yt-dlp kept current automatically in Docker (refreshed on container start)
- Uploads up to 2GB via a bundled local Telegram Bot API server (vs. 50MB on the standard API)
- Audio-only sources (e.g. SoundCloud) are auto-detected and always fetched as tagged MP3
- Direct media URLs (e.g. an imageboard
.webm) are transcoded to a streamable MP4 (H.264/AAC,moovat the start) so Telegram plays them inline instead of attaching as a file - Audio results are a single post: MP3 with embedded cover art, album-art thumbnail, and title/artist/duration
- Each result post shows the original source URL as plain (non-linked) text
- Authenticated downloads via browser cookies (bare Python) or a mounted
cookies.txt(Docker) - Handles split video+audio sources (HLS/DASH) that have no single muxed stream
- Optional headless-browser fallback: when no yt-dlp extractor can handle a page whose player builds the media URL in JavaScript, the page is loaded in headless Chromium, the media request is captured, and that URL is handed back to yt-dlp
- Unit-tested with pytest
tg-media-bot/
├── main.py # Entry point: builds Bot/Dispatcher, starts polling
├── docker-compose.yml # Bot + local Telegram Bot API server
├── Dockerfile # Bot image (installs ffmpeg + yt-dlp)
├── flake.nix # Nix package, dev shell, checks (see packaging/nix/)
├── requirements.txt # Python dependencies
├── .env.example # Configuration template
├── src/
│ ├── bot/
│ │ ├── handlers.py # URL extraction, download/upload orchestration
│ │ └── router.py # Dispatcher, command routing, auth middleware
│ ├── commands/
│ │ └── handlers.py # /start, /help, /audio, /video, /status, /minimal, /topic, etc.
│ ├── config/
│ │ └── settings.py # Env-based settings (singleton)
│ ├── downloaders/
│ │ └── ytdlp.py # yt-dlp subprocess wrapper
│ ├── queue/
│ │ └── manager.py # Async queue + per-user rate limiting
│ ├── services/
│ │ ├── chat_store.py # Persistent group-chat allowlist
│ │ ├── cleanup.py # Temp file cleanup
│ │ ├── media_cache.py # file_id cache for instant resends
│ │ ├── minimal_store.py # Per-chat minimal-UI toggle
│ │ ├── topic_lock.py # Per-chat forum-topic restriction
│ │ └── uploader.py # Telegram upload (video/audio/document)
│ ├── types/
│ │ └── download.py # DownloadTask, DownloadStatus, MediaFormat
│ └── utils/
│ ├── logger.py # Structured logging
│ └── sanitizer.py # Filename sanitization
├── packaging/ # aur/ (PKGBUILD + unit), nix/ (package + home-manager module)
└── tests/ # pytest suite (see Testing)
Running via Docker Compose brings up the bot and a local Telegram Bot API server, which raises the upload limit from 50MB to 2GB.
- Bot token from @BotFather (
/newbot). - API ID + API hash from my.telegram.org/apps (needed by the local Bot API server).
cp .env.example .env
nano .envSet at minimum BOT_TOKEN, TELEGRAM_API_ID, TELEGRAM_API_HASH, and ALLOWED_USERS.
docker compose up -d # pulls the prebuilt image from GHCR
docker compose logs -f botdocker-compose.yml references the published image ghcr.io/antlis/tg-media-bot:latest, so no local build is needed. To build from source instead (e.g. for unreleased changes), use docker compose up -d --build.
To stop: docker compose down.
Port note: the local Bot API server publishes on host port
8082(8082:8081indocker-compose.yml). The bot reaches it over the internal Compose network ashttp://telegram-bot-api:8081, so the host port only matters if another service already occupies8081. Adjust if8082is also taken.
This path uses the standard Telegram Bot API (50MB upload limit) and supports Firefox cookies.
# System packages (Arch)
sudo pacman -S yt-dlp ffmpeg
# Python deps
python -m venv venv && source venv/bin/activate
pip install -r requirements.txt
# Configure
cp .env.example .env # set BOT_TOKEN, ALLOWED_USERS
# Run
python main.pySee Installation Guide for systemd service setup and Firefox cookie configuration.
Arch users can install a packaged build with a systemd service:
paru -S tg-media-bot # or: yay -S tg-media-bot
sudoedit /etc/tg-media-bot/.env # set BOT_TOKEN, ALLOWED_USERS
sudo systemctl enable --now tg-media-botPackaging files and publishing notes live in packaging/aur/.
nix run github:antlis/tg-media-bot # BOT_TOKEN etc. from the environment
nix build .#with-browser # + headless-Chromium fallback (large)
nix develop # dev shell with deps + pytestyt-dlp and ffmpeg are put on the bot's PATH by the wrapper; yt-dlp comes
from nixpkgs, so bump it with nix flake update. For home-manager, import
homeManagerModules.default and enable services.tg-media-bot — it runs a user
service, keeps state under ~/.local/state/tg-media-bot, and reads BOT_TOKEN
from environmentFile (extra env vars go in settings). nix flake check
builds the package and runs the test suite.
The local Bot API server (2 GB uploads) is a separate service — keep it in the
compose file (docker compose up -d telegram-bot-api) or run nixpkgs'
telegram-bot-api, then set API_SERVER_URL. Don't run the Docker bot and
the Nix bot at once — they'd poll the same token.
All settings are loaded from .env (see src/config/settings.py).
| Variable | Required | Default | Purpose |
|---|---|---|---|
BOT_TOKEN |
yes | — | Telegram bot token from @BotFather |
TELEGRAM_API_ID |
Docker only | — | From my.telegram.org/apps; used by the local Bot API server |
TELEGRAM_API_HASH |
Docker only | — | From my.telegram.org/apps; used by the local Bot API server |
ALLOWED_USERS |
recommended | empty (open) | Comma-separated Telegram user IDs allowed to use the bot |
API_SERVER_URL |
no | empty | Local Bot API base URL; set automatically in Docker |
TEMP_DIR |
no | /tmp/tg-media-bot |
Working directory for downloads |
MAX_PARALLEL_DOWNLOADS |
no | 3 |
Global concurrent download limit |
RATE_LIMIT_PER_USER |
no | 2 |
Concurrent downloads per user |
DOWNLOAD_TIMEOUT |
no | 3600 |
Per-download timeout in seconds |
LOG_LEVEL |
no | INFO |
DEBUG/INFO/WARNING/ERROR |
LOG_FILE |
no | empty (/data/... in Docker) |
Persist logs to a rotating file for a durable download record |
USE_BROWSER_COOKIES |
no | true |
Use browser cookies (forced off in Docker) |
BROWSER_NAME |
no | firefox |
Browser to read cookies from |
COOKIES_FILE |
no | empty | Path to a Netscape cookies.txt for authenticated downloads; takes precedence over browser cookies when present (the Docker way to auth) |
ALLOWED_CHATS_FILE |
no | empty | Path to a JSON file persisting group chats an allowed user has activated the bot in |
TOPIC_LOCK_FILE |
no | empty | Path to a JSON file persisting per-chat forum-topic locks set via /topic lock |
MEDIA_CACHE_FILE |
no | empty | Path to a JSON file caching file_ids so repeat URLs are resent instantly |
PROXY_URL |
no | empty | Proxy used only as a fallback retry when a download fails with a geo/region block (socks5h://… or http://…) |
ENABLE_BROWSER_FALLBACK |
no | true |
Try the headless-browser fallback when no yt-dlp extractor can handle a page (needs Chromium in the image — see INSTALL_BROWSER) |
BROWSER_FALLBACK_TIMEOUT |
no | 45 |
Seconds the fallback waits for the page to load and start playing |
BOT_API_HOST_PORT |
no | 8082 |
Docker only: host port for the local Bot API server |
YTDLP_AUTO_UPDATE |
no | true |
Docker only: refresh yt-dlp to the latest release on container start |
The headless-browser fallback needs a Chromium binary in the image, which is opt-in (it adds ~450 MB). Build with it by setting INSTALL_BROWSER=true in .env before docker compose build (the compose file forwards it as the INSTALL_BROWSER build arg), or docker build --build-arg INSTALL_BROWSER=true. On the bare-Python/AUR install, run playwright install chromium once. Without Chromium the fallback simply no-ops and the bot reports the original failure.
The bot is gated by ALLOWED_USERS. An outer_middleware on every message (src/bot/router.py) checks from_user.id against the allowlist before any handler runs:
- Empty / unset → open to everyone.
- Set → only listed IDs are served; others get a denial reply and are logged.
ALLOWED_USERS is read at startup. To add a user, append their ID and restart:
docker compose up -d bot # no rebuild needed — .env is read on startTo find a user's numeric ID, have them message @userinfobot.
The bot also works in group chats. When an allowed user uses it inside a group, that group is activated — its other members can then use the bot there too, without being individually allowlisted. Set ALLOWED_CHATS_FILE to persist activated groups across restarts (in Docker this defaults to /data/allowed_chats.json on the bot-logs volume); leave it unset to keep them in memory only.
Required one-time setup: Telegram bots ship with "group privacy" enabled, which stops the Bot API from forwarding plain messages (like a pasted link) sent in a group — only commands, @mentions, and replies to the bot get through. To let people just paste a link, disable it: message @BotFather → /mybots → your bot → Bot Settings → Group Privacy → Turn off, then remove and re-add the bot to any group it's already in (the change doesn't apply retroactively to existing memberships).
Forum topics: in a group with topics enabled, every reply (status messages, the downloaded file) is posted into whichever topic the request came from — never "General" — so different topics can be used for different purposes without their results bleeding into each other. To confine the bot to one topic per group (e.g. a "bots" topic, ignoring everything posted elsewhere), send /topic lock from inside that topic. /topic unlock lifts the restriction, and /topic status shows the current lock. This is per-chat, so different groups can each lock to their own topic (or not lock at all). /topic itself always works regardless of the current lock, so a chat can't get stuck; set TOPIC_LOCK_FILE to persist locks across restarts.
Some sources (Instagram, age-restricted videos, etc.) need a logged-in session. Two options:
- Bare Python: set
USE_BROWSER_COOKIES=trueandBROWSER_NAMEto pull cookies from your local browser. - Docker: export a Netscape
cookies.txt, drop it in./cookies/, and it's used per-download (COOKIES_FILE=/cookies/cookies.txt, mounted bydocker-compose.yml). A presentcookies.txttakes precedence over browser cookies.
By default the bot logs to stdout (docker compose logs bot), which resets when the container is recreated. Set LOG_FILE to also persist logs to a rotating file. In Docker this is wired by default to /data/tg-media-bot.log on the named bot-logs volume, so the record of every download (timestamp, user ID, URL, platform, filename, size) survives restarts and rebuilds.
# Tail the persistent log
docker compose exec bot tail -f /data/tg-media-bot.log
# Just the completed downloads
docker compose exec bot grep "Download completed" /data/tg-media-bot.logRotation keeps ~110 MB of history (10 × 10 MB files). For a privacy-minded setup, set LOG_LEVEL=WARNING to stop recording URLs/user IDs.
Send the bot any media URL (or up to 3 URLs in one message) and it downloads and returns the file. The active format mode (video/audio) applies to each download.
| Command | Description |
|---|---|
/start |
Welcome message |
/help |
Show help and the list of supported platforms |
/audio |
Switch to audio-only mode — downloads are converted to MP3 |
/video |
Switch to video mode (default) — includes video when available |
/formats <url> |
Show inline buttons to pick a download quality (Best / 1080p / 720p / 480p / Audio) |
/status |
Show your queued/active downloads and their task IDs |
/cancel <task_id> |
Cancel one of your active downloads (get the ID from /status) |
/minimal on|off |
Toggle minimal UI for this chat — no status/progress messages, no caption on media |
/topic lock|unlock|status |
Restrict the bot to one forum topic in this group (see Use in groups) |
Notes:
/audioand/videoset a per-user preference that persists until changed.- A download is queued per URL;
/statusreports each one's task ID, which/cancelconsumes. - Per-user concurrency is bounded by
RATE_LIMIT_PER_USER; the global cap isMAX_PARALLEL_DOWNLOADS. - Audio posts are a single message: the MP3 with embedded cover art, an album-art thumbnail in the player, title/artist/duration tags, and the source URL in the caption. (Telegram doesn't allow a standalone photo and an audio file in one post, so the cover rides along as the player thumbnail.)
- SoundCloud links are always audio — no need to send
/audiofirst. - Every post includes the original source URL as monospace, non-linked text — copyable, but Telegram won't turn it into a link or fetch a preview.
See Command Reference for full examples and sample responses.
The test suite uses pytest (with pytest-asyncio) and covers the pure-logic units — config parsing, sanitization, URL extraction, platform detection, yt-dlp command building, the caption builder, the allowlist middleware, the queue, and ffmpeg thumbnail resizing. No network or Telegram access is required; downloads and get_info are mocked.
# In a virtualenv with dev deps
pip install -r requirements-dev.txt
pytest
# Or, without managing a venv (uses uv)
uv run --with pytest --with pytest-asyncio --with aiogram --with structlog \
--with python-dotenv --with aiohttp pytestThe thumbnail tests need ffmpeg on PATH; they're skipped automatically if it's missing. yt-dlp is not required — the version probe is patched out in tests.
- Architecture Overview
- Installation Guide
- Command Reference
- Troubleshooting
- Contributor guide for AI agents
Downloads that use HLS/DASH fragments are fetched in parallel
(CONCURRENT_FRAGMENTS, default 16). This is faster and lets a download finish
before sites that expire segment URLs shortly after issuing them invalidate
them; if a high-resolution file still can't complete in that window, the bot
automatically retries at progressively lower quality.
Some sites build their media URL in JavaScript or hide it behind a site-specific API, so neither yt-dlp nor the headless-browser fallback can reach it. You can teach the bot about such a site with a plugin: a small Python file that turns a page URL into a media URL yt-dlp can download.
Drop .py files into a plugin directory — the ./plugins folder mounted into the
container by default (PLUGIN_DIR=/plugins), or any directory named by the
PLUGIN_DIR env var on a bare-metal install. Each plugin exposes two callables:
def match(url: str) -> bool: ... # claim the URLs you handle
async def resolve(url: str): ... # -> media_url | (media_url, referer) | ResolveResult | NoneFor a URL a plugin claims, its resolver runs before yt-dlp; the returned URL
then goes through the normal download / recode / upload path. A plugin that
returns None or raises is skipped, falling through to yt-dlp and the
headless-browser fallback. See examples/plugin_example.py
for a complete template.
A resolver may also assemble a playlist itself and return a local file://
URL (e.g. after rewriting a site's rotating segment hosts) — the bot enables
yt-dlp to read it automatically.
Plugins are not committed — the plugins/ directory is gitignored, so
site-specific extractors stay private to your deployment. Set ENABLE_PLUGINS=false
to ignore the directory entirely.
