All-round Telegram spam firewall for your personal account
Private message blocking & group moderation โ all in one.
TeleShield is a full-featured Telegram spam defense system covering private DM blocking and group moderation. Unlike Bot API bots, it logs in as you โ handling the personal-account protection that bots simply cannot.
| Feature | Command | Description |
|---|---|---|
| DM scan | --scan |
Scans recent non-contact conversations, matches spam patterns, blocks |
| Group scan | --group-scan |
Scans recent group messages, kicks ad senders (admin required) |
| Live listener | --listen |
Runs in the background, watching DMs + groups with second-level response |
| Scoring engine | built-in | Multi-signal spam scoring (tiered regex / link density / @ mentions / weak account traits / frequency) โ automatically decides block / flag / pass |
| Group behavior analysis | built-in | New members posting links instantly, message-flood ads โ auto kick |
| Dry run | --dry-run |
Safe preview โ shows results without blocking or kicking |
| Image OCR | built-in | Image-only ads โ local Tesseract text recognition โ pattern match; data never leaves your machine |
| Learn mode | --learn <text> |
Manually flag spam; extracts keywords and generates regex patterns automatically |
| Block reports | --report [day|week], --report-html |
Daily/weekly summaries + HTML visual reports (category stats / trends / details) |
| Blacklist | --blacklist add|remove|list|import|export [id|file] |
Auto block/kick on sight; JSON exchange for community lists |
| Whitelist | --whitelist add|remove|list|import|export [id|file] |
Whitelisted users are never scanned, blocked or kicked |
| Status panel | --status |
Overview of blocks, kicks, lists and learn-mode state |
- Python 3.9+
- Telegram API credentials from my.telegram.org/apps
- (Optional) Tesseract OCR for image-ad detection
Option 1: one-click install (recommended)
git clone https://github.com/c92d58/TeleShield.git
cd TeleShield
bash install.sh # checks dependencies + pip install + creates data directoryOption 2: install via pip
pip install "teleshield[ocr]"
# system dependency for OCR (optional but strongly recommended)
apt install tesseract-ocr tesseract-ocr-chi-simteleshield --setupEnter, in order:
API IDโ from my.telegram.org/appsAPI Hashโ same sourcePhone numberโ with country code, e.g.+852****5931Verification codeโ Telegram sends it to your phone
Security note: credentials can also be written to
~/.teleshield/.env(TELESHIELD_API_ID/TELESHIELD_API_HASH/TELESHIELD_PHONE), and--setupreads them automatically. Never pass credentials as command-line arguments โ they would leak into your shell history.
After a successful login the session is saved under ~/.teleshield/; you won't need to log in again.
# โโโ DM protection โโโ
# dry run first to preview the results
teleshield --dry-run
# actually scan recent DM spam
teleshield --scan
# start the live listener (background; full DM + group protection)
teleshield --listen
# โโโ Group moderation โโโ
# scan all groups you moderate, kick ad senders
teleshield --group-scan
# โโโ Learning & reports โโโ
# manually flag spam text so the program learns new patterns
teleshield --learn "ๅ ๅพฎไฟก abc123 ๆ่ณ็ฉฉ่ณบๆฅๅ
ฅ้่ฌ"
# view block summaries
teleshield --report # last 24 hours
teleshield --report week # last 7 days + trend
teleshield --report-html # generates an HTML report (~/.teleshield/report_week.html)
# โโโ List management โโโ
# whitelist (never blocked)
teleshield --whitelist add 12345678
teleshield --whitelist list
# blacklist (blocked on sight)
teleshield --blacklist add 87654321
# community list exchange (JSON)
teleshield --blacklist export my_list.json
teleshield --blacklist import community.json
teleshield --blacklist remove 87654321
# full status
teleshield --status| Command | Description |
|---|---|
--setup |
First-time setup (credentials via environment variables TELESHIELD_API_ID / TELESHIELD_API_HASH / TELESHIELD_PHONE or interactive input โ not arguments) |
--scan |
Scans non-contact DMs and blocks spam |
--dry-run |
Dry-run scan (does not actually block) |
--listen |
Live listener โ DM blocking + group kicking + behavior analysis at once |
--group-scan |
Scans moderated groups and kicks ad senders |
--status |
Full status panel |
--report [day|week] |
Block summary (default: day) |
--report-html [day|week] |
Generates an HTML visual report (~/.teleshield/report_<period>.html) |
--learn <text> |
Manually flags spam text and learns new patterns |
--whitelist add|remove|list|import|export [user_id|file] |
Whitelist management (import/export via JSON) |
--blacklist add|remove|list|import|export [user_id|file] |
Blacklist management (import/export via JSON) |
TeleShield can automatically moderate any group where you have admin rights:
| Scenario | Behavior |
|---|---|
--listen running |
New group message โ auto-detect โ kick the ad sender |
--group-scan |
Scans the last 20 messages โ batch kicks |
| Admins skipped | Group admins and the creator are never affected |
| Whitelist skipped | Whitelisted users are never kicked |
| 3-day window | Only messages from the last 3 days are examined |
Kicks use ChatBannedRights(view_messages=True) โ Telegram's "block + remove", so the user cannot rejoin.
When you meet a new spam pattern, use --learn to teach TeleShield:
# Example: flag an ad containing a URL
teleshield --learn "https://bit.ly/3XabcDe ๅ
่ฒป้ ๅ BTC"
# Example: flag a LINE/WeChat promotion
teleshield --learn "โๅฎๆนLINE๏ผ@free888 ๆฏๆฅๆจ่ฆ้ฃ่ก"How learning works:
| Step | Description |
|---|---|
| ๐ Extract keywords | filters stopwords, extracts high-value 2โ6 character keywords |
| ๐งฉ Generate regex | builds reusable patterns from URLs, IDs and other structures |
| ๐พ Persistent storage | saved in config.json, loaded on every start |
| ๐ Immediate effect | is_spam() uses the new patterns right away |
View accumulated learning results via --status.
Replaces single-layer regex decisions โ multi-dimensional weighted scoring that automatically picks block / flag / pass:
| Signal | Points |
|---|---|
| Tiered regex hit (high risk / medium / low) | +3 / +2 / +1 (semantic clusters stack independently) |
| Learned patterns (keywords / custom regex) | +2 |
| Link density (โฅ3 URLs) | +1 |
| Heavy @ mentions (โฅ2) | +1 |
| Weak account traits (no username/avatar/bio โ at least 2) | +1 |
| Message burst (โฅ5 in a short window) | +1 |
Decision thresholds:
- โฅ5 points โ
block(block in DMs / kick in groups) - โฅ3 points โ
flag(record only, no action โ watch it in--report) - <3 points โ
pass(let through, avoiding false positives)
Design note: a single weak signal (one t.me link, one "ๆ่ณ/invest" word) will not cause a false positive; only combined signals escalate. Every verdict prints its score.
In --listen mode, suspicious behavior patterns inside groups are monitored automatically:
| Behavior | Threshold | Action |
|---|---|---|
| New member posts a link after joining | within 5 minutes of joining | auto kick |
| Flooding ads (including links) | โฅ3 messages in 120 s | auto kick |
| Mass @ mention flooding | โฅ3 messages in 120 s | auto kick |
Behavior analysis is independent of text scoring โ even when link text matches no pattern (short URLs / images), the behavior still triggers. Records are marked with the reason [behavior].
# daily report
teleshield --report
# weekly report (includes daily trend)
teleshield --report weekSample report:
๐ Block summary โ last 24 hours
โโโโโโโโโโโโโโโโโโโโโโโโโโโโ
Total blocked: 12
Source:
โข DMs: 10
โข Groups: 2
Top 5 ad categories:
โข Investment scams: 5
โข Part-time job scams: 3
โข Adult content: 2
โข Gambling: 1
โข English spam: 1
Daily trend:
2026-07-14: 12
TeleShield ships with 30+ tiered regexes (covering both Traditional and Simplified Chinese), infinitely extendable through learn mode:
| Level | Category | Examples |
|---|---|---|
| ๐ด High risk | Traffic funneling | ๅ ๆๅพฎไฟกใๅ VใVไฟกใvx |
| ๐ด High risk | Adult content | ่ฃธ่ใ็ด็ฎใๆดไบคใๆไบบ |
| ๐ด High risk | Gambling | ่ณญๅใๅ ญๅๅฝฉใไธๆณจใcasinoใbetting |
| ๐ด High risk | Job-scam | ๅ ผ่ทใๅทๅฎใๆฅๅ ฅใ่บบ่ณบใๅจๅฎถๅทฅไฝ |
| ๐ Medium | Investment | ๆ่ณใๅธถๅฎใ่ทๅฎใ้ๅใ็ฉฉ่ณบใ้ซๅๅ ฑ |
| ๐ Medium | Selling | ๅบๅฎใๆน็ผใไปฃ่ณผใไปฃ็ผใๆธ ๅ |
| ๐ Medium | Fake offers | ่จปๅ้ใๅ ่ฒป้ ใ็ด ๅ ใๅชๆ ็ขผ |
| ๐ Medium | Engagement farming | ้ป่ฎใๅท็ฒใๅท่ฎใๆผฒ็ฒ |
| ๐ก Low | Weak signals | t.me linksใ@ mentionsใtg accountsใclick here |
Both scripts covered: every category matches Traditional and Simplified Chinese (e.g. ่ณญๅ/่ตๅ, ็ฉฉ่ณบ/็จณ่ต) โ ads from HK, TW and CN are all caught. False-positive protection: single characters (ๅบ/ๅ/ๅฝฉ/ๅฎ) were removed in favor of semantic clusters; a lone weak signal never triggers a verdict.
- Logs in via MTProto (Telegram's official protocol), not Bot API
- The session file (
~/.teleshield/user.session) is stored with Telethon's internal encryption and auto-chmod 600 (protects login credentials from other users on the same machine) - API credentials are stored only in local
~/.teleshield/config.json(atomic write + chmod 600) or.env - Credentials are never accepted as command-line arguments (no shell-history leaks) โ environment variables or interactive input only
| Feature | Required permission |
|---|---|
| DM blocking | none (any account can block others) |
| Group kicking | group admin (needs ban_users) |
| Image OCR | local Tesseract, no network permission needed |
- The session file is your Telegram identity โ it is chmod 600 automatically; never delete or share it
- Group kicks are irreversible โ preview with
--group-scan dryfirst - All sensitive files (config.json / block_log.json / learned_patterns.json / .env) are chmod 600 automatically
TeleShield/
โโโ teleshield/ # Python package
โ โโโ __init__.py # version definition (single source)
โ โโโ __main__.py # python -m teleshield entry point
โ โโโ cli.py # command parsing & dispatch
โ โโโ commands.py # core actions (scan/listen/report/lists)
โ โโโ config.py # paths/.env/storage (atomic writes + chmod 600)
โ โโโ patterns.py # tiered spam patterns (severe/moderate/low, both scripts)
โ โโโ scoring.py # spam scoring engine (v0.10.0)
โ โโโ behavior.py # group behavior analysis (v0.10.0)
โ โโโ ocr.py # local Tesseract OCR (data stays local)
โ โโโ client.py # Telethon client factory
โ โโโ systemd.py # v0.11.0 โ service unit + logrotate generation
โ โโโ updater.py # v0.11.0 โ GitHub Release check + checksum verify
โ โโโ ml.py # v0.11.0 โ local Naive Bayes (char-bigram, no deps)
โ โโโ dashboard.py # v0.11.0 โ localhost web dashboard (http.server)
โ โโโ cloudsync.py # v0.11.0 โ optional list sync to Cloudflare KV
โโโ tests/ # 378 pytest cases (all offline; no network in CI)
โโโ .github/workflows/ # CI (ruff + pytest on 3 versions + build + auto Release)
โโโ pyproject.toml # packaging (pip install teleshield)
โโโ install.sh # one-click install script
โโโ .env.example # environment variable example
โโโ README.md
โโโ LICENSE
~/.teleshield/ # auto-generated at runtime (chmod 600)
โโโ user.session # Telegram login session (encrypted + 600)
โโโ config.json # settings + learned patterns + lists
โโโ learned_patterns.json # learn-mode patterns, separate storage
โโโ block_log.json # block records + blocked message text (for local ML)
โโโ ml_model.json # trained classifier (delete it to disable that layer)
โโโ cloud_backup_*.json # pre-sync snapshots (rollback for --cloud)
โโโ .env # credentials (optional)
โโโ report_*.html # HTML reports (--report-html)
โโโ logrotate.conf # user-level rotation config (system mode uses /etc)
teleshield --systemd install # user-level (starts on login)
teleshield --systemd install --system # system-level (starts on boot, needs root)
teleshield --systemd status
teleshield --systemd uninstallThe generated unit sets Restart=always with RestartSec=5, and a
StartLimitIntervalSec/StartLimitBurst pair so a permanently broken config
backs off instead of restarting in a tight loop. It is also hardened:
NoNewPrivileges, PrivateTmp, ProtectSystem=full, ProtectHome=read-only,
with ReadWritePaths limited to the data directory.
Log rotation uses copytruncate on purpose: the daemon keeps its log file open,
and plain rotation would leave it writing to the moved inode.
teleshield --update --check # only report
teleshield --update # download, verify sha256, installThe download is verified against the asset's sha256 digest before it is
installed, and a failed check deletes the file. If the release publishes no
digest, the update still proceeds but says so.
teleshield --ml train # train from block_log + labelled fixtures
teleshield --ml statsA Multinomial Naive Bayes classifier in pure Python โ no scikit-learn, no
numpy, no tokenizer package. Chinese has no word delimiters, so it uses
character bigrams instead: ๅ ๆๅพฎไฟก yields ๅ ๆ, ๆๅพฎ, ๅพฎไฟก. URLs and long
digit runs are normalized to <URL> / <NUM>, otherwise every ad's unique link
becomes a one-off feature and nothing is learned.
Where it sits in the pipeline, and what it is not allowed to do:
severe regex โ scoring โ Jev (semantic) โ thresholds โ local ML
The ML layer only ever upgrades review to block. It never downgrades a
block, and it never turns an allow into a block โ false positives are the
expensive direction, so a weak model is only trusted to be more conservative.
It also gets a say when the semantic layer is unreachable, which is otherwise a
blind spot.
Measured with tools/calibrate.py --judge typesafe on the labelled set:
blocked 21 vs 20 without it, with zero false positives and zero false allows,
and the safe auto_block plateau widens from [0.30, 0.50] to [0.30, 0.80].
That is one extra catch on 23 cases โ real, small, and free. Delete
~/.teleshield/ml_model.json to switch the layer off entirely.
teleshield --dashboard # http://127.0.0.1:8787
teleshield --dashboard --port 9000
teleshield --dashboard --allow-remote # bind 0.0.0.0 โ no auth, trusted nets onlyServes block statistics (by day, by hour, by source, top reasons), the recent
block log, and black/white list management. Built on http.server with zero
dependencies, bound to localhost by default, and it re-reads the config before
writing lists so it cannot clobber what the daemon just saved.
teleshield --cloud status
teleshield --cloud push whitelist
teleshield --cloud pull all --merge
teleshield --cloud sync allKeeps black/white lists in Cloudflare KV so several machines share one list.
Requires all three TELESHIELD_CF_* variables; without them it says which one
is missing instead of failing obscurely. Every destructive operation snapshots
the previous state to cloud_backup_<list>.json first, so a wrong direction is
recoverable. The API token is never written to a backup, a log, or an error
message.
Done (v0.10.0):
- Phase 1 engineering: modular refactor, .env configuration, pytest framework, CI/CD, pip packaging, install.sh
- Phase 2 features: tiered rule engine, spam scoring, group behavior analysis, HTML reports, community list import/export
- Security-audit fixes: session/config 600 permissions, env-based credentials, both-script coverage, false-positive regression tests
Done (v0.11.0):
- Phase 3: systemd one-click deployment (daemon + log rotation + auto-restart)
- Auto-update (checks GitHub Release + checksum)
- ML classifier (local Naive Bayes trained on block_log)
- Web dashboard (view block stats + manage lists)
- Cloud list sync (optional, black/white lists โ CF KV)
Planned:
- Richer labelled set โ the calibration set needs real cases, especially ones the regex layer misses but the semantic layer catches
- Re-run
tools/calibrate.pyonce enough real traffic accumulates, and re-check theml_blockthreshold against the enlarged set - Scheduler for recurring scans (cron/timer), so
--scanneed not be manual
