Skip to content

About

๐Ÿ›ก๏ธ Telegram personal-account spam firewall โ€” real-time blocker for DMs & groups

Resources

Stars

4 stars

Watchers

0 watching

Forks

Latest commit

ย 

History

29 Commits

Folders and files

NameName
Last commit message
Last commit date
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 

Repository files navigation

๐Ÿ›ก๏ธ TeleShield

All-round Telegram spam firewall for your personal account
Private message blocking & group moderation โ€” all in one.

Python License Telethon Release

๐ŸŒ Product Page ยท ๐Ÿ“ฆ Download ยท ๐Ÿš€ Quick Start


๐Ÿ“‹ Overview

TeleShield is a full-featured Telegram spam defense system covering private DM blocking and group moderation. Unlike Bot API bots, it logs in as you โ€” handling the personal-account protection that bots simply cannot.


โœจ Features

Feature Command Description
DM scan --scan Scans recent non-contact conversations, matches spam patterns, blocks
Group scan --group-scan Scans recent group messages, kicks ad senders (admin required)
Live listener --listen Runs in the background, watching DMs + groups with second-level response
Scoring engine built-in Multi-signal spam scoring (tiered regex / link density / @ mentions / weak account traits / frequency) โ€” automatically decides block / flag / pass
Group behavior analysis built-in New members posting links instantly, message-flood ads โ†’ auto kick
Dry run --dry-run Safe preview โ€” shows results without blocking or kicking
Image OCR built-in Image-only ads โ†’ local Tesseract text recognition โ†’ pattern match; data never leaves your machine
Learn mode --learn <text> Manually flag spam; extracts keywords and generates regex patterns automatically
Block reports --report [day|week], --report-html Daily/weekly summaries + HTML visual reports (category stats / trends / details)
Blacklist --blacklist add|remove|list|import|export [id|file] Auto block/kick on sight; JSON exchange for community lists
Whitelist --whitelist add|remove|list|import|export [id|file] Whitelisted users are never scanned, blocked or kicked
Status panel --status Overview of blocks, kicks, lists and learn-mode state

๐Ÿš€ Quick Start

Prerequisites

  • Python 3.9+
  • Telegram API credentials from my.telegram.org/apps
  • (Optional) Tesseract OCR for image-ad detection

Install

Option 1: one-click install (recommended)

git clone https://github.com/c92d58/TeleShield.git
cd TeleShield
bash install.sh          # checks dependencies + pip install + creates data directory

Option 2: install via pip

pip install "teleshield[ocr]"
# system dependency for OCR (optional but strongly recommended)
apt install tesseract-ocr tesseract-ocr-chi-sim

First-time setup

teleshield --setup

Enter, in order:

  1. API ID โ€” from my.telegram.org/apps
  2. API Hash โ€” same source
  3. Phone number โ€” with country code, e.g. +852****5931
  4. Verification code โ€” Telegram sends it to your phone

Security note: credentials can also be written to ~/.teleshield/.env (TELESHIELD_API_ID / TELESHIELD_API_HASH / TELESHIELD_PHONE), and --setup reads them automatically. Never pass credentials as command-line arguments โ€” they would leak into your shell history.

After a successful login the session is saved under ~/.teleshield/; you won't need to log in again.

Basic usage

# โ”€โ”€โ”€ DM protection โ”€โ”€โ”€

# dry run first to preview the results
teleshield --dry-run

# actually scan recent DM spam
teleshield --scan

# start the live listener (background; full DM + group protection)
teleshield --listen

# โ”€โ”€โ”€ Group moderation โ”€โ”€โ”€

# scan all groups you moderate, kick ad senders
teleshield --group-scan

# โ”€โ”€โ”€ Learning & reports โ”€โ”€โ”€

# manually flag spam text so the program learns new patterns
teleshield --learn "ๅŠ ๅพฎไฟก abc123 ๆŠ•่ณ‡็ฉฉ่ณบๆ—ฅๅ…ฅ้Ž่ฌ"

# view block summaries
teleshield --report         # last 24 hours
teleshield --report week    # last 7 days + trend
teleshield --report-html    # generates an HTML report (~/.teleshield/report_week.html)

# โ”€โ”€โ”€ List management โ”€โ”€โ”€

# whitelist (never blocked)
teleshield --whitelist add 12345678
teleshield --whitelist list

# blacklist (blocked on sight)
teleshield --blacklist add 87654321

# community list exchange (JSON)
teleshield --blacklist export my_list.json
teleshield --blacklist import community.json
teleshield --blacklist remove 87654321

# full status
teleshield --status

๐Ÿ“– Full Command Reference

Command Description
--setup First-time setup (credentials via environment variables TELESHIELD_API_ID / TELESHIELD_API_HASH / TELESHIELD_PHONE or interactive input โ€” not arguments)
--scan Scans non-contact DMs and blocks spam
--dry-run Dry-run scan (does not actually block)
--listen Live listener โ€” DM blocking + group kicking + behavior analysis at once
--group-scan Scans moderated groups and kicks ad senders
--status Full status panel
--report [day|week] Block summary (default: day)
--report-html [day|week] Generates an HTML visual report (~/.teleshield/report_<period>.html)
--learn <text> Manually flags spam text and learns new patterns
--whitelist add|remove|list|import|export [user_id|file] Whitelist management (import/export via JSON)
--blacklist add|remove|list|import|export [user_id|file] Blacklist management (import/export via JSON)

๐Ÿ‘ฅ Group Moderation in Detail

TeleShield can automatically moderate any group where you have admin rights:

Scenario Behavior
--listen running New group message โ†’ auto-detect โ†’ kick the ad sender
--group-scan Scans the last 20 messages โ†’ batch kicks
Admins skipped Group admins and the creator are never affected
Whitelist skipped Whitelisted users are never kicked
3-day window Only messages from the last 3 days are examined

Kicks use ChatBannedRights(view_messages=True) โ€” Telegram's "block + remove", so the user cannot rejoin.


๐Ÿง  Learn Mode in Detail

When you meet a new spam pattern, use --learn to teach TeleShield:

# Example: flag an ad containing a URL
teleshield --learn "https://bit.ly/3XabcDe ๅ…่ฒป้ ˜ๅ– BTC"

# Example: flag a LINE/WeChat promotion
teleshield --learn "โž•ๅฎ˜ๆ–นLINE๏ผš@free888 ๆฏๆ—ฅๆŽจ่–ฆ้ฃ†่‚ก"

How learning works:

Step Description
๐Ÿ” Extract keywords filters stopwords, extracts high-value 2โ€“6 character keywords
๐Ÿงฉ Generate regex builds reusable patterns from URLs, IDs and other structures
๐Ÿ’พ Persistent storage saved in config.json, loaded on every start
๐Ÿ”„ Immediate effect is_spam() uses the new patterns right away

View accumulated learning results via --status.


๐Ÿง  Spam Scoring Engine (v0.10.0)

Replaces single-layer regex decisions โ€” multi-dimensional weighted scoring that automatically picks block / flag / pass:

Signal Points
Tiered regex hit (high risk / medium / low) +3 / +2 / +1 (semantic clusters stack independently)
Learned patterns (keywords / custom regex) +2
Link density (โ‰ฅ3 URLs) +1
Heavy @ mentions (โ‰ฅ2) +1
Weak account traits (no username/avatar/bio โ€” at least 2) +1
Message burst (โ‰ฅ5 in a short window) +1

Decision thresholds:

  • โ‰ฅ5 points โ†’ block (block in DMs / kick in groups)
  • โ‰ฅ3 points โ†’ flag (record only, no action โ€” watch it in --report)
  • <3 points โ†’ pass (let through, avoiding false positives)

Design note: a single weak signal (one t.me link, one "ๆŠ•่ณ‡/invest" word) will not cause a false positive; only combined signals escalate. Every verdict prints its score.


๐Ÿ” Group Behavior Analysis (v0.10.0)

In --listen mode, suspicious behavior patterns inside groups are monitored automatically:

Behavior Threshold Action
New member posts a link after joining within 5 minutes of joining auto kick
Flooding ads (including links) โ‰ฅ3 messages in 120 s auto kick
Mass @ mention flooding โ‰ฅ3 messages in 120 s auto kick

Behavior analysis is independent of text scoring โ€” even when link text matches no pattern (short URLs / images), the behavior still triggers. Records are marked with the reason [behavior].


๐Ÿ“Š Block Reports

# daily report
teleshield --report

# weekly report (includes daily trend)
teleshield --report week

Sample report:

๐Ÿ“Š Block summary โ€” last 24 hours
โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
   Total blocked: 12

   Source:
     โ€ข DMs: 10
     โ€ข Groups: 2

   Top 5 ad categories:
     โ€ข Investment scams: 5
     โ€ข Part-time job scams: 3
     โ€ข Adult content: 2
     โ€ข Gambling: 1
     โ€ข English spam: 1

   Daily trend:
     2026-07-14: 12

๐Ÿ” Spam Patterns

TeleShield ships with 30+ tiered regexes (covering both Traditional and Simplified Chinese), infinitely extendable through learn mode:

Level Category Examples
๐Ÿ”ด High risk Traffic funneling ๅŠ ๆˆ‘ๅพฎไฟกใ€ๅŠ Vใ€Vไฟกใ€vx
๐Ÿ”ด High risk Adult content ่ฃธ่Šใ€็ด„็‚ฎใ€ๆดไบคใ€ๆˆไบบ
๐Ÿ”ด High risk Gambling ่ณญๅšใ€ๅ…ญๅˆๅฝฉใ€ไธ‹ๆณจใ€casinoใ€betting
๐Ÿ”ด High risk Job-scam ๅ…ผ่ทใ€ๅˆทๅ–ฎใ€ๆ—ฅๅ…ฅใ€่บบ่ณบใ€ๅœจๅฎถๅทฅไฝœ
๐ŸŸ  Medium Investment ๆŠ•่ณ‡ใ€ๅธถๅ–ฎใ€่ทŸๅ–ฎใ€้‡ๅŒ–ใ€็ฉฉ่ณบใ€้ซ˜ๅ›žๅ ฑ
๐ŸŸ  Medium Selling ๅ‡บๅ”ฎใ€ๆ‰น็™ผใ€ไปฃ่ณผใ€ไปฃ็™ผใ€ๆธ…ๅ€‰
๐ŸŸ  Medium Fake offers ่จปๅ†Š้€ใ€ๅ…่ฒป้ ˜ใ€็ด…ๅŒ…ใ€ๅ„ชๆƒ ็ขผ
๐ŸŸ  Medium Engagement farming ้ปž่ฎšใ€ๅˆท็ฒ‰ใ€ๅˆท่ฎšใ€ๆผฒ็ฒ‰
๐ŸŸก Low Weak signals t.me linksใ€@ mentionsใ€tg accountsใ€click here

Both scripts covered: every category matches Traditional and Simplified Chinese (e.g. ่ณญๅš/่ตŒๅš, ็ฉฉ่ณบ/็จณ่ตš) โ€” ads from HK, TW and CN are all caught. False-positive protection: single characters (ๅ‡บ/ๅš/ๅฝฉ/ๅ”ฎ) were removed in favor of semantic clusters; a lone weak signal never triggers a verdict.


โš™๏ธ Security & Permissions

Authentication

  • Logs in via MTProto (Telegram's official protocol), not Bot API
  • The session file (~/.teleshield/user.session) is stored with Telethon's internal encryption and auto-chmod 600 (protects login credentials from other users on the same machine)
  • API credentials are stored only in local ~/.teleshield/config.json (atomic write + chmod 600) or .env
  • Credentials are never accepted as command-line arguments (no shell-history leaks) โ€” environment variables or interactive input only

Required permissions

Feature Required permission
DM blocking none (any account can block others)
Group kicking group admin (needs ban_users)
Image OCR local Tesseract, no network permission needed

Risk notes

  • The session file is your Telegram identity โ€” it is chmod 600 automatically; never delete or share it
  • Group kicks are irreversible โ€” preview with --group-scan dry first
  • All sensitive files (config.json / block_log.json / learned_patterns.json / .env) are chmod 600 automatically

๐Ÿ—‚๏ธ Project Structure

TeleShield/
โ”œโ”€โ”€ teleshield/            # Python package
โ”‚   โ”œโ”€โ”€ __init__.py        # version definition (single source)
โ”‚   โ”œโ”€โ”€ __main__.py        # python -m teleshield entry point
โ”‚   โ”œโ”€โ”€ cli.py             # command parsing & dispatch
โ”‚   โ”œโ”€โ”€ commands.py        # core actions (scan/listen/report/lists)
โ”‚   โ”œโ”€โ”€ config.py          # paths/.env/storage (atomic writes + chmod 600)
โ”‚   โ”œโ”€โ”€ patterns.py        # tiered spam patterns (severe/moderate/low, both scripts)
โ”‚   โ”œโ”€โ”€ scoring.py         # spam scoring engine (v0.10.0)
โ”‚   โ”œโ”€โ”€ behavior.py        # group behavior analysis (v0.10.0)
โ”‚   โ”œโ”€โ”€ ocr.py             # local Tesseract OCR (data stays local)
โ”‚   โ”œโ”€โ”€ client.py          # Telethon client factory
โ”‚   โ”œโ”€โ”€ systemd.py         # v0.11.0 โ€” service unit + logrotate generation
โ”‚   โ”œโ”€โ”€ updater.py         # v0.11.0 โ€” GitHub Release check + checksum verify
โ”‚   โ”œโ”€โ”€ ml.py              # v0.11.0 โ€” local Naive Bayes (char-bigram, no deps)
โ”‚   โ”œโ”€โ”€ dashboard.py       # v0.11.0 โ€” localhost web dashboard (http.server)
โ”‚   โ””โ”€โ”€ cloudsync.py       # v0.11.0 โ€” optional list sync to Cloudflare KV
โ”œโ”€โ”€ tests/                 # 378 pytest cases (all offline; no network in CI)
โ”œโ”€โ”€ .github/workflows/     # CI (ruff + pytest on 3 versions + build + auto Release)
โ”œโ”€โ”€ pyproject.toml         # packaging (pip install teleshield)
โ”œโ”€โ”€ install.sh             # one-click install script
โ”œโ”€โ”€ .env.example           # environment variable example
โ”œโ”€โ”€ README.md
โ””โ”€โ”€ LICENSE

~/.teleshield/             # auto-generated at runtime (chmod 600)
โ”œโ”€โ”€ user.session           # Telegram login session (encrypted + 600)
โ”œโ”€โ”€ config.json            # settings + learned patterns + lists
โ”œโ”€โ”€ learned_patterns.json  # learn-mode patterns, separate storage
โ”œโ”€โ”€ block_log.json         # block records + blocked message text (for local ML)
โ”œโ”€โ”€ ml_model.json          # trained classifier (delete it to disable that layer)
โ”œโ”€โ”€ cloud_backup_*.json    # pre-sync snapshots (rollback for --cloud)
โ”œโ”€โ”€ .env                   # credentials (optional)
โ”œโ”€โ”€ report_*.html          # HTML reports (--report-html)
โ””โ”€โ”€ logrotate.conf         # user-level rotation config (system mode uses /etc)

๐Ÿ› ๏ธ Operations (v0.11.0)

Run it as a service

teleshield --systemd install      # user-level (starts on login)
teleshield --systemd install --system   # system-level (starts on boot, needs root)
teleshield --systemd status
teleshield --systemd uninstall

The generated unit sets Restart=always with RestartSec=5, and a StartLimitIntervalSec/StartLimitBurst pair so a permanently broken config backs off instead of restarting in a tight loop. It is also hardened: NoNewPrivileges, PrivateTmp, ProtectSystem=full, ProtectHome=read-only, with ReadWritePaths limited to the data directory.

Log rotation uses copytruncate on purpose: the daemon keeps its log file open, and plain rotation would leave it writing to the moved inode.

Update in place

teleshield --update --check   # only report
teleshield --update           # download, verify sha256, install

The download is verified against the asset's sha256 digest before it is installed, and a failed check deletes the file. If the release publishes no digest, the update still proceeds but says so.

Local ML layer

teleshield --ml train         # train from block_log + labelled fixtures
teleshield --ml stats

A Multinomial Naive Bayes classifier in pure Python โ€” no scikit-learn, no numpy, no tokenizer package. Chinese has no word delimiters, so it uses character bigrams instead: ๅŠ ๆˆ‘ๅพฎไฟก yields ๅŠ ๆˆ‘, ๆˆ‘ๅพฎ, ๅพฎไฟก. URLs and long digit runs are normalized to <URL> / <NUM>, otherwise every ad's unique link becomes a one-off feature and nothing is learned.

Where it sits in the pipeline, and what it is not allowed to do:

severe regex โ†’ scoring โ†’ Jev (semantic) โ†’ thresholds โ†’ local ML

The ML layer only ever upgrades review to block. It never downgrades a block, and it never turns an allow into a block โ€” false positives are the expensive direction, so a weak model is only trusted to be more conservative. It also gets a say when the semantic layer is unreachable, which is otherwise a blind spot.

Measured with tools/calibrate.py --judge typesafe on the labelled set: blocked 21 vs 20 without it, with zero false positives and zero false allows, and the safe auto_block plateau widens from [0.30, 0.50] to [0.30, 0.80]. That is one extra catch on 23 cases โ€” real, small, and free. Delete ~/.teleshield/ml_model.json to switch the layer off entirely.

Dashboard

teleshield --dashboard              # http://127.0.0.1:8787
teleshield --dashboard --port 9000
teleshield --dashboard --allow-remote   # bind 0.0.0.0 โ€” no auth, trusted nets only

Serves block statistics (by day, by hour, by source, top reasons), the recent block log, and black/white list management. Built on http.server with zero dependencies, bound to localhost by default, and it re-reads the config before writing lists so it cannot clobber what the daemon just saved.

Cloud list sync (optional)

teleshield --cloud status
teleshield --cloud push      whitelist
teleshield --cloud pull      all --merge
teleshield --cloud sync      all

Keeps black/white lists in Cloudflare KV so several machines share one list. Requires all three TELESHIELD_CF_* variables; without them it says which one is missing instead of failing obscurely. Every destructive operation snapshots the previous state to cloud_backup_<list>.json first, so a wrong direction is recoverable. The API token is never written to a backup, a log, or an error message.


๐Ÿงฉ Roadmap

Done (v0.10.0):

  • Phase 1 engineering: modular refactor, .env configuration, pytest framework, CI/CD, pip packaging, install.sh
  • Phase 2 features: tiered rule engine, spam scoring, group behavior analysis, HTML reports, community list import/export
  • Security-audit fixes: session/config 600 permissions, env-based credentials, both-script coverage, false-positive regression tests

Done (v0.11.0):

  • Phase 3: systemd one-click deployment (daemon + log rotation + auto-restart)
  • Auto-update (checks GitHub Release + checksum)
  • ML classifier (local Naive Bayes trained on block_log)
  • Web dashboard (view block stats + manage lists)
  • Cloud list sync (optional, black/white lists โ†’ CF KV)

Planned:

  • Richer labelled set โ€” the calibration set needs real cases, especially ones the regex layer misses but the semantic layer catches
  • Re-run tools/calibrate.py once enough real traffic accumulates, and re-check the ml_block threshold against the enlarged set
  • Scheduler for recurring scans (cron/timer), so --scan need not be manual

๐Ÿ“„ License

MIT


TECXIA | DIGITAL STUDIO

Made with โค๏ธ by TECXIA ยท Keep Telegram clean

About

๐Ÿ›ก๏ธ Telegram personal-account spam firewall โ€” real-time blocker for DMs & groups

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages