Real-time harassment detection and stalking defense for Gmail, X/Twitter, Instagram and Facebook, as a Chrome extension.
Text is classified by a model running entirely on your own machine. No cloud inference, no API key, and no message content leaves the computer.
- 🔍 Local toxicity detection — Laya, a System One decision model, running locally via ONNX Runtime
- 🛡️ Stalking & threat signals — separate probes for stalking, threats, doxxing, sexual coercion, self-harm and targeted harassment, not just a single toxicity score
- 🖼 Image moderation via the Sightengine API (the one part still remote)
- 🔐 Nothing leaves the machine for text — no model provider, no key, no per-token cost
- ♻ Graceful degradation — if Redis is down, moderation still runs uncached rather than leaving you unprotected
- ⚡ Manifest V3 Chrome extension
- 🔒 Local-only API — CORS accepts extension origins only, so no website can drive it
| Component | Technology | Role |
|---|---|---|
| Extension | Chrome MV3 | Captures page content, renders badges |
| Text moderation | Laya via @receptron/laya (ONNX) |
8-question classification, on-device |
| Image moderation | Sightengine API | Nudity, violence, gore, self-harm |
| Backend | Node.js 20, Express | Orchestration, caching, health |
| Cache | Redis (optional) | SHA-256 keyed result cache, 5 minute TTL |
git clone https://github.com/akshayvibe/SafeSphere.ai.git
cd SafeSphere.ai/backend
cp .env.example .env # add Sightengine keys if you want image moderation
docker compose up --buildThen load the extension from the extension/ folder in this repo and open Gmail.
⚠️ Load the extension from this repository, not a downloaded copy. ASafeSphere.ai-mainfolder in your Downloads directory will look identical but runs stale code, and Chrome will happily keep using it. The console line[SafeSphere] content script loaded (extension v1.1.0)confirms you have the current version.
Running without Docker, and other detail
- Node.js v20+ — required by the ONNX runtime
- ~2 GB RAM — the model stays resident once loaded
- ~2 GB disk — the ONNX bundle downloads from Hugging Face on first run
- Chrome
- Sightengine credentials — only if you want image moderation; text needs nothing
cd backend
npm install
npm startonnxruntime-node and @huggingface/tokenizers ship native prebuilds, so their
install scripts must run. They are allow-listed under allowScripts in
package.json; if your npm blocks them:
npm approve-scripts --no-allow-scripts-pin onnxruntime-node @huggingface/tokenizersRedis is optional natively too — it defaults to redis://127.0.0.1:6379 and the
backend simply runs uncached if it is absent.
- The image is
node:20-slim, not alpine: the ONNX and tokenizer bindings are glibc-only and fail on musl. - Give Docker Desktop ≥ 4 GB of memory or the container is OOM-killed mid-load.
- First start downloads ~1.7 GB into the named
laya-modelsvolume. It is not resumable — an interrupted first run starts over. - Helpers:
npm run docker:up,npm run docker:logs.
/app/node_modules is an anonymous volume, kept separate from the host so the
container uses Linux ONNX binaries instead of macOS ones. Compose reuses an
existing anonymous volume when recreating a container, so a stale one can survive a
rebuild and shadow new packages.
npm run docker:refreshRemoves containers and their anonymous volumes, then rebuilds. The named
laya-models volume is left alone, so the model is not re-downloaded.
-
The content script captures visible text and image URLs.
-
It messages the background service worker, which POSTs to
http://127.0.0.1:3000/moderation/analyzeComment. Content scripts cannot call the API directly — in MV3 theirfetch()is CORS-checked against the page origin, so Gmail would block the request. -
The controller checks Redis (
moderation-laya-text:<sha256>, 5 min TTL). -
On a miss, Laya answers eight questions in one forward pass — seven yes/no probes (
noul) plus a five-levelseverityscore:is_toxic·is_harassment·is_threat·is_stalking·is_sexual·is_self_harm·is_doxxing·severity -
Each probe is weighted, and the strongest signal sets the rating — one credible threat is not diluted by six quiet categories. Severity can raise a score but never lower a concrete signal.
-
Results are cached and returned; the extension badges each item and blurs anything rated above 7.
Images are still sent to Sightengine. There is no local image model in this stack, so image URLs leave the machine. If that matters for your threat model,
extension/contentScript_image.jsis the piece to replace.
- Text never leaves the machine. Classification runs in-process via ONNX Runtime. No third-party inference, no API key.
- Results are cached for 5 minutes only, keyed by SHA-256 of the text.
- The local API accepts extension origins and localhost only — an arbitrary website cannot post to it.
- Sightengine credentials live in
backend/.env, never in the extension bundle. - There is no cloud storage, no telemetry and no audit log in this codebase. Nothing is persisted beyond the 5-minute Redis cache.
Redis is treated as optional: a cache miss and a cache outage are
indistinguishable to the caller, so losing Redis costs a slower response, not your
protection. A circuit breaker skips the cache for REDIS_BACKOFF_MS (30s default)
so an outage costs one log line instead of a failed round trip per request.
The model is treated differently, because no fallback keeps a user safe. If it
cannot load, the endpoint returns 503 and the extension shows an error rather
than implying the content was cleared.
Read this before relying on it for an actual safety decision.
- The rating weights are policy, not calibration.
SIGNAL_WEIGHTSinlayaClient.jsis hand-authored and was tuned against a small set of examples after finding that an earlieris_sexualwording mislabelled benign mail. It has not been calibrated on a real labelled dataset. Treat the cutoffs as provisional. - Question wording materially changes results. The same probe reworded can move
a score from 0.62 to 0.27 on identical input. The
is_sexualquestion carries a comment explaining this — do not reword it casually. - State is truncated at 512 tokens, so very long messages are only partly classified.
- The English checkpoint is the default. For other languages set
LAYA_SUBFOLDER=multilingual; the English checkpoint is substantially weaker there. - ~600 ms per classification warm, so the extension deliberately scans a small number of items per page rather than everything.
- Image moderation is remote and dependent on Sightengine.
| Method | Endpoint | Body | Response |
|---|---|---|---|
GET |
/health |
– | Status |
POST |
/moderation/analyzeComment |
{ text } or { texts: [] } |
Moderation result |
POST |
/moderation/analyzeImage |
{ imageUrl } |
Moderation result |
A result contains rating (1-10), toxic, message, contentTypes, confidence,
category, per-signal signals, and the raw model output.
GET /health reports textModel.state as idle / loading / ready / error;
only ready is ok, and cache as connected / connecting / bypassed.
{
"status": "ok",
"textProvider": "laya-local",
"imageProvider": "sightengine",
"textModel": { "state": "ready", "error": null },
"cache": "connected"
}
Issues and pull requests are welcome. If you change the questions in
layaClient.js, please re-test against benign text — a wording change can silently
turn a working classifier into a false-positive machine.
- Laya by Convai Innovations (Apache 2.0)
- @receptron/laya for the ONNX Node runtime
- Sightengine for image moderation
Let's make the internet safer for everyone, starting with protecting women online.