Does a TypeSafe Jev-labelled customer-support stream spike before a brand admits an outage? Somewhat: at the same false-alarm rate it catches more incidents than raw tweet volume, with more lead time — but a good outage keyword list gets most of the way.
170,400 customer tweets to seven support accounts (Spectrum, Comcast, Hulu, Activision, Cox, Sprint,
Verizon) from Kaggle's Customer Support on Twitter
(Oct–Nov 2017), each labelled by TypeSafe Jev (typesafe/jev-1.13-20260917,
20 tweets per call, $1.84 total) with two typed questions: what is this tweet about (service down /
product bug / account & billing / question / other complaint / praise / other) and does it suggest the
problem is widespread (yes/no). Baselines on the same tweets: tweet volume, an outage keyword regex,
and VADER negative sentiment.
Like the release-note ground truth in the GitHub
study, the event is defined by the company's own words: a reply that says a problem is known and
shared — "there is an outage in your area", "a known issue we're investigating", "our technicians are
working to restore service". Generic help offers ("I'd like to help with any service issues, please DM")
and denials ("we do not have any reports of an outage") are excluded (src/jsp/ack.py).
- Precision of that rule, checked by reading 40 random matches: 37/40 (92.5 %). A first, looser version scored ~22 % and was discarded — most matches were "sorry for the service issues, please DM".
- An incident is an hour in which a brand posts ≥ 3 such acknowledgements and ≥ 4× its median hourly rate over the previous week (bursts within 12 h merged); its time is the brand's first acknowledgement in that hour. 62 incidents: Hulu 35, Activision 15, Spectrum 7, Cox 3, Comcast 1, Sprint 1, Verizon 0. The single Comcast incident is 2017-11-06 19:00 UTC — the nationwide Comcast internet outage reported that day, which is a useful sanity check on the definition.
A built-in circularity: brands acknowledge because complaints arrive, so every customer signal rises before acknowledgements to some extent. The question is not "does anything rise first" but which signal separates incident hours from ordinary busy hours — and at what false-alarm cost.
Every signal is an hourly z-score against the same brand's previous 7 days. Positives are the 3 hours before an incident's first acknowledgement (151 hours); negatives are hours ≥ 24 h from any incident (7,650 hours).
| signal | AUC | 95 % CI |
|---|---|---|
| Jev: widespread | 0.616 | [0.557, 0.671] |
| volume | 0.612 | [0.554, 0.668] |
| VADER negative | 0.598 | [0.543, 0.653] |
| Jev: service down | 0.585 | [0.526, 0.644] |
| Jev: down × widespread | 0.585 | [0.525, 0.645] |
| keyword | 0.572 | [0.509, 0.633] |
No signal ranks hours meaningfully better than another. All are weak (≈ 0.6).
Threshold = the 99th percentile of each signal in that brand's quiet hours (so every signal raises 1.76 false alarms per week); an incident is caught if the signal crosses it in the 12 h before the first acknowledgement.
| signal | incidents caught (of 62) | median lead before the brand's first "we know" |
|---|---|---|
| Jev: down × widespread | 17 | 4.1 h |
| Jev: service down | 16 | 3.6 h |
| keyword | 12 | 1.7 h |
| Jev: widespread | 12 | 2.4 h |
| volume | 10 | 2.4 h |
| VADER negative | 9 | 1.7 h |
Robustness (results/robustness.md) — the same comparison at three alarm
budgets, with an exact McNemar test on the same incidents:
| alarm threshold | Jev down × wide | volume | keyword | Jev vs volume (only-Jev / only-volume, p) | Jev vs keyword (p) |
|---|---|---|---|---|---|
| 95 % (≈ 8.8 alarms/week) | 36 | 25 | 31 | 13 / 2, p = 0.007 | 8 / 3, p = 0.23 |
| 99 % (≈ 1.8/week) | 17 | 10 | 12 | 9 / 2, p = 0.065 | 8 / 3, p = 0.23 |
| 99.5 % (≈ 0.9/week) | 11 | 6 | 11 | 6 / 1, p = 0.13 | 3 / 3, p = 1.00 |
The advantage over volume holds at every threshold and in both of the larger brands (Hulu 20/35 vs 12/35, Activision 7/15 vs 3/15 at 95 %). The advantage over a hand-built outage keyword list is consistent in direction but never significant.
A weak label (a brand may acknowledge in reply to one tweet and not another about the same outage), over 128,750 answered tweets, 2,051 of them acknowledged:
| signal | AUC |
|---|---|
| Jev: widespread | 0.630 |
| Jev: down × widespread | 0.607 |
| Jev: service down | 0.589 |
| keyword | 0.531 |
| VADER negative | 0.501 |
- Volume is a poor alarm. Support accounts get busy for many reasons (a game launch, a billing change, a TV event); Jev's service down × widespread filter removes most of that background, which is why it catches 1.7× the incidents at the same false-alarm rate, ~1.7 h earlier.
- A careful keyword list is almost as good. Outage tweets use a small vocabulary ("down", "no internet", "anyone else?"). Jev adds cases the list misses ("WTH is up with my wifi today?!? Had to use cell data") but not enough to be significant with 62 incidents.
- Sentiment is the wrong tool — the same finding as the GitHub-issues study: people reporting a broken service are often not "negative" in a lexicon's sense, and angry tweets are mostly about other things.
- With the GitHub-issues study this makes two aggregate-signal tests: where fixes are faster than complaints accumulate (claude-code, ~25 h), nothing works; where the organisation lags its customers (a 2017 support desk), a Jev-filtered stream gives a few hours of warning over raw volume.
- Multiple comparisons. Six signals × three thresholds; the single p < 0.01 result should be read together with the consistent direction across thresholds, not on its own.
- Acknowledgement ≠ outage start. The brand's first "we know" is a lower bound on when the outage began, so lead times are measured against the acknowledgement, not the failure.
- Customer stream coverage. Tweets are included if they mention the brand's handle or the brand replied to them; the dataset anonymises some handles to numbers, so a few unanswered complaints are missing.
- Hulu dominates (35 of 62 incidents), and several Hulu incidents are product bugs (live playback, a single show) rather than full outages.
- 2017 data, possible model familiarity with the big incidents (e.g. the Comcast outage).
The dataset has no explicit license on Kaggle and contains other people's tweets, so it is not
redistributed: data/labels.jsonl has tweet ids, brand, time and Jev's answers only.
uv venv --python 3.12 .venv && uv pip install --python .venv/bin/python -e ".[dev]"
.venv/bin/pytest -q
# download thoughtvector/customer-support-on-twitter into data/raw/ (kaggle CLI), then build the pickle:
.venv/bin/python -c "import pandas as pd; df=pd.read_csv('data/raw/twcs/twcs.csv'); df['t']=pd.to_datetime(df.created_at, format='%a %b %d %H:%M:%S %z %Y', utc=True); df.to_pickle('data/raw/twcs.pkl')"
OPENROUTER_API_KEY=... .venv/bin/python -m jsp.label # ~$1.8, or keep the committed labels
.venv/bin/python -m jsp.analyze && .venv/bin/python -m jsp.robustjev-search-rerank-eval · jev-cold-start-prior · jev-issue-pulse · jev-news-cold-start
MIT (code). Not affiliated with TypeSafe or any of the brands.