Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

jev-support-pulse

Does a TypeSafe Jev-labelled customer-support stream spike before a brand admits an outage? Somewhat: at the same false-alarm rate it catches more incidents than raw tweet volume, with more lead time — but a good outage keyword list gets most of the way.

170,400 customer tweets to seven support accounts (Spectrum, Comcast, Hulu, Activision, Cox, Sprint, Verizon) from Kaggle's Customer Support on Twitter (Oct–Nov 2017), each labelled by TypeSafe Jev (typesafe/jev-1.13-20260917, 20 tweets per call, $1.84 total) with two typed questions: what is this tweet about (service down / product bug / account & billing / question / other complaint / praise / other) and does it suggest the problem is widespread (yes/no). Baselines on the same tweets: tweet volume, an outage keyword regex, and VADER negative sentiment.

Ground truth: the brand's own "we know"

Like the release-note ground truth in the GitHub study, the event is defined by the company's own words: a reply that says a problem is known and shared — "there is an outage in your area", "a known issue we're investigating", "our technicians are working to restore service". Generic help offers ("I'd like to help with any service issues, please DM") and denials ("we do not have any reports of an outage") are excluded (src/jsp/ack.py).

  • Precision of that rule, checked by reading 40 random matches: 37/40 (92.5 %). A first, looser version scored ~22 % and was discarded — most matches were "sorry for the service issues, please DM".
  • An incident is an hour in which a brand posts ≥ 3 such acknowledgements and ≥ 4× its median hourly rate over the previous week (bursts within 12 h merged); its time is the brand's first acknowledgement in that hour. 62 incidents: Hulu 35, Activision 15, Spectrum 7, Cox 3, Comcast 1, Sprint 1, Verizon 0. The single Comcast incident is 2017-11-06 19:00 UTC — the nationwide Comcast internet outage reported that day, which is a useful sanity check on the definition.

A built-in circularity: brands acknowledge because complaints arrive, so every customer signal rises before acknowledgements to some extent. The question is not "does anything rise first" but which signal separates incident hours from ordinary busy hours — and at what false-alarm cost.

Results

Every signal is an hourly z-score against the same brand's previous 7 days. Positives are the 3 hours before an incident's first acknowledgement (151 hours); negatives are hours ≥ 24 h from any incident (7,650 hours).

Ranking hours (AUC, bootstrap 95 % CI over incidents)

signal AUC 95 % CI
Jev: widespread 0.616 [0.557, 0.671]
volume 0.612 [0.554, 0.668]
VADER negative 0.598 [0.543, 0.653]
Jev: service down 0.585 [0.526, 0.644]
Jev: down × widespread 0.585 [0.525, 0.645]
keyword 0.572 [0.509, 0.633]

No signal ranks hours meaningfully better than another. All are weak (≈ 0.6).

Early warning at a fixed false-alarm budget

Threshold = the 99th percentile of each signal in that brand's quiet hours (so every signal raises 1.76 false alarms per week); an incident is caught if the signal crosses it in the 12 h before the first acknowledgement.

signal incidents caught (of 62) median lead before the brand's first "we know"
Jev: down × widespread 17 4.1 h
Jev: service down 16 3.6 h
keyword 12 1.7 h
Jev: widespread 12 2.4 h
volume 10 2.4 h
VADER negative 9 1.7 h

Robustness (results/robustness.md) — the same comparison at three alarm budgets, with an exact McNemar test on the same incidents:

alarm threshold Jev down × wide volume keyword Jev vs volume (only-Jev / only-volume, p) Jev vs keyword (p)
95 % (≈ 8.8 alarms/week) 36 25 31 13 / 2, p = 0.007 8 / 3, p = 0.23
99 % (≈ 1.8/week) 17 10 12 9 / 2, p = 0.065 8 / 3, p = 0.23
99.5 % (≈ 0.9/week) 11 6 11 6 / 1, p = 0.13 3 / 3, p = 1.00

The advantage over volume holds at every threshold and in both of the larger brands (Hulu 20/35 vs 12/35, Activision 7/15 vs 3/15 at 95 %). The advantage over a hand-built outage keyword list is consistent in direction but never significant.

Per tweet — did the brand answer this tweet with an acknowledgement?

A weak label (a brand may acknowledge in reply to one tweet and not another about the same outage), over 128,750 answered tweets, 2,051 of them acknowledged:

signal AUC
Jev: widespread 0.630
Jev: down × widespread 0.607
Jev: service down 0.589
keyword 0.531
VADER negative 0.501

Reading it

  1. Volume is a poor alarm. Support accounts get busy for many reasons (a game launch, a billing change, a TV event); Jev's service down × widespread filter removes most of that background, which is why it catches 1.7× the incidents at the same false-alarm rate, ~1.7 h earlier.
  2. A careful keyword list is almost as good. Outage tweets use a small vocabulary ("down", "no internet", "anyone else?"). Jev adds cases the list misses ("WTH is up with my wifi today?!? Had to use cell data") but not enough to be significant with 62 incidents.
  3. Sentiment is the wrong tool — the same finding as the GitHub-issues study: people reporting a broken service are often not "negative" in a lexicon's sense, and angry tweets are mostly about other things.
  4. With the GitHub-issues study this makes two aggregate-signal tests: where fixes are faster than complaints accumulate (claude-code, ~25 h), nothing works; where the organisation lags its customers (a 2017 support desk), a Jev-filtered stream gives a few hours of warning over raw volume.

Caveats

  • Multiple comparisons. Six signals × three thresholds; the single p < 0.01 result should be read together with the consistent direction across thresholds, not on its own.
  • Acknowledgement ≠ outage start. The brand's first "we know" is a lower bound on when the outage began, so lead times are measured against the acknowledgement, not the failure.
  • Customer stream coverage. Tweets are included if they mention the brand's handle or the brand replied to them; the dataset anonymises some handles to numbers, so a few unanswered complaints are missing.
  • Hulu dominates (35 of 62 incidents), and several Hulu incidents are product bugs (live playback, a single show) rather than full outages.
  • 2017 data, possible model familiarity with the big incidents (e.g. the Comcast outage).

Reproduce

The dataset has no explicit license on Kaggle and contains other people's tweets, so it is not redistributed: data/labels.jsonl has tweet ids, brand, time and Jev's answers only.

uv venv --python 3.12 .venv && uv pip install --python .venv/bin/python -e ".[dev]"
.venv/bin/pytest -q
# download thoughtvector/customer-support-on-twitter into data/raw/ (kaggle CLI), then build the pickle:
.venv/bin/python -c "import pandas as pd; df=pd.read_csv('data/raw/twcs/twcs.csv'); df['t']=pd.to_datetime(df.created_at, format='%a %b %d %H:%M:%S %z %Y', utc=True); df.to_pickle('data/raw/twcs.pkl')"
OPENROUTER_API_KEY=... .venv/bin/python -m jsp.label    # ~$1.8, or keep the committed labels
.venv/bin/python -m jsp.analyze && .venv/bin/python -m jsp.robust

Related

jev-search-rerank-eval · jev-cold-start-prior · jev-issue-pulse · jev-news-cold-start

MIT (code). Not affiliated with TypeSafe or any of the brands.

About

Does a Jev-labelled support-tweet stream spike before a brand admits an outage? At equal false alarms it catches 17 vs 10 incidents (volume), ~4h ahead; a good keyword list is almost as good.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages