Fix Prometheus latency histograms mis-scaled by 1000x#693
Merged
Conversation
Every latency Histogram in FaustMetrics (events_runtime_latency, producer_send_latency, producer_error_send_latency, assign_latency, rebalance_done_consumer_latency, rebalance_done_latency, http_latency, consumer_commit_latency) was created with no explicit `buckets=`, so prometheus_client defaulted to its second-scale DEFAULT_BUCKETS (0.005-10.0). Every one of these histograms is fed millisecond values via Monitor.ms_since() (see PrometheusMonitor.on_*), so real latency observations (typically tens to hundreds of milliseconds) mostly land in the `+Inf` overflow bucket instead of a meaningful bucket -- silently mis-scaling every latency dashboard/alert built on Faust's Prometheus sensor by 1000x. Add MS_LATENCY_BUCKETS, a millisecond-scale bucket tuple with the same shape as Histogram.DEFAULT_BUCKETS (scaled x1000), and pass it explicitly to all 8 latency histograms. Fixes #260. Tests: parametrized check that every latency histogram's bucket boundaries equal MS_LATENCY_BUCKETS, plus a direct check that a 250ms observation lands in the 250ms bucket rather than the overflow bucket. Verified both fail without the fix (missing MS_LATENCY_BUCKETS import; KeyError on the 250.0 bucket, since the old default buckets only had a 0.25 boundary). flake8/black/isort clean; full tests/unit/sensors suite passes (excluding the pre-existing, unrelated test_datadog.py errors caused by the datadog package not being installed in this sandbox). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HHPL4VFWQRQPpjR1gXSKyL
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## master #693 +/- ##
=======================================
Coverage 94.15% 94.15%
=======================================
Files 104 104
Lines 11136 11137 +1
Branches 1201 1201
=======================================
+ Hits 10485 10486 +1
Misses 550 550
Partials 101 101 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
wbarnha
added a commit
that referenced
this pull request
Jul 19, 2026
Per review, the v0.12.0 changelog/release notes should describe only what is already on master, not work still in open PRs. - Remove the not-yet-merged items: the offset-commit data-loss fixes (#606/#707, #316/#692), the optional OpenTracing/OpenTelemetry extras (#685/#686, #688/#681), web_application_options (#704), and the reported-issue fix stack (#693-#703, #705). These will be added back as they merge. - Add a Dependencies section noting the current runtime/client libraries: mode-streaming >= 0.4.0, aiokafka >= 0.10.0 (compatible with recent 0.13/0.14 releases), the new confluent-kafka >= 2.0.0 for faust[ckafka], and the faust-cchardet fork replacing unmaintained cchardet. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HHPL4VFWQRQPpjR1gXSKyL
wbarnha
added a commit
to SpencerWhitehead7/faust
that referenced
this pull request
Jul 21, 2026
…st-streaming#708) * docs: prepare v0.12.0 release notes and changelog Resume the Keep a Changelog format (dormant since v0.8.10) with a v0.12.0 section, and add standalone GitHub release notes covering the changes since v0.11.3 plus the pending fix stack. Highlights: two offset data-loss fixes (faust-streaming#606/faust-streaming#707, faust-streaming#316/faust-streaming#692), the re-added confluent-kafka driver, Python 3.14 support (3.8/3.9 dropped), OpenTracing/OpenTelemetry made optional, and a live-broker CI harness. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HHPL4VFWQRQPpjR1gXSKyL * docs: drop closed codecov.yml PR (faust-streaming#683) from v0.12.0 notes PR faust-streaming#683 (codecov.yml with a 1% coverage threshold) was closed without merging, so remove it from the changelog and release notes to keep the v0.12.0 change list accurate. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HHPL4VFWQRQPpjR1gXSKyL * docs: derive Sphinx version from the package instead of a stale constant `docs/conf.py` hardcoded `version_dev='1.1'` / `version_stable='1.0'` - robinhood-era values that never matched faust-streaming's 0.x line, so the published GitHub Pages docs advertised the wrong version. Derive the documented major.minor from `faust.__version__` (which setuptools_scm resolves from the git tag), so the docs always report the real version and this can't silently drift between releases. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HHPL4VFWQRQPpjR1gXSKyL * docs: scope v0.12.0 notes to merged work; note dependency updates Per review, the v0.12.0 changelog/release notes should describe only what is already on master, not work still in open PRs. - Remove the not-yet-merged items: the offset-commit data-loss fixes (faust-streaming#606/faust-streaming#707, faust-streaming#316/faust-streaming#692), the optional OpenTracing/OpenTelemetry extras (faust-streaming#685/faust-streaming#686, faust-streaming#688/faust-streaming#681), web_application_options (faust-streaming#704), and the reported-issue fix stack (faust-streaming#693-faust-streaming#703, faust-streaming#705). These will be added back as they merge. - Add a Dependencies section noting the current runtime/client libraries: mode-streaming >= 0.4.0, aiokafka >= 0.10.0 (compatible with recent 0.13/0.14 releases), the new confluent-kafka >= 2.0.0 for faust[ckafka], and the faust-cchardet fork replacing unmaintained cchardet. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HHPL4VFWQRQPpjR1gXSKyL * docs: set v0.12.0 changelog date to 2026-07-19 Replace the UNRELEASED placeholder with the release date and drop the now-satisfied "set the date at tag time" note. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HHPL4VFWQRQPpjR1gXSKyL * Delete RELEASE_NOTES_v0.12.0.md --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Note: Before submitting this pull request, please review our contributing guidelines.
Description
Fixes #260.
Every latency
HistograminFaustMetrics(events_runtime_latency,producer_send_latency,producer_error_send_latency,assign_latency,rebalance_done_consumer_latency,rebalance_done_latency,http_latency,consumer_commit_latency) was created with no explicitbuckets=, soprometheus_clientdefaulted to its second-scaleDEFAULT_BUCKETS(0.005–10.0).Every one of these histograms is fed millisecond values via
Monitor.ms_since()(see the variousPrometheusMonitor.on_*methods). Real latency observations — typically tens to hundreds of milliseconds — mostly land in the+Infoverflow bucket instead of any meaningful bucket, silently mis-scaling every latency dashboard/alert built on Faust's Prometheus sensor by 1000x.Fix
Add
MS_LATENCY_BUCKETS, a millisecond-scale bucket tuple with the same shape asHistogram.DEFAULT_BUCKETSscaled x1000, and pass it explicitly to all 8 latency histograms.Tests
MS_LATENCY_BUCKETS.250ms observation lands in the250ms bucket rather than the overflow bucket.ImportErroron the missingMS_LATENCY_BUCKETSsymbol;KeyErroron the250.0bucket boundary, since the old default buckets only had a0.25boundary) — confirming the tests aren't vacuous.flake8,black --check,isort --check-onlyclean.tests/unit/sensorssuite passes (118 passed), excluding the pre-existingtest_datadog.pyerrors caused by thedatadogpackage not being installed in this sandbox — unrelated to this change.🤖 Generated with Claude Code
Generated by Claude Code