The prometheus handler's exposition has changed. Nothing fails to compile,
but the series it publishes are different, so anyone already scraping this
package will see renamed metrics and different staleness behaviour on upgrade.
The datadog, influxdb, otlp and veneur handlers are untouched.
What changed, and why:
-
Counters are now suffixed with
_total.Incr("requests")publishedapp_requestsand now publishesapp_requests_total. This is not only Prometheus naming convention: the OpenMetrics encoder keys the type line on the suffix, so a counter without it was published asunknown. A name that already ends in_totalis left alone — as judged on the name once exposed, not as received: a field is joined to its scope by an_, soIncr("requests.total")reports the bare fieldtotaland already publishesapp_requests_total, and.renders as_besides. Where the suffix would land on a name a sibling field already occupies — a counterhitsbeside a fieldhits_totalin the same scope — the suffix is dropped rather than the two being merged onto one series.svc.hitspublishessvc_hitsandsvc.hits_totalpublishessvc_hits_total; merging them put two samples with identical labels under one# TYPEline, of which a scraper keeps one. Which names the handler has seen decides this — not the order they arrived in, and not which are still live, so a counter does not change the name it publishes under when its sibling stops being reported and is swept by theMetricTimeoutcleanup. Queries and dashboards referring to the old names need updating. This includes the metrics the library reports about itself:go_version_valueandstats_version_valuebecomego_version_value_totalandstats_version_value_total. -
Histograms always emit a
+Infbucket. The handler allocated exactly one bucket per registered boundary and never appended an overflow bucket, so observations above the highest boundary were counted in_sumand_countbut landed in no bucket at all.histogram_quantile()returnsNaNunless the highest bucket is+Inf, so any histogram whose registered boundaries did not already end inmath.Inf(+1)could not be evaluated. -
Histograms with no registered boundaries fall back to
prometheus.DefaultBuckets.stats.Bucketsis empty by default and a miss returned a nil slice with no error, so such a histogram published_sumand_countwith no_bucketseries and nothing looked wrong. The defaults are the reference Prometheus client's, suited to latencies in seconds; they are a floor, not a substitute for choosing boundaries. This adds bucket series for histograms that previously published none: a histogram that published 2 series (_sum,_count) now publishes 14 (11 boundaries,+Inf,_sumand_count), per label set. This applies to every histogram the lookup finds nothing for; the registrations shipped byhttpstatsandnetstatsnow resolve (below), so those get their own boundaries rather than these. Counters and gauges are unaffected — they publish one series each, as before. -
Bucket
lelabels sort numerically. They compared as raw strings, which put+Inffirst and10ahead of2. -
# TYPEis declared once per scope, not once per field name. The dedup discarded the scope, so same-named fields coming from different engine prefixes looked like repeats and every one after the first was published untyped. Sub-engines derived withWithPrefixexist precisely so subsystems can reuse short field names, so this fired readily. The handler now tracks every family it has declared rather than comparing each metric against the previous one, so a family interrupted in the sorted output is no longer declared twice. A repeated# TYPEmakes the text format parser reject the entire exposition, so this also fixes scrapes that failed outright when a histogram shared a scope with a field sorting between its_bucketand_sumseries —latencynext tolatency_size, say. -
Timestamps are no longer exposed. The field is optional, and a series carrying one opts out of Prometheus stale-marker handling — the scraper kept serving the last value for five minutes after a series stopped being exported. The scraper now assigns scrape time.
MetricTimeoutis unaffected. -
Bucket registrations made by
httpstatsandnetstatsnow resolve. They never had.HistogramBuckets.Setsplits its argument on the last., but these registrations are written"http.message:body.bytes"— the:formsplitMeasureFieldused before b45dd38 ("fix typo insplitMeasureField()", Aug 2019) changed the separator. The strings were never updated, so each has been keyed on a name nothing looks up since. The string form cannot express these keys at all:Setcuts at the last., and every one of these fields carries one of its own (body.bytes,header.size,rtt.seconds). A second mismatch sat behind the first — the lookup uses the measure name after the engine prefix is attached, and a package registering frominit()cannot know that prefix. This changes the series published for anyone usinghttpstatsornetstats: those histograms published no_bucketseries before, and would otherwise have taken the seconds-scaleDefaultBucketsabove — eleven boundaries no byte count can reach. They now publish the byte and duration boundaries those packages declare.procstatsregisters"go.memstats:gc_pause.seconds"for fields declaredtype:"gauge", so that entry remains inert and is left alone. -
New:
HistogramBuckets.SetKey,SetUnprefixedandLookup.SetKeytakes theMeasureandFieldhalves directly, reaching keysSetcannot express.SetUnprefixedregisters a measure named without the prefix an engine will add to it, for packages registering frominit(), andLookupresolves those by dropping leading segments from the measure name after an exact lookup fails. Exact registrations always win, and only registrations made throughSetUnprefixedare matched that way: matching every registration by suffix cannot tell a derived measure from an unrelated one ending the same way.Setis unchanged, and theotlphandler — which reads the registry directly — is untouched. -
New:
Engine.SetBuckets(name, buckets...).Observetakes a name relative to the engine, whileHistogramBuckets.Setneeds the fully-qualified name, so registering buckets meant restating the engine prefix — and a mismatch was an ordinary map miss, indistinguishable from no registration at all.SetBucketsderives the key from the engine's own prefix, so callers pass the same string they pass toObserve. An ancestor can register for a sub-engine by naming the path to it —root.SetBuckets("db.latency", ...)coversroot.WithPrefix("db").Observe("latency", ...)— so oneinitfunction covers a whole tree. Buckets are not inherited: a sub-engine resolves only what was registered for its own prefix. AHandlerholding its own non-nilBucketsnever reads the global registry, soSetBucketshas no effect on it.Buckets.Setis unchanged.
The minimum supported Go version is now 1.26. The golang.org/x/* modules
(net, sys, sync, text) all declare go 1.26.0 as of their latest
releases, and stats depends on them both directly and transitively through
gRPC, so the whole module requires 1.26.
Update all dependencies. Most notably, gRPC moves to v1.83.2, which addresses GHSA-vp52-pcj8-j9qc, GHSA-2v4p-qf9q-27wj, and GHSA-qc2q-p7wx-3px3; the OpenTelemetry SDK and exporters move to v1.46.0.
The minimum supported Go version is now 1.25. The otlp package, previously
its own nested module, has been folded back into github.com/segmentio/stats/v5
so that go test ./..., go build ./..., and lint all cover it and dependency
bumps only need to happen in one place. Because the OpenTelemetry SDK and gRPC
both declare go 1.25.0, every consumer of stats now requires Go 1.25 and
transitively pulls in the OTel SDK and gRPC, even if otlp is unused.
go get github.com/segmentio/stats/v5/otlp is no longer needed (or valid) as a
separate module fetch; otlp is available as soon as you depend on
github.com/segmentio/stats/v5.
Apply 'go fix ./...' on the codebase. Several references to interface{} have
been replaced by any; there should be no API or performance differences.
Add full OpenTelemetry OTLP exporter support with official SDK integration.
New Feature: OpenTelemetry OTLP Exporter (SDKHandler)
The otlp package now includes a production-ready SDKHandler that uses the
official OpenTelemetry SDK with comprehensive support for modern observability
requirements:
- Dual Transport Support: Both gRPC and HTTP/Protobuf protocols
- Environment Variables: Support for the standard
OTEL_*environment variables includingOTEL_EXPORTER_OTLP_ENDPOINT,OTEL_EXPORTER_OTLP_PROTOCOL(and theOTEL_EXPORTER_OTLP_METRICS_PROTOCOLoverride), andOTEL_RESOURCE_ATTRIBUTES - Resource Detection: Host and process metadata plus environment attributes
by default; cloud (AWS, GCP, Azure) and Kubernetes detection is opt-in via the
contrib/detectors/*packages, with examples in the package README - All Metric Types: Counter, Gauge, and Histogram with proper semantics
- Exponential Histograms: Optional support for exponential histogram aggregation with configurable bucket size and scale
- Temporality Configuration: Configurable metric temporality (cumulative or delta) with cumulative as the default for Prometheus compatibility
- Tag Preservation: Automatic conversion of stats tags to OpenTelemetry attributes
- Production Ready: Thread-safe instrument caching, proper context handling, and comprehensive error handling
Usage Example:
import (
"context"
"github.com/segmentio/stats/v5"
"github.com/segmentio/stats/v5/otlp"
)
// Simple usage with environment variables
handler, err := otlp.NewSDKHandlerFromEnv(ctx)
if err != nil {
log.Fatal(err)
}
defer handler.Shutdown(ctx)
stats.Register(handler)
// Or with explicit configuration
handler, err := otlp.NewSDKHandler(ctx, otlp.SDKConfig{
Protocol: otlp.ProtocolGRPC,
EndpointURL: "http://localhost:4317",
})Implementation Details:
- Gauges use native
Float64Gaugeinstrument for instantaneous value recording - Background context for metric recording to prevent context cancellation issues
- Efficient two-level locking pattern for instrument caching (read locks in hot path)
- Cumulative temporality by default (Prometheus-compatible)
- Comprehensive documentation including cloud resource detector examples
SDKConfig.EndpointURLtakes a full URL including thehttp://orhttps://scheme- SDK defaults are used when unset (ExportInterval: 60s, ExportTimeout: 30s)
Deprecated:
otlp.Handleris now deprecated in favor ofotlp.SDKHandler(will be removed in v6.0.0)otlp.HTTPClientis now deprecated in favor ofotlp.SDKHandlerwithProtocolHTTPProtobuf(will be removed in v6.0.0)otlp.NewHTTPClient()is now deprecated (will be removed in v6.0.0)
The legacy Handler has been marked as Alpha since 2022 and has minimal to zero usage.
Migration is straightforward - see deprecation notices in code for examples.
See the otlp package documentation for complete details and examples.
When reporting go/stats versions, ensure that any user provided tags are included with the go-version and stats version reporting tags, to ensure better correlation on the Datadog side.
At the same time, don't report a timestamp with this metric, to avoid problems where Prometheus says that the metric is too old (we only report it one time).
More lenient sanitization for tag values, which can contain commas, slashes, and other characters that are not allowed in a metric name.
The sanitization process introduced in the v5.6.0 release did not properly handle characters in the extended Latin-1 supplement, e.g. "÷". This issue has been fixed in this release.
Remove golang.org/x/exp from the list of dependencies, in favor of the "slices" standard library package. This bumps the minimum supported Go version to 1.23.
The Datadog client should have faster performance, by copying less metric data before writing it to the socket.
Remove outdated README content and add debugstats examples.
Fix an error in the v5.6.0 release related to metric names with a longer buffer size.
-
In the
datadoglibrary: invalid characters in metric names, field names, or tag keys/values will be replaced with underscores. Accents and other diacritics will be removed (e.g. é will be replaced with 'e'). This change also improves performance of HandleMeasures by about 15-20%. #192 -
External calls to github.com/segmentio/objconv were replaced by imports of github.com/segmentio/stats/v5/util/objconv, which is a fork of the library (it has since been archived). This allowed to to substantially reduce the surface we import.
influxdb: calls to objconv/json were replaced with github.com/segmentio/encoding/json (a library with substantially more production experience and test coverage). #193
-
prometheus: Fix a deadlock from concurrent calls to collect() and cleanup(). Thank you Matthew Hooker for this contribution. #194
- Add logic to replace invalid unicode in the serialized datadog payload with '\ufffd'.
- Fix a regression in configured buffer size for the datadog client. Versions 5.0.0 to 5.3.1 would ignore any configured buffer size and use the default value of 1024. This could lead to smaller than expected writes and contention for the file handle.
- Fix version parsing logic.
-
Add
debugstatspackage that can easily print metrics to stdout. -
Fix file handle leak in procstats in the DelayMetrics collector.
go_version.valueandstats_version.valuewill be emitted the first time any metrics are sent from an Engine. Disable this behavior by settingGoVersionReportingEnabled = falsein code, or setting environment variableSTATS_DISABLE_GO_VERSION_REPORTING=true.
Add support for publishing stats to Unix datagram sockets (UDS).
In the httpstats package, replace misspelled http_req_content_endoing
and http_res_content_endoing with http_req_content_encoding and
http_res_content_encoding, respectively. This is a breaking change; any
dashboards or queries that filter on this tag must be updated.