This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
MX Map is a DNS-based email provider classifier for European municipalities (92 countries, ~20,000 municipalities). It runs a 3-stage async pipeline that produces data.json, which powers an interactive Leaflet.js map showing where municipalities host their official email. Forked from mxmap.ch (Swiss municipalities). Covers all 27 EU member states plus 20 non-EU European countries.
The classifier above (92 countries) is the inherited infrastructure; since 2026 the active focus is Italy, as the data engine behind the Osservatorio Nazionale Sovranità Digitale. Read this section before touching any public-facing or KPI logic — it captures the framing, the analytical model, and the decisions/assumptions/constraints accumulated so far, so a new contributor can pick up the project.
- MxMap (this repo) = the data engine. Classifies the email provider of ~22,987 Italian PA entities (from IndicePA) →
data.json+ public static artifacts at the deploy root:kpi.json(aggregate KPIs) andreport.json(the structured report), plus the pagesstatistiche.html,report.html,storia.html(gated),anomalie.html,methodology.html. - Osservatorio Nazionale Sovranità Digitale = the presentation/advocacy layer. A separate Hugo repo (
fpietrosanti/osservatorio-nazionale-sovranita-digitale; localC:\Users\admin\osservatorio-nazionale-sovranita-digitale). Stakeholder-oriented; it fetches our artifacts (aupdate-kpi.ymlAction downloadskpi.json; same pattern planned forreport.json) and renders them with its own styling. - Contract: MxMap produces, the Osservatorio consumes. Artifacts are static files at the deploy root — no API, no auth, CC BY-SA 4.0. Measurement belongs here; editorial/presentation belongs to the Osservatorio. If you change an artifact's schema, update
docs/STATS_KPI.mdand tell the Osservatorio side. - Custom domain:
mxmap.it(CNAME). Canonical URLs arehttps://mxmap.it/...; oldfpietrosanti.github.io/mxmap.it/...redirect there. - Public brand: the site presents publicly as "Osservatorio Nazionale Sovranità Digitale" — the visible
<title>, header,og:site_name,application-nameand JSON-LDnameall use it. "MxMap Italia" / "MxMap.it" is kept only as the repo/tool name, themxmap.itdomain, and an SEOalternateName/keyword (so "mxmap" searches still resolve). Note: the separate advocacy siteosservatorio.mxmap.itshares the same brand — they are two fronts of the one observatory (data/map here, reports/advocacy there).
Upstream — official successor project: mxmap/secassure2026
- The original upstream
davidhuser/mxmap(mxmap.ch) is officially "succeeded by mxmap/secassure2026" (its repo description says so). The successor is DACH-scoped (DE/AT/CH), titled "Email Provider Dependencies and Email Security in Municipalities", with a refactored architecture:src/mail_municipalities/{core,domain_resolver,provider_classification,security_analysis,analysis}+src/security_test(Kotlin/Docker DANE/DNSSEC + SPF/DMARC scanner, needs outbound port 25) + a publicsecurity.html. - Policy: follow, never merge. Our fork diverged (IT pipeline, IndicePA, SEO pages). New upstream work is reviewed and features are cherry-picked/re-implemented where they fit. The weekly workflow
.github/workflows/upstream-watch.ymldigests new upstream commits into an issue labeledupstream-watch(same auto-issue pattern as the nightly). Dev convenience: the laptop clone hasgit remote upstream→git fetch upstreamto inspect/diff. - Candidates spotted at first review (2026-07-26): (a) DMARC/SPF policy evaluation (
security_analysis, DNS-only — cheap to port and very relevant for a PA observatory); (b) the DANE/DNSSEC scanner (port 25 — could run on the Scaleway server); (c) the probe/signature classifier refactor (low priority). NB: our google._domainkey TXT fix (#17) was implemented downstream and validated; upstream ticket mxmap/mxmap-ch#28 sits on the OLD repo — if we engage upstream again, re-file against secassure2026.
sovereignty_of(provider)+material_row(entity)inhistoricize.pyare canonical —stats,kpi,reportall reuse them. Never re-derive sovereignty elsewhere.- 6 MxMap buckets:
USA (CLOUD Act),Altri provider esteri,Italia — Cloud sovrano,Italia — Provider commerciali,Italia — Infrastruttura autonoma,Sconosciuto. - 4 Osservatorio buckets (
kpi.provider_to_sov4):extra_eu(USA + non-European foreign, e.g. Zoho/Yandex),eu_non_it(European non-IT providers — OVH, Hetzner, IONOS, Scaleway, Gandi, Infomaniak; CH/UK counted as European for simplicity — viaEU_NON_IT_PROVIDERS, #21),it(the 3 Italian buckets),unknown. In the 6-bucket model these European providers sit inAltri provider esteri; theeu_non_it/extra_eusplit happens only in the 4-bucket. - ISD — Indice di Sovranità Digitale: % entities under Italian jurisdiction, computed over classified (unknowns excluded), on provider sovereignty (legal control).
mx_jurisdiction(where the MX physically sits) is a complementary technical indicator — the gap between the two is itself a finding.
- By GROUP (works): 15 citizen-friendly clusters keyed on the
bfscategory code. NB: thebfsuses the project's own codes (COM=Comuni,PRO=Province,CMM,REG), not IndicePAL6/L5— the 54 real codes are mapped instats.CLUSTERS(full coverage, noother). - By AREA (active): every IT entity now carries
regione/provincia/comune/macroarea, injected byscripts/enrich_geo.pyvia the structural geographic enrichment — it resolves the seed's cleanipa_codice_comune_istat(the comune-sede, present 100%) against the official ISTAT crosswalk (data/istat_comuni.json) ingeo.py. Coverage 20/20 regions, 100% (Sardegna uses pre-2016 legacy province prefixes 112–119, mapped to current provinces; seegeo.SARDEGNA_LEGACY_PROV). This bypasses IndicePA's dirtyregionfield, not derived from it. Consumed bystats.compute_by_region→stats_by_region.json+ the report "Analisi per aree" section (regional ISD league table + macroarea summary + most-sovereign/most-exposed extremes) + the Statistiche page (regional stacked bars). - GOTCHA — the map's region/province grouping uses
regione/provincia, NOT the legacy seed fields. The interactive map (index.html) historically grouped the Regioni level bycantonand the Province level bydistrict— both stale seed fields that the geo work left only ~33–34% populated (andcantonis full of garbage like assembly/entity names). Reading them directly silently broke both choropleths (72 fake "regions", mostly-empty provinces).build_frontend.pynow keys IT regions byregione(20) and provinces by theprovinciacar-plate sigla (107);index.html'sgetGroupKeyreads the clean fields andmatchGroupFeaturematches the topo polygons byname:it(bilingual regions: Sardegna/Valle d'Aosta/Trentino) and by sigla (ISO3166-2/short_name) for provinces. Guard:build_frontend._assert_it_geo_coverageexits non-zero ifregione/provinciacoverage isn't 100% (CI-smoke + nightly catch a regression before deploy). Known topo gap: the province TopoJSON lacksAO(Aosta) andSU(Sud Sardegna) polygons (pre-2016 naming) — those 2 provinces render uncolored though their entities are still counted. Note: enrich_geo gives region/province but not building-level geometry, so this powers analytics, not new polygons.
The IndicePA constraint (#2)
IndicePA is not a clean source: email domains are incoherent/incomplete. The whole pipeline exists to reprocess it; our data is not a direct read. Continuous remediation is a core functional dependency (issue #2), disclosed in the methodology. The same dirtiness affected the territorial field — now worked around structurally: instead of IndicePA's incomplete/wrong region, the geo enrichment resolves the clean ipa_codice_comune_istat against the ISTAT crosswalk (see the AREA axis above).
Manual PA additions (not in IndicePA). Some real public bodies have no standalone IndicePA record but their own mail domain — e.g. the Armed Forces (Esercito/Marina/Aeronautica, subdomains of *.difesa.it with their own MX). Curate these in data/pa_manual_additions.json (same schema as a seed entry: id = IT-{ipa_codice_categoria}-{ipa_codice_ipa}, domain, ipa_codice_comune_istat for geo, …). fetch_indicepa.py merges them into the seed after the IndicePA fetch (load_manual_pa_additions()), so they survive the nightly refresh. Dedup is by id and domain: if IndicePA ever starts covering one, the manual entry is skipped automatically. They then flow through preprocess/classify/geo like any other entity. Add more by editing the JSON — no code change.
- Historicization is gated: the
historicize/build_dcatsteps innightly.ymlare commented out andstoria.html/per-entity timelines are empty until run #1 — the first clean scan after the ~700 anomalies (#4) are fixed. Do not activate before then. - One reality: no "reality vs methodology" distinction. Methodology freezes at run #1; everything after is real change.
- Fotografia anticipata: the current snapshot KPIs (
kpi.json,report.json,stats_current) are live now (non-gated); only the time-series wait for run #1.
- Lead with segmented extremes, not the national average (citable/actionable findings).
- Report style = management consulting: answer-first titles, exhibits (pie charts), recommendations with owners, methodology in calce, both site links. The "andamento" section is pre-built but shows "just started" until historicization is live.
- Tone down politically-sensitive, small-N segments. PA Centrale (ministries, ~52 entities) is deliberately kept out of the headline spotlight (
SPOTLIGHT_EXCLUDEinreport.py): small numbers + a charged "attack-the-state" reading. The spotlight features robust, citizen-data sectors (Istruzione, Sanità, Ricerca). Security/defense segments, if surfaced, go via policy framing, not loud percentages. Frame as protecting citizens, not accusing institutions.
Per-entity & geographic SEO pages (#15)
scripts/build_entity_pages.pyemits, every nightly, ~53k static pages: one per entity (/ente/{provincia-sigla}/{nome-ente}/) with full scan data + sovereignty verdict (6/4-bucket, reusingsovereignty_of/provider_to_sov4— never re-derive) + reliability + nearby entities (reputational nudge) + a "Riporta un errore" link that becomes an emphasised "Aiutaci a risolvere l'anomalia" CTA for anomalous/low-confidence entities; plus geographic hubs (/aree/{regione}/{sigla}/{comune}/), category facets (/categoria/{cluster}/) and lightweight domain aliases (/dominio/{dominio}/, canonical → entity).- Pure URL/slug logic in
src/mail_sovereignty/pages.py(ruff-gated, unit-tested intests/test_pages.py): deterministic, collision-free slugs (full entity name, province namespace, stable per-bfs token on the rare namesake; 0 collisions on the real 22,987). Generator runs an integrity assert (#entity-pages == #entities, unique URLs, ISD in range). build_entity_pages.pyis the single sitemap authority: writessitemap.xml(a sitemap index) +sitemap-core/aree/categorie+sitemap-enti-{regione}.xml(×20). It replaced the standalonebuild_sitemap.pystep in the nightly (the latter is kept only as an importable lib for the core page list).<lastmod>=kpi.json:generated_at.- Artifact-only, never committed. The generated
/ente /aree /categoria /dominio+sitemap*.xmlare git-ignored; they ship inside the Pages artifact (deploy is decoupled from git). Covered by thesmokejob (subset via--limit/--solo-regione). ~53k files build in ~50s. ⚠️ GitHub Pages MUST stay in "GitHub Actions" mode (build_type=workflow). Since the #15 tree is git-ignored, if Pages flips to legacy/branch mode it serves only the committed files onmain→ the entire/ente /aree /sitemap.xmltree 404s site-wide (happened 2026-06-22). Bothdeploy.yml(on push) andnightly.ymlgenerate the pages and publish them viaupload-pages-artifact+deploy-pages, which only takes effect in workflow mode. Guard:deploy.ymlrunsactions/configure-pages@v5(enablement: true) to re-assert the mode every deploy. If pages 404 site-wide, checkgh api repos/mxmap-it/mxmap.it/pages→build_typeand reset withgh api -X PUT repos/mxmap-it/mxmap.it/pages -f build_type=workflow, then re-run a deploy.⚠️ deploy-pagescan fail transiently on the large artifact. With ~53k files the Pagessyncing_filesphase occasionally returns "Deployment failed, try again later" (seen 2026-07-03; intermittent — the same artifact deploys fine most nights). It's retriable and a failed deploy does not take down the previous one (the site keeps serving the last good deploy). Two mitigations are in place: (1) bothnightly.yml(deploy job) anddeploy.ymldo one automatic retry ofdeploy-pagesafter a 45s pause (self-heals normal transients); (2) both invokebuild_entity_pages.pywith--no-domain-aliases— the ~22k/dominio/pages were canonical-redirects, not in the sitemap, ~zero SEO value, and ~42% of the files; dropping them takes the artifact from ~53k to ~31k files for a far more reliable sync. If a deploy still fails both retry attempts (GitHub Pages having a bad window), the site keeps serving the last good deploy — just re-run later (gh run rerun <id> --failed). To re-enable aliases, remove the--no-domain-aliasesflag from both workflows.- Honest per-URL
<lastmod>+ tiered sitemaps.build_entity_pages.pykeeps a content-hash→date state indata/page_lastmod.json(committed by the nightly): a URL's<lastmod>bumps only when its page content actually changes — declaring "everything changed nightly" teaches Google to distrust the sitemap. Entity sitemaps are split into 3 measurable GSC segments:sitemap-enti-territoriali/scuole/istituzioni.xml. The IndexNow submitter reads these lastmods and submits the real daily delta first, then the rotating slice. - IndexNow (daily "ping").
scripts/indexnow_submit.py+indexnow.ymlsubmit a deterministic rolling slice (~1000 URLs/day, full ~30k cycle ≈ monthly) of the live sitemap toapi.indexnow.org— one submission is shared with Bing/Yandex/Seznam/Naver/Yep. The key is public by design, hosted as<key>.txtat the site root (committed). Google does not support IndexNow and retired sitemap pings in 2023: for Google, the levers are the maintained sitemaplastmod, internal links, and authority — never add a "Google ping", it doesn't exist. GSC note: Googlebot occasionally fetches URL-looking strings from inline JS template literals → a stray "Not found (404)" in GSC with Source=Website is expected noise, not a broken link (self-check crawl of real internal links passes 200).
- Logic in
src/mail_sovereignty/(importable, ruff-gated, coveragefail_under=84); thin CLI inscripts/; viewer HTML + public artifacts at root. Every KPI generator follows the numbers-tested rule (below). - Driven from a Windows laptop without
uv: forstats-style logic (stdlib +mail_sovereigntyonly),pip install ruff==0.15.5 pytestinto the system Python to format/lint/test locally; else round-trip via the server. Push from the laptop (server deploy key is read-only); PowerShell forgit push(HTTPS creds), Bash tool for commits. - See
docs/ROADMAP.mdfor the issue-driven roadmap. - Keep the docs in sync (mandatory). On every feature commit or significant change, update the relevant sections of
README.md(Italian, contributor-facing: how it works, corner cases, artifacts) anddocs/ROADMAP.md(phases/issues). Treat docs drift as a bug. The README A-to-Z + roadmap are how a new contributor (or Claude) onboards.
Never use system python3 directly. Always use uv run python3 or uv run for all Python commands. The system Python may be an older version (e.g., 3.9) that doesn't support the type annotations and features used in this codebase.
CI (.github/workflows/ci.yml) runs uv run ruff format src tests --check and fails the entire CI (exit 1, pytest skipped) if any file under src/ or tests/ is not ruff-formatted. Hand-written Python is almost never compliant (string wrapping, @parametrize layout, trailing commas), so it breaks CI every time.
Mandatory before committing any change to src/ or tests/ Python:
uv run ruff format src tests # auto-format (mutates files)
uv run ruff check src tests # lint (separate CI gate, also blocks)Then git add the reformatted files. There is a committed .pre-commit-config.yaml (ruff hooks) — run pre-commit install once to enforce this automatically on git commit.
Dev-machine note: if the working machine has no uv (e.g. a Windows laptop driving a remote server), run the format on the server, then copy the formatted files back before committing. CI checks only src tests, not scripts/ — do not bulk-commit a repo-wide reformat of scripts/.
The nightly (.github/workflows/nightly.yml) broke twice on the data commit/push and stopped the public site from updating. The fix is structural and must be preserved — never patch around it:
-
The deploy NEVER depends on a git commit/push.
update-databuilds the data and publishes it withactions/upload-pages-artifact(path.); thedeployjob only runsactions/deploy-pageson that artifact. Do not re-introduceactions/checkout+ref: mainin thedeployjob — that re-couples the site to a successful push and is exactly the bug that recurred. The site must update from the in-run artifact regardless of git. -
Data commits are best-effort and non-blocking. The "Commit and push data" step in
update-dataiscontinue-on-error: truewithid: commit. A commit/push failure must never fail the job or block the deploy. Keep the git history of data as a side-effect, never on the critical path.⚠️ Gotcha (bit us 2026-06-19): because this commit is best-effort and currently fails (#10, branch protection blocks the bot push),maincan lag the live (artifact) data. Sincedeploy.yml(on push) regenerates the #15 entity pages from the COMMITTEDdata.json/kpi.json, a push after a nightly-with-failed-commit republishes stale data/pages and silently reverts the live stats. Recovery:gh run download <nightly-run> -n github-pages, extract the freshdata.json/kpi.json/report.json/dist, commit them. Real fix = make the nightly data commit reliable somain== live. -
Every script invoked by
nightly.ymlMUST be covered by thesmokejob inci.yml. Thesmokejobpy_compiles every nightly script (catches import/syntax breakage, including the network-only ones) and runs the deterministic pipeline tail (compute_confidence,report_confidence,report_anomalies,validate,build_frontend,build_public_dataset,historicize,build_dcat) on the committeddata.json. If you add or rename a step innightly.yml, add/update it in thesmokejob in the SAME PR. This is what guarantees a code change can never silently break the 04:00 run — it goes red in CI first. -
Failures auto-open a GitHub issue (auto-detection). A commit failure opens/updates an issue labeled
nightly-commit; any hard step failure (build/pipeline) opens/updates one labelednightly-failurevia a finalif: failure()step. Do not remove these and do not let the labels go missing (the steps create them idempotently). -
Before merging any change that touches the nightly or its scripts: confirm
ci.yml'ssmokejob is green. Locally/on the server you can dry-run the tail with the same commands the smoke job uses (seeci.yml). Network steps (fetch_indicepa,preprocess,postprocess,recover/finalize,enrich) can't run in CI — they are onlypy_compiled; their runtime failures are transient (network) and surface via thenightly-failureissue, never by silently dropping the site.
Any statistic or KPI we publish (Indice di Sovranità Digitale, CLOUD Act share, per-category breakdowns, coverage, market concentration, …) MUST be both unit-tested and self-verified at build time. A wrong public number is worse than a missing one — this is a transparency observatory; the figures are the product.
Two mandatory layers for every KPI/statistics generator:
-
Unit tests with hand-computed expected values. The compute logic lives in
src/mail_sovereignty/(importable, coverage-gated byfail_under), not only inscripts/. Atests/test_*.pyexercises it on a small synthetic fixture whose totals are worked out by hand, asserting exact values (counts, shares, index, segmentation). Reference:src/mail_sovereignty/stats.py+tests/test_stats.py. -
Runtime integrity assertions on real data. The module exposes an
assert_integrity()that checks internal consistency — counts sum to the population, shares sum to ~100%, the headline index equals its definition, the segmentation covers everything with no oversizedother, noNaN/out-of-range — and the build script calls it on every run, exiting non-zero on any violation. Reference:stats.assert_integrity(), called byscripts/build_stats.py. Each invariant has a test proving it actually fires on corrupted input.
Wiring (so it runs automatically): unit tests run in the test CI job (pytest --cov); the build script with its integrity check runs in the nightly and in the smoke CI job (see the rule above). A KPI that isn't covered by both a unit test and a runtime invariant must not ship.
Dev-machine note: src/ + tests/ are ruff-gated. The Windows laptop has no uv, but ruff==0.15.5 and pytest can be pip installed into the system Python to format/lint/test stats-style modules locally (logic that only imports stdlib + mail_sovereignty), avoiding the server round-trip. Match the pinned ruff version (see uv.lock).
uv sync # Install dependencies
uv sync --group dev # Install with dev dependencies
# Pipeline (run in order, each reads/writes data.json)
uv run preprocess # DNS lookups + classification (~30s for small countries)
uv run preprocess DE # Single country
uv run preprocess DE:BY # Single Bundesland (Bavaria)
uv run postprocess # Overrides, SMTP banners, scraping (~5 min)
uv run validate # Confidence scoring + quality gate
# TopoJSON split (requires mapshaper: npm install -g mapshaper)
uv run python3 scripts/split_topo.py # Splits monolithic TopoJSON -> topo/
# Seed data fetching
uv run python3 scripts/fetch_wikidata.py DE # Fetch Gemeinden from Wikidata
uv run python3 scripts/fetch_boundaries.py DE # Fetch boundaries from Overpass
# Tests
uv run pytest # All tests
uv run pytest tests/test_classify.py # Single file
uv run pytest tests/test_classify.py::test_name # Single test
uv run pytest --cov --cov-report=term-missing # With coverage (90% threshold)
# Lint
uv run ruff check src tests
uv run ruff format src tests
# Local frontend
python -m http.serverAll three stages operate on data.json at the repo root:
-
Preprocess (
preprocess.py) — Loads municipalities fromdata/municipalities_{cc}.jsonseed files (92 countries) +data/overrides.json. For each municipality: extracts domain (or guesses from name with diacritics transliteration), performs async MX/SPF/CNAME/ASN/autodiscover/DKIM/TXT-verification/tenant DNS lookups via 3 resolvers (system, Google, Cloudflare) with shared cache, classifies provider, detects gateways. Concurrency: 20. Supports sub-country filtering (DE:BYscans only Bavaria). -
Postprocess (
postprocess.py) — Four sub-steps: (a) applyMANUAL_OVERRIDESdict with DNS re-lookup for domain-only overrides, (b) retry DNS for unknowns that have a domain, (c) SMTP banner check on primary MX of independent/unknown entries (deduplicated, concurrency 5), (d) scrape municipality websites for email addresses on remaining unknowns (concurrency 10). Includes TYPO3 Caesar cipher decryption for obfuscated mailto: links. -
Validate (
validate.py) — Scores each entry 0–100 based on DNS data quality (has domain, MX, SPF, provider match, etc.). Quality gate: average score ≥ 70 and ≥ 80% of entries above 80 confidence. Writesvalidation_report.jsonandvalidation_report.csv. Exits 1 on failure.
classify() returns tuple[str, str] — (provider, reason).
Priority order:
- Direct MX match — MX hostname contains provider keyword
- CNAME resolution — MX host's CNAME target matches a provider
- Known gateway look-through — MX matches a
GATEWAY_KEYWORDSentry (SeppMail, Barracuda, FortiMail, SecMail, D-Fence, Cisco IronPort, MailAnyone, Comendo, Heimdal, StaySecure, edelkey, ippnet, garmtech, etc.) → check SPF (only if exactly one main provider found) → autodiscover → DKIM → TXT verification → MS365 tenant (viagetuserrealm.srf) for the actual backend provider. If no backend identified, returns "independent" with reason mentioning the gateway. - Self-hosted gateway detection — MX exists but doesn't match any provider or gateway → check DKIM for a hidden backend provider (e.g.,
mail.muhu.eeon Radicenter but DKIM →*.onmicrosoft.com= Microsoft) - Local ISP — MX ASN matches known ISP ASNs (
LOCAL_ISP_ASNSin constants.py) - Independent — MX exists but doesn't match any known provider and no DKIM backend found
- Unknown — No MX records found
SPF vs DKIM for backend detection:
- SPF is only used in step 3 (known gateways), and only when exactly one main provider is found in SPF. If multiple providers appear (e.g., Microsoft + Google), SPF is ambiguous — municipalities often include
spf.protection.outlook.comfor shared calendars or hybrid sending without hosting mailboxes on Microsoft. In ambiguous cases, fall through to autodiscover/DKIM. - DKIM is used in both steps 3 and 4. DKIM CNAMEs (
selector1._domainkey.domain → *.onmicrosoft.com) prove a Microsoft 365 tenant is configured to sign mail for that domain — this is definitive proof of mail hosting. DKIM is the most reliable signal for identifying the actual backend provider.
Provider values: microsoft, google, aws, zone, telia, tet, elkdata, local-isp, independent, unknown.
All provider detection is keyword-based. To add a new provider:
- Add
*_KEYWORDSlist toconstants.py - Add to
PROVIDER_KEYWORDSdict - Add to
SMTP_BANNER_KEYWORDSif applicable - Add to the two provider-matching loops in
classify()(steps 1 and 2) - Add display name mapping + color in
index.html
Per-country file-based DNS cache in data/dns_cache/. Domain-scoped: all DNS queries for a domain stored together. TTL: 7 days.
Partitioned caches for large countries: DnsCache("DE", partition="09") → de_09.json. Configured via PARTITIONED_COUNTRIES in constants.py. Currently only Germany is partitioned (16 files, one per Bundesland).
The preprocess CLI supports CC:STATE syntax for scanning subsets of large countries:
uv run preprocess DE:BY # Bavaria only (abbreviation)
uv run preprocess DE:09 # Bavaria only (state code)
uv run preprocess DE:BY,NW # Bavaria + Nordrhein-Westfalen
uv run preprocess DE:BY IT # Bavaria + all of ItalyState codes are in DE_STATES dict in constants.py. When filtering, only the scanned entries are replaced in data.json — other states/countries are preserved.
Single-page app with three admin-level views (Region/District/Municipality) and per-country lazy-loaded TopoJSON.
Data loading: Fetches data-summary.json + topo/manifest.json on startup. data-detail.json loaded in background. The manifest maps each country × level to a TopoJSON file (or an object of per-state files with bboxes for viewport loading). Files are fetched on demand and cached in memory (topoCache). Default view is "Districts".
Multi-level toggle: Three-button segmented control (top-left) switches between Region, District, and Municipality views. Each level loads different TopoJSON files per country.
Viewport-based loading: For countries with many municipalities (currently DE), the manifest municipality entry is an object mapping filenames to bboxes. Only files whose bbox intersects the visible viewport are fetched. On moveend, new files are loaded as needed.
Per-country layers: Each country has its own L.geoJSON layer stored in countryLayers Map. Country filter buttons add/remove layers and restyle them (active = provider-colored, inactive = gray).
Aggregation: At Region/District levels, multiple municipalities map to one polygon. computeAggregation() groups municipalities by region name or district key (AT: first 6 chars of ID, BE: first 5, DE: first 8). matchGroupFeature() matches dissolved TopoJSON features to groups by name, name_en, or name:en property.
Popups: Municipality level shows individual DNS data (MX, SPF, DKIM, autodiscover, TXT verifications, MS365 tenant status). Region/District level shows aggregated view: dominant provider badge, stacked provider bar chart, scrollable municipality list with provider dots.
Statistics panel: Always shows municipality-level data — unaffected by the level toggle.
Gateway markers: Shield icons only shown at municipality level.
Splits baltic-municipalities.topo.json (monolithic source) into per-country per-level files in topo/. Some countries have standalone TopoJSON files generated by scripts/fetch_boundaries.py.
topo/
manifest.json # { CC: { levels, files, sizes } }
{cc}_municipality.topo.json # Per-country municipality boundaries (simplified 15%)
{cc}_region.topo.json # Dissolved by region field (simplified 8%, quantization 5k)
{cc}_district.topo.json # Dissolved by district key (AT, BE, DE)
de_municipality_XX.topo.json # Per-Bundesland DE files (16 files, viewport-loaded)
Manifest format: For most countries, files.municipality is a string filename. For DE, it's an object mapping filenames to bounding boxes:
"municipality": {
"de_municipality_01.topo.json": [8.3, 53.3, 11.3, 55.0],
"de_municipality_09.topo.json": [8.9, 47.3, 13.8, 50.6]
}District key extraction: AT → first 6 chars of ID (AT-101), BE → first 5 (BE-11), DE → first 8 (DE-01001).
Run: uv run python3 scripts/split_topo.py (requires mapshaper CLI).
Splits data.json into two files for faster initial page load:
data-summary.json— Loaded immediately. Contains fields needed for map rendering, legend, and stats.data-detail.json— Loaded in background after map renders. Contains popup-only fields:mx,spf,reason,autodiscover,dkim,txt_verifications,tenant,smtp_software.
Run: uv run python3 scripts/build_frontend.py
All lookups use 3 resolvers (system, Google, Cloudflare) sharing a single dns.resolver.Cache to avoid redundant queries. The core function resolve_robust(qname, rdtype) provides universal multi-resolver fallback — all higher-level functions delegate to it. Key functions: lookup_mx(), lookup_txt() (returns SPF + TXT verification tokens in one query), lookup_spf(), resolve_spf_includes() (recursive BFS with loop detection), resolve_mx_cnames(), resolve_mx_asns(), resolve_mx_countries() (both via Team Cymru DNS — ASN + country code from same query), lookup_autodiscover(), lookup_dkim() (checks selector1/selector2/google._domainkey CNAMEs — definitive proof of mail hosting, e.g. CNAME to *.onmicrosoft.com = Microsoft 365), lookup_tenant() (queries Microsoft's getuserrealm.srf endpoint to detect MS365 tenants — returns Managed or Federated).
Tests use pytest-asyncio (auto mode) and respx for HTTP mocking. DNS is mocked via AsyncMock on resolver objects. Fixtures in conftest.py provide sample_municipality, sovereign_municipality, sample_data_json. Coverage threshold is 90%.
GitHub Actions nightly workflow (.github/workflows/nightly.yml) runs preprocess → postprocess → validate → build (frontend, public dataset, stats, kpi, report) → upload Pages artifact → deploy, with a separate best-effort data commit. The deploy is decoupled from the git commit (see "CRITICAL: Nightly must never break"). Default branch is main. The custom domain is mxmap.it (CNAME at repo root — keep it in the deploy artifact).
DE nightly rotation: Germany's 11K Gemeinden are too many to scan every night. The workflow rotates 3 Bundesländer per night (6-day cycle), while all other countries are scanned every night.
Policy (project owner, 2026-09-29): open-source/community basemap, no API key, no account, no paid plan. The map background must stay a zero-cost, fully open dependency.
- Incident timeline (why the map "suddenly" broke with no commit): the keyless CARTO config dated from the original fork (untouched since 2026-03). CARTO started watermarking keyless requests to
basemaps.cartocdn.comaround 2026-08-28 and enforced the API-key requirement on 2026-09-23 — tiles kept answering HTTP 200 but carried an "API KEY REQUIRED" watermark (2049 fixed bytes): the map broke silently and server-side. No git rollback can ever fix this class of breakage — every commit in our history contains the same keyless URLs the provider no longer honors. Upstreammxmap/secassure2026reacted (2026-08-27, "add carto api key to hide watermark") by injecting aCARTO_KEYGitHub secret at deploy; we deliberately did NOT follow (the "secret" is world-readable in the served JS, quota burnable by anyone, forks/local dev stay broken, billing relationship). - Current provider (owner's decision, 2026-09-29): OSM standard
https://tile.openstreetmap.org/{z}/{x}/{y}.png— OSMF community, open data, keyless, native up to z19, slippy axis order/{z}/{x}/{y}. Respect the OSMF tile usage policy (moderate traffic, real attribution© OpenStreetMap contributors). - Toponym overlay (labels above polygons): OSM raster has labels baked in, so entity polygons would cover them. Solution: a second tileLayer with the SAME URLs on a dedicated pane
toponimi(createPane, zIndex 450 = above overlayPane/polygons 400, below markers 600) blended via CSSmix-blend-mode: darken+ a filter that pushes midtones to white — only dark pixels (labels, admin borders) emerge above the polygons. Same URLs ⇒ browser HTTP cache ⇒ zero extra requests to OSM. Hidden-by-default +@supports (mix-blend-mode: darken)guard: browsers without blending just show the base map (labels under polygons), never opaque tiles over data. - Documented fallback (if OSM is ever unreachable/degraded): Esri World Light Gray
Base+Referenceatserver.arcgisonline.com— keyless but proprietary; axis order/tile/{z}/{y}/{x}(inverted vs slippy! a mismatch shows the wrong place with zero errors) and native max zoom 16 (maxNativeZoom: 16). - Battery:
tests/test_basemap.py— structural checks (2 layers: base with attribution + toponym overlay on panetoponimi; axis-order-per-host rule; no keyless CARTO; preconnects aligned;foreign-faded+ blend CSS contract) + functional checks (real land tiles for Rome & Milan: 200,image/*, minimum size, and the anti-placeholder check: two different coordinates MUST return different bytes — the only check that catches watermark-behind-200 breakage). Runs on every push/PR (CItestjob) and daily in the standalonebasemap-batteryjob of nightly.yml (provider drift detected ≤24h, auto-issue labelbasemap, never blocks the data pipeline).
Create a JSON array of municipalities:
[{"id": "XX-001", "name": "City Name", "country": "XX", "region": "Region", "domain": "city.xx", "osm_relation_id": 12345}]Sources: national statistics office API for official list, Wikidata SPARQL for OSM relation IDs + domains (P856 website, P402 OSM relation), Nominatim as fallback for OSM IDs.
preprocess.py: Add"XX": "municipalities_xx.json"toSEED_FILES. Add diacritics transliteration pairs. Add"XX": [".xx"]totld_mapinguess_domains(). Add name suffixes to strip (e.g., " kommun", " kunta").constants.py: Add country-specific ISP ASNs toLOCAL_ISP_ASNS. Use Team Cymru DNS (origin.asn.cymru.com) to identify ASNs of municipalities classified as "independent" — many will be local ISPs. Add gateway keywords if the country uses local email security appliances (FortiMail, SecMail, etc.). Municipal IT cooperatives (e.g., Norwegian IKT companies like Hedmark IKT, Lofoten IKT) often act as gateways — add them toGATEWAY_KEYWORDSorLOCAL_ISP_ASNSas appropriate.constants.py: Add"example.xx"toSKIP_DOMAINS. Add country-specific contact page paths toSUBPAGESif different from existing ones.
Preferred: scripts/fetch_boundaries.py — Fetches from Overpass API, converts via Python GeoJSON converter, annotates with region/country, creates TopoJSON via mapshaper.
- Add country to
COUNTRY_CONFIGwith admin_level and ISO code - Run:
uv run python3 scripts/fetch_boundaries.py XX
Alternative: monolithic file — For countries in baltic-municipalities.topo.json, run scripts/split_topo.py.
Feature IDs must be relation/XXXXX matching osm_relation_id in seed data.
- Add country code to
COUNTRY_LISTandFLAGSmap - Add country button in
.country-filtersdiv - Add country name to
countryNamesand color tocountryColorsin stats
- Add country to
COUNTRIESandLEVEL_MAPinscripts/split_topo.py - Re-run
uv run python3 scripts/build_frontend.pyto rebuild data-summary.json and data-detail.json
- Update
test_loads_all_countriesexpected countries set - Update
test_no_country_generates_all_tldsexpected TLDs - Add diacritics test case in
TestGuessDomains
Run preprocess → check "independent" municipalities → look up their MX ASNs → add missing local ISPs to LOCAL_ISP_ASNS → re-run. Typical pattern: first run has too many "independent", iteratively adding ISP ASNs and gateway keywords brings it down to a handful of genuinely self-hosted servers. Also check for gateway patterns in MX hostnames (e.g., iphmx.com = Cisco IronPort, comendosystems.com = Comendo) and add to GATEWAY_KEYWORDS.
For countries with >2,000 municipalities (currently DE with ~11K):
- Partitioned DNS cache — Add to
PARTITIONED_COUNTRIESinconstants.pywith a lambda extracting the partition key from the municipality ID. - Sub-country filtering —
uv run preprocess CC:STATEto scan subsets. - Per-state TopoJSON — Manifest uses object format
{filename: [bbox]}for viewport-based loading. - Nightly rotation — Scan N states per night instead of all at once.
- Dissolved district topo — Generate
{cc}_district.topo.jsonby dissolving municipality features by district key prefix.
When fetching boundaries from Overpass:
- Area ID format:
3600000000 + relation_id(not string concatenation3600{id}) - Rate limiting: Overpass returns 429 after rapid queries. Use 10-15s delays between states, 90s backoff on 429.
- Timeouts (504): Large states may timeout. Retry up to 3 times with 30s waits.
- City-states: Berlin, Hamburg, Bremen use different admin levels (4/6/9 instead of 8). Try multiple levels for states with ≤5 expected municipalities.
- osmtogeojson strips properties: The npm
osmtogeojsontool doesn't preserve OSM tags as flat GeoJSON properties. Use the Pythonconvert_osm_to_geojson_simple()fallback for boundary fetching. - osm_id format: Feature IDs must be
relation/XXXXX(with prefix), not bare integers. Theconvert_osm_to_geojson_simplefunction sets this correctly, but verify after mapshaper processing.
Municipality domains are the most error-prone part of the data. Always verify domains via web search — do not trust automated guessing alone.
Many municipality domains do NOT match the municipality name:
- Corporate namesakes:
nokia.fi(phone company),outokumpu.fi(mining company),noo.ee(meat factory). Cities usenokiankaupunki.fi,outokummunkaupunki.fi,nvv.ee. - Tourism/portal sites:
hiiumaa.ee(tourism portal, municipality is atvald.hiiumaa.ee),peipsi.ee(tourism NGO, municipality ispeipsivald.ee),rouge.ee(community portal, municipality isrougevald.ee). - Gaming/unrelated sites:
siauliu.ltwas a Counter-Strike gaming site; the actual municipality domain issiauliuraj.lt.
Swedish-speaking Finnish municipalities consistently use their Swedish name for domains: Kruunupyy→kronoby.fi, Luoto→larsmo.fi, Maalahti→malax.fi, Vöyri→vora.fi, Kristiinankaupunki→krs.fi.
After Latvia's 2021 municipal reform, many novads (counties) have their own domains distinct from the main city: bauskasnovads.lv (not bauska.lv), valmierasnovads.lv (not valmiera.lv), ventspilsnd.lv (not ventspils.lv which is the city). The seed data domain with no MX causes the pipeline to guess the city domain instead — always set the correct novads domain in seed data.
Norwegian municipalities mostly use name.kommune.no, but exceptions exist: ha.no (Hå), sarpsborg.com (Sarpsborg), voss.herad.no (Voss herad), mgk.no (Midtre Gauldal), ahk.no (Aurskog-Høland). Sami-language municipalities may use .suohkan.no for their website but .kommune.no for email (e.g., Kautokeino). Post-merger municipalities sometimes retain stale "nye" (new) domains — always verify the current domain has MX records.
Some municipalities use different domains for their website and email. When the seed data domain has no MX records, the pipeline falls back to guessing from the municipality name, which may find a wrong domain. Always check MX records for the seed data domain; if empty, search for the actual email domain.
For each country, web-search every municipality to verify domains. Use dig +short domain MX to confirm MX records exist. The MANUAL_OVERRIDES dict in postprocess.py handles cases where the guessed domain is wrong but the seed data domain is correct for the website (overrides trigger DNS re-lookup on the corrected domain).
Municipalities often use local email security gateways (FortiMail, SecMail, D-Fence, Barracuda, etc.) that relay to cloud providers. The pipeline detects these via GATEWAY_KEYWORDS in constants.py. When a gateway is detected, the pipeline checks SPF → autodiscover → DKIM to identify the backend provider.
Small local IT companies can also act as gateways (e.g., edelkey.net for Helsinki, ippnet.fi for Parkano, garmtech.com for Saulkrasti). Add these to GATEWAY_KEYWORDS when discovered — otherwise they get classified as "independent" instead of the actual backend provider.
Gateway SPF ambiguity: When looking through a gateway, SPF is only trusted if exactly one main provider keyword is found. Many municipalities have multiple providers in SPF (e.g., Microsoft for mailboxes + Google for transactional email), making SPF ambiguous. In those cases, the pipeline falls through to autodiscover, DKIM, TXT verification, and MS365 tenant detection for a definitive answer. If none identify a backend, the municipality is classified as "independent" with a reason mentioning the gateway name.
MS365 tenant detection: lookup_tenant() queries Microsoft's login.microsoftonline.com/getuserrealm.srf endpoint. A Managed or Federated response proves an MS365 tenant exists for that domain. Used as a last-resort signal in gateway look-through (after DKIM and TXT verification) — not used for self-hosted MX since having a tenant doesn't prove mailboxes are hosted there.
DKIM is the most reliable signal for identifying the backend provider. A CNAME at selector1._domainkey.domain pointing to *.onmicrosoft.com is definitive proof of Microsoft 365, even when MX and SPF point elsewhere.
Norwegian IKT cooperatives: Many Norwegian municipalities share IT infrastructure via regional IKT companies (Hedmark IKT, Lofoten IKT, IKT Sunnmøre, etc.). These appear as shared MX hosts or DKIM tenants (e.g., lofotenikt.onmicrosoft.com). They typically relay to Microsoft 365.
Per-country implementation guides are in docs/countries/. These were used during initial setup and may be outdated, but contain useful context about domain pitfalls, ISP discovery, and admin level choices. Available: Andorra, Australia, Austria, Belgium, Czechia, Denmark, Germany, Luxembourg, New Zealand, Norway, Sweden. Use these as examples when adding new countries.