Postal Codes Dataset for United States, US

558
Updated:
Files:1
Size:3.02 MB
Rows:41,488
Formats:csv

Postal Codes Dataset for United States, US including name of the city, town, or place, various administrative divisions and alternative city names.

Part of a Data Solution

  • Postal Codes Solution

    Every worldwide postal-code dataset in one bundle — the ultimate global reference for precise postal codes.

    Explore →

Premium

You're viewing a free sample

Get the complete dataset — all rows and files — delivered instantly after checkout, with lifetime access to the latest version.

  • Secure checkout via Stripe
  • Instant download after payment
  • Lifetime access to the latest version
  • Creative Commons Attribution 4.0 (GeoNames) license
20% off
$49.90$39.92
one-time payment

API Access

Access dataset files directly from scripts, code, or AI agents.

Browse dataset files
Dataset Files

Each file has a stable URL (r-link) that you can use directly in scripts, apps, or AI agents. These URLs are permanent and safe to hardcode.

/logistics/postal-codes-us/
https://datahub.io/logistics/postal-codes-us/_r/-/.licensing-notes.md
https://datahub.io/logistics/postal-codes-us/_r/-/.migration-notes.md
https://datahub.io/logistics/postal-codes-us/_r/-/ATTRIBUTION.txt
https://datahub.io/logistics/postal-codes-us/_r/-/README.md
https://datahub.io/logistics/postal-codes-us/_r/-/datapackage.yaml
Key Files

Start with these files — they give you everything you need to understand and access the dataset.

datapackage.yaml— metadata & schema
https://datahub.io/logistics/postal-codes-us/_r/-/datapackage.yaml
README.md— documentation
https://datahub.io/logistics/postal-codes-us/_r/-/README.md
Typical Usage
  1. 1. Fetch datapackage.yaml to inspect schema and resources
  2. 2. Download data resources listed in datapackage.yaml
  3. 3. Read README.md for full context

Data Files

postal-codes-us-sample


About this dataset

United States (US): migration notes

Internal record of the P5 migration from the old scripts to the producer: the discrepancies found, their causes, and who decided what. This is a dated record, so the numbers are as of the migration and not kept current. It is not shipped in the zip.

  • Issue: pc-28e.7.56 (closed 2026-10-05)
  • Commit: 72c273f (worktree commit 44bb982, same content), landed in batch 6 on 2026-10-03
  • Parity verdict: explained, row_order only; signed off under the standing rule
  • Producer: custom, refresh disabled

Sources

InputURLNotes
GeoNames postal exporthttps://download.geonames.org/export/zip/US.zip
Census 2020 ZCTA-to-County relationship, via the zctaCrosswalk R packagehttps://raw.githubusercontent.com/arilamstein/zctaCrosswalk/main/data/zcta_crosswalk.rdaThe same mirror the original pipeline used, because census.gov blocks automated downloads. datapackage.yaml keeps the census.gov reference page as the origin and lists the mirror as the copy the producer fetches

The .rda is read by a small stdlib reader for R's XDR format inside produce.py. pyreadr isn't in the venv, and adding it would have meant touching files outside datasets/us/.

Which builder was ported

Not the in-repo scripts. datasets/us/scripts/{build_base,fetch_source}.py were a 2026-08-19 reconstruction, written on the belief that no US build script survived (it only checked origin/main and a local copy). On 2026-10-01 the user pulled the live full file and sample and asked for a comparison (session a51a99b8). The reconstruction did not match live:

  • 41,490 rows vs 41,488: two extra FPO AA rows (96860, 96863), which live carries only in alternative_city_name.
  • 33,468 admin_name2 cells: bare GeoNames names (Douglas) vs Census names (Douglas County). 32,152 differ only by the legal suffix.
  • 335 admin_code2 cells, 267 of them Connecticut: current GeoNames planning-region codes (091xx) vs live's Census 2020 county codes (090xx).
  • 1 row with admin_name3/admin_code3 filled (Monroe Township / 47280) where live is blank.

The live file was built by an unmerged pipeline on origin/feat/us-consolidated-r2 (commit ae1dc64, 2026-07-27, by amautadev): scripts/us/{geonames_base,fetch_crosswalk,census_enrich,integrity_checker}.py, run monthly by .github/workflows/run-monthly-us-consolidated-r2-script.yml. That pipeline explains every difference above. produce.py ports it.

Parity against the live zip (2026-10-03)

Live reference: s3://postal-codes/us/consolidated/latest/us.zip#us.csv. Rows: 41,488 live and 41,488 fresh, none only on one side, 0 changed cells. The only class is row_order.

Cause: GeoNames reordered its export (upstream drift). 137 row positions differ. The first is at line 24847, where Nevada 89044 now comes before 89046. The sorted key lists are identical. Today's US.txt has 41,490 lines with 2 repeated codes, which dedup to 41,488 rows. A --no-fetch rerun is byte-identical.

The informational sample check against https://postal.datahub.io/us/us.csv reports leading_zero_restored ×100 and float_suffix_dropped ×100. That is the known bug in the live sample (create_samples_and_stats.py read the file without dtype=str, e.g. 2020.0 for 02020), not a difference in the full file.

Decisions

DecisionByBasis
Port the origin/feat/us-consolidated-r2 pipeline, not the in-repo reconstructionOrchestrator instruction (bead description and agent prompt)Finding of session a51a99b8 (2026-10-01), where the agent proposed the port. No explicit user answer to that proposal is recorded. The user approved launching batch 6 on 2026-10-03 with "us came from the unmerged feat/us-consolidated-r2 branch" in the question (session 7df0c5be)
Accept the row_order differenceStanding rule (row_order/CRLF only, user, 2026-10-02)Recorded in the sign-off section of docs/parity/p5.md
Fetch the crosswalk from the zctaCrosswalk GitHub mirror and list it honestly in sourcesOrchestrator instructionSame mirror the original used; census.gov blocks bots
Read the .rda with a built-in reader instead of adding pyreadrAgentNo new dependency; keeps the commit inside datasets/us/
Replace the old minimum-row-count check (≥40,000 ZIPs) with structural checksAgentBrief: no row-count asserts
Rewrite the README's "~1.2% PO-box / point ZIPs" line to name the rows that actually lack a county code (military APO/FPO/DPO, Marshall Islands)AgentThe old line was wrong
Add a dated correction to .licensing-notes.md saying the 2026-08-19 reconstruction is supersededAgentNot asked for by the brief

Transforms carried over (unchanged)

  • geonames_base.py: one row per postal_code; the first row wins and each further distinct place_name is appended to alternative_city_name, pipe-separated. Blank-code rows dropped. admin_name3/admin_code3 blanked.
  • census_enrich.py, per ZIP: if GeoNames' state+county code is one of the ZCTA's counties, or the ZCTA lies in a single county, admin_code2 = the 5-digit GEOID and admin_name2 = the Census name through Python str.title(). That's why live has Prince George'S County, and it is reproduced on purpose. Otherwise the GeoNames name stays, with code = state FIPS + its 3-digit county code. Rows with no state FIPS keep GeoNames' code: 509 blank-state military rows, and 2 MH rows that keep 020.
  • integrity_checker.py: hard checks are now ProducerErrors (5-digit codes, unique postal_code), plus all 50 states and DC present and admin_code1 a known code (a state, DC, MH or empty).
  • Cells copied verbatim (csv, no pandas).

Open questions

  • Resolved (user, 2026-10-05, pc-28e.7.59): the README now says "ZIP codes across the 50 states, DC and the Marshall Islands".
  • .licensing-notes.md still has the 2026-08-19 counts (41,490 / 41,488, 513 lookup misses) above the correction. Left for the pc-hfz sweep.

Provenance

Reconstructed 2026-10-05 from the agent transcript 7df0c5be-a523-4080-87f9-6045a6a41515/subagents/agent-aea04d7e06039ee77.jsonl, the orchestrator session 7df0c5be-…, the earlier session a51a99b8-ac96-4db5-ba52-284359b960ad (the 2026-10-01 live comparison and the discovery of the unmerged branch), the commit message, the bd close reason and docs/parity/p5.md.