Postal Codes Dataset for France, FR

456
Updated:
Files:1
Size:4.49 MB
Rows:51,589
Formats:csv
License:CC-BY-4.0

Postal Codes Dataset for France, FR including name of the city, town, or place, various administrative divisions and alternative city names.

Part of a Data Solution

  • Postal Codes Solution

    Every worldwide postal-code dataset in one bundle — the ultimate global reference for precise postal codes.

    Explore →

Premium

You're viewing a free sample

Get the complete dataset — all rows and files — delivered instantly after checkout, with lifetime access to the latest version.

  • Secure checkout via Stripe
  • Instant download after payment
  • Lifetime access to the latest version
  • Creative Commons Attribution 4.0 license
20% off
$49.90$39.92
one-time payment

API Access

Access dataset files directly from scripts, code, or AI agents.

Browse dataset files
Dataset Files

Each file has a stable URL (r-link) that you can use directly in scripts, apps, or AI agents. These URLs are permanent and safe to hardcode.

/logistics/postal-codes-fr/
https://datahub.io/logistics/postal-codes-fr/_r/-/.licensing-notes.md
https://datahub.io/logistics/postal-codes-fr/_r/-/.migration-notes.md
https://datahub.io/logistics/postal-codes-fr/_r/-/ATTRIBUTION.txt
https://datahub.io/logistics/postal-codes-fr/_r/-/README.md
https://datahub.io/logistics/postal-codes-fr/_r/-/datapackage.yaml
Key Files

Start with these files — they give you everything you need to understand and access the dataset.

datapackage.yaml— metadata & schema
https://datahub.io/logistics/postal-codes-fr/_r/-/datapackage.yaml
README.md— documentation
https://datahub.io/logistics/postal-codes-fr/_r/-/README.md
Typical Usage
  1. 1. Fetch datapackage.yaml to inspect schema and resources
  2. 2. Download data resources listed in datapackage.yaml
  3. 3. Read README.md for full context

Data Files

postal-codes-fr-sample

About

Last updated
3 October 2026
Total rows
100
Format
CSV
File size
8.07 kB

About this dataset

France (FR): migration notes

Internal record of the P5 migration from the old scripts to the producer: the discrepancies found, their causes, and who decided what. This is a dated record, so the numbers are as of the migration and not kept current. It is not shipped in the zip.

  • Issue: pc-28e.7.22 (closed 2026-10-03)
  • Commit: f34ad50 (worktree commit e5beae7, which amended the first attempt a6e2f02); the docs/parity/p5.md row and sign-off landed separately in 23d7af9
  • Parity verdict: unexplained (the gate blocks on rows only in fresh and on unclassified cells); accepted by the user's primary-key decision
  • Producer: custom, refresh disabled

Sources

InputURLNotes
GeoNames postal exporthttps://download.geonames.org/export/zip/FR.zipThe only input, as in the old builder

Despite the multi-script builder, fr is single-source. The four old scripts (fetch_source.py, build_base.py, integrity_checker.py, package.py) were a 2026-08-24 corrected republish of GeoNames, written to fix the defects of the earlier pandas-based shared parser.

Parity against the live zip

First attempt, a6e2f02 (2026-10-03, 08:00 UTC)

Rows: 51,585 live and 51,585 fresh, none only on one side. 6 rows changed, 14 cells (latitude 6, longitude 6, admin_name3 1, admin_code3 1); classed as coordinate_precision 1, unclassified 13.

KeyLiveFresh
63700 MontaigutAmbert / 631, 45.615, 3.449Riom / 634, 46.179, 2.8088
14113 Cricquebœuf49.4015, 0.146849.4022, 0.1456
89520 Saints-en-Puisaye47.621, 3.262247.6208, 3.2617

Cause: first-wins dedup over GeoNames' row order. FR.txt has 51,611 rows with 24 repeated (postal_code, place_name) groups (26 extra rows). The old builder sorted by (admin_code1, admin_code2, postal_code, place_name) and kept the first row of each group, so on a tie the winner was whichever row GeoNames listed first. The agent checked all 24 groups: no rule (first, last, min or max coordinate) reproduces the live picks. It concluded GeoNames had reordered rows inside 6 groups since the 2026-08-24 build. The 2026-08-24 input wasn't available to compare, so this cause was inferred, not traced to an old copy of FR.txt.

The same check found that 4 of the 24 groups are different communes, not coordinate-only duplicates as the old builder's docstring and README said. The old dedup dropped one of each:

  • 62760 Thièvres: Pas-de-Calais (62/621) and Somme (80/804)
  • 73670 Saint-Pierre-d'Entremont: Isère (38/381) and Savoie (73/732)
  • 30130 Pont-Saint-Esprit: Gard (30/302, two copies) and Vaucluse (84/843)
  • 63700 Montaigut: arrondissement Riom (634, two copies) and Ambert (631)

Landed version, e5beae7 (2026-10-03, 13:10 UTC)

Rows: 51,585 live and 51,589 fresh. 0 only in live, 4 only in fresh: the second commune of each of the 4 groups above (62760 Thièvres 80/804, 63700 Montaigut 63/634, 73670 Saint-Pierre-d'Entremont 73/732, 30130 Pont-Saint-Esprit 84/843). 9 rows changed, 18 cells (latitude 9, longitude 9); classed as coordinate_precision 3, unclassified 15.

KeyLiveFresh
59242 Templeuve-en-Pévèle50.5267, 3.17550.5234, 3.1781
89520 Saints-en-Puisaye47.621, 3.262247.6208, 3.2617
22210 Plémet48.1771, -2.594348.1769, -2.595

Other changed rows in the report: 70240 La Villeneuve-Bellenoye-et-la-Maize, 66150 Corsavy, 66160 Le Boulou, 63700 Buxières-sous-Montaigut. The report shows only the first examples, so the last 2 of the 9 rows aren't in the records.

Cause: the new tie-break. Every changed row is a coordinate-only twin (same widened key, ~100–200 m apart). The producer now keeps the twin with the smallest (latitude, longitude), and in these 9 the live file had kept the other one.

The informational sample check against https://postal.datahub.io/fr/fr.csv was also unexplained. Nobody looked into it.

Decisions

DecisionByBasis
Widen the primary key to (postal_code, place_name, admin_code2, admin_code3), as for bg, keeping the 4 communesUser, 2026-10-03 (session d33e1b7e), choosing "Widen primary key" over "Accept drift, keep key" and "Hold fr"The option said it keeps all 4 communes, stops the output depending on GeoNames' row order, and that parity would show added rows
Collapse only rows equal on the widened key, deterministically and not by upstream order; stop on any other collisionOrchestrator instruction, following the user's choice
On a tie keep the smallest (latitude, longitude); a repeat differing in anything but latitude, longitude and accuracy is a ProducerErrorAgentAn order-independent rule; no non-coordinate collision exists in today's file. Reported to the user after landing, with no recorded reply
Accept the 4 added rows and the 18 coordinate cellsUser (the key decision above)The option said parity would show added rows and that the output would stop depending on row order, which is where the 18 cells come from. The agent was told drift and added rows were pre-approved. After landing, the orchestrator told the user the 18 cells came from a rule it picked, not from the source, and invited objections. The user's next reply ("fine, keep it. start batch 5") answered a pl question and didn't mention fr, so there is no explicit fr answer, only no objection. Recorded in the sign-off section of docs/parity/p5.md
Extend the sort to (admin_code1, admin_code2, postal_code, place_name, admin_code3, latitude, longitude)AgentMakes the order independent of GeoNames' file order
Keep admin_code1 as supplied (no ISO crosswalk)Agent: a faithful portGeoNames already carries the official post-2016 INSEE region codes; the old builder did the same
producer: custom, not geonamesAgentThe builder has fixes and guards on top of the raw export
Remove row, distinct-code, shape-breakdown and duplicate counts and the "Last updated" line from README; standard accuracy descriptionOrchestrator instruction (P5 brief doc rules, which the brief labels user decisions)
Rewrite the README duplicates sections for the widened keyAgent, after the user's decisionThe old text said every dropped repeat was a coordinate-only duplicate, which was wrong for 4 groups

Transforms carried over

Folded from fetch_source.py → build_base.py → integrity_checker.py → package.py:

  • FR.txt read with the stdlib csv module as text (QUOTE_NONE), values copied verbatim. That alone avoids the defects of the pre-2026-08-24 legacy parser: postal codes zero-padded to 15 characters, admin_code1 as 11.0, literal nan cells, punctuation stripped from place names.
  • postal_code is text in the French shape: 5 digits plus an optional CEDEX / CEDEX N / SP NN / AIR / CITYSSIMO suffix.
  • Rows GeoNames leaves without a region (98799 Clipperton Island) sort first.
  • alternative_city_name is empty (not supplied by this source).
  • The old integrity asserts are now ProducerErrors: every region code known and all 13 present, one name per region code, French postal shape, no .0 codes, no nan/none/null cells, the (widened) primary key unique, at most 10 coordinates outside metropolitan France and Corsica. No row-count asserts.
  • Old builder, scripts/README.md and publish_r2.py removed. ATTRIBUTION.txt unchanged.

Open questions

  • Resolved (user, 2026-10-05, pc-28e.7.59): the smallest-(latitude, longitude) tie-break is kept.
  • Resolved (2026-10-05, pc-28e.7.60): added a dated correction to .licensing-notes.md naming the 4 distinct communes. Was: .licensing-notes.md is now wrong in places. It still says the 26 dropped rows are 24 groups of coordinate-only duplicates and gives 51,585 published rows. 4 of those groups are distinct communes, which are now kept. It was left alone and isn't filed under pc-hfz (pc-hfz has an es note but none for fr).
  • Admin-unit counts ("96 departments", "13 regions") stay in the README by rule; any remaining counts belong to the pc-hfz sweep.

Provenance

Reconstructed 2026-10-05 from the agent transcript d33e1b7e-5027-498a-a97f-ba914ee2524e/subagents/agent-a8631f0e94344e8e3.jsonl (both hand-backs), the orchestrator session d33e1b7e-… (the AskUserQuestion answer, the SendMessage to the agent and the landing), commits f34ad50 and 23d7af9, the bd notes and close reason, and docs/parity/p5.md.