A calibrated, retrieval-augmented postal-address parser.
Live demo · Docs & blog · Getting started
Mailwoman turns free-text postal addresses into structured components (house number, street, locality, region, postcode, country) and resolves them to coordinates against an open gazetteer. It is a small transformer encoder (~30M params) that performs BIO token classification over a 33-label schema. It is not an LLM, and no part of it is generative. It is conventional named-entity recognition, which suits short, structured strings.
It runs in Node.js and the browser, without Elasticsearch or a multi-gigabyte libpostal install. The model is about 30 MB. It resolves coordinates from a SQLite gazetteer at the finest tier available: admin, postcode, street, or rooftop (US and French national registers).
npx mailwoman parse "1600 Amphitheatre Parkway, Mountain View, CA 94043"npm install mailwoman @mailwoman/neural @mailwoman/neural-weights-en-usThat installs the CLI, the neural runtime, and the US-English model weights. For French
addresses, add @mailwoman/neural-weights-fr-fr. For coordinate resolution, add
@mailwoman/resolver-wof-sqlite.
Important
Requires Node.js ≥ 24.18.0.
# Parse an address into components
mailwoman parse "350 5th Ave, New York, NY 10118"
# Other output formats
mailwoman parse "350 5th Ave, New York, NY 10118" --format tuple # or: xml
# Resolve components to a Who's On First place + coordinate (needs a gazetteer DB)
mailwoman parse "350 5th Ave, New York, NY 10118" --resolve --resolve-db ./wof.sqlite
# Geocode (admin/postcode coordinate)
mailwoman geocode "1600 Amphitheatre Pkwy, Mountain View, CA 94043"import { decodeAsJSON } from "@mailwoman/core/decoder"
import { createRuntimePipeline } from "mailwoman"
import { NeuralAddressClassifier } from "@mailwoman/neural"
const classifier = await NeuralAddressClassifier.loadFromWeights({ locale: "en-US" })
const parse = createRuntimePipeline({ classifier })
const { tree } = await parse("1600 Amphitheatre Parkway, Mountain View, CA 94043")
console.log(decodeAsJSON(tree))
// { region: "CA", locality: "Mountain View", street: "Amphitheatre",
// house_number: "1600", street_suffix: "Parkway", postcode: "94043" }The mailwoman package README and Getting
started document the
full library API: confidence, the per-stage pipeline result, resolution, browser loading, and
configuration.
If you already run a geocoding stack, three HTTP servers speak the wire formats your clients
use today. They need neither PostgreSQL, Elasticsearch, nor an osm2pgsql import:
| Package | Speaks | Start it |
|---|---|---|
@mailwoman/nominatim |
Nominatim — /search, /reverse, /status |
npx @mailwoman/nominatim serve |
@mailwoman/photon |
Photon autocomplete — /api, /reverse (GeoJSON) |
npx @mailwoman/photon serve |
@mailwoman/libpostal |
libpostal — /parse, /expand (no gazetteer needed) |
npx @mailwoman/libpostal serve |
Point geopy's Nominatim(domain="localhost:8080") at the first one, and forward and reverse
geocoding keep working. The API returns an OpenCage-style annotations block with the
IANA timezone, UN/LOCODE, EU NUTS codes, coordinate formats, sun times, and currency.
@mailwoman/annotations composes that block.
This table compares what each system needs to run. It does not compare accuracy. Mailwoman began as a fork of Pelias Parser and ships wire-compatible drop-ins for the other three.
| Mailwoman | libpostal | Nominatim | Photon | |
|---|---|---|---|---|
| Footprint | ~30 MB model + SQLite gazetteer | multi-GB data blob | PostgreSQL + planet import | Elasticsearch/OpenSearch index |
| Runs in the browser | ✓ (WebGPU / WASM) | ✗ | ✗ | ✗ |
| Parse → labeled components | ✓ with calibrated confidence | ✓ | — | — |
| Forward + reverse geocoding | ✓ (Who's On First gazetteer) | ✗ (parse only) | ✓ | ✓ |
| Autocomplete / type-ahead | ✓ (Photon-compatible API) | ✗ | ✗ | ✓ |
| Annotations (tz, currency, …) | ✓ OpenCage-style block | ✗ | ✗ | ✗ |
If you are migrating off Nominatim, Photon, or libpostal, each has a compatible server. The server keeps the same wire interface and reads a SQLite file instead of a cluster. A hosted Photon-compatible endpoint is live for evaluation:
curl "https://photon.mailwoman.ai/api?q=berlin&limit=3" # hosted — nothing to install
npx @mailwoman/photon serve # Photon-compatible /api, /reverse
npx @mailwoman/nominatim serve # Nominatim-compatible /search, /reverse, /status
npx @mailwoman/libpostal serve # libpostal-compatible /parse, /expandOpenAPI 3.1 specifications for all three ship in each package and at mailwoman.ai/openapi.
The work splits into two parts:
- The model learns the grammar. A sequence labeler trained from scratch on a diverse corpus of real and synthetic addresses decides which span is a street, a locality, or a postcode.
- The gazetteer supplies place knowledge. A provenance-tracked Who's On First database resolves parsed components to real-world places and coordinates.
Gazetteer knowledge reaches the model at inference as soft input features (anchors). The
features inform the model's decision and never override it. If you know RAG from the LLM
world, this is RAG for token classification. The confidence numbers the parser returns are
calibrated probabilities. When it reports 0.88, it is right about 88%
of the time.
For the longer version, read What Mailwoman Is.
A locale appears here only when a coordinate-graded eval backs it.
| Tier | Locales | What backs the claim |
|---|---|---|
| 1 — first-class, floor-enforced | US, FR | Per-tag eval floors block every release. FR coordinate panel (n=3000): 100% resolve, resolved-p90 6.6 km |
| 2 — trained + coordinate-paneled | IT, PT, PL, AT, CZ, DE, AU, BE, ES, NL, CH, HR, DK, FI | Per-locale coordinate panels (n=1000 each), resolved-p90 ≤ 10 km across the set; NL resolved-p50 0.05 km |
| 3 — trained, thinly measured | NO, SE | Coordinate panels exist, but residual misses are not yet fully characterized. Claims beyond this: unverified |
The browser demo carries the same coverage. Full receipts live in the scope declaration and the eval reports.
Not every query is an address. "coffee near Honolulu" is a category search. Mailwoman
distinguishes it from an address, splits the category from the location constraint, and
resolves it against poi.db, a sealed spatial layer built from Overture Places (US, Canada,
Mexico, France). Categories without a permissively licensed source, such as fire hydrants and
post boxes, abstain intentionally instead of returning an empty result that looks like an
answer. You can build that layer yourself from your own OSM extract.
@mailwoman/mcp exposes the whole toolset through an MCP server over stdio for any
MCP-compatible agent: parse, geocode, POI search, and an OverpassQL export. Both features ship
on npm as of mailwoman 7.2.1.
mailwoman poi "gas station near Springfield, IL" --db poi.dbmailwoman is the entry point to 33 published packages. The rest of the toolkit includes:
| Package | What it does |
|---|---|
@mailwoman/match |
Geocode-first record matcher: block → score → cluster (Fellegi-Sunter) |
@mailwoman/registry |
Resolve messy address records to geocoded entities, export GeoJSON |
@mailwoman/address-id |
Stable address primary key (<state>.<H3-cell>.<hash>) for joins + dedup |
@mailwoman/annotations |
The OpenCage-style annotation composer behind the drop-in servers |
@mailwoman/timezone-lookup |
Coordinate → IANA timezone (point-in-polygon, node:sqlite) |
@mailwoman/un-locode-lookup |
Place → UN/LOCODE trade-location codes |
@mailwoman/nuts-lookup |
EU coordinate → NUTS statistical regions |
@mailwoman/codex |
Postal address formatter (components → a per-country string) + the reference data behind it |
@mailwoman/mcp |
MCP server exposing parse/geocode/POI-search/OverpassQL-export to agents over stdio |
@mailwoman/poi-taxonomy |
Category lexicon behind POI-query detection: Overture taxonomy snapshot + synonym table |
Mailwoman is dual-licensed:
- AGPL-3.0-only for open-source use. You may use, modify, and redistribute the software, but you must share your modifications and, for network services, your source.
- A commercial license for closed-source/commercial use without the AGPL's source-sharing
obligation. Contact
teffen@sister.software.
Release notes live on the GitHub releases page.
The privacy policy
states what we do and do not collect. Funders and sponsors can read our machine-readable
funding.json.
Report security vulnerabilities privately per SECURITY.md.
Portions of Mailwoman derived from Pelias Parser remain
under the MIT license, and Mailwoman bundles third-party data under its own terms. See
THIRD_PARTY_NOTICES.md for the full attribution list.
Note
This section is for working on Mailwoman in this repository. If you only want to use Mailwoman, the published packages above are all you need. You do not need to clone the repo, build anything, or read further.
Mailwoman is a Yarn 4 monorepo: one root package (mailwoman) plus the scoped
@mailwoman/* workspaces that compose it. Start with AGENTS.md for the
orientation map (workspaces, where to read next, the release pipeline) and
docs/records/plan/ for the design record.
git clone https://github.com/sister-software/mailwoman.git
cd mailwoman
yarn install
yarn compile # tsc -b across all workspaces
yarn test # vitest (runs from source, no precompile)Mailwoman began as a TypeScript fork of Pelias Parser, a
rule-based engine: a tokenizer, a set of dictionary/pattern classifiers, and an
ExclusiveCartesianSolver that enumerated consistent solutions. As the neural sequence
labeler matured into the primary parse path, the rule engine was retired, and v7.0.0 deleted it
from the tree (@mailwoman/classifiers and the @mailwoman/core/{solver,classification}
implementation). The last standalone release is the frozen @mailwoman/classifiers@6.x.
Consumers of the published package use only the neural pipeline.
Fork and open a pull request against main on a feature branch. Please include unit tests.
The model-work runbook (which evals a change must pass, how to add a training-data subset) is
docs/engineering/CONTRIBUTING_MODEL_WORK.mdx.
This project stands on the work of the Pelias community, OpenStreetMap, Who's On First,
OpenAddresses, and the wider open-geo ecosystem. See
THIRD_PARTY_NOTICES.md.