Points of interest are the quiet backbone of countless geospatial products. Delivery routing, site selection, footfall analysis, insurance risk models, navigation apps and retail planning all depend on knowing what is where: the pharmacy on the corner, the fuel station by the motorway exit, the coffee chain that opened three new branches last month. Commercial POI datasets exist, but they are expensive, uneven across regions and not always fresh. Many GIS teams end up building or enriching their own from public web sources. Doing that well is harder than it first appears.
What a useful POI record contains
A minimal record includes a name, a category, an address and coordinates. A genuinely useful one adds opening hours, contact details, brand or chain affiliation, a source reference and a last-verified date. That last field matters more than most, because POI data decays constantly as businesses open, close, move and rebrand.
Where the data comes from
Public POI data is spread across many kinds of source, each with strengths and weaknesses:
- Open data. OpenStreetMap and government open data portals provide broad coverage and clear licences, though completeness varies by region.
- Official registries. Business registers, licensing databases and health inspection records are authoritative for certain categories but often lack coordinates.
- Brand store locators. Chains publish their own branch lists, usually the most accurate source for those brands.
- Business websites and directories. Useful for hours, contact details and categories, but quality and terms of use vary widely.
The hard parts
Collecting records is the easy step. The difficult work comes afterwards.
- Entity resolution. The same cafe can appear in five sources with slightly different names, addresses and coordinates. Deciding which records describe the same place requires fuzzy matching on names, normalised addresses and spatial proximity.
- Geocoding quality. Addresses that geocode to a postcode centroid rather than a rooftop can put a POI hundreds of metres from its real position, which ruins proximity analysis.
- Category taxonomies. Every source categorises differently. Mapping them to one consistent schema is tedious but essential.
- Freshness. Without regular re-collection, a dataset accumulates closed businesses and misses new ones.
- Licensing. Open data licences such as OpenStreetMap’s ODbL carry attribution and share-alike obligations, and in the EU, databases can be protected by database rights. Know what you may use and how.
Collecting from store locators and websites
Brand store locators are often the richest source for chains, but they come with technical quirks. Many return only the branches near a searched location, so complete coverage means querying a grid of points across each country. Some serve different content or block requests depending on the visitor’s country, especially when a brand runs separate sites per market. And repeated queries from a single server address quickly run into rate limits, leaving gaps that look like missing branches rather than missing data.
Good collection practice addresses these problems directly. Query at a moderate pace, cache responses, log every failed request so it can be retried, and compare counts against any totals the brand publishes. For sites that respond differently by country, the requests need to come from that country. Residential proxies are a common solution: they route traffic through IP addresses that internet service providers have assigned to real households, so a locator serves the same results a local visitor would see. Providers offering residential proxies, such as Proxy-Cheap, support country-level targeting across a wide range of markets, which helps when a dataset spans many countries.
Quality assurance
A POI dataset is only as trustworthy as its validation. Sample records in each region and check them against imagery, official sources or ground truth. Track a confidence score per record based on how many independent sources agree. Flag records that have not been confirmed recently. And measure coverage against known benchmarks, such as the number of branches a chain reports in its annual report. These checks turn a pile of scraped rows into a dataset analysts can rely on.
Responsible collection
Respect the terms of use of every source, honour robots.txt guidance and keep request rates low enough that no website is burdened by your collection. Prefer official APIs and open data where they exist. POI datasets describe businesses rather than people, but some sources include personal information, such as sole traders’ names or home-based business addresses; minimise and protect that data in line with privacy law.
Scheduling refresh cycles
Not every record needs refreshing at the same rate. Fast-changing categories such as restaurants, cafes and retail turn over far more quickly than hospitals, schools or fuel stations. A practical approach is to set refresh frequencies by category and region, re-collect high-churn categories monthly and stable ones quarterly, and trigger extra checks when a source shows signs of change, such as a sudden drop in a chain’s branch count. Tracking first-seen and last-seen dates for every record makes closures detectable: if a POI disappears from all sources across two refresh cycles, it has probably closed.
Choosing the right storage
POI data benefits from a spatial database such as PostGIS, which makes proximity queries, deduplication by distance and coverage analysis straightforward. Keep raw source records separate from the merged, cleaned layer, so every final POI can be traced back to the evidence behind it and matching can be re-run when your rules improve. That traceability also helps when customers or licensors ask where a record came from.
Bringing it together
Building POI data from public sources is a pipeline, not a one-off scrape: collection, normalisation, geocoding, deduplication, categorisation, validation and scheduled refresh. Teams that treat it this way end up with datasets that are fresher and better suited to their region and use case than many off-the-shelf alternatives. The investment is real, but so is the payoff for every spatial analysis that depends on knowing what is where.