How a moat gets built
Three layers, three properties, and a disclosure posture that publishes the rulebook on purpose.
A data moat is not a dataset. It is a position: the only complete, normalized, continuous record of a market that nobody else is watching. Building one has the same shape every time.
The three layers, built in reverse order of glamour
One — the rulebook. The governing rules are public. They are also scattered across dozens of jurisdictions, written for humans, published as PDFs, and revised without notice. Normalizing them to machine-readable form comes first because everything downstream is uninterpretable without it.
Two — the index. The recorded documents that attach a rule to a specific asset, keyed to whatever identifier the jurisdiction actually uses, and durable for decades.
Three — the events. Listings, sales, transfers, expirations. The layer everyone wants to start with, and the layer that is worth nothing without the two beneath it. An event you cannot interpret is a row.
The three properties
Accumulated normalization. Thirty jurisdictions, thirty schemas, one output. The thirty-first is nearly free. A new entrant redoes all thirty. Their cost is linear in jurisdictions; ours is asymptotically zero.
A time series that cannot be back-bought. History is the only component that cannot be reconstructed at any price once the window has passed. Every other component — the rules, the documents, the software — can be acquired late. Start the clock early. Never sell the history.
Work that is tedious rather than hard. There is no single scraping target and no one large engineering problem. There are thirty small interpretation problems, each slightly different, none of them interesting. That is precisely the shape a large vendor is structurally worst at and a small operator is best at.
The disclosure posture
Most of this gets published, and the reasoning is not generosity.
Publish the rulebook in full, with quotes and sources. It derives from public documents. Withholding it captures almost no value, and publishing it makes SocioClimate the reference point for reading them. It also opens a correction loop: thirty published jurisdictions is thirty free proofreaders.
Publish existence, sell depth. Whether a parcel carries a restriction is the public good. The terms, the expiration date, and the portfolio view are the product.
Never publish the time series. Publish findings derived from it instead. Nobody becomes a reference by holding a file.
What makes a market a candidate
Four conditions, and a market has to meet all four.
- A recurring stream of records nobody else sees. Not a one-off dataset. A stream, or there is no time series to accumulate.
- Governing rules that are public but scattered. Public, or the rulebook cannot be published. Scattered across jurisdictions, or normalizing them is not worth anything.
- A segment too small for a national vendor to bother with. If the market is large enough to warrant a real budget, it already has one.
- An administrative rather than market-set price. Where the number comes out of a formula or a committee rather than a negotiation, there is no public price record to compete with — and no reason for one to appear.
If this describes what you already have
Door one
Partner with us
You already generate the raw material. Every day your business produces records that exist nowhere else and that nobody has normalized — the by-product of doing the work. We turn that stream into an asset you own a share of, and we do the part you have no reason to be good at: schema design, normalization across jurisdictions, productization, and finding the buyer.
You keep operating. Nothing about the day job changes.