The wine data behind this site, as JSON

Every figure on the VinSip reference pages had to be read out of a primary document before it was allowed onto a page. That compilation — the figure, the regulator, the citation and the page that explains what it means — is published here as 98 facts and five other datasets. No key, no rate limit, CORS open, CC BY 4.0.

What is published

Eight files under https://vinsip.app/api/v1/. Each one is a static file regenerated on every site build from the same module or table the pages render from, so an endpoint cannot disagree with a page — there is no second copy of anything to drift.

The endpoints under /api/v1/, with what each returns
FileSizeWhat it carries
wine-facts.json 98 facts Every regulated figure the reference pages state — duty rates, labelling thresholds, sulphite limits, permitted bottle volumes, alcohol units — quoted with the body that published it and linked to the document it was read from. Drawn from 20 pages and 28 primary sources.
country-rules.json 6 countries The legal purchase age, the excise duty regime, the VAT treatment and the trading-hours rule for each country the directory covers, with the legislation or tax authority behind each.
wine-regions.json 29 regions Wine regions with the grape varieties planted and the styles made, plus a centroid for mapping and the parent region where there is one.
directory.json 60 cities How many wine shops, wine bars, wine-focused restaurants and wineries are indexed in each city and country, out of 3172 venues, with the listing page for every row. Counts only — see below.
guides.json 63 pages The whole guide corpus: slug, title, which of the 30 locales each page was actually built in, and the markdown twin's address.
vinsip.json one object What the VinSip app does, what the free tier includes, and an explicit list of what it does not do.
status.json freshness When each dataset was last rebuilt and how many rows it carries, so a client can decide whether to refetch without downloading the large payloads again.
openapi.json OpenAPI 3.1 A machine-readable description of every endpoint above.

The licence, and what it asks of you

Everything here is CC BY 4.0. Take it, redistribute it, build a product on it, commercially or not. The one condition is attribution, and there is a better way to satisfy it than naming this site: credit the page the figure came from. Each fact carries a page.url as well as a source.url, and the page is where the number is explained, dated and put in context. A reader who follows that link can check the claim in one hop, which is the difference between a citation and a footnote.

The endpoints are files behind a CDN. There is no key to request, no quota to exceed and nothing to rate-limit, so cache them rather than polling: status.json carries the rebuild timestamp and the row counts, and it is a few hundred bytes.

What the fields mean

A JSON schema can say that claim is a string. It cannot say that the string is a sentence with its own attribution inside it and must not be split. The semantics below are the part implementers get wrong, and every one of them is a mistake somebody can make while reading the payload correctly.

claim — a sentence, not a value

Each fact is written the way it should appear: the number and the body that published it, in one sentence. “The United Kingdom charges £30.64 per litre of pure alcohol on wine between 8.5% and 22% ABV from 1 February 2026” is one claim. Extracting 30.64 from it and presenting the figure bare drops the alcohol band, the tax base and the date — three qualifications that decide whether the number is right. If you need a numeric field, parse the claim and keep the sentence next to it.

source versus page

source.url is the regulator's own document: HMRC's duty rates, the European Commission's excise tables, the eCFR, the OIV, the NIAAA. It is the authority. page.url is our page, which explains what the figure means and how it interacts with the others. Cite the source for the number and the page for the explanation; citing us for the number alone is weaker than the data deserves.

verified — a date, and a warning

Tax and labelling figures go stale on a schedule. The UK alcohol duty rate is superseded every February; the EU excise duty tables are reissued; two of the US labelling measures in this dataset are proposals under an extended comment period and are not rules. verified is the date the figure was last read at its source, not the date of the build that emitted the file. Quote the date with the figure and a reader can tell how much weight it takes.

purchase_age — the age to buy, not the age to drink

Several countries set a lower age for drinking wine with a meal or in a private home than for buying a bottle off a shelf, and the two are genuinely different rules. The field is the purchase age. Where the distinction matters, the country's statements array says so in a sentence; do not infer a drinking age from the number.

Why duty is prose and not a number

The obvious shape for country-rules.json would be a duty_per_litre column, and it would be wrong. The regimes differ in kind, not only in value: Britain charges per litre of pure alcohol, so the duty on a bottle depends on its strength; France charges per hectolitre of wine, so it does not; Germany charges nothing at all on still wine and a sparkling-wine tax above 6% ABV. One column forces all three into a shape none of them has, and every consumer of it then compares quantities that are not comparable. The regimes are therefore statements, and the figures live in wine-facts.json with their sources.

locales — what was built, not what could be

A guide's locales array lists the locales that page actually exists in. The site is thirty-lingual on the families that earn organic traffic and English and French on the venue directory and the wine regions, because directory prose is computed from the listings rather than written. A locale missing from the array has no page at that path. Requesting one anyway returns the English page rather than a 404 — the navigation downgrades a path to its English original whenever the locale does not build that family — so an absent locale is a fact about the translation, not about the URL.

venues — a count of our corpus, not a census

A city's venues figure is how many venues of that kind VinSip has indexed there. It is not how many exist. Coverage is city-first: 60 cities are indexed because a specialist wine trade exists in them at a scale worth mapping, and a listing with fewer than six venues of a kind is left out entirely, because it cannot fill a page honestly and padding it would be the dishonest fix. A city absent from the file is an editorial threshold, not a finding about the city.

The dataset that is deliberately incomplete

The obvious thing to publish from a site with 3172 venue listings is the venue table. It would be the most linkable file here, and it is not ours to publish. Those records — names, addresses, phone numbers, ratings, review counts, opening hours — were retrieved from the Google Places API, whose terms permit displaying that content in our own interface and do not permit redistributing it as a file. Publishing it keyless under a CC licence would be licensing content we do not own.

So directory.json carries counts and page URLs. A count is our own measurement of our own corpus and carries none of Google's content. If you need the listings, the page on each row is where they are shown, under the licence they came with. The omission is stated in the payload itself, in a not_published block, because an unexplained gap reads as an oversight and the next person to look at the file would otherwise “fix” it.

Markdown twins

Every indexable page on this site has a markdown copy beside it at <path>/index.md, and the HTML URL returns the same markdown to a request carrying Accept: text/markdown. Agent stacks split roughly evenly between the two conventions, so both work. The saving is not marginal: a city listing is around 180 KB of HTML, because this site inlines its whole stylesheet into <head> and embeds the venue list twice — once as cards and once as a JSON blob for the map — and it is a few kilobytes as markdown.

/llms.txt indexes the site and this API. /llms-full.txt carries the full text of the hubs, the documents and every reference page. The programmatic families are too large for one file and are reachable through their own twins.

MCP and A2A

A read-only Model Context Protocol server runs at https://vinsip.app/mcp (Streamable HTTP, no authentication), described by its server card. It exposes the same data as six tools: the tax and purchase rules for a country, a search over the sourced figures, a region lookup, venue coverage for a city, a guide lookup, and what the app does and does not do. Every tool returns the source URL alongside the figure, so an agent that used a tool still has a page to cite.

An A2A endpoint answering the same questions in prose runs at https://vinsip.app/a2a, described by its agent card. Neither endpoint writes anything, charges anything or reads anything about a user — there is nothing here to authenticate, because there is nothing here to protect.

Four agent skills are published for stacks that load them: wine tax and duty, reading a wine label, venue coverage by city, and what a wine app can and cannot tell you. They spend most of their length on the ways of reading the data wrongly, which is the part a schema cannot carry.

Discovery, without parsing any HTML

Every response from this origin carries Link headers pointing at the API catalogue, the OpenAPI description, this page and the AI catalogue, so an agent that has fetched anything at all can find the machine surface from the response alone. RFC 9727's catalogue and the ARD manifest list the same things in documents.

RFC 9728 Protected Resource Metadata is published for an unprotected resource on purpose: it declares an empty authorization_servers list rather than being absent, so a client can tell “declared none” from “not published” and learns in one request that there is no authentication to do. For the same reason there is no oauth-authorization-server document and no openid-configuration: both would have to name a token endpoint and a JWKS URL that answer, and inventing either would teach every agent that reads it that this site misdescribes itself. /auth.md says the same in words.

What this API will not do

It does not identify a wine. Scanning happens inside the VinSip app, on the user's own device, and there is no endpoint on this origin that takes a photograph. It does not score wines, value them, authenticate them or say whether a particular bottle is sound — those are properties of the liquid, and no dataset holds them. It sells nothing, so there is no checkout to integrate and no commerce protocol to negotiate. And it reads nothing about anyone: there are no accounts on this site and no request to these endpoints is associated with a person.

Corrections

A figure here that is wrong is a worse problem than a figure that is missing, because the whole point of the compilation is that somebody opened the source. If you find one that has been superseded or misread, write to [email protected] with the document you read it in. Corrections land on the page and in the endpoint in the same build.

VinSip
VinSip Scan wine labels. Get recommendations.
Get the App