PremiumTally

Price index as of June 2026

How PremiumTally builds its numbers

The site’s working documentation: the pipeline behind every number, the rules constraining it, and what we found wrong in the sources.

Where every number comes from

Nothing here is estimated or written by hand: every figure is computed at build time from a committed copy of a government publication.

The premium layer comes from 13 state regulators, each publishing the sample premiums insurers filed for hypothetical drivers it defines. We store the artefact with a SHA-256 hash and parse it; no two share a layout, so each parser is written to its own source.

The statutory layer covers all 51 jurisdictions, read from statute text rather than a summary. The price layer is the Bureau of Labor Statistics motor-vehicle-insurance index, series CUUR0000SETE on base 1982-84=100, whose latest reading is June 2026, when the twelve-month change fell to −4.1%.

Derived aggregates only, which is why no insurer is named

For every state, geography and source-published profile we publish the count of filed rates, the minimum, the lower quartile, the median, the upper quartile and the maximum. No average, because one figure hides the spread and the spread is the finding: in California, in San Francisco - Glen Park, one published profile draws 48 filed rates running from $5,258 to $220,416 on annual premiums, a range of 41.9×.

Per-insurer premiums exist only in the ingest run and the archived sources; they never enter the committed datasets or reach a page. A ranking site answers which company; this site answers what the distribution looks like.

How the quartiles are computed

Percentiles use the exclusive method, the default of Python’s statistics.quantiles function with four cut points. On samples this small it diverges from the inclusive method, and every pinned fixture reproduces only under the exclusive one.

Aggregates compute in integer cents from whole-dollar filed rates and round once, at the end, to whole dollars, with ties going to the even dollar.

Where a cell holds only two filed rates the exclusive method can place a quartile outside the range observed; that quartile is clamped to the lowest and highest filed rate in the cell, so nothing published falls outside the data behind it.

Sample size travels with every figure: 102 filed rates behind the largest cell, 14 behind the smallest.

Where the driver profiles come from

Bands come from the source’s own published profiles. If a regulator rates six drivers we publish six, and we never interpolate one it did not rate. Every page is keyed to a profile identifier that must exist in the ingested data, and the build fails on a missing identifier or on a profile dropped without a written reason.

One regulator rates each lettered example four times, once per pairing of vehicle and liability option, at four different premium levels, so all four ship. Another writes its drivers as narrative scenarios, so the band label is its own scenario text. A third publishes mileage bands that differ per persona, including two a hundred miles apart that are real rather than a typo.

The two-source rule for statutory figures

Every statutory figure is read from the legislature’s own statute text and corroborated against the state insurance department or motor-vehicle agency, and ships only when the primary text confirms it. Where they disagree, or only a secondary reproduction can be reached, the jurisdiction publishes no figure and links the statute.

Today 45 of the 51 jurisdictions meet that bar, and 6 render a not-yet-verified block instead. Commercial statute mirrors are never authorities on currency: one still showed a superseded liability minimum for a state, stamped current as of a date years past, after that legislature had raised it. One indexed figure is withheld for the same reason, its bulletin being unreachable.

The accuracy gate has three legs

The first is a source-snapshot diff: every artefact is stored with a SHA-256 in a manifest and every build recomputes it. Drift fails the build; a new edition becomes a new dated record, never a mutation of the old.

The second is extracted-value fixtures: hand-verified values pinned in test files and recomputed on every run, plus a full geography column re-derived by hand and asserted equal to its committed cell.

The third is the strongest, because the Bureau of Labor Statistics publishes both the index levels and its own percent changes. We recompute the change from the levels and must reproduce the Bureau’s published rounded figure exactly before the site builds; it currently reproduces the figure published for June 2026 of −4.1%.

Six-month and annual premiums are never mixed

Six-month premiums come from 6 regulators: Arizona, District of Columbia, North Dakota, Nevada, Oklahoma and Utah. Annual premiums come from 7: California, Florida, Hawaii, Maryland, New Hampshire, Texas and West Virginia. Every figure names its term.

We never convert between them, because doubling a six-month premium is not the same claim as an annual one, and we never place two states’ figures in one table. One regulator’s heading reads as annual while describing its publication cycle; its policy term is six months, and reading it otherwise would double every figure in that state.

Undated sources

Of the 13 premium sources, 3 publish no rates-effective date anywhere: Florida, Nevada and Utah. An undated figure cannot carry a freshness stamp, so those pages say the effective date is not published by the source, in those words, show a retrieval date instead, and carry no year in the title.

Where a date exists it is the document’s own: the oldest survey is Arizona, rates effective March 2023, the newest Maryland at August 2026. Where rates are more than eighteen months old a banner sits above the table, so an old survey ships as what it is, not as current.

The jurisdictions we hold out

Some 38 jurisdictions publish nothing usable, and the reasons are recorded so nobody re-discovers them. Two regulators’ domains time out from this build connection, so both are held out rather than filled from a stale cache. One publishes a real dataset behind a form needing a driven browser. One last compiled its guide in 2018 and another in 2013. One state retired the country’s most completely specified dataset: its page redirected to an article and its file returns not-found.

The rest publish complaint indexes or filing-search tools with no per-insurer premiums; those states still get a requirements page that says so.

What we found in the sources

Reading these documents closely turns up things worth recording. They are consumer publications made by small teams; finding the defect is the republisher’s job.

  • One regulator prints the same premiums three times: a full alphabetical table, then a county subset twice. Only the master table may be parsed; the repeats triple-count the largest insurers. An edition also disagrees with itself on one insurer’s speeding-conviction cell, and the master table wins.
  • One survey prints a figure so far below every sibling figure in its edition that it fails a plausibility floor. It is excluded rather than corrected to a guess, and that cell’s sample size is one lower than its neighbours.
  • One regulator publishes two versions of the same grid, a dated PDF and an undated web page, differing on about a third of their rows. The dated artefact ships: an undated revision cannot carry a freshness claim, and two divergent rows price its youngest hypothetical driver below its older one.
  • One regulator’s current edition states no policy term. Its preceding edition said in terms that rates were six-month, and the two price the same examples at the same order of magnitude, so the term is carried forward as recorded provenance.
  • One regulator publishes no rates-effective date anywhere in its tool or data responses, describing the figures only as the most recent rate filings its office approved.
  • Smaller defects are handled the same way: scenario tables duplicated with their second halves missing, a contents page labelling drivers with an age its hypotheticals do not use, and cells printed as zero that are non-responses.
  • Retrieval is a finding too: one regulator’s domain answers scripted clients with an empty body, so a state library’s mirror is the route; one blocks on connection fingerprint rather than user agent; one blocks every client we have, so its file comes from a byte-exact public archive capture. One has discontinued its survey without saying so, logged as a freshness event.

Licensing decisions, stated plainly

The National Association of Insurance Commissioners publishes the report most often quoted as the national average, prohibits reproduction without written permission, releases its data roughly two years late, and prints its own caution against direct state-by-state comparison; it is cited in words on one explainer page and is a data source nowhere. A widely quoted industry institute permits non-commercial use only, and this site carries advertising, so it is not a source, a citation or a cross-check. One regulator permits reproduction of its guide only in its entirety, so none of it is reproduced.

Questions

Why publish quartiles instead of an average?
Because for this data an average is close to meaningless. Filed rates for one defined driver in one place routinely run several times apart from lowest to highest, and one number hides the fact a reader needs.
Why does the sample size change from page to page?
Insurers decline to quote some hypothetical drivers and regulators print the decline rather than a premium. A distribution built on 14 filed rates is a weaker statement than one built on 102.

Written and maintained by PremiumTally Editorial. Last reviewed 10 August 2026. Every figure on this page is filled from a committed dataset at build time; the build fails on any figure that does not reconcile to it.