← back to dashboard  ·  method  ·  Türkçe

What we measure, and how

This page is not marketing — it is the product. For a closed score to be trusted in this field you need either a BlackRock-sized brand or an auditable methodology. The second is within our reach.

The core claim

An official gazette is not what a state says, but what it does. It is published before it reaches the news, its date is exact, and it is binding. A speech shows intent; a regulation shows a committed resource. This system counts the second kind.

Role Gap  Δᵢ = Kᵢ − Rᵢ Capacity minus assigned role. Δ>0 revisionist pressure · Δ<0 proxy fatigue.

The binary choice was dropped: separate K and R per document

The first schema forced K/R as a binary: a document was either a capacity signal or a role signal. That was the model's most contestable assumption, and it was wrong. A “basing agreement” is both — it grows the host state's actual power and imposes an obligation on it. Under a binary schema one of those two opposed readings disappears.

each document   k ∈ [−5,+5]  ·  r ∈ [−5,+5]  ·  Δ contribution = k − r A basing agreement can now be k=+3, r=+4 → Δ contribution −1: it confers power but commits more.

Δ flow is a flow, not a stock. The Δ = K − R in the thesis layer is a stock: an actor's accumulated position, assigned by hand. The Δ measured on the dashboard is the net pressure produced within a window. They are not the same unit and sit in separate boxes.

Only their signs are compared. If the flow's sign matches the stock assumption it reads “supports”; if it is opposite, “contradicts”. This is exactly where the thesis is tested — a contradiction says either the assumption is wrong or the trend is turning.

Δ is hidden when label coverage is low. Δ is computed only from labelled documents; an unlabelled document contributes zero. At low coverage this does not merely make Δ incomplete — it makes it incomparable across time. If August is 100% labelled and March 20%, the series shows a fake August spike, precisely where fracture detection looks.

So below 85% coverage the figure is not shown as “small” — it is treated as invalid. The volume layer is unaffected: it needs no labels and is always complete, which is why fracture detection works regardless of labelling progress.

But coverage alone is not enough. Measured: in the United Kingdom's 90-day window, 16 of 16 documents were labelled — 100% coverage — yet only 9 of them were relevant. 100% coverage does not turn nine documents into a hundred. Δ now requires a second, independent condition: at least 12 relevant documents in the window.

The referee — Δ's error margin

The rest of the system measures states; this layer measures the instrument. It exists because of one finding: of 99 relevant labels, 49 carry r=0 and only 20 carry k=0. The labeller assigned zero on the role axis to half the documents, and on the capacity axis to far fewer. Since Δ = k − r, this is a systematic bias that makes every state look like a capacity builder — and indeed all four countries came out with positive Δ.

This is not a model error but a behaviour: “capacity” is something visible (an authority, a budget, a permit), whereas “role” is a theory-laden judgement. Writing zero on the axis you are unsure about is a reasonable reflex — but done 99 times in a row it produces an artefact, not a measurement. A human cannot do this; they get tired and ask “am I always writing zero in this field?” A machine does not ask.

On 20 August 2026 the same 40 documents were labelled blind twice more: once by the same model (test–retest), once by a different model (inter-rater). The sample is seeded, so all three rounds cover exactly the same documents.

κ (quadratic weighted)test–retestinter-rater
relevance+0.754+0.441
field+0.847+0.908
k (capacity)+0.773+0.779
r (role)+0.902+0.798
k drift (per document)−0.45−0.44

The expectation was that the role axis would come out weak, being a theory-laden judgement. Measurement said the opposite — r is more reliable than k in both rounds.

But this does not solve the r=0 problem — it changes its meaning. Across all four measurements, 55–60% of documents received r=0, and two different models write zero on the same documents. So r=0 is not noise but a stable default.

A caveat: two instruments from the same model family reading the same instruction text is not independent evidence. Whether the zeros come from the documents or from the definition of r can only be settled by a human round. Reliability is not validity: an instrument that makes the same error consistently scores high.

On the capacity axis both rounds scored 0.45 points lower per document than the base round. Two independent instruments drifting the same way suggests the problem lies not in the noise of a round but in the base round itself. That is not scatter but directional drift, and across a 45-document window its total exceeds Δ itself. So the two failure modes are tested separately:

band (random) = Δ recomputed 400 times with drift-removed disagreement → a symmetric 95% interval around the point estimate
drift (directional) = where Δ lands when the mean per-document difference is applied Δ's sign must survive both. If it does not, neither the magnitude nor the direction is reported, and no thesis test is run in that window.

Combining the two into one band was the first version's flaw, and measurement exposed it: Turkey's 90-day band came out as [−18.33 … −5.45] — the band excluded its own point estimate (+2.41). The figure was not wrong, its meaning was: the band was asserting “Δ's true value is −12”, and there is no basis for that claim; we do not know which round is right. Separating them is more conservative: a narrow band plus a large drift means the instrument is consistently unstable, and a reading that looks only at the band would call it sound. That is the most misleading case of all.

The band is not produced by a second implementation of Δ, but by calling the production aggregation function again with perturbed labels. A separate formula would silently drift from the very figure it claims to measure — the same decision as in the time machine.

The weakest link is not the scores but the relevance decision. With a different labeller, 8 of the 33 documents that entered Δ in the base round dropped out entirely (24%). That is a far larger lever than ±1 point shifts: a shifted score moves Δ by one unit, a dropped document erases its entire contribution. In the first version this sat on the “not modelled” list; a measured but unmodelled figure is more dangerous than an unmeasured one, because it is assumed to have been accounted for. It is now inside the band.

The band is fed by the inter-rater round when one exists; the test–retest round is a known lower bound and serves only as a fallback, labelled as such in the dashboard.

What is not modelled (the irrelevant→relevant direction, disagreement over field assignment, documents the filter never surfaced, and systematic bias shared by both labellers) widens the band further, so the published band is a lower bound on reality.

The pipeline

  1. Ingest — daily cron, raw documents from the source. Deduplicated by hash, written to a permanent archive.
  2. Normalise — HTML/PDF → plain text. Scanned PDFs go through OCR; confidence is stored.
  3. Rule filter — weighted key terms. A free layer; it removes the obviously irrelevant.
  4. Labelling — per document: field, orientation, k and r scores, target actor, rationale. Every label is stored with its author; click a document in the panel to see its rationale.
  5. Aggregation — no model here. Code produces the score, by a deterministic formula.
  6. Publish — daily static JSON. No server-side computation.

The archive rule

Fetch a source once, never fetch it again. The raw document, its download timestamp and its hash are kept permanently. The reason: when the taxonomy changes the archive is reprocessed, the source is not refetched. If a source site deletes its history or moves behind a paywall — as happened with Japan's Kanpō archive — this is all that remains. The archive is this project's real asset, not the model output.

Role archetypes

Country colours on the globe are role archetypes: hegemon, revisionist, rising, bloc member, proxy, double-bound, isolated, buffer, periphery. If the thesis claims that states are positioned by their assigned role rather than by their power, then that is what the map should carry.

These assignments are assumptions, not measurements. The measured layer is separate and looks different on the globe: amber pillars (decision volume) and targeting links.

A single assignment is a simplification. India is both rising and double-bound, Iran both revisionist and isolated, Ukraine both buffer and proxy. Only the dominant one is shown — the same limit the binary K/R choice carried, and it must be stated as plainly.

Two images are produced: one to be looked at (archetype colours) and one to be read (each country filled with a unique code colour, no antialiasing). Clicking resolves the country by reading a pixel at the hit point's UV — running a point-in-polygon test in the browser across 177 countries would be both slow and wrong at the antimeridian.

Sources

countrysourceclasslicence
United StatesFederal Register (JSON API)APublic domain — 17 U.S.C. §105
TürkiyeResmî Gazete (HTML/PDF)CUnclear — legal opinion needed before republication
United Kingdomlegislation.gov.uk (Atom)AOpen Government Licence v3.0
PolandDziennik Ustaw (ELI API)AOfficial legal text — not copyrightable

Class A: official API or bulk download · B: regular structure, no API · C: scraping + PDF/OCR · D: restricted access. Turkish content is archived and enters the measurement, but its full text is not republished here; only the title, source link and derived score are shown.

Why not The Gazette

The Gazette was tried first for the UK, and measured: of the 516 notices published on 19 August 2026, 484 were corporate insolvencies, personal insolvencies and probate notices. The wrong source for measuring state behaviour. The real counterpart of the Federal Register is the statutory instrument stream — legislation.gov.uk.

A second trap found there: the feed returns 20 records per page and hides the rest behind a rel="next" link. Without pagination, every day with more than 20 items silently lost the remainder — more insidious than returning nothing, because partial data looks entirely normal. It was caught by noticing that the daily maximum was exactly 20 on nine separate days: a number repeating at the top of a distribution is a ceiling, not data.

The rule filter

This layer's job is not to decide correctly. Its job is to keep the obviously irrelevant half of the hundreds of daily documents away from the expensive layer. Its threshold is therefore deliberately loose: a wrong elimination is expensive, a wrong pass costs a few cents.

score = Σ(positive term weight) − Σ(negative term weight) A term in the title carries full weight, in the body 45%. Threshold 1.0; if no positive term appears in the title the threshold rises to 2.2. Every matched term is recorded — which decision was made and why stays auditable.

Negative terms do not eliminate a document, they lower its score. A strong positive term can beat a negative one.

The raised body threshold was measured, not guessed: in the first version only 39% of passing documents were genuinely relevant, and the dominant failure was documents whose body mentioned a topic term while the title was about something else entirely — a car-rental regulation passed because the word “sanction” appeared in its text. Gazette titles are deliberately descriptive.

Calibration — the one place the system can audit itself

There is no ground truth; the hand-labelled gold set is the only truth we have. The filter is measured against it at every version. The two error types are not equal: a miss is expensive (that document never reaches the labelling layer), a false pass is cheap.

versionprecisionrecallwhat changed
r139%100%first version — passes everything
r290%56%negative list + body threshold; cut too much
r389%95%gaps in the positive list closed

r2's collapse was not caused by the threshold but by omissions in the positive list: antidumping existed only in its hyphenated form (anti-dumping) and the Federal Register does not hyphenate it. A single hyphen missed ten documents. Without the gold set this would have been invisible — which is the whole point of having one.

Live calibration figures are in the dashboard's system tab, alongside the values actually in production.

OCR — and a silent error we measured

In Türkiye, presidential decisions, international treaties and board rulings are published as scanned PDFs. Their content is an image; they appear to have no text layer. Our first threshold was simple: “a PDF yielding fewer than 180 characters is scanned.” It was wrong.

We measured the distribution across 83 PDFs and it was bimodal: Poland's genuine text PDFs yield 3,400–4,200 characters per page, Türkiye's scans 4–206. Nothing in between.

206 is not a coincidence. A scanned page does carry a text layer — but it is not content, it is the masthead: “13 August 2026 THURSDAY · Resmî Gazete · No: 33339”.

A total-character threshold mistook that boilerplate for content. A 22-page scanned court ruling counted as “has text” on the strength of 459 characters and never entered the OCR queue. Silently. Of 83 PDFs, 49 were effectively scanned while only 37 were flagged.

is it scanned = (extracted characters ÷ page count) < 400 Measured per page, not in total. A wrong OCR is cheap (a few seconds of processing); a missed OCR silently becomes a wrong label.

Whose confidence is the confidence?

Human review was initially triggered by the document average. Looking at the output showed that this was wrong: the Türkiye–Saudi Arabia visa-exemption agreement runs to 17 pages with an average confidence of 58.8. But the decision text is on the first page and is clean — the low score comes from the maps and coordinate tables on the following 16 pages.

human review = first-page confidence < 70 Not the document average. In a gazette the decision text is always on the first page; failing to read the annexes is not failing to read the decision.

OCR output does not alter the raw archive; it is derived data, stored separately with its confidence. If the engine or language changes it is regenerated without returning to the source. The Turkish language pack is mandatory: without ğ, ş, ı, İ, ö, ü, ç the output silently breaks keyword matching.

Aggregation

intensity(country, window) = Σ [ score × field weight × 0.5(age/90) ] 90-day half-life: a new decision outweighs an old one, but the old one is not zeroed.
fracture = (7-day average ÷ 30-day average) ≥ 1.8  and  ≥ 3 documents in 7 days Not a prediction — a deviation measurement: “more decisions than usual are being produced here.” It does not say why.

Windows are divided by the number of days available, not by their nominal length, and fracture is not computed until the archive is at least 30 days deep. The earlier 21-day rule produced an artefact visible in the time series: Turkish fractures began exactly on the archive's 21st day, because a 30-day baseline computed from 21 days of data and still divided by 30 is roughly 30% too low, inflating the ratio by about 43%.

Against silent breakage

Gazette sites change structure without notice. When the ingest layer breaks it must not silently return empty: “zero documents today” is an alarm, not a normal day.

But some sources genuinely do not publish. The Federal Register does not publish at weekends or on federal holidays; legislation.gov.uk produces 0–9 documents a day. If we cannot tell the two apart we either miss real failures or get false alarms twice a week and stop reading alarms — the second being more dangerous. The calendar is therefore source-specific.

For Türkiye no fixed holiday calendar is hard-coded. On religious holidays the index page returns HTTP 200 with a notice instead of documents: “pursuant to Presidential Decree No. 10, the Official Gazette is not published today.” That notice is read directly — if the calendar changes the code need not, and “not published” never gets confused with “the parser broke”.

Known limits

  1. The gold set was also produced by a model. The first 92 labels were assigned document by document, by hand — but the hand was a language model's. Better than no reference; not the same as a human-verified one. The real job of weekly human calibration is to correct this set, not to enlarge it.
  2. A labeller cannot grade its own measure. The daily agent may not add its output to the gold set; if it did, the system would confirm its own error as reference and drift silently.
  3. The K/R distinction is still interpretive — the binary was removed, but whether a document deserves k=+3 or k=+4 is judgement. Consistency of the severity scale matters more than individual accuracy: Δ is a time series, and a drifting criterion corrupts it.
  4. Ease of access is inversely correlated with thesis relevance. The actors that build the containment triangles — China, Russia, India, Pakistan — are all in the hard-access class. Taking the easy path would build a Western-centric index while the thesis's actual claim stays unmeasured. Source class is therefore visible on every country card.
  5. Unsupervised loops amplify their own errors. Hence the target is not “fully autonomous”: daily automatic production plus weekly human-approved calibration.

Positioning

The claim is not “I measure the world correctly.” The claim is: I convert state behaviour into a traceable unit, and I show exactly how I convert it.

The taxonomy, weights and decay coefficient on this page are the values actually used in production — not a separate marketing text. Back to dashboard · thesis model · Türkçe