Skip to content

Research · OBS-005

OBS-005: Attested model-weight mining on Pearl: the weight-provenance scan

A pre-registered scan asked one question of every candidate block on Pearl: was the operand it committed to byte-identical to a tensor from a published checkpoint? For a set of blocks in the chain’s first weeks the answer is yes, recomputable by anyone from raw block bytes. The commitments then stop, and nothing after them matched. This page is the result, its bounds, and the commands to check it.

RecordedRecorded 2026-08-09
Matched blocks

1,110

Blocks whose committed operand was byte-identical to a published checkpoint tensor, out of 53,718 candidate blocks: 53,717 recorded by the scan as shape-eligible, heights 1 to the frozen tip h96,774, plus 1 added by the coverage-closure round, which carries its further match at h68,332.

Unit of analysis: blocks, resolving to far fewer coinbase addresses and fewer operators still. Block counts are not independent observations.

Control matches

0of 840,000 pairs

Nothing in the reused control ever matched: 2,130,000 pairs across the runs, and 0 of 2,372,384 in the era-stratified redraw.

The control is an implementation canary against gross false positives. It has no power against false negatives. A single control match would have voided the run.

Last attested block

h68,332

05 Jun 2026. Nothing matched in the scanned candidate blocks after it, through the frozen tip h96,774; blocks above that bound were never scanned.

Basis: Combined corpus, the union of the 4 hashing runs fixed by the OBS-005 section 10 errata.

Who mined them, and when

F0a · Who mined these blocks, and when

Matched blocks by calendar day and by coinbase address. Basis: DS-002-weight-scan-v1, 1,066 matched blocks.

Exploratory (post-hoc). The calendar-day framing is post-hoc: the pre-registration bins by height, not by date.

By calendar day

2026-04-282026-04-292026-05-25

656 of 1,066 matched blocks (61.5%) were mined on 2026-04-29, by 7 addresses, over 21 days with any match at all.

By coinbase address

rank 1 · 432 blocksrank 46 · 1 block

46 addresses, 4.27 effective as 1/HHI. The top address holds 432 blocks and the top three hold 75.6% between them. Addresses are not identities, and one operator can hold many.

Matched blocks by calendar day and by coinbase address, on the basis named on the chart. The unit is blocks and the population is a handful of operators during the chain’s launch weeks, not the network; addresses are not identities, and one operator can hold many. Data tier: frozen. Scanned to h96,774; blocks above that bound were never hashed, and the live tip is higher and moving.

F0b · Match rate by declared batch dimension

Every rate carries its own numerator and denominator. Strata are the declared m, which is a claim; the match is what is proven.

  • Varying m (not 8,192 and not 32,768)86.4%1,036 of 1,199
  • Constant m = 8,1923.0%73 of 2,469
  • Constant m = 32,7680.0%0 of 50,049

    Measured zero: 50,049 blocks scanned, none matched. Not an absence of data.

Over every stratum together, 1,109 of 53,717 scanned candidate blocks matched (2.06%). The two constant strata are declared-shape mass points: n 57,344 k 8,192 at 31,831 blocks, n 28,672 k 4,096 at 16,530 blocks, n 10,240 k 8,192 at 1,688 blocks. Blocks from the DS-006 closure cells the frozen classifier omitted; scanned and matched, with no decoded declared parameters in the DS-002 extract.

Match rate within each declared batch-dimension stratum, each row labelled with its own numerator and denominator. Shape declarations are claims and the match is what is proven; a stratum reading zero means those exact published bytes were not committed, not that nothing was computed. Data tier: frozen. Scanned to h96,774; blocks above that bound were never hashed, and the live tip is higher and moving.

F1 · When it happened, and when it stopped

Matched blocks per height bin, over the denominator you choose. The peak is bin-width dependent: 9.31% of scanned candidates at 5,000-block bins against 24.27% at 1,000.

  • Matched blocks in bin
  • Anchors and consensus changes
One cell per 200 heights, genesis to the frozen tip h96,774. 153 of 484 cells hold a matched block; the last is h68,200 to h68,399, standing alone well after the cluster ends. Cell height scales with the square root of the count, so the rug shows where matches are, not how many.

Share of scanned candidate blocks whose committed operand was byte-identical to a published tensor, in 5,000-block bins, showing when such commitments occurred and when they stopped. Rates depend strongly on which population is in the denominator, so read the stratified view first; this shows neither that a model was served nor who mined. Denominator in this view: scanned blocks in the varying-m stratum. Data tier: frozen. Scanned to h96,774; blocks above that bound were never hashed, and the live tip is higher and moving.

What a match proves, and what it does not

Read this before any figure below. The scan proves one thing precisely, and the distance between that thing and what a reader might assume it means is the whole reason this page exists.

What this proves

proven
  • The operand committed by that block was byte-identical to a specific published tensor, under a key derived from that block’s own on-chain header and mining config.
  • The sampled output entries consensus verified were computed from strips of that tensor, after a publicly derivable perturbation.
  • The block was accepted by mainnet consensus, so the commitment is on chain and anyone can recompute it from raw block bytes.

What this does not prove

not proven
  • Not that a model was run end to end, and not that any request was served.
  • Not that the computation was fresh rather than replayed or precomputed.
  • Not that the full declared arithmetic was performed; consensus opens a sample of the result, never all of it.
  • Not the identity, affiliation or intent of any operator: a match identifies weights, and the same public weights are available to anyone.
  • Not anything about private, fine-tuned or locally requantized weights, which are indistinguishable from noise on chain.

Unit of analysis: blocks, resolving to far fewer coinbase addresses and fewer operators still. Block counts are not independent observations.

Pre-registered · frozenRecorded 2026-08-09PREREG-002@d4cb529b (external link)PREREG-003@d2b83c55 (external link)Tool 609cf9f90191-wb58f7a2Decoder v1docs/research/datasets/DS-002-weight-scan-v1 (external link)sha256 6d28d8f4e1abScan tip h96,774Basis 1,110 matched blocksFrozen before the scan ran; corrections ship as a new numbered document, never as an edit.

What the matched blocks were

F2 · The two extinctions

Eligibility fell from 99.98% of blocks to 4.00%; the matched share peaked at 9.30% and reached a measured zero. The two are different quantities and only one is a behaviour.

Exploratory (post-hoc). The eligibility series is post-hoc: PREREG-002 registers the matched-share series, not the shape-filter series drawn beside it.
  • Matched share of all blocks
  • Eligible candidate share of all blocks
  • Measured zero (open circle, strip below)

Two different quantities against the same block axis: the matched share of all blocks, and the share of all blocks eligible to be candidates at all. The second is a property of the shape filter, not a behaviour rate, and neither series is evidence about work that is not a published checkpoint. Data tier: frozen. Scanned to h96,774; blocks above that bound were never hashed, and the live tip is higher and moving.

F3 · What the matched blocks committed to

13 of 18 searched candidate shapes matched at least once; Llama-3.3-70B / gate_up_fused / tp1 alone holds 64.9% of the matched set.

By model
Llama-3.3-70B1,034 · 93.2%
Gemma-4-31B44 · 4.0%
Llama-3.1-8B32 · 2.9%
By layer family
gate_up_fused868 · 78.2%
o145 · 13.1%
qkv_fused97 · 8.7%
By tensor-parallel degree
tp1989 · 89.1%
tp2121 · 10.9%

Matched blocks by the model, layer family and tensor-parallel degree of the tensor they committed to. This is the composition of the matched set, not a ranking of models on the chain, and a family absent here was searched and matched nothing rather than never searched. Data tier: frozen. Scanned to h96,774; blocks above that bound were never hashed, and the live tip is higher and moving.

F4a · How often each tensor was matched

328 distinct tensors, median 2 matches each, maximum 16; 154 were matched exactly once.

How often each distinct tensor was matched, and which layer indices were reached in each family. Broad layer coverage rules out a single tile being replayed; it does not establish that a forward pass ran, and layers with no buffer are drawn as unsearched rather than as unmatched. Data tier: frozen. Scanned to h96,774; blocks above that bound were never hashed, and the live tip is higher and moving.

F4b · Which layer indices were reached

One cell per transformer layer index. Llama-3.3-70B / gate_up_fused / tp1 matched at 80 of 80 layers.

Llama-3.3-70B / gate_up_fused / tp180/80
Llama-3.3-70B / o / tp164/80
Llama-3.3-70B / gate_up_fused / tp250/80
Llama-3.3-70B / qkv_fused / tp135/80
Gemma-4-31B / gate_up_fused / tp124/60
Llama-3.3-70B / o / tp220/80
Llama-3.1-8B / gate_up_fused / tp117/32
Llama-3.3-70B / qkv_fused / tp29/80
Gemma-4-31B / gate_up_fused / tp27/60
Llama-3.1-8B / qkv_fused / tp13/32
Gemma-4-31B / o / tp22/60
Gemma-4-31B / o / tp11/60
Llama-3.1-8B / o / tp11/32
Llama-3.1-8B / gate_up_fused / tp20/32
Llama-3.1-8B / o / tp20/32
Llama-3.1-8B / qkv_fused / tp20/32
Llama-3.3-70B / o / tp40/80
Llama-3.3-70B / o / tp80/80
  • 1 match
  • 2 to 4
  • 5 to 9
  • 10 or more
  • searched, no match
  • never searched

How often each distinct tensor was matched, and which layer indices were reached in each family. Broad layer coverage rules out a single tile being replayed; it does not establish that a forward pass ran, and layers with no buffer are drawn as unsearched rather than as unmatched. Data tier: frozen. Scanned to h96,774; blocks above that bound were never hashed, and the live tip is higher and moving.

F5 · The declared batch dimension

631 distinct values over 1,109 matched blocks. Modes: 16,384 (152), 8,192 (73), 14,000 (46).

  • Matched blocks
  • vLLM default max_num_batched_tokens (16,384)

The batch dimension each matched block declared, with the reference stack’s default of 16,384 marked. The batch dimension is operator-configurable and is a declaration, so no configuration and no workload is inferred from it here. Data tier: frozen. Scanned to h96,774; blocks above that bound were never hashed, and the live tip is higher and moving.

Declared m above 16,384

15.4%

171 of 1,109 matched blocks declared a batch dimension above the reference stack’s default.

Denominator: matched blocks with decoded declared parameters. vLLM default max_num_batched_tokens is the reference default, not a consensus rule: any m in range is legal.

What would have shown a false positive, and what was never looked at

F6 · The negative control, at two scales

0 matches in 2,130,000 control pairs on the reused control, and 0 in 2,372,384 on the era-stratified redraw.

Row 1: true scale, every hashed pair

1,110 matching pairs · 0.0162%

6,832,618 (block, buffer) pairs hashed across the 4 runs, each run’s own control pairs included. One matched block is one matching pair, because a committed root can equal at most one distinct-bytes buffer.

Row 2: legible scale, each population on its own track

  • Scanned candidate blocks1,110 of 53,718

    2.07% of the scanned candidate population.

  • Reused control, h41,237 to h52,2800 of 1,000

    2,130 distinct buffers, 2,130,000 pairs, 7 declared shapes, 824 of 1,000 of them at n 16,384 k 1,024.

  • Era-stratified control (DS-007-stratified-control-v1)0 of 952

    7 buffer sets, 2,372,384 pairs, drawn across 5 era strata: pre-boundary 200, post-boundary 200, moe 200, dense-only 200, rank-penalty 152.

Each track is its own population normalised to full width, so the widths are not comparable between rows. The block counts are in the table.

The control is an implementation canary against gross false positives. It has no power against false negatives. A single control match would have voided the run.
PREREG-002 section 1, verbatim, frozen at d4cb529b8602.

The reused control covers h41,237 to h52,280 only, so it says nothing about the region after the matched cluster ends. That is why the control was redrawn: DS-007-stratified-control-v1 draws 952 blocks across 5 era strata spanning h1 to h96,405, with 45 declared shapes excluded from the draw, and matched 0 of them.

The negative control at both its true scale and a legible one: the same strata as pairs and as blocks. The control is an implementation canary against gross false positives with no power against false negatives, and the pair count is not a count of independent trials. Data tier: frozen. Scanned to h96,774; blocks above that bound were never hashed, and the live tip is higher and moving.

F7 · What was scanned, and what was not

The scan froze at h96,774. The live tip could not be read, so the ruler ends at the snapshot tip h97,590 recorded in the datasets.

Scanned
h1h96,774
53,718 candidate blocks hashed against 4,244 buffers over 4 runs, 6,832,618 pairs.
Never scanned
h96,775h97,590
Blocks above the frozen tip were never hashed against any buffer. Nothing on this page is a statement about them.

The scanned height range against the chain, with everything after the frozen tip drawn as unscanned. Every zero on this page is bounded by this range, by the buffer sets listed beside it, and by the layouts and weight variants that were never searched at all. Data tier: frozen. Scanned to h96,774; blocks above that bound were never hashed, and the live tip is higher and moving.

F7 · The search space, and its edges

PREREG-003 enumerated 31 candidate cells, of which 18 were already covered and 3 were scanned to close the gap. 0 recorded candidates remain unhashed.

Recorded but unhashed at first, since closed

Llama-3.3-70B / o / tp4
338 blocks
n 8,192, k 2,048. Closed by DS-002b-o-tp48-v1: 0 matched.
Gemma-4-31B / gate_up_fused / tp1
32 blocks
n 43,008, k 5,376. Closed by DS-002b-gemma-v1: 32 matched.
Gemma-4-31B / gate_up_fused / tp2
9 blocks
n 21,504, k 5,376. Closed by DS-002b-gemma-v1: 9 matched.
Llama-3.3-70B / o / tp8
5 blocks
n 8,192, k 1,024. Closed by DS-002b-o-tp48-v1: 0 matched.
Gemma-4-31B / o / tp2
2 blocks
n 5,376, k 4,096. Closed by DS-002b-gemma-v1: 2 matched.

386 blocks in total, of which 43 matched once a buffer existed for their shape. Until then they were counted as candidates and never tested, which is a gap in coverage, not a negative result.

Never searched at all

FP8 layer families
down_proj on both Llama checkpoints, and qkv layers 0 to 39 on the 70B and 0 to 15 on the 8B, are FP8 and were excluded by PREREG-002 section 3 before the run.
Private, fine-tuned and locally requantized weights
Only public static checkpoints are externally verifiable. A private model, a customer fine-tune or a local requantization is indistinguishable on chain from synthetic noise.
Padding and transposition conventions beyond the registered variants
Three alternative layouts were registered and probed to zero (column-major byte order, and two gate_up tp2 fusion orders). Any convention outside that set remains unsearched.
MoE expert-stacked layouts
Never in the candidate universe: PREREG-002 declares the extraction problem open for expert-stacked tensors.
The activations operand
hash_a is not attested by anything here. The match is one-sided toward the weights operand.

Buffer sets by run: DS-002-weight-scan-v1 840 · DS-002b-gemma-v1 330 · DS-002b-o-tp48-v1 960 · DS-006-coverage-closure-v1 2,114. Every zero on this page means “not these exact bytes, in this range, under these layouts”, and never that nothing was computed.

Who this does not identify

What a match does not tell you about who

Frozen text, verbatim

A match identifies weights, not a miner. These three readings were written down before the data was read, so that whichever the data supports is not a post-hoc rescue.

  1. 01Operator bootstrap fleet. Matches concentrated early, in the unattributed stratum, are consistent with Pearl Research Labs (or an affiliate) mining its own published checkpoints to bootstrap the chain. The published checkpoint’s own repository timeline is recorded in the dataset manifest for exactly this reason, and it cuts both ways: a checkpoint available from genesis is equally available to anyone.

  2. 02Third-party miners running the published stack on public weights.

  3. 03Mimicry. Anyone can download the same weights and mine them; a match proves the operand, never the identity or intent of the operator.

Keshi does not name a party. No coinbase-based attribution beyond the existing pool-label ladder enters this analysis, and the unattributed stratum stays unattributed. Where the strata cannot distinguish these readings, OBS-005 says so explicitly rather than choosing one.

Unit of analysis: blocks, resolving to far fewer coinbase addresses and fewer operators still. Block counts are not independent observations. OBS-005 section 11 carries this text in full.

Attributed to a labelled pool

0of 1,109

Every matched block with decoded attributes is unattributed: 1,109 of 1,109 carry no pool label. Shares are never redistributed onto labelled pools.

Denominator: the matched blocks whose declared attributes the frozen extract decoded. 1 further matched block from the coverage-closure round has no decoded extract record and is counted in neither column.

The matched set, by stratum

pre-moe
1,109 · 100.0%
OFFICIAL_CONSISTENT
1,055 · 95.1%
MODEL_SHAPED_CUSTOM
54 · 4.9%

Era and certificate class over the same 1,109 blocks. Every one of them predates the MoE fork.

Check it yourself

F9 · The anatomy of one match

Four steps, all of them recomputable from raw block bytes and a published checkpoint.

header[0:76]mining_config[0:52]both halves on chain, fixed before the resultblake3job_key (32 bytes)blake3(tensor_bytes, key = job_key)tensor bytes extracted from the public checkpointcomputed root=hash_bcommitted in the certificatebyte-identical, no tolerance
  1. 01Take the block’s own bytes

    The first 76 bytes of the wire header, before the parts that depend on the result, followed by the 52-byte mining configuration from the certificate. Both halves are on chain and neither is chosen after the fact.

  2. 02Derive the per-block key

    BLAKE3 over those two halves gives job_key, a 32-byte key unique to this block. Because it is per block, the same weights hashed for a different block produce a different root, and reuse of weights across blocks is undetectable.

  3. 03Hash the published tensor under that key

    Take the tensor bytes as extracted from the public checkpoint and compute keyed BLAKE3 with job_key as the key. Nothing here is Keshi-specific: the buffer comes from the published checkpoint and the hash is the reference implementation.

  4. 04Compare with what the block committed

    The certificate carries hash_b, the commitment to the weight-side operand. If the two are byte-identical, the operand that block committed to was those exact published bytes. That is the whole claim.

How one match is checked: the block’s own header and mining config derive a per-block key, the published tensor is hashed under that key, and the result is compared to the certificate’s committed value. Equality proves which bytes were committed, and nothing about what was computed with them afterwards. Data tier: frozen. Scanned to h96,774; blocks above that bound were never hashed, and the live tip is higher and moving.

Verify it yourself

Block h22,816 is one of the matched blocks, recorded in OBS-005 section 8 as independently reproduced from raw chain bytes. Reproducing it needs three things and one command.

  1. 01The raw block

    Any Blockbook instance you trust serves it. The script defaults to a loopback node, so pass --blockbook to point it somewhere else.

  2. 02blake3

    pip install blake3. The script uses the reference Python binding, not Keshi code.

  3. 03The tensor buffer

    Extracted from the public checkpoint with scripts/extract-weights.py, using the layout PREREG-002 section 3 fixes. The dataset manifest pins its sha256.

The command

python3 scripts/verify-hashb-match.py \
  --height 22816 \
  --buffer gate_up_fused.tp1.s0.L000.bin \
  --blockbook https://<a-blockbook-you-trust>

The script defaults to a loopback Blockbook, so the last flag is the one to change. Source: scripts/verify-hashb-match.py (external link). The buffer is extracted with scripts/extract-weights.py; the dataset MANIFEST pins the sha256 of every one of the 843 files.

What it prints on a match

height 22816  block <64-hex block id>
certificate v1   declared m <m>  n <n>  k <k>  rank <rank>
header self-check       : SHA256d(header) == block id  ✓
job_key                 : <64-hex>
buffer                  : gate_up_fused.tp1.s0.L000.bin (<n·k> bytes)
on-chain hash_b         : <64-hex>
blake3(buffer, job_key) : <64-hex>

MATCH: the operand this block committed IS these weight bytes.
Proves committed-operand identity only: not that inference ran,
that a customer was served, or that the computation was fresh.

The last three lines are the tool’s own, printed on a successful match. It trusts no Keshi code: it fetches the raw block, splits the certificate from the header, checks that SHA256d of the header reproduces the block id, derives the key from those bytes and hashes the buffer with the reference blake3 binding. A mismatch exits non-zero. The bracketed values differ per block and per tensor and are not transcribed here; the three closing lines are the tool’s own, printed exactly as shown.

The certificate fields this uses are on every block page: block 22,816 shows the same declared parameters and the same committed hash_b the command compares against.

Data availability

Nothing on this page downloads the dataset, and no link here starts a download. The two small files are enough to check every headline number; the full corpus of per-pair results is a clone or a Zenodo record away.

Dataset
docs/research/datasets/DS-002-weight-scan-v1
Size
169.46 MiB over 843 files
Manifest sha256
6d28d8f4e1ab0eb7
Tool · decoder
609cf9f90191-wb58f7a2 · v1

Route 1: the small files, pinned

Summaries, manifests and the committed sidecars. Enough to recompute every headline number on this page without downloading a single run output. Each link is pinned to the commit its digest was taken at.

Route 2: the full corpus

Every hashed pair, not only the hits. The tag pins the DS-002 dataset as published; the later coverage and control runs are pinned by the commits listed beside their files above.

git clone --branch ds-002-v1 https://github.com/InverseAltruism/keshi-research.git

Route 3: the archived record

The exact input to every figure on this page is one committed file, obs005.generated.json, generated from the datasets above by build-obs005.mjs v1. Every chart here is reproducible from it without the dataset.


Companion: who mines Pearl now

Different population, different period. Every matched block is pre-MoE and unattributed, months before the concentration measured below. Nothing here is linked to the weight-provenance result, no axis, legend or colour scale is shared with it, and no operator above is identified with any pool here.

Mining pool share

Attribution by coinbase output address against Keshi's label table.

  • Labelled pool
  • Unattributed

Blocks whose coinbase address carries no label are reported as one “unattributed” row. They are never hidden and never redistributed across known pools.

24h window

7d window

30d window

Report the Nakamoto coefficient with its basis: attributed-only blocks and all blocks give different numbers, and quoting either bare has already produced one error in this project. Read the 24h figure as a spike, not a level, which is why the three windows sit side by side. The address-level and entity-level bases are on the network telemetry page.

PoUW work rate

Live, rolling 30-day window. Included as context for the concentration figures above, not as a measure of anything in the weight-provenance result.

matmul attempts per second, a difficulty-derived consensus measure, not a measure of AI computation performed.

Frozen named-pool audit

Held for right of reply

The frozen named-pool audit row is held back pending right of reply. Publishing a per-pool measurement on this site is publication, so the audited figures ship once the notified operators have had their window, not before.

The frozen weight-provenance result above is unaffected either way: all 1,110 matched blocks are unattributed and pre-MoE, and no figure on this page names a party.


The deposited note

The frozen document every figure above reads from, rendered unedited from the source repository at the commit shown. Corrections to it ship as new dated blocks, never as edits, which is why its section 10 errata restates several figures rather than replacing them. The pre-registration fixes what could be reported before any of it ran, and PREREG-003 does the same for the coverage-closure and control-redraw rounds behind the errata.

The pre-registered weight-provenance scan: 1,110 blocks commit byte-exactly to published checkpoint tensors across three model species, with zero hits against 1,000 control blocks (2.13 million pairs), and consensus verification provably computed the sampled outputs from strips of those tensors after a publicly derivable perturbation (the attestation is one-sided toward the weights operand). Resolved to operators, the finding is launch-week adoption by a handful of operators that wound down inside the chain’s first six weeks; the coverage-completion run then matched all 43 Gemma-candidate blocks nobody had ever hashed. Only public static checkpoints are externally verifiable at all, so every zero is bounded by that ceiling.

RecordedRecorded 2026-08-09docs/research/OBS-005-weight-provenance-scan.md (external link)@7f9d428c396eRendered from the source repository at the commit shown. The content is committed alongside this site, never fetched live.

Dated observation note, grading the pre-registered weight-provenance scan. Recorded 2026-08-09. This note grades PREREG-002 (frozen at d4cb529b8602f1f73bab681f4ae428ba6cfd3fff, compiled into the scan tool as prereg2Commit) over dataset DS-002-weight-scan-v1; the DS-002b coverage completion (the phase P2 remainder plus phase P3) is folded in as §9. Every number in this note regenerates from the committed datasets via scripts/ds002-reproduce.sh (sections 1, 2, 2b, 3 need only the dataset; 4 and 5 need a Keshi API or an emitted address artifact), and each table below names the script that produced it. Definitions: metrics.md#weight-provenance-match, #matched-share-series, #tensor-coverage, #attested-model-weight-mining.

Label discipline: a byte-exact hash_b match is attested model-weight mining, never "certified inference". A match proves facts about the committed operand and the sampled computation (§0.2); it does not prove inference occurred end to end, was fresh, or served anyone.

The DS-002b coverage runs (completing phase P2's remaining o tp4/tp8 grid and running phase P3, Gemma) completed 2026-08-09 the same day and are graded in §9; their datasets are DS-002b-o-tp48-v1 and DS-002b-gemma-v1, every file sha256-verified against its MANIFEST after copy.

0. What is knowable at all (read before any number)

0.1 The verifiability ceiling

job_key = blake3(header[0:76] ‖ mining_config[0:52]) includes the coinbase-dependent merkle root, so hash_b is per-block: identical weights hash to different values in different blocks, and an observer can verify a weight commitment only by guessing the exact bytes and recomputing the keyed root per (candidate, block). Consequences, stated before the result so the result cannot be over-read:

  1. Only public, static checkpoints are externally verifiable. Private models, customer fine-tunes, locally requantized variants, evolving training weights, and synthetic noise are mutually indistinguishable on chain. This scan therefore measures verifiable public-checkpoint mining, a lower bound on whatever useful work exists.
  2. No negative result in this note supports a "no useful work" claim. A zero in a stratum means "not these exact published bytes", nothing more. A locally requantized 70B, or a training run, is not excluded by any zero below.
  3. Weight reuse across blocks is undetectable by construction. Nothing here fingerprints an operator's model over time.

0.2 What a match proves (source-verified claim ladder)

Verified 2026-08-09 against the pinned Pearl source at v1.2.1 and v1.3.0 (the zk-pow/src/v1 verifier that governs every matched-era block is byte-identical between the two tags). Mainnet consensus accepts a block only after a Plonky2 recursion over the Pearl STARK verifies against public inputs taken from the header-committed certificate (node/blockchain/validate.go:333). Inside that proof the opened strips of A and B are authenticated against the certificate's hash_a/hash_b by in-circuit keyed-BLAKE3 Merkle recomputation (v1/circuit/pearl_air.rs:91-103), and the same strip values, plus noise the verifier derives as b_noise_seed = blake3(job_key ‖ hash_b), feed the folded transcript that must hash below target. Therefore a consensus-accepted block whose hash_b equals the keyed root of a published tensor supports, beyond operand identity:

the sampled output entries verified by consensus were computed from strips of that tensor, after a publicly-derivable rank-r perturbation.

Residuals bounding the strong form: the multiplied operands are the noised strips, not the pristine ones the root commits to; each strip enters products over dot_product_length = k - (k mod r) of its entries; values are int7 range-checked; row/column semantics are a labeling outside consensus; and nothing binds hash_a to real activations, so the attestation is one-sided toward the weights operand. The matched era is verified by the V1 path, which enforces h·w ≥ 32 sampled output entries with no explicit upper bound (the 32-256 bound is V2-only); per-block h·w is read from the on-chain mining config. Full cites in docs/pearl-notes.md §Certificate internals.

0.3 The unit of analysis is operators, not blocks

The 1,066 matched blocks resolve to 46 distinct coinbase addresses with an effective operator count (1/HHI) of 4.27, and 656 of 1,066 matches (62%) fall on a single calendar day (§4). N = 1,066 is block-level pseudo-replication; every claim below is therefore phrased as what a handful of operators did during launch weeks, not as 1,066 independent observations. Addresses are not entities either: the co-spend clustering that would bound entity counts is open work (roadmap Phase 14.1).

1. What was run

keshictl scan-weights (binary 609cf9f90191-wb58f7a2, PREREG-002 compiled in, preregFrozen: true in summary.json) over the certificate corpus frozen at snapshot tip 96,774: 53,717 candidate blocks (every T2-passing block, both OFFICIAL_CONSISTENT and, via --include-model-shaped, MODEL_SHAPED_CUSTOM) plus the 1,000-block negative control fixed by §2 of the prereg (first 1,000 CUSTOM blocks in (height, hash) order). Candidate buffers: 840 tensors from the two published Llama checkpoints (revisions and byte counts pinned in the MANIFEST; 102.0 GB by the manifest totals). 4,401,184 (block, buffer) pairs hashed (summary.json), which is ~1.47 PB of keyed BLAKE3 computed as pairs times manifest buffer bytes, over 25.5 h on 2026-08-07/08 by the run log timestamps.

Per-block derivation proofs ran at extract time for all 54,717 blocks (SHA256d header reconstruction equals the block id; the 52-byte mining_config rebuilt from decoded params byte-equals the raw certificate): zero failures, so PREREG-002 §6's "derivation check failure" row is not applicable, which is itself positive evidence that the decoder and the population are sound. Determinism across worker counts and across resume is a tested property of the tool.

Coverage gap declared by the run: 386 candidate blocks were recorded but never hashed (343 shaped as 70B o_proj tp4/tp8 shards, 43 shaped as Gemma-4-31B tensors), because those buffers were not in the DS-002 weights root. §9 closes this gap as DS-002b (the phase P2 remainder plus phase P3).

2. The outcome table, graded

Copied verbatim from PREREG-002 §6 and graded in place. Applicable rows in bold; the interpretations are the frozen text, unedited.

Outcome Pre-registered interpretation Grade
≥1 match, control clean Cryptographic evidence of attested model-weight mining, the first we are aware of on any chain: the committed operand equals published tile T. Report per pool/era with counts. Still not proof that inference occurred or was served. The attribution alternatives below apply and must be stated alongside any such result. APPLIES: 1,066 matched blocks (DS-002; 1,109 combined with DS-002b, §9), 0 hits in 840,000 control pairs (2,130,000 across the three runs). §11 states the alternatives.
Matches concentrated in one pool/era, control clean As above, plus a measured heterogeneity; report shares with the unattributed remainder always shown, never redistributed. APPLIES for era and operator structure (§3, §4, §9): all matches are pre-MoE, heights 21,068-60,881 (DS-002; combined to h62,319), ~4.3 effective operators (48 addresses combined). Per-pool shares are not testable (§6): 100% of matched blocks are unattributed.
Zero matches, control clean, full phase completed No block in the scanned set committed to any searched tile. Coverage-bounded per §5. This is not proof of absence, and it does not establish that mining was synthetic. Publish the exact search space alongside the result. Applies per-stratum: 0 of 50,049 scanned m=32,768 blocks matched any searched tile (§3), within PREREG-002 §5's coverage bounds and the §0.1 ceiling.
Zero matches, phases incomplete Report as inconclusive with the exact phases run; no inference about the unscanned space. Does not apply: P1 and the bulk of P2 ran as DS-002; DS-002b completed P2's remaining o tp4/tp8 grid and ran P3, Gemma (§9), with exact populations of 343 and 43 candidate blocks.
Any control match Implementation defect. The run is void, not a finding. Fix, re-run, and record the void run in the errata. Did not occur: 0 of 840,000 control pairs (DS-002), 0 of 330,000 (DS-002b Gemma), 0 of 960,000 (DS-002b o tp4/tp8).
Derivation check failure Run aborts; treated as a data/decoder defect and investigated before any scan result is reported. Did not occur: 0 failures across 54,717 blocks.

3. The finding, in three claims with all denominators

The candidate universe is not one population. Bucketing the 53,331 scanned candidates by their declared batch dimension m (scripts/ds002-strata.py):

Stratum Scanned candidates Matched Rate
varying m (m ∉ {8,192, 32,768}) 1,179 1,016 86.17%
m = 8,192 constant 2,103 50 2.38%
m = 32,768 constant 50,049 0 0.00%
total 53,331 1,066

An earlier internal headline ("an extinction curve of matched share vs height") is largely a composition artifact of these strata: the per-bin rate mostly tracks each bin's varying-m share. The claims that survive, each with its denominator:

  1. Separation. 86.17% of blocks with live-traffic-like varying batch dimensions committed byte-exactly to published checkpoint tensors, while 0 of 50,049 constant-m=32,768 blocks did. Two structurally different workloads coexisted on the chain from its first days, and only one of them is attested model-weight mining (the other is "not these exact published bytes", §0.1).
  2. Within-stratum decline. Inside the fixed m=8,192 stratum, the match rate falls from 49/1,532 (3.20%) at or below the boundary h54,972 to 1/571 (0.18%) above it; two-sided Fisher exact P = 2.96e-06. Caveat: addresses churn across the boundary too (28 distinct below, 9 above, only 4 in common), so this is not a clean within-operator comparison, and no causal reading is attached. (A review-session figure of "P ≈ 1e-7" did not reproduce from the committed script and is superseded by the value here.)
  3. The population disappeared. The varying-m population itself vanishes: its last scanned candidate is h56,473 (DS-002 population; DS-002b's Gemma-shaped candidates extend the combined tail to h62,319, §9). The scan cannot distinguish "these operators stopped committing published weights" from "these operators stopped producing blocks"; both are consistent with every number here.

The table above is the DS-002 population; the Gemma-shaped stratum that DS-002b added (phase P3) is graded in §9 and matched at 43/43, including 23/23 at m = 8,192, so the m = 8,192 row here is a fact about Llama-shaped declarations, not about the batch dimension itself.

Struck claims, recorded so they are never reused: "zero matches in ~36,000 blocks to the tip" (the truth, all datasets combined, is 3,556 scanned candidates above h54,972 with 3 matches, §9; candidate eligibility is height-correlated, ~100% of the first 20,000 heights vs ~4% near the tip, so matched/all-blocks measures the eligibility filter, not a match rate). The headline peak is bin-width dependent and is never quoted without one: 9.31% at 5,000-block bins (h20,000+), 24.27% at 1,000-block bins (h24,000+).

4. Who and when (scripts/ds002-operator-structure.py, ds002-deep-checks.py)

All 1,066 matched blocks are unattributed (no coinbase pool tag). They resolve to 46 distinct addresses; the top address mined 432 of them (40.5%), the top three 76%; effective operators (1/HHI) 4.27. The match-days tell the story: 21 distinct days (calendar days in the API's UTC+2 block-time offset), 656 of 1,066 matches (62%) on 2026-04-29, the chain's third day (genesis 2026-04-27); steady decay through mid-May; last cluster 2026-05-17 (h54,972); one straggler 2026-05-25 (h60,881). Height range 21,068-60,881; h21,068 is 2026-04-28, roughly genesis plus one day (the chain's first ~20,000 blocks compressed into its first ~36 hours and match nothing).

The m=32,768 stratum is structurally different on addresses too: a 600-block sample resolves to 457 distinct addresses (diffuse) vs 46 in 1,066 (concentrated), and the address overlap between the two populations is zero. Two disjoint fleets.

Reading, within §0.3's limits: this is launch-week adoption of the reference stack by a handful of operators, wound down within three weeks with a thin tail to 2026-05-27 (§9), not a network-wide practice that decayed. The checkpoint repositories were created on genesis day and never modified (revision shas pinned in the MANIFEST), which cuts both ways per the frozen attribution text (§11).

5. Pre-registered analyses of the matched population (PREREG-002 §6.1-6.4)

5.1 Extinction curve and correlates (ds002-aggregate.py, ds002-strata.py)

Matched blocks per 5,000-block bin, combined across all three datasets (ds002-strata.py --combine over DS-002 plus both DS-002b runs): the numerator is the union of matches (1,109), the denominator all 53,717 recorded candidate blocks, and across the three runs every recorded candidate has been hashed against its family's buffers, so no known match sits in a zero row. The stratified §3 table is the honest headline; this series is reported with both its numerator and denominator per bin:

Height bin Matched / candidates Rate
0-19,999 0 / 19,999 0.0%
20,000+ 465 / 4,995 9.3%
25,000+ 221 / 4,954 4.5%
30,000+ 72 / 4,843 1.5%
35,000+ 105 / 4,183 2.5%
40,000+ 71 / 4,362 1.6%
45,000+ 122 / 4,141 2.9%
50,000+ 50 / 2,692 1.9%
55,000+ 1 / 1,333 0.1%
60,000+ 2 / 610 0.3%
65,000 - tip 0 / 1,605 0.0%

Correlates at the boundary, as the prereg requires:

  • Consensus changes are refuted as a confound, the strongest defensive result here. All three of Pearl's consensus changes (MoE fork 71,935; dense-only 91,630; rank-penalty 96,251; registry/PCCR.md) post-date the boundary h54,972 by 16,963 blocks, and even the last matched block (h62,319, combined) by 9,616. The clean pre-fork window h54,973-71,934 contains, combined across the three datasets, 2,507 recorded candidates with exactly 3 matches (h56,217, h60,881, h62,319; DS-002 alone: 2,505 scanned candidates, 1 match). Nothing in consensus changed at or near the boundary. (A review-session count of 2,612 for the DS-002 window did not reproduce; 2,505 is the committed script's value.)
  • Calendar: boundary 2026-05-17; DS-002's straggler 2026-05-25; the combined last attested block 2026-05-27 (h62,319, §9). The entire matched era is the chain's first month.
  • The second extinction: candidate eligibility itself (blocks whose declared shape matches any published-model tensor) falls from ~100% of blocks in the first 20,000 heights to ~4% near the tip (per-bin candidates over bin width, table above). Both extinctions are reported; they are different facts with different denominators.
  • Difficulty (live API on the recording date): 755 at the onset h21,068, 2,796,565 at the boundary h54,972, 3,372,839 at the last varying-m candidate h56,473, 5,962,295 at the straggler h60,881. The matched era spans a ~3,700-fold difficulty rise; the population faded as the work got expensive. Correlate, not cause.
  • First appearance of each labeled pool (scripts/pool-first-appearance.py, full sweep at tip 97,605 on the recording date; pools are named here as neutral public chain facts under the roadmap 9.4 disclosure split, and nothing in this row is a conduct claim): PearlHash h46,963 (2026-05-08), Pearl Fortune h51,473 (2026-05-13), LuckyPool h57,963 (2026-05-22), Hero Miners h66,836 (2026-06-03), Kryptex h67,675 (2026-06-04). The first labeled pools arrive in the two weeks straddling the boundary: attested model-weight mining ended as public pool infrastructure emerged. Correlate, not cause; the matched blocks themselves are all unattributed, before and after.

5.2 Corrected layer mixture (ds002-mixture-tv.py)

Over matched blocks, observed layer-family frequencies against MAC-share expected vectors built from the checkpoints' own quantization config (gate_up and o over all 80 layers, qkv over layers 40-79 only, down excluded as FP8, per PREREG-002 §3):

  • Llama-3.3-70B, N = 1,034: observed gate_up 0.7737 / o 0.1364 / qkv 0.0899 vs expected 0.8116 / 0.1159 / 0.0725. TV = 0.0379, MIXTURE_CONSISTENT (threshold 0.10).
  • Llama-3.1-8B, N = 32: INSUFFICIENT_SAMPLE (N < 200; the frozen verdict is reported, not the statistic).

The matched population looks like full forward passes, not a single reused tile. This does not settle the stronger alternative in §12 (iterating a downloaded checkpoint without inference); §0.2's language is held either way.

5.3 Batch-size trace (ds002-deep-checks.py)

Matched-block m: modes 16,384 (152 blocks), 8,192 (50), 14,000 (46), then a long irregular tail (1,281, 6,689, 1,111, 8,932, 7,305, 14,138, ...). 171 of 1,066 matched blocks (16.0%) declare m > 16,384, so the pilot's "matched population is 100% m ≤ 16,384" claim did not replicate and is recorded in §10. m is operator-configurable; no configuration is inferred.

5.4 Tensor coverage (ds002-tensor-coverage.py)

1,066 matches land on 293 distinct tensors: matches per tensor median 2, mean 3.64, max 16; 127 tensors matched exactly once. 70B gate_up tp1 alone covers 80 of 80 layer indices. Per-family distinct layers: 70B gate_up tp1 80, o tp1 64, gate_up tp2 50 (both shards), qkv tp1 35, o tp2 20, qkv tp2 9; 8B gate_up 17, qkv 3, o 1. Nine of the ten populated DS-002 candidate families matched (8B gate_up tp2, 11 candidates, did not); with DS-002b the count is twelve of fifteen (three Gemma families matched completely, o tp4 and tp8 scanned to zero, §9). The "one tile reused indefinitely" caveat is refuted for this population.

6. H2 (per-pool concentration) is NOT_TESTABLE, with a premise erratum

PREREG-002 §2 chose the scan set expecting labeled-pool traffic inside MODEL_SHAPED_CUSTOM. The data contradict the premise: all 53,717 candidate blocks are unattributed, so no per-pool grading is possible, and H2 is NOT_TESTABLE rather than silently dropped. (The controls do carry labels, PearlHash 69 and Pearl Fortune 24 among the 1,000, so the attribution join itself works; the candidates simply have no labeled members. A 2,400-block API sample across five labeled pools, recorded 2026-08-08, found 0% model-shaped blocks.)

7. The 14 MODEL_SHAPED_CUSTOM matches, presented unsmoothed

Fourteen matched blocks fail PREREG-001's T1 (tile patterns outside the unmodified official stack) while matching published weights byte-exactly (ds002-deep-checks.py): heights 48,782-54,950 plus the straggler 60,881; 12 of the 14 declare m = 8,192 exactly, the other two m = 5,806 and m = 6,858; all rank 128 with official minimum dims; families gate_up tp1 (11) and qkv tp1 (3). The frozen classifier and the cryptography disagree about these blocks, and the disagreement is the finding: tile patterns identify mining software, not honesty. Custom or reconfigured software committed genuine published weights. (This also seeds roadmap 14.4's software census; it is not smoothed into either neighboring class.)

8. Controls, specificity, and the positive control

  • Negative control: 0 matches in 840,000 control pairs (1,000 CUSTOM blocks × 840 buffers). Scope honestly stated: the control blocks span heights 41,237-52,280 only, and §1 of the prereg itself limits the control to an implementation canary; it has no power against false negatives and no power in the post-boundary region.
  • Post-boundary positive control (spike-in), closing that gap: scripts/ds002-spikein-control.py doctors the first three post-boundary candidate records (h54,973, h54,980, h54,981) in a scratch copy of the extract, setting each hashB to the keyed BLAKE3 of a real buffer (three distinct tensors) under that block's own on-chain-derived job_key, then runs the production Go hash stage (release binary eb3b7989dfa1-wb73d662) over the copy: 3 of 3 detected, 0 false extras (run 2026-08-09). The pipeline detects post-boundary matches by construction when they exist. The spike-in exercises the hash stage (manifest loading, keyed BLAKE3, match comparison, result emission), not the extract stage, whose derivation proofs are checked per block in §1.
  • Specificity framing (binding): "every matched block commits to exactly one of 293 distinct tensors" is a data-integrity check, not statistical evidence (one hash_b can only ever match one distinct-bytes buffer; the MANIFEST records zero byte-identical buffers). The evidence for any single match is the keyed-BLAKE3 equality, the clean control, and the independent reproduction path: scripts/verify-hashb-match.py re-derives everything from raw Blockbook bytes and pip blake3, trusting no Keshi code. Reproduced on the recording date: block 22,816 (70B gate_up L000, DS-002), block 56,217 (Gemma-4-31B gate_up L007) and block 62,319 (Gemma-4-31B gate_up L032), the last attested block on the chain (DS-002b, §9), all green with the header self-check passing.

9. DS-002b closes the coverage gap (P2 remainder + P3), and finds a third species

Two matched coverage runs, executed 2026-08-09 on release binary eb3b7989dfa1-wb73d662 (which emits the §10 coverage fields; both summaries record snapshot tip 97,590, scan bound toHeight 96,774, and their unscanned counts, and neither has any assigned candidate id without a buffer), both pinned to DS-002's snapshot tip with --to 96774, both with the same deterministic 1,000-block control (the prereg fixes control selection, so this is the same control as DS-002, not a second one):

  • Run 3a, 70B o_proj tp4/tp8 (960 sliced buffers, byte-checked against the DS-002-era tp1 buffers and golden-checked against the reference miner before launch): 343 candidate blocks (338 tp4-shard shaped, 5 tp8), all below the boundary. 1,071,360 pairs, 0 matches, 0 control hits. This independently confirms a reviewer's out-of-band hash of the same 343 blocks (also 0), and closes the last unvalidated layout family with a clean zero: every populated 70B family is now either matched on-chain (tp1, tp2) or scanned to zero (o tp4/tp8).
  • Run 3b, Gemma-4-31B gate_up + o at tp 1,2 (330 buffers at pinned revision f1dfba688ce6343b0433de57ca4dc0f3d1c5baa5; the checkpoint's 10 attention-variant layers carry o at k = 16,384, a shape outside the frozen tables, excluded as a declared exclusion): 43 candidate blocks. 333,200 pairs; all 43 of 43 matched; 0 control hits; no block matched more than one tensor. Every family matched completely: gate_up tp1 32/32, gate_up tp2 9/9, o tp2 2/2, across 34 distinct tensors.

The Gemma result upgrades the coverage story into a finding of its own:

  • A third model species, at a 100% rate over its searched universe: every block whose declared shape matched a searched Gemma tensor committed byte-exactly to the published Gemma-4-31B weights, 43 of 43. The searched universe is bounded by PREREG-002 §5's declared exclusions (the 10 attention-variant o layers and the FP8 families are not searched, so a block declaring those shapes would not be a candidate here). Within the 43: 20 of 20 varying-m and 23 of 23 m = 8,192 blocks matched (contrast the Llama-shaped m = 8,192 stratum at 2.38%, §3).
  • The tail extends: matched heights run 26,883 (2026-04-29, the launch-spike day) through 62,319 (2026-05-27), which replaces h60,881 as the last attested model-weight block on the chain. Two of the 43 sit above the boundary: h56,217 (2026-05-19) and h62,319. Both are independently reproduced from raw chain bytes (§8).
  • The operator cluster is the same one, not a new phenomenon: the 43 blocks resolve to 4 addresses, effective operators (1/HHI) 1.15, with one address mining 40 of 43 (h39,763-47,350). Two of the four also appear in DS-002's matched set: the address behind the §7 MODEL_SHAPED_CUSTOM cluster (15 Llama-matched blocks, h48,782-54,950) mined Gemma h56,217, and an early two-block Llama address mined the first Gemma match h26,883. The combined matched population is 1,109 blocks over 48 addresses; the addresses were resolved per block from the API on the recording date (same method as §4, ds002-operator-structure.py).
  • The classifier disagreement recurs, stronger: 40 of the 43 are MODEL_SHAPED_CUSTOM under the frozen classifier and 3 OFFICIAL_CONSISTENT; with §7's 14, the cryptography now contradicts the tile-pattern class on 54 blocks across two model families. Tile patterns identify software, never honesty.

Combined tail, all datasets, stated with its denominators: above the boundary h54,972 there are 3,556 scanned candidates (3,554 DS-002 + 2 DS-002b Gemma + 0 o tp4/tp8) and exactly 3 matched blocks (h56,217, h60,881, h62,319); zero matches from h62,320 to the frozen tip 96,774. The attested era runs 2026-04-28 to 2026-05-27, the chain's first month.

Even after DS-002b, "zero unscanned" remains false and is never claimed: zero records in DS-002's extract carry a Qwen3 candidate id (checked by grep over the committed dataset on the recording date), so Qwen3 needs a sentence, not a scan; MoE expert-stacked layouts were never in the candidate universe (PREREG-002 declares the extraction problem open); and §0.1's ceiling bounds everything.

10. Errata, dated

  • 2026-08-09: DS-002 violates PREREG-002 §7.2 and §2: summary.json and MANIFEST.md record neither the snapshot tip nor the unscanned count. The dataset is append-only and is not edited; the values are recorded here (snapshot tip 96,774; 386 unscanned candidates) and the tool now emits both (plus per-family unhashed accounting) in every later run, DS-002b included. Policy: ERRATA-POLICY.md.

  • 2026-08-09: the pilot claim "matched population is 100% m ≤ 16,384" did not replicate on the full run (171/1,066 above the bound), §5.3.

  • 2026-08-09: the review-session statistics "P ≈ 1e-7" (m=8,192 decline) and "2,612 candidates in the clean pre-fork window" (DS-002 population) did not reproduce from the committed scripts; the values in §3 and §5.1 (P = 2.96e-06; 2,505 for DS-002, 2,507 combined) supersede them.

  • Struck and never reused: "zero matches in ~36,000 blocks to the tip"; any matched/all-blocks rate quoted as a match rate; any peak rate without its bin width.

  • 2026-08-11: Consolidated post-audit errata (PREREG-003 coverage closure and control redraw; datasets DS-006 and DS-007). An audit dated 2026-08-10 found never-scanned certificate cells and a mis-specified negative control in the original OBS-005 / DS-002 work. The registered coverage-closure and control-redraw scans are now committed as DS-006 (coverage closure) and DS-007 (stratified control), and the corrections they support are recorded here in one dated block per ERRATA-POLICY.md. The deposited version-1 copy retains the original text; the external artifact gains a version naming this block (ADR-0014 decision 3). Section references are to this note.

    1. The last attested block is h68,332 (2026-06-05), not h62,319. §9's statements "zero matches from h62,320 to the frozen tip 96,774" and "h62,319 replaces h60,881 as the last attested model-weight block" are superseded. The h68,332 match (Gemma attention-variant o, n 5,376, k 16,384, layer 53) was found by the post-publication shape sweep, recorded first as a roadmap erratum, and has now been re-derived into a committed results file under PREREG-003 R1 (DS-006-coverage-closure-v1; 1 match in 10 pairs; independently reproduced from raw block bytes). §5.1's "65,000 - tip | 0 / 1,605" row is likewise superseded: with the DS-006 closure scan folded in, the band holds 1 attested block among 1,606 scanned, the added scanned block being h68,332 itself. The two era sentences that read the old end date as a calendar month are superseded the same way: §5.1's "the entire matched era is the chain's first month" and §9's "the attested era runs 2026-04-28 to 2026-05-27, the chain's first month". On the combined basis the attested era runs 2026-04-28 (h21,068) to 2026-06-05 (h68,332), 39 days after the 2026-04-27 genesis, so the era is the chain's first six weeks, not its first month. Wherever this note and the sections disagree on the era, this note governs.

    2. Three never-scanned certificate populations are now closed (PREREG-003 R1, DS-006): two cells the frozen classifier omitted against PREREG-002 §3's implied-mineable table (the Qwen3-30B attention o tp4 cell and the Gemma attention-variant o cell), plus one clearly labeled exploratory alternative-sharding cell (the 70B o column-parallel tp2 layout hypothesis). The counts, over DS-001's span (heights 0 to 96,405): the Qwen3-30B attention o tp4 cell (2,048, 1,024), 10 blocks, 0 matches in 1,920 pairs; the 70B o column-parallel tp2 cell (4,096, 8,192), 371 blocks, 0 matches in 59,360 pairs; the Gemma attention-variant o cell, 1 block, 1 match (item 1); 28 of the 31 enumerated cells hold 0 corpus blocks. Both zeros are layout-conditional in the PREREG-002 §5 sense.

    3. The negative control's composition is disclosed and its stated null was wrong. PREREG-002 §1 described control blocks as committing to differently-shaped operands. In fact 10 of the 1,000 control blocks declare the Qwen3 o tp4 shape (a published checkpoint shape), the control has exactly 7 distinct (n, k) with 824 of 1,000 at one shape, and the same 1,000 blocks were reused across all three runs, so "0 hits in 2,130,000 control pairs across the three runs" describes 1,000 distinct blocks against 2,130 distinct buffers, never 2.13 million independent trials. The counts themselves stand; item 2's scan establishes that the 10 checkpoint-shaped members do not commit to that checkpoint's tensors, so no contamination occurred, and the count "0 in 840,000 control pairs" remains literally correct (those pairs never included a Qwen3 attention buffer). A redrawn, era-stratified control with a corrected shape-exclusion null is registered as PREREG-003 R5 and committed as DS-007. R5 drew 952 blocks stratified by era (200 each from the pre-boundary, post-boundary, MoE-window, and dense-only strata; 152 from the shorter rank-penalty stratum), each block's declared shape absent from every searched cell, and hashed them with the shape filter disabled against all seven searched buffer sets (the 70B, 8B, and Gemma reference sets plus the o tp4/tp8, Qwen3 attention, column-parallel, and Gemma attention-variant sets): 0 matches in 2,372,384 pairs. The corrected control reproduces the clean canary the original was meant to provide, now with a null that holds by construction.

    4. The o-family byte-order hypothesis for the unmatched varying-m residue: refuted for the registered population. 162 of the 163 unmatched varying-m candidates are o-family shapes; PREREG-003 R2 probed the column-major byte-order variant (V1) with a built-in positive control: the 20 control blocks, drawn by the registered seed from the 144 matched o blocks, matched 20 of 20 under the standard convention, each at exactly the tensor DS-002 had recorded for it, and 0 of 1,680 pairs under V1, which validates the variant pipeline in both directions. Over the population, V1 produced 0 matches in 3,040 pairs for the 95 8B o tp1 blocks and 0 matches in 10,160 pairs for the 67 70B o blocks (7 tp1, 60 tp2). A transposed byte-order commitment to these checkpoints' o tensors is excluded for all 162 blocks at the pinned revisions; the residue remains unmatched and unexplained. PREREG-003 R3 likewise probed the two registered gate_up tp2 layout variants (P1, up|gate concatenation order; P2, slicing the already-fused tensor per rank) over the 1,905 unmatched gate_up tp2 blocks, with the same built-in positive control (20 of 20 matched gate_up tp2 blocks re-matched under the standard convention, 0 of 20 under each variant): 0 matches in 304,800 pairs for P1 and 0 in 304,800 for P2. Neither alternative fusion or shard layout explains that residue either.

    4a. No block committed to either third-party int8 artifact that was scanned (PREREG-003 R4). The population is the 2,216 unmatched candidates outside the m 32,768 stratum, which is not a non-constant-m residue: 2,053 of the 2,216 declare constant m 8,192 and only 163 vary (counts over DS-002's extract). It was scanned against a published third-party W8A8 quantization of the 70B architecture (0 matches in 326,080 pairs) and an 8B candidate (0 in 3,744), at the revisions pinned in DS-006. Verified: the 496 buffers built for R4 (400 for the 70B artifact, 96 for the 8B) are byte-distinct from every buffer any earlier run searched, 0 sha256 in common with the 2,130 DS-002 and DS-002b reference buffers and 0 with all 3,148 non-R4 buffers recorded in DS-006, so R4 put new operands in front of these blocks rather than re-searching the pearl-ai bytes. Inferred: different quantization grids produce different int8 codes, which is the reason to expect a third-party requantization of the same architecture to be a distinct operand rather than a re-encoding. The only checkpoints any block attested to remain the published pearl-ai ones.

    What R4 does not cover. PREREG-003 R4 named two 70B candidates as verified eligible before the freeze. The RedHat W8A8 artifact was retrieved and scanned, as above; the second named candidate was not retrieved in the session that produced this note and so was never scanned. Whether its int8 buffers are byte-equivalent to the scanned artifact is unverified, and no claim is made either way. The R4 zero therefore bounds the two artifacts named in DS-006 at their pinned revisions, one 70B and one 8B, and says nothing about third-party int8 quantization in general.

    1. Effective sample size, restated at the honest unit. The combined 1,109 matched blocks resolve to 48 addresses; effective addresses (1/HHI) 4.59; effective match-days 2.73; effective (address, day) units 5.40; 657 of 1,109 (59.2%, combined basis) fall on 2026-04-29 via 7 addresses, and the top address mined 430 of its 432 matches on that one day. §0.3's "62%" is the DS-002-only basis (656 of 1,066) and §0.3's 4.27 is the DS-002-only effective-address figure; both reproduce on the DS-002 basis, and the basis for each is stated here so the DS-002-only and combined figures are not conflated. Including item 1's h68,332, whose address is new, the attested population is 1,110 blocks over 49 addresses (effective addresses 4.60).

    2. "0 of 50,049" is a three-shape statement. The constant-m 32,768 stratum comprises exactly three declared shapes from two checkpoints at tp1: (57,344, 8,192) x 31,831, (28,672, 4,096) x 16,530, and (10,240, 8,192) x 1,688. The zero excludes those exact published bytes for those shapes and nothing broader.

    3. §9's phrase "closes the last unvalidated layout family with a clean zero" is withdrawn. A zero cannot validate a layout. The 8B gate_up tp2, 70B o tp4, and 70B o tp8 layouts remain golden-checked only, exactly as PREREG-002 §5 warned; their zeros are layout-conditional. Separately, the layouts that did match are now also conformance-tested by a second, from-specification extractor implementation (six buffers spanning the six matched layout kinds, gate_up, qkv, and o at tp1 and tp2, on the 70B checkpoint, byte-exact against the recorded dataset hashes; scripts/independent-extract.py).

    4. §3 claim 2's Fisher test is address-confounded. One address mined 811 of the 1,483 below-boundary unmatched m 8,192 candidates (811 of the stratum's 2,053 unmatched candidates overall; heights 33,777 to 49,675) and has exactly one match of its own (h37,772), so both arms of the 3.20% versus 0.18% comparison are dominated by a handful of addresses and the two-sided P of 2.96e-06 over blocks overstates the evidence. Full-population address counts for the unmatched stratum are 43 distinct below the boundary and 9 above, 5 in common; §3's parenthetical (28, 9, 4) was computed on seeded 300-block samples and is superseded by these. The crosstab stands; the P is withdrawn as the headline statistic.

    Every figure above was produced by commands against the committed datasets on 2026-08-11, with the combined-basis address and day figures resolved per block from the API: items 5 and 8 need the per-block minerAddress that the committed operator sidecar does not carry at the 1,110 basis. The PREREG-003 runs are frozen at commit d2b83c55626865a1c32254e893fb731307da889c and their datasets carry that hash. Corrections to this block, if ever needed, ship as a further dated block, never as edits.

11. Attribution alternatives (frozen text, verbatim)

A match identifies weights, not a miner. These readings are written down now so that whichever the data supports is not a post-hoc rescue:

  1. Operator bootstrap fleet. Matches concentrated early, in the unattributed stratum, are consistent with Pearl Research Labs (or an affiliate) mining its own published checkpoints to bootstrap the chain. The published checkpoint's own repository timeline is recorded in the dataset manifest for exactly this reason, and it cuts both ways: a checkpoint available from genesis is equally available to anyone.
  2. Third-party miners running the published stack on public weights.
  3. Mimicry. Anyone can download the same weights and mine them; a match proves the operand, never the identity or intent of the operator.

Keshi does not name a party. No coinbase-based attribution beyond the existing pool_labels ladder enters this analysis, and the unattributed stratum stays unattributed. Where the strata cannot distinguish these readings, OBS-005 says so explicitly rather than choosing one.

What remains measurable regardless of which reading is true: when real-weight mining occurred, in what volume, in which strata, and, if it declines, when it stopped. That trajectory is the finding; the identity of the miner is not ours to assert.

12. Threats to validity, and what would falsify this

  • Composition, not decline: §3 exists because the aggregate curve mostly tracks stratum shares. Any reader quoting a single rate from this note without its stratum and denominator is misquoting it.
  • Eligibility filter: candidate share falls ~100% to ~4% with height (§5.1); nothing here says what ineligible blocks computed.
  • Address churn: the §3 claim-2 comparison is across different address sets (28/9/4); it is a population statement, not an operator statement.
  • Pseudo-replication: §0.3; the effective sample is ~4.3 operators.
  • Control scope: negative control spans h41,237-52,280 (canary only); post-boundary detection power is demonstrated by the §8 spike-in, not by the control.
  • Coverage bounds: §9's exact search space, plus §0.1's ceiling; every zero is "not these exact bytes".
  • The stronger alternative left open: a miner iterating a downloaded checkpoint without serving inference is consistent with §5's mixture, batch, and coverage results. Nothing in this note distinguishes real serving from replaying checkpoint layers; §0.2's claim is exactly what is proven, no more.
  • What would falsify the finding: any control match (voids the run by the frozen table); failure of the independent raw-bytes reproduction on any published match; byte-identical buffers in the manifest (would void the specificity integrity check); a derivation-proof failure on a matched block; the spike-in failing to detect doctored records. As of the recording date, all five checks pass.

13. Prior art

The one adjacent measurement paper (arXiv:2606.04819, retrieved 2026-08-08) analyzes miner economics on the dominant pool binary; it never touches certificates, hash_b, or checkpoint matching, and it lists workload provenance as unsolved future work requiring an external PKI. Its artifact repository is unreachable (404) and it has no academic citations we could find. On that basis, and per the frozen outcome table's own wording, this is the first cryptographic evidence of attested model-weight mining on any chain that we are aware of. External review (roadmap 8.6) is invited before publication; the review window and invitations are logged in OUTREACH-log.md.

14. Reproduction

  • Dataset-only: KESHI_DS002_DIR=<dir> scripts/ds002-reproduce.sh regenerates every table above (sections 1-3 offline; 4-5 with a Keshi API or an emitted address artifact). Datasets under docs/research/datasets/DS-002-weight-scan-v1/ (and DS-002b-* per §9), each with sha256 MANIFESTs; licence and field-by-field schema in DATA-LICENCE.md.
  • Any single match, trusting no Keshi code: verify-hashb-match.py --height H --buffer <tensor.bin> (raw Blockbook block, pip blake3, header self-check).
  • The positive control: ds002-spikein-control.py --out <scratch> --run.