EXP-01---- traces committed---- incidents---- reproducibility---- bondedSolana devnet

EXP-01 / Collimate / Verification tiers

Tiers

Attested
Sampled
Proven

A collimator does not amplify a beam. It absorbs everything travelling at the wrong angle so that what leaves the far side is parallel, and therefore comparable. These three tiers do that to a verification claim: each one states the evidence it carries and, in the same breath, the evidence it does not, so that a record can never borrow the strength of a tier it did not use.

This page is a specification, not a pitch. Where a figure exists it is published with the conditions it was measured under. Where one does not exist yet, the gap is published instead.

The three tiers, side by side

Cost rises down the table and so does what the tier is willing to state. The right-hand columns are the ones that matter: a tier that carries no reproducibility weight is not a weaker version of one that does, it is a different kind of claim.

Collimation

DivergentPlatesResolved

Two of the five rays never reach the far side. That is the point of the instrument, and it is the point of the tiers: most of what a provider could assert about its own output is stopped at the plate, and only what survives is written to the record.

AxisAttestedSampledProven
Track shapeShort and faint. The track stops where the evidence stops.Comparison beads along the track, one for each re-execution that matched.Bent by the field and cut hard into the plate.
ProvesA TEE signature shows that the declared binary executed on the declared hardware.A randomly selected rerun produced the same output hash as the committed one.A zero-knowledge proof of the core operations, checked mathematically rather than by re-execution.
Does not proveThat the model weights match the ones registered on the plate. The attestation is only as good as the trust placed in the enclave vendor.Anything about the inferences that were not sampled. The guarantee is statistical, not per-inference.That a full large model was proved end to end. Whole-model zkML is not currently feasible; the proved operation set is published per plate and everything outside it is unproved.
CostLowestMiddleHighest
Reproducibility weight0.00.61.0
Committed at this tier22693

Indexer ok / read through slot 507,602,399 / 2026-10-05 03:59 UTC

What KORTX does NOT prove

This section is required and it is not moved, shortened, or folded away. A verification system that publishes only its guarantees has published half a specification, and the missing half is the half a reader needs in order to be wrong about something.

Outside every tier

  1. That a model is good.

    KORTX checks whether a computation can be reproduced or certified. A reproducible model can still be wrong, biased, or unsafe. Reproducibility is a property of the execution, not of the answer.

  2. That a committed output is correct.

    A committed hash binds an output to a plate, an input, a fingerprint and a seed. It says the output came out of that model under those conditions. It says nothing at all about whether the output is right.

  3. That the declared weights are the weights that ran, except where the tier says so.

    Hardware attestation measures the GPU driver and the VBIOS. It does not measure model weights. At the attested tier the plate's weights hash is a declaration, not a measurement.

  4. That a prover spent computation matching the advertised model size.

    A zero-knowledge proof certifies that an equation holds. It does not certify an amount of work. Published research demonstrates a model that presents itself as twelve layers while computing six, at no additional serving cost.

  5. Anything about an inference that was not sampled.

    The sampled tier gives a probability and the sampling rate that produced it is printed on the record. It does not extend to the traces that were never drawn, and no amount of matched samples turns that probability into a guarantee.

  6. A zero-knowledge proof of a full large language model forward pass.

    No system does this today. The published ceiling for full-inference proofs is 13B parameters, at 803 seconds on one A100 for a single 2,048-token sequence, producing a 188 kB artifact. A Solana transaction holds 1,232 bytes.

  7. Survival of a compromise at the hardware vendor's root of trust.

    At the attested tier, trust reduces to the silicon vendor and the vendor's attestation service. If the vendor's root of trust lies, the tier tells you nothing and cannot tell you that it is telling you nothing.

  8. Detection of a provider who is honest on sampled traces and dishonest elsewhere.

    Beyond the probability the sampling rate gives, that behaviour is invisible to this system. The probability is stated on the record rather than implied by a badge.

  9. Reproduction of an inference that was never deterministic.

    Above temperature 0 the output is drawn from a distribution, so byte-identical reproduction is not merely hard, it is undefined. Even at temperature 0 the same prompt run 1,000 times produced 80 distinct outputs: the first 102 tokens were identical every time and divergence began at token 103. That is why a plate pins temperature, top-p, seed bound and backend, and why a trace whose plate does not pin them cannot be re-run against.

Tier specifications

Each tier below is stated as a mechanism, a coverage claim, an exclusion, the precondition without which it means nothing, and the trust that is left over when it has done its work. The last of those is the one nobody prints.

Attested

Lowest cost
A short faint condensation track fading a third of the way across the vessel
Reproducibility weight
0.0contributes no evidence to a score
Counts toward score
No
Committed here
22 traces
Track
Short and faint. The track stops where the evidence stops.

ProvesA TEE signature shows that the declared binary executed on the declared hardware.

Does not proveThat the model weights match the ones registered on the plate. The attestation is only as good as the trust placed in the enclave vendor.

Mechanism
The provider runs inference inside a confidential virtual machine with a confidential-computing GPU. The GPU produces a signed report over its driver and VBIOS measurements, checked against the vendor's signed reference values. The report is bound to the trace and its hash goes on chain.
Covers
That measured firmware and driver ran on genuine silicon, in a machine whose memory was not readable by the host operating system across the PCIe boundary.
Does not cover
Model weights are not measured. On-package HBM is not encrypted; the protection covers what crosses PCIe. And NVIDIA GPU attestation reports are not bound to the identity of the confidential VM that presents them, so a report from one machine can in principle be presented by another.
Trust assumption
The hardware vendor's root of trust, the vendor's attestation service, and the cloud operator's platform firmware. Physical attacks are explicitly outside the threat model that both major CPU vendors publish, and interposer attacks against DDR5 memory have been demonstrated for under 1,000 USD in parts.
Measured cost
  • Hardware list price premium: none. Azure prices the confidential H100 size and the non-confidential H100 size of the same generation identically at 6.98 USD per hour, Linux, East US 2, checked against the Azure retail price API on 2026-08-20.
  • Throughput: confidential-computing overhead on H100 falls between minus 0.13 percent and 21 percent across three independent measurements, varying with model size, sequence length and serving mode. Cost per token rises by roughly the same range.
Scorer noteQuoted from the index
Attestation is not a rerun. An attested trace contributes no evidence to the reproducibility score unless it is challenged and re-run.

Sampled

Middle cost
A condensation track carrying bright comparison beads at intervals along its length
Reproducibility weight
0.6per matched comparison
Counts toward score
Yes
Committed here
69 traces
Track
Comparison beads along the track, one for each re-execution that matched.

ProvesA randomly selected rerun produced the same output hash as the committed one.

Does not proveAnything about the inferences that were not sampled. The guarantee is statistical, not per-inference.

Mechanism
A fraction of committed traces is drawn at random and re-executed by an independent bonded node under a pinned environment. The re-run output is hashed and compared byte for byte. A mismatch opens an incident.
Covers
That the drawn trace reproduces exactly. With the execution environment pinned, 10,000 runs on two hosts produced identical hashes with zero bit-level drift, and a batch-size swing of plus or minus 20 percent did not break that.
Does not cover
Traces that were not drawn. Detection follows 1 - (1-p)^K for sampling rate p and K dishonest attempts. At p = 0.05 the expected number of dishonest traces before the first catch is 20.
Precondition
Five properties of the execution must be pinned and recorded: hardware SKU, exact weights and quantization format, parallelism topology, software and kernel versions, and the batch size of each forward pass. One of these cannot be pinned across a heterogeneous provider network. An A100 and an H100 running the same weights on the same input agree 0.0 percent of the time at the bit level. KORTX therefore draws re-runs from architecture-separated verifier pools; comparing across architectures would slash honest providers at a rate near 100 percent. The pool key is the architecture each node declares in the verifier registry, behind its bond. The on-chain Verifier account carries no architecture field, so this separation rests on the assignment code and on the stake standing behind that declaration, not on the program. It is also conditional: the filter runs only when the provider's architecture is supplied, and a draw made without it is an unrestricted pool rather than a safe default.
Trust assumption
That at least one independent verifier is running and is not colluding with the provider, and that both bonds exceed the gain from cheating. A verifier paid only for finding errors earns nothing while the system works correctly. This is the verifier's dilemma, and it produced a real Bitcoin fork in July 2015.
Measured cost
  • Re-execution: the sampling rate times the inference cost. At p = 0.05 that is a 5 percent surcharge on served inference.
  • Determinism: between 1.8 percent and 133 percent depending on the serving stack. The stack is named wherever that figure is quoted, because the two ends of that range are different systems rather than a margin of error.
Scorer noteQuoted from the index
A matched sample is direct reproducibility evidence, weighted below a proof because it compares outputs rather than the computation.

Proven

Highest cost
A condensation track curving through the magnetic field and terminating in a sharp bright point
Reproducibility weight
1.0per matched comparison
Counts toward score
Yes
Committed here
3 traces
Track
Bent by the field and cut hard into the plate.

ProvesA zero-knowledge proof of the core operations, checked mathematically rather than by re-execution.

Does not proveThat a full large model was proved end to end. Whole-model zkML is not currently feasible; the proved operation set is published per plate and everything outside it is unproved.

Mechanism
A zero-knowledge proof that the declared circuit, evaluated on the committed input, produces the committed output. Verification requires no trust in the provider.
Covers
The arithmetic of the declared circuit, exactly.
Does not cover
The amount of work. A proof certifies that private weights exist which map the input to the output through the public architecture. It does not certify that the prover spent computation matching the advertised model size.
Precondition
The model has to be in the accepted set below. A workload outside that set is rejected rather than quietly downgraded to a weaker tier, because a silent downgrade leaves the requester believing they received a proof.
Trust assumption
The soundness of the proving system, the correctness of the circuit compiler, and the claim that the declared architecture is the one the provider advertises. The last of these is not cryptographic.
Measured cost
  • Classical models are inside a practical range: 0.118 to 6.161 seconds in the framework author's own published benchmark, at 19.4 MB to 383 MB of memory.
  • Neural networks are not. A third-party benchmark puts MobileNetV2 at roughly four hours and 204 GB of RAM, and language-model-scale proofs are measured in tens of minutes to days. The full table is below.
Scorer noteQuoted from the index
The strongest evidence class carried by this index, within the published proved operation set.

Why attested is worth zero

The reproducibility score is a Wilson lower bound over matched comparisons, and every comparison enters it with a weight set by the evidence that produced it. One of those weights is zero, and that zero is the single most load-bearing number on this site.

Weight per comparison

Read from the index
Incident rerun1.0

A re-execution forced by a bonded challenge. Adversarial, so it is weighted like a proof rather than like a routine sample.

Proven1.0

A proof of the declared circuit, within the accepted model set. The strongest evidence class this index carries.

Sampled0.6

A matched sample is direct reproducibility evidence, weighted below a proof because it compares outputs rather than the computation.

Attested0.0

Attestation is not a rerun. It contributes nothing to the score unless the trace is challenged and re-run.

The scorer's own sentence

Attested traces carry zero reproducibility weight: TEE attestation proves a binary ran, not that its output reproduces. They are reported under traces_untested.

An attestation answers the question "did this binary run on this hardware". Reproducibility answers a different question: "does this output come back if the work is done again". The first cannot be evidence for the second, so folding attested traces into the score at any weight above zero would inflate every provider that never submits to a re-run. They are counted, and they are counted as untested.

The consequence is deliberate and it is unflattering: a provider can commit thousands of attested traces and still be reported as unrated. That is the correct answer, because nobody has checked their work.

Coverage across this index

Traces committed

94

Carrying rerun evidence

21

Untested

73

22 of these are attested and have never been challenged.

Coverage

22.3%

Share of committed traces that any rerun has ever touched.

Committed traces by tier

Attested
2223%
Sampled
6973%
Proven
33%

Read against the coverage figure above: the largest tier on this network is also the one that carries no reproducibility weight at all.

What proven actually accepts

The proven tier refuses work it cannot prove instead of quietly serving it at a lower tier. A silent downgrade is worse than a rejection, because the requester keeps believing they asked for a proof and got one.

Accepted

O

Linear and logistic regression

0.118 s measured, 19.4 MB

O

Support vector machines

0.318 s measured, 23.7 MB

O

Decision trees and tree ensembles, random forests included

0.308 s to 6.161 s measured, 23.7 MB to 383 MB

O

Small convolutional networks

Practical only off the response path; see MobileNetV2 below

Refused

X

Transformer language models, full inference

The framework author states neural networks are outside benchmarking scope

X

Autoregressive multi-token generation

The circuit grows with the square of the token count

X

Diffusion pipelines

SDXL end to end measured at 68,950 s and a 145.84 MB proof

X

3D volumetric networks

A 19M-parameter 3D-UNet measured at 568,543 s, with 4 hours of verification alone

X

Variable-length sequence input

Circuits are fixed-shape; every length needs its own setup

X

Anything on a real-time response path

None of the above fits inside a request the user is waiting on

Documented circuit constraints

  • Non-unit dilation in convolution and deconvolution is refused by the compiler outright.
  • Division is supported only by constants.
  • Exponents are supported only as constants.
  • An ONNX operator the circuit does not know fails at circuit layout rather than being approximated silently.

There is no complete list of unsupported operators and this page does not pretend to publish one. The compiler routes ONNX through an intermediate representation that folds more than a hundred operators into roughly twenty primitives, so an operator missing from the source's match arms is not necessarily unsupported. What is documented is the failure mode: an unknown operator is a hard error at circuit layout, not a quiet approximation.

Approximated, not refused

Non-linear functions are replaced by lookup tables. The circuit is proved correctly; the function it proves is a near neighbour of the one the model ran.

FunctionMean errorP99 error
GELU0.026%0.158%
SiLU0.021%0.158%
RMSNorm0.003%0.034%
Softmax, end to endNot publishedNot published

Every proving system in this class quantizes to integers. One reports cosine similarity of at least 99.6 percent against the unquantized model, and applies no outlier smoothing to the largest model it supports. Quantization error is not a rounding artefact of the proof; it changes the model that was proved.

Published proving cost

Every row below was measured and published by somebody else, on their hardware, under conditions printed in the same row. None of it is a KORTX measurement and none of it is an estimate. Where a source did not measure a quantity, the cell says so.
WorkloadInputProve timeMemoryProof sizeHardware
Linear regressionNot statedFramework author's benchmark, 2024-01-28One inference0.118 s19.4 MBNot measuredNot stated
Support vector machineNot statedFramework author's benchmark, 2024-01-28One inference0.318 s23.7 MBNot measuredNot stated
Tree ensemble regressionNot statedFramework author's benchmark, 2024-01-28One inference0.308 s23.7 MBNot measuredNot stated
The same benchmark states directly that verification time and proof size were not measured, and that neural networks are outside its scope.
Random forest classificationNot statedFramework author's benchmark, 2024-01-28One inference6.161 s383 MBNot measuredNot stated
A competing zkVM took 173.4 s on the same model and needed up to 10.2 GB. A third framework failed this case with out-of-memory errors on a machine holding 1,000 GB of RAM.
MobileNetV2Not statedThird-party benchmarkOne imageAbout 4 h204 GBNot measuredXeon E5-2665, 16 threads, 350 GB
GPT-2124MDeepProve, published benchmark512 tokens7.64 minNot published10.71 MiB24 cores, 504 GB, EPYC 9254
Gemma 31BDeepProve, published benchmark512 tokens18.95 minNot published21.73 MiB24 cores, 504 GB, EPYC 9254
The only published end-to-end figure for a 512-token sequence at 1B parameters. The proof is 21.73 MiB.
LLaMa-213BzkLLM, published paperOne 2,048-token sequence803 s23.1 GB188 kBOne A100 SXM4 40 GB, EPYC 7413
The paper states a semi-honest verifier assumption; its completeness and soundness results rest on that premise. This is not a guarantee against a malicious verifier, and the scheme is not an on-chain verification target.
LLaMA-27BZKTorch, published paperOne token2,645 sNot published22.85 MBXeon Platinum 8358, 64 threads, 4 TB
One token. Verification alone takes 100.14 s.
3D-UNet19MZKTorch, published paper1 x 1 x 128 cubed568,543 s (6.6 days)Not published1.4 MBXeon Platinum 8358, 64 threads, 4 TB
Verification alone takes 4 hours.
  • Proof artifacts in the table run from 188 kB to 22.85 MB. A Solana transaction holds 1,232 bytes, so those two exceed the on-chain limit by 152 times and 18,500 times. On Solana today only Groth16-class proofs verify on chain directly, at 78,293 to 108,762 compute units against a 1,400,000 unit transaction ceiling.
  • Proving a production workload cost 75 times the inference it certified.
  • Every figure above is reproduced from a published source. None is a KORTX measurement and none is an estimate. Where a source did not measure something, the cell says so instead of carrying a plausible number.

Not yet measured

Not yet measured
Proving time per trace on this network at the proven tier
Not yet measured
Verifier re-execution wall clock at the sampled tier
Not yet measured
Incident to resolution wall clock
Not published by the vendor
Attestation call latency, rate limit, SLA and pricing
No project in this field publishes one
False-slash rate against honest providers
No project in this field publishes one
Challenge period length used by comparable systems

These are published when they are measured, not before. Filling them with an estimate would make this page indistinguishable from the marketing it exists to replace, and an estimate that is later contradicted costs more than an empty field ever did.

On-chain cost of one commit

Rent exemptionA 342-byte trace account. Returned in full when the account is closed.
3,271,200 lamports
Signature feeThe only part of a commit that is actually consumed.
5,000 lamports
Total locked per commitDeposit plus fee. Writing this as the cost of a trace overstates it by a factor of 655.
0.0032762 SOL

No USD conversion is printed here. The SOL price was not obtained from a primary source in the measurement record, and a hardcoded conversion would be a number this page cannot defend.

What sampling actually catches

The sampled tier is a probability and this is the probability. It is published in full because a reader who has this table cannot be surprised later, and a reader who does not have it will assume something more generous than the truth.
Sampling rate pK=1K=10K=20K=50K=100K=200
1%1.00%9.56%18.21%39.50%63.40%86.60%
2%2.00%18.29%33.24%63.58%86.74%98.24%
5%5.00%40.13%64.15%92.31%99.41%99.99%
10%10.00%65.13%87.84%99.48%99.99%99.99%
20%20.00%89.26%98.85%99.99%99.99%99.99%

Probability at least one dishonest trace is caught. 1 - (1-p) ^ K

  • Read the first column, not the last. A provider who cheats once at a 5 percent sampling rate is caught 5 percent of the time. Sampling is a deterrent priced against a bond, not a guarantee.
  • 99.99 percent means at least 99.99 percent. It is never 100. For any finite K the expression stays below 1, so cells that would round up to 100.00 are written as 99.99 instead.

Choosing a tier

There is no recommended tier. There is a workload, a set of constraints it imposes, and the tier that fits inside them. The last row is the one most likely to apply and the least likely to be printed anywhere else.
SituationTierReason
Serving a transformer language model on a live response pathAttested or sampledProven is structurally unavailable here, not merely expensive.
Batch inference over classical models: regression, SVM, treesProvenProving time of 0.118 to 6.161 seconds is inside a practical range.
A provider network with mixed GPU architecturesAttestedSampled requires architecture-separated verifier pools first. An A100 and an H100 agree 0.0 percent of the time at the bit level.
A single-architecture pool that can pin its execution environmentSampledWith five properties pinned, 10,000 runs produced identical hashes and zero drift.
A small convolutional network with no latency requirementProvenMobileNetV2 costs roughly 4 hours and 204 GB. It works, off the response path.
A workload whose correctness matters more than its provenanceNone of themNo tier here evaluates whether an output is right. Verification and evaluation are different problems and this system only does the first.

Tiered verification is not a KORTX invention. Other projects offer a choice of verification strength. The narrower position here is that the tier travels with the output, on chain, as part of the record rather than as a claim in a dashboard.

Where the tiers appear

A tier is not a page, it is a field on a record. These are the surfaces that carry it.
Trace ViewerOne inference: input hash, output hash, fingerprint, seed, its tier, and whether any rerun ever compared against it.
Model PlatesThe registered model behind a trace, including the determinism fields that decide whether reproduction is even defined.
Track BoardProvider reproducibility, weighted by the table above, with unrated providers held apart from low-scoring ones.