EXP-01 / Collimate / Verification tiers
Tiers
A collimator does not amplify a beam. It absorbs everything travelling at the wrong angle so that what leaves the far side is parallel, and therefore comparable. These three tiers do that to a verification claim: each one states the evidence it carries and, in the same breath, the evidence it does not, so that a record can never borrow the strength of a tier it did not use.
This page is a specification, not a pitch. Where a figure exists it is published with the conditions it was measured under. Where one does not exist yet, the gap is published instead.
The three tiers, side by side
Collimation
Two of the five rays never reach the far side. That is the point of the instrument, and it is the point of the tiers: most of what a provider could assert about its own output is stopped at the plate, and only what survives is written to the record.
| Axis | Attested | Sampled | Proven |
|---|---|---|---|
| Track shape | Short and faint. The track stops where the evidence stops. | Comparison beads along the track, one for each re-execution that matched. | Bent by the field and cut hard into the plate. |
| Proves | A TEE signature shows that the declared binary executed on the declared hardware. | A randomly selected rerun produced the same output hash as the committed one. | A zero-knowledge proof of the core operations, checked mathematically rather than by re-execution. |
| Does not prove | That the model weights match the ones registered on the plate. The attestation is only as good as the trust placed in the enclave vendor. | Anything about the inferences that were not sampled. The guarantee is statistical, not per-inference. | That a full large model was proved end to end. Whole-model zkML is not currently feasible; the proved operation set is published per plate and everything outside it is unproved. |
| Cost | Lowest | Middle | Highest |
| Reproducibility weight | 0.0 | 0.6 | 1.0 |
| Committed at this tier | 22 | 69 | 3 |
Indexer ok / read through slot 507,602,399 / 2026-10-05 03:59 UTC
What KORTX does NOT prove
Outside every tier
That a model is good.
KORTX checks whether a computation can be reproduced or certified. A reproducible model can still be wrong, biased, or unsafe. Reproducibility is a property of the execution, not of the answer.
That a committed output is correct.
A committed hash binds an output to a plate, an input, a fingerprint and a seed. It says the output came out of that model under those conditions. It says nothing at all about whether the output is right.
That the declared weights are the weights that ran, except where the tier says so.
Hardware attestation measures the GPU driver and the VBIOS. It does not measure model weights. At the attested tier the plate's weights hash is a declaration, not a measurement.
That a prover spent computation matching the advertised model size.
A zero-knowledge proof certifies that an equation holds. It does not certify an amount of work. Published research demonstrates a model that presents itself as twelve layers while computing six, at no additional serving cost.
Anything about an inference that was not sampled.
The sampled tier gives a probability and the sampling rate that produced it is printed on the record. It does not extend to the traces that were never drawn, and no amount of matched samples turns that probability into a guarantee.
A zero-knowledge proof of a full large language model forward pass.
No system does this today. The published ceiling for full-inference proofs is 13B parameters, at 803 seconds on one A100 for a single 2,048-token sequence, producing a 188 kB artifact. A Solana transaction holds 1,232 bytes.
Survival of a compromise at the hardware vendor's root of trust.
At the attested tier, trust reduces to the silicon vendor and the vendor's attestation service. If the vendor's root of trust lies, the tier tells you nothing and cannot tell you that it is telling you nothing.
Detection of a provider who is honest on sampled traces and dishonest elsewhere.
Beyond the probability the sampling rate gives, that behaviour is invisible to this system. The probability is stated on the record rather than implied by a badge.
Reproduction of an inference that was never deterministic.
Above temperature 0 the output is drawn from a distribution, so byte-identical reproduction is not merely hard, it is undefined. Even at temperature 0 the same prompt run 1,000 times produced 80 distinct outputs: the first 102 tokens were identical every time and divergence began at token 103. That is why a plate pins temperature, top-p, seed bound and backend, and why a trace whose plate does not pin them cannot be re-run against.
Tier specifications
Attested

- Reproducibility weight
- 0.0contributes no evidence to a score
- Counts toward score
- No
- Committed here
- 22 traces
- Track
- Short and faint. The track stops where the evidence stops.
ProvesA TEE signature shows that the declared binary executed on the declared hardware.
Does not proveThat the model weights match the ones registered on the plate. The attestation is only as good as the trust placed in the enclave vendor.
- Mechanism
- The provider runs inference inside a confidential virtual machine with a confidential-computing GPU. The GPU produces a signed report over its driver and VBIOS measurements, checked against the vendor's signed reference values. The report is bound to the trace and its hash goes on chain.
- Covers
- That measured firmware and driver ran on genuine silicon, in a machine whose memory was not readable by the host operating system across the PCIe boundary.
- Does not cover
- Model weights are not measured. On-package HBM is not encrypted; the protection covers what crosses PCIe. And NVIDIA GPU attestation reports are not bound to the identity of the confidential VM that presents them, so a report from one machine can in principle be presented by another.
- Trust assumption
- The hardware vendor's root of trust, the vendor's attestation service, and the cloud operator's platform firmware. Physical attacks are explicitly outside the threat model that both major CPU vendors publish, and interposer attacks against DDR5 memory have been demonstrated for under 1,000 USD in parts.
- Measured cost
- Hardware list price premium: none. Azure prices the confidential H100 size and the non-confidential H100 size of the same generation identically at 6.98 USD per hour, Linux, East US 2, checked against the Azure retail price API on 2026-08-20.
- Throughput: confidential-computing overhead on H100 falls between minus 0.13 percent and 21 percent across three independent measurements, varying with model size, sequence length and serving mode. Cost per token rises by roughly the same range.
- Scorer noteQuoted from the index
- Attestation is not a rerun. An attested trace contributes no evidence to the reproducibility score unless it is challenged and re-run.
Sampled

- Reproducibility weight
- 0.6per matched comparison
- Counts toward score
- Yes
- Committed here
- 69 traces
- Track
- Comparison beads along the track, one for each re-execution that matched.
ProvesA randomly selected rerun produced the same output hash as the committed one.
Does not proveAnything about the inferences that were not sampled. The guarantee is statistical, not per-inference.
- Mechanism
- A fraction of committed traces is drawn at random and re-executed by an independent bonded node under a pinned environment. The re-run output is hashed and compared byte for byte. A mismatch opens an incident.
- Covers
- That the drawn trace reproduces exactly. With the execution environment pinned, 10,000 runs on two hosts produced identical hashes with zero bit-level drift, and a batch-size swing of plus or minus 20 percent did not break that.
- Does not cover
- Traces that were not drawn. Detection follows 1 - (1-p)^K for sampling rate p and K dishonest attempts. At p = 0.05 the expected number of dishonest traces before the first catch is 20.
- Precondition
- Five properties of the execution must be pinned and recorded: hardware SKU, exact weights and quantization format, parallelism topology, software and kernel versions, and the batch size of each forward pass. One of these cannot be pinned across a heterogeneous provider network. An A100 and an H100 running the same weights on the same input agree 0.0 percent of the time at the bit level. KORTX therefore draws re-runs from architecture-separated verifier pools; comparing across architectures would slash honest providers at a rate near 100 percent. The pool key is the architecture each node declares in the verifier registry, behind its bond. The on-chain Verifier account carries no architecture field, so this separation rests on the assignment code and on the stake standing behind that declaration, not on the program. It is also conditional: the filter runs only when the provider's architecture is supplied, and a draw made without it is an unrestricted pool rather than a safe default.
- Trust assumption
- That at least one independent verifier is running and is not colluding with the provider, and that both bonds exceed the gain from cheating. A verifier paid only for finding errors earns nothing while the system works correctly. This is the verifier's dilemma, and it produced a real Bitcoin fork in July 2015.
- Measured cost
- Re-execution: the sampling rate times the inference cost. At p = 0.05 that is a 5 percent surcharge on served inference.
- Determinism: between 1.8 percent and 133 percent depending on the serving stack. The stack is named wherever that figure is quoted, because the two ends of that range are different systems rather than a margin of error.
- Scorer noteQuoted from the index
- A matched sample is direct reproducibility evidence, weighted below a proof because it compares outputs rather than the computation.
Proven

- Reproducibility weight
- 1.0per matched comparison
- Counts toward score
- Yes
- Committed here
- 3 traces
- Track
- Bent by the field and cut hard into the plate.
ProvesA zero-knowledge proof of the core operations, checked mathematically rather than by re-execution.
Does not proveThat a full large model was proved end to end. Whole-model zkML is not currently feasible; the proved operation set is published per plate and everything outside it is unproved.
- Mechanism
- A zero-knowledge proof that the declared circuit, evaluated on the committed input, produces the committed output. Verification requires no trust in the provider.
- Covers
- The arithmetic of the declared circuit, exactly.
- Does not cover
- The amount of work. A proof certifies that private weights exist which map the input to the output through the public architecture. It does not certify that the prover spent computation matching the advertised model size.
- Precondition
- The model has to be in the accepted set below. A workload outside that set is rejected rather than quietly downgraded to a weaker tier, because a silent downgrade leaves the requester believing they received a proof.
- Trust assumption
- The soundness of the proving system, the correctness of the circuit compiler, and the claim that the declared architecture is the one the provider advertises. The last of these is not cryptographic.
- Measured cost
- Classical models are inside a practical range: 0.118 to 6.161 seconds in the framework author's own published benchmark, at 19.4 MB to 383 MB of memory.
- Neural networks are not. A third-party benchmark puts MobileNetV2 at roughly four hours and 204 GB of RAM, and language-model-scale proofs are measured in tens of minutes to days. The full table is below.
- Scorer noteQuoted from the index
- The strongest evidence class carried by this index, within the published proved operation set.
Why attested is worth zero
Weight per comparison
A re-execution forced by a bonded challenge. Adversarial, so it is weighted like a proof rather than like a routine sample.
A proof of the declared circuit, within the accepted model set. The strongest evidence class this index carries.
A matched sample is direct reproducibility evidence, weighted below a proof because it compares outputs rather than the computation.
Attestation is not a rerun. It contributes nothing to the score unless the trace is challenged and re-run.
The scorer's own sentence
Attested traces carry zero reproducibility weight: TEE attestation proves a binary ran, not that its output reproduces. They are reported under traces_untested.
An attestation answers the question "did this binary run on this hardware". Reproducibility answers a different question: "does this output come back if the work is done again". The first cannot be evidence for the second, so folding attested traces into the score at any weight above zero would inflate every provider that never submits to a re-run. They are counted, and they are counted as untested.
The consequence is deliberate and it is unflattering: a provider can commit thousands of attested traces and still be reported as unrated. That is the correct answer, because nobody has checked their work.
Coverage across this index
Traces committed
94
Carrying rerun evidence
21
Untested
73
22 of these are attested and have never been challenged.
Coverage
22.3%
Share of committed traces that any rerun has ever touched.
Committed traces by tier
- Attested
- 2223%
- Sampled
- 6973%
- Proven
- 33%
Read against the coverage figure above: the largest tier on this network is also the one that carries no reproducibility weight at all.
What proven actually accepts
Accepted
Linear and logistic regression
0.118 s measured, 19.4 MB
Support vector machines
0.318 s measured, 23.7 MB
Decision trees and tree ensembles, random forests included
0.308 s to 6.161 s measured, 23.7 MB to 383 MB
Small convolutional networks
Practical only off the response path; see MobileNetV2 below
Refused
Transformer language models, full inference
The framework author states neural networks are outside benchmarking scope
Autoregressive multi-token generation
The circuit grows with the square of the token count
Diffusion pipelines
SDXL end to end measured at 68,950 s and a 145.84 MB proof
3D volumetric networks
A 19M-parameter 3D-UNet measured at 568,543 s, with 4 hours of verification alone
Variable-length sequence input
Circuits are fixed-shape; every length needs its own setup
Anything on a real-time response path
None of the above fits inside a request the user is waiting on
Documented circuit constraints
- Non-unit dilation in convolution and deconvolution is refused by the compiler outright.
- Division is supported only by constants.
- Exponents are supported only as constants.
- An ONNX operator the circuit does not know fails at circuit layout rather than being approximated silently.
There is no complete list of unsupported operators and this page does not pretend to publish one. The compiler routes ONNX through an intermediate representation that folds more than a hundred operators into roughly twenty primitives, so an operator missing from the source's match arms is not necessarily unsupported. What is documented is the failure mode: an unknown operator is a hard error at circuit layout, not a quiet approximation.
Approximated, not refused
Non-linear functions are replaced by lookup tables. The circuit is proved correctly; the function it proves is a near neighbour of the one the model ran.
| Function | Mean error | P99 error |
|---|---|---|
| GELU | 0.026% | 0.158% |
| SiLU | 0.021% | 0.158% |
| RMSNorm | 0.003% | 0.034% |
| Softmax, end to end | Not published | Not published |
Every proving system in this class quantizes to integers. One reports cosine similarity of at least 99.6 percent against the unquantized model, and applies no outlier smoothing to the largest model it supports. Quantization error is not a rounding artefact of the proof; it changes the model that was proved.
Published proving cost
| Workload | Input | Prove time | Memory | Proof size | Hardware |
|---|---|---|---|---|---|
| Linear regressionNot statedFramework author's benchmark, 2024-01-28 | One inference | 0.118 s | 19.4 MB | Not measured | Not stated |
| Support vector machineNot statedFramework author's benchmark, 2024-01-28 | One inference | 0.318 s | 23.7 MB | Not measured | Not stated |
| Tree ensemble regressionNot statedFramework author's benchmark, 2024-01-28 | One inference | 0.308 s | 23.7 MB | Not measured | Not stated |
| The same benchmark states directly that verification time and proof size were not measured, and that neural networks are outside its scope. | |||||
| Random forest classificationNot statedFramework author's benchmark, 2024-01-28 | One inference | 6.161 s | 383 MB | Not measured | Not stated |
| A competing zkVM took 173.4 s on the same model and needed up to 10.2 GB. A third framework failed this case with out-of-memory errors on a machine holding 1,000 GB of RAM. | |||||
| MobileNetV2Not statedThird-party benchmark | One image | About 4 h | 204 GB | Not measured | Xeon E5-2665, 16 threads, 350 GB |
| GPT-2124MDeepProve, published benchmark | 512 tokens | 7.64 min | Not published | 10.71 MiB | 24 cores, 504 GB, EPYC 9254 |
| Gemma 31BDeepProve, published benchmark | 512 tokens | 18.95 min | Not published | 21.73 MiB | 24 cores, 504 GB, EPYC 9254 |
| The only published end-to-end figure for a 512-token sequence at 1B parameters. The proof is 21.73 MiB. | |||||
| LLaMa-213BzkLLM, published paper | One 2,048-token sequence | 803 s | 23.1 GB | 188 kB | One A100 SXM4 40 GB, EPYC 7413 |
| The paper states a semi-honest verifier assumption; its completeness and soundness results rest on that premise. This is not a guarantee against a malicious verifier, and the scheme is not an on-chain verification target. | |||||
| LLaMA-27BZKTorch, published paper | One token | 2,645 s | Not published | 22.85 MB | Xeon Platinum 8358, 64 threads, 4 TB |
| One token. Verification alone takes 100.14 s. | |||||
| 3D-UNet19MZKTorch, published paper | 1 x 1 x 128 cubed | 568,543 s (6.6 days) | Not published | 1.4 MB | Xeon Platinum 8358, 64 threads, 4 TB |
| Verification alone takes 4 hours. | |||||
- Proof artifacts in the table run from 188 kB to 22.85 MB. A Solana transaction holds 1,232 bytes, so those two exceed the on-chain limit by 152 times and 18,500 times. On Solana today only Groth16-class proofs verify on chain directly, at 78,293 to 108,762 compute units against a 1,400,000 unit transaction ceiling.
- Proving a production workload cost 75 times the inference it certified.
- Every figure above is reproduced from a published source. None is a KORTX measurement and none is an estimate. Where a source did not measure something, the cell says so instead of carrying a plausible number.
Not yet measured
- Not yet measured
- Proving time per trace on this network at the proven tier
- Not yet measured
- Verifier re-execution wall clock at the sampled tier
- Not yet measured
- Incident to resolution wall clock
- Not published by the vendor
- Attestation call latency, rate limit, SLA and pricing
- No project in this field publishes one
- False-slash rate against honest providers
- No project in this field publishes one
- Challenge period length used by comparable systems
These are published when they are measured, not before. Filling them with an estimate would make this page indistinguishable from the marketing it exists to replace, and an estimate that is later contradicted costs more than an empty field ever did.
On-chain cost of one commit
- Rent exemptionA 342-byte trace account. Returned in full when the account is closed.
- 3,271,200 lamports
- Signature feeThe only part of a commit that is actually consumed.
- 5,000 lamports
- Total locked per commitDeposit plus fee. Writing this as the cost of a trace overstates it by a factor of 655.
- 0.0032762 SOL
No USD conversion is printed here. The SOL price was not obtained from a primary source in the measurement record, and a hardcoded conversion would be a number this page cannot defend.
What sampling actually catches
| Sampling rate p | K=1 | K=10 | K=20 | K=50 | K=100 | K=200 |
|---|---|---|---|---|---|---|
| 1% | 1.00% | 9.56% | 18.21% | 39.50% | 63.40% | 86.60% |
| 2% | 2.00% | 18.29% | 33.24% | 63.58% | 86.74% | 98.24% |
| 5% | 5.00% | 40.13% | 64.15% | 92.31% | 99.41% | 99.99% |
| 10% | 10.00% | 65.13% | 87.84% | 99.48% | 99.99% | 99.99% |
| 20% | 20.00% | 89.26% | 98.85% | 99.99% | 99.99% | 99.99% |
Probability at least one dishonest trace is caught. 1 - (1-p) ^ K
- Read the first column, not the last. A provider who cheats once at a 5 percent sampling rate is caught 5 percent of the time. Sampling is a deterrent priced against a bond, not a guarantee.
- 99.99 percent means at least 99.99 percent. It is never 100. For any finite K the expression stays below 1, so cells that would round up to 100.00 are written as 99.99 instead.
Choosing a tier
| Situation | Tier | Reason |
|---|---|---|
| Serving a transformer language model on a live response path | Attested or sampled | Proven is structurally unavailable here, not merely expensive. |
| Batch inference over classical models: regression, SVM, trees | Proven | Proving time of 0.118 to 6.161 seconds is inside a practical range. |
| A provider network with mixed GPU architectures | Attested | Sampled requires architecture-separated verifier pools first. An A100 and an H100 agree 0.0 percent of the time at the bit level. |
| A single-architecture pool that can pin its execution environment | Sampled | With five properties pinned, 10,000 runs produced identical hashes and zero drift. |
| A small convolutional network with no latency requirement | Proven | MobileNetV2 costs roughly 4 hours and 204 GB. It works, off the response path. |
| A workload whose correctness matters more than its provenance | None of them | No tier here evaluates whether an output is right. Verification and evaluation are different problems and this system only does the first. |
Tiered verification is not a KORTX invention. Other projects offer a choice of verification strength. The narrower position here is that the tier travels with the output, on chain, as part of the record rather than as a claim in a dashboard.