EpochCore

Evidence · IBM Quantum

1,605 Heron jobs. Every one enumerated.

The complete IBM Quantum export for this account between 2026-02-09 and 2026-07-02 — job ids, devices, circuits, gate counts, shots and timestamps. Not a selection. Failures included.

What runs on quantum hardware

The pharma build runs on real Heron. Its batched end-to-end validation is job d8sghksbp3hs738332ug on ibm_kingston — six circuits in one job, 2,048 shots each, physically correct results.

The clinical chain routes to a QPU on every mint. Trial-arm prioritisation runs QAOA on IBM Heron and seals the job id, the device, the shot count and the queue state it chose from into the record.

And the circuits are not all small. Six ran at 24 logical qubits or above, to completion — including 110 qubits at depth 415 with 1,204 two-qubit gates on ibm_kingston. Every one is a job id in the file below.

What does not, and the limits of this file

The three sealed discovery steps run on GPU. Fold → generate → predict affinity are NVIDIA NIM, and those records correctly carry no quantum field. That is the only thing “classical” describes.

This corpus stops on 2026-07-02. It is a complete export up to that date, not a running feed, so every job run since is outside it — including the clinical QAOA. Absence from this file is not evidence a thing never ran.

The corpus

What is in the file.

1,605
jobs
3,436
circuits parsed
18,822,898
shots
2.54 h
QPU time
156
qubits, every circuit
23,856
deepest circuit
devicejobs
ibm_fez513
ibm_boston365
ibm_kingston332
ibm_marrakesh213
ibm_pittsburgh171
ibm_torino11

Status: done 1,565 · error 26 · cancelled 14. Primitives: sampler 1,459, estimator 135. Account instances, labelled rather than named because instance ids are credentials: A 1,129 · B 259 · C 193 · D 24. Circuit depth runs 1 to 23,856, median 107.

Verification · tier 1

What you can check with no account at all.

IBM job ids are scoped to the instance that submitted them. A stranger who queries one of these gets RuntimeJobNotFound, so publishing the ids does not by itself make them checkable, and this page does not pretend otherwise.

The gate histogram is different. It is checkable by anyone against IBM’s public device documentation, in about five minutes, with no account and no cooperation from us.

Heron’s native two-qubit gate is CZ, realised through tunable couplers. Eagle and every earlier IBM device use ECR. Heron R2 is 156 qubits. Here is what 3,436 transpiled circuits in this corpus actually contain:

156 is the transpiled device width — every circuit is laid out across the full chip. The logical width is separate and it is in the same file: 0q×135 · 1q×1,879 · 2q×168 · 4q×507 · 6q×629, running all the way to 156. Both numbers are worth reading, and neither is a proxy for the other.

gatecountwhat it is
cz365,431two-qubit gate — Heron
ecr0two-qubit gate — Eagle and earlier
cx0two-qubit gate — not a Heron basis gate
sx800,631single-qubit
rz669,335single-qubit, virtual
x56,389single-qubit
delay156,310idle instruction
measure9,304readout
barrier4,611scheduling
reset16mid-circuit reset

The check

365,431 CZ gates. Zero ECR. Zero CX. Width 156 on every one of 3,436 circuits, across six named Heron devices. That is the transpiler output of Heron hardware and of nothing else — not Eagle, not a generic simulator, not a spreadsheet.

Verification · tiers 2 and 3

The rest, honestly.

  1. Anyone, no account. The gate histogram above, plus internal consistency: shot counts against result cardinality, queue timestamps against QPU seconds, device names against IBM’s fleet.
  2. A reviewer we grant instance read access. Every job id below resolves in IBM’s own console, with IBM’s timestamps, not ours. Ask and we will arrange it under NDA — that is the offer, and it costs us nothing because the jobs are real.
  3. IBM. Can confirm the account and the jobs directly.

The record

Every job, in one file.

Job id, device, primitive, status, queue and run timestamps, QPU seconds, per-circuit width, depth, size and full gate counts, and per-result shot counts. Sixteen rows of it, evenly spaced through the corpus rather than chosen:

job iddevicedepth czshotscreated
d6mk5ls3pels73a0i5b0fez3661,0242026-03-08
d6mkstgfh9oc73eo5k20fez912026-03-08
d6n2o6ofh9oc73eoknv0boston912026-03-09
d7ddlcr0g7hs73drqla0fez4742254,0962026-04-11
d7dfga15a5qc73dpq010kingston4971884,0962026-04-12
d7dgat15a5qc73dpqse0boston4431734,0962026-04-12
d7fbh421u7fs739lqai0fez4632254,0962026-04-14
d7pbnnm25ies7395hfegkingston146274,0962026-04-30
d7pp4m1oagoc73fiic50fez271814,0962026-04-30
d7q0se1oagoc73fisce0pittsburgh2564,0962026-05-01
d7q2tcc3lfgs73fglg5gfez2464,0962026-05-01
d8c88c47avuc73dqj6pgpittsburgh4621704,0962026-05-28
d8c8fqqjki0s73aqs5l0marrakesh4851904,0962026-05-28
d8iebl66983c73ds1uegboston4491704,0962026-06-07
d8nnkro32u0s73fcmi60fez116324,0962026-06-15
d8oa5o832u0s73fddvs0fez56114,0962026-06-16
Download all 1,605 · JSON, 1021 KB

sha256 5e58a35392141bafc2c09eccbf348d3ae5f827dd9267f5a08e428186b223b88c

IBM instance identifiers are credentials and are replaced with the labels A–D. Nothing else is removed. The build refuses to write the file if any CRN, UUID or account-id-shaped string survives into it.

Completeness

Three gaps, named.

A published dataset that hides its holes is worth less than one that points at them.

  • 11 jobs (all ibm_torino) return no API payload. They appear from the CSV export only, with device and status but no circuit detail.
  • 24 jobs returned a payload our 2026-09-18 retrieval could not decode. That is our bug, not IBM’s. They carry no circuit or result detail here and are excluded from every aggregate on this page.
  • 1,416 of 1,605 carry result data. The remainder either failed, were cancelled, or completed without our retrieving their results.

Every number on this page is computed from the file at build time. None of them is typed by hand.

Cross-substrate experiment · R13

The A/B, and why it is not a result yet.

On 2026-05-16 we ran one experiment across five substrates: GPU statevector, GPU MPS, our router, IBM with the flash_sync carrier, and IBM without it as the control. Four proprietary algorithms, two variants each, forty runs.

All sixteen of its IBM job ids are in the file above. Twelve of them returned RuntimeJobFailureError; four completed. None returned shots. The GPU arms returned 1,024 shots per run; the eight that failed outright were all on GPU MPS.

So

The rig is real, it is genuinely cross-substrate, and it reached real hardware with a real control arm. It has not yet produced a hardware A/B result, and no claim on any EpochCore page rests on one. When it does, it will appear here with its job ids.

Width, on hardware

The wide circuits.

Most of this corpus is small circuits — that is what a fingerprinting and calibration workload looks like. It is not the whole of it. 4 circuits ran at 24 logical qubits or above with at least 100 two-qubit gates, on real Heron, to completion. The gate floor is deliberate: it excludes wide-register readouts, which carry a large clbit count without being a wide computation.

logical qdepthcz job iddevicedate status
120480119d8t1eodbh0os73eott6gkingston2026-06-23DONE
1104151,204d8t0pjcbp3hs7383qmogkingston2026-06-23DONE
1085561,353d8t12f5bh0os73eot7ogmarrakesh2026-06-23DONE
249261,420d8vacs0pknjs73a0ikpgkingston2026-06-26DONE

A 110-qubit circuit at depth 415 with 1,204 two-qubit gates, completed on ibm_kingston. A 108-qubit one at depth 556 with 1,353. These are not toy circuits and they are not simulations — each row is a job id in the file above, on a named device, with IBM’s own status.

What they are, and why they are batched

These are φ-weighted swarm-consensus circuits from the deployed quantum_consensus tool, which runs an exact state-vector below a calibrated threshold and escalates to hardware above it. The wide runs are the escalations — and the design is the interesting part:

job iddevicepubs in one session
d8t0pjcbp3hs7383qmogkingston 4q parity → top 1111 at p 0.375  +  110q escalation, depth 415, 1,204 cz
d8t1eodbh0os73eott6gkingston 2q parity → top 00 at p 0.512  +  120q escalation, depth 480
d8t12f5bh0os73eot7ogmarrakesh 108q escalation, depth 556, 1,353 cz, 2,048 shots
d7pbnne25ies7395hfd0kingston 17-agent φ-weighted consensus, 4q, depth 534, 174 cz, 4,096 shots

The small arm and the wide arm ride in the same job. Same chip, same calibration window, one SamplerV2 session — so the parity circuit is not a separate experiment you have to trust, it is a control that experienced exactly what the escalation experienced. The manifest shows it directly: those jobs carry n_pubs 2 with both widths inside.

The project’s own dossier reports that for the 4-qubit consensus circuit, the exact state-vector top outcome was 1111 and real Heron agreed, at Hellinger fidelity 0.920 raw and un-mitigated. This corpus confirms the hardware half of that independently — pulled from IBM’s API, not from the receipt — showing top 1111 at p 0.375 on that pub.

The router

A fit, not a live loop.

The routing weights were fitted on this corpus. That is what “measured on 1,605 real IBM Heron jobs” means — a fit against recorded outcomes, not a credential and not a real-time telemetry feed.

Here is why a table is worth having at all, straight out of the file. Wall time from submission to finish, completed jobs:

devicejobsp50 s p90 sp99 s
fez5127.1176.13210.6
boston35717.4783.5319377.0
kingston3307.5114.94024.4
marrakesh2097.1234.43349.7
pittsburgh1536.461.32002.2

The median tells you nothing — every device answers in seconds, within a few seconds of each other. The tail is everything. At p90 the spread is 13x, from pittsburgh at 61 s out to 783 s. The gap between wall time and QPU seconds is almost entirely queue, which is why a scheduling decision moves an order of magnitude while the physics does not move at all.

We backtested the policy rather than asserting it

A counterfactual backtest over this corpus (15 September 2026) scored four routing policies against the operator’s real choices on the 1,333 jobs after a 30-day training window, with coverage accounted for so no policy can look good by being scored on the easy half:

policymedian sp90 sresult
operator (actual)6.9320.7the baseline
frozen table, first 30d6.9963.0loses — 3.0× worse at the tail
trailing 6h median6.5553.6loses
trailing 24h median6.5483.6loses
oracle (upper bound)5.88.8not a policy — the headroom

Both of our own candidate policies lost to the human, and we publish that. A frozen calibration-style table is 3× worse than the operator at p90, and a trailing-median proxy is better than the table and still worse than the operator. What the backtest does establish is headroom: an oracle reaches p90 8.8 s, 36× below the operator. The signal is there; a trailing median is not it.

Month to month the best-performing backend held its position in 0 of 3 transitions. With n=3 that cannot establish non-persistence — it establishes that the persistence a frozen table assumes is unevidenced here, which is enough to stop trusting one.

Two signals the backtest could not use, because a historical export does not carry them, are live queue depth at submit and the backend’s operational flag and recent error state. The dispatcher reads both today and publishes every candidate’s value on each run; recording them per job and re-running this exact backtest against the same baselines is the next piece of work.

Reproduced independently

That table is not asked for on trust. The method is reimplemented from the raw IBM export in tools/routing_backtest.py and re-run. Three of the five policies reproduce their scored-job counts exactly1,102, 1,051 and 1,104 — with p90 within 1.8%, 3.7% and 2.3%. It independently selects the same frozen-table pick, finds the same two rankable backends in the training window, and lands both headline ratios: the frozen table 3.1× worse than the operator — the report says 3.0×, and the difference is the reproduction, not a disagreement — and 36× of headroom to the oracle. The two trailing-window policies differ more — 13 and 15 fewer jobs scored, p90 13% and 18% lower — and both still lose to the operator, which is the finding. The full report is available on request.

The threshold that does exist

Inside the discovery chain the gate that actually decides is the drug-likeness pre-screen, and it is a number:

score = qed − 0.10 × lipinski_violations − 0.05 × n_alerts
rule: Lipinski Ro5 (≤1 violation) + Veber + verdict != deprioritize

Measured behaviour, which is the interesting part: it passes 93% of the candidates our own generator proposes, but only 3.1% of 97 known potent KRAS G12D compounds, 65.1% of 415 EGFR, and 59.1% of 232 MAT2A. The KRAS number is not a bug — real G12D inhibitors are beyond-Rule-of-5 by design, and a Ro5 screen rejects them. It is published here because it is the honest limit of the screen.

The discovery chain has no automated GPU→QPU escalation today — that specific chain, not the product. The quantum_consensus tool above escalates on a calibrated threshold and has done so on hardware. Sealing the discovery chain's routing decision as a visible step, including when it says no, is the next build.

On hardware · the pharma build

One job, six circuits, checkable three ways.

The Quantum Pharma build’s batched end-to-end validation ran on ibm_kingston on 2026-06-22 as a single Qiskit Runtime job holding six pubs — four circuits plus two zero-noise-extrapolation folds — fired only once the queue drained to 15 pending. Its results are physically correct: GHZ-3 splits 000/111, QFT-3 concentrates on 000, Bell gives 00+11.

job iddevicepubs shots eachdate
d8sghksbp3hs738332ugibm_kingston 62,0482026-06-22

Three independent records agree on it: the build’s own captured artifact, its evidence dossier, and the manifest on this page — which lists the same job with six circuits at depths 10, 9, 32, 7, 10, 10 and the CZ counts those circuits imply. None of the three was produced from the others. That is what a checkable claim looks like, and it is the reason the whole export is published rather than a summary.

On hardware · the clinical chain

The clinical chain routes to a QPU.

Step 03 of the clinical-evidence chain prioritises trial arms with QAOA, executed on IBM Heron through Qiskit Runtime SamplerV2. It is not a simulator and not a stand-in. The sealed record carries what a reviewer needs:

fieldsealed value
hardware_qputrue
ibm_job_iddano0jg2fm4c73f4g6m0
backendibm_pittsburgh
job_statusDONE · 4,096 shots
verifyQiskitRuntimeService().job('dano0jg2fm4c73f4g6m0').result()
backend_choiceibm_pittsburgh (0 pending) from 6 eligible of 6 · candidates boston 311, phoenix 338, fez 1, pittsburgh 0, marrakesh 0, kingston 0 · rule least-queue among eligible

That last row is the routing decision itself, written down. Six candidate devices, each with its queue depth at the moment of submission, the eligibility filter, and the rule that picked one. It is published on every run, so the claim that the dispatcher reads live backend state is checkable per mint rather than asserted.

The job id is not in the corpus above, and that is not a discrepancy: the export ends 2026-07-02 and this ran on 2026-09-20. Same instance-scoping caveat applies — the verify line is the call a reviewer with access runs.

Not claimed

What this page does not say.

  • That any discovery step ran on quantum hardware. None has. The clinical trial-design step does, and is evidenced above.
  • That this corpus is current. It ends 2026-07-02 and is not a live feed.
  • That the seal proves execution. It proves integrity — that our key signed a given content and nothing changed after. It does not prove which machine ran it.
  • That a stranger can resolve these job ids unaided. They cannot; see tier 2.
  • That any affinity number here or elsewhere is a measurement. They are predictions.
  • That the flash_sync experiment has a hardware result. It does not, yet.