Quick answer: A PQC pilot scorecard needs measurable, pre-defined criteria across nine dimensions to produce a genuine go or no-go signal rather than a subjective impression: compatibility (pass rate across real clients and middleboxes), latency (handshake and signing time under load), throughput (operations per second sustained), certificate and message size (actual bytes versus the classical baseline), failure rate (outright failures, not just degraded performance), observability (whether you can actually see which algorithm was negotiated in production), recovery (rollback time if something breaks), control evidence (whether the pilot produces audit-usable documentation), and business impact (user-facing effect, if any). This guide provides a scorecard across all nine dimensions, building on the lab design covered in our companion guide.
A pilot that ends with “it seemed to work fine” is not a pilot; it’s an anecdote. The difference between a pilot that produces a defensible production decision and one that produces a shrug is whether the scoring criteria were defined and measured before the pilot started, not assembled after the fact to justify whatever happened.
Key Takeaways
- Nine scoring dimensions, compatibility, latency, throughput, size, failure rate, observability, recovery, control evidence, and business impact, produce a defensible go/no-go signal.
- Observability and control evidence are the two dimensions most pilots skip, and the two that matter most for a defensible production decision, not just a technical one.
- Recovery time should be measured directly, not assumed, since a pilot that never tests its own rollback path has not actually validated the rollback plan.
- Scores should be captured against pre-defined thresholds set before the pilot runs, not evaluated retroactively against whatever result the pilot produced.
The Nine Scoring Dimensions
| Dimension | What to measure | Example threshold |
|---|---|---|
| Compatibility | Pass rate across real clients, applications, and middleboxes in the pilot population | ≥99% successful negotiation across tested population |
| Latency | Handshake and signing time under realistic load, not idle single-operation time | Within an agreed percentage of classical baseline under peak concurrency |
| Throughput | Sustained operations per second under production-representative volume | Meets or exceeds current peak issuance/signing volume |
| Certificate/message size | Actual measured bytes for certificates, handshakes, and signed artifacts versus classical baseline | Confirmed within infrastructure MTU and storage capacity limits |
| Failure rate | Outright failures distinct from degraded performance, tracked by failure type | Zero unexplained failures; all failures traced to a known, documented cause |
| Observability | Whether negotiated algorithm and failure telemetry is actually visible in production monitoring | 100% of pilot traffic has negotiated-algorithm data captured |
| Recovery | Measured time to roll back to classical-only operation if triggered | Rollback executes within the defined change-control window |
| Control evidence | Whether the pilot produces documentation usable for audit or compliance purposes | Evidence package reviewed and accepted by compliance/audit stakeholder |
| Business impact | Any user-facing or operational effect observed during the pilot | No unplanned user-facing incidents attributable to the pilot |
Observability and Control Evidence: The Two Dimensions Most Pilots Skip
Compatibility, latency, throughput, and size get measured in almost every pilot, since they’re the obvious technical questions. Observability and control evidence get skipped far more often, and they’re exactly the two dimensions that determine whether a pilot’s result is usable for anything beyond internal technical confidence. A pilot with no negotiated-algorithm telemetry cannot tell you, after the fact, whether the traffic you thought was testing PQC actually negotiated it, or silently fell back to classical algorithms the whole time. A pilot with no audit-usable documentation produces a result your compliance team cannot actually use to support a regulatory filing or an executive report, no matter how technically sound the underlying work was.
Why Recovery Time Needs to Be Measured, Not Assumed
A rollback plan that exists only on paper is not a validated rollback plan. Trigger an actual rollback during the pilot, deliberately, and measure how long it genuinely takes to return to classical-only operation, including the time to detect the trigger condition, execute the change, and confirm the rollback succeeded. This is the step most pilots skip, since it feels like manufacturing a problem rather than solving one, but a pilot that never exercises its own rollback path has left the single most consequential failure mode entirely untested.
What We’d Actually Recommend
Set specific, numeric thresholds for every dimension before the pilot starts, using the lab testing baseline covered in our PQC lab design guide as the starting reference point. Explicitly include observability and control evidence in the scorecard rather than treating them as implicit byproducts of the technical work. Deliberately trigger and measure a rollback during the pilot itself, not just document a theoretical rollback plan, and require every scorecard dimension to clear its threshold, not just a majority, before authorizing broader production rollout.
How Encryption Consulting Can Help
Our PQC Advisory Services design and run pilots against this nine-dimension scorecard, setting thresholds specific to your environment and producing the control evidence your compliance and audit stakeholders actually need to accept the pilot’s results.
CBOM Secure establishes the classical baseline this scorecard measures against, giving the latency, throughput, and size comparisons real production data to score against rather than a generic industry figure.
A Scorecard, Not an Impression
“The pilot went well” is not a production decision; it’s a feeling. Nine measured dimensions, with thresholds set before testing begins, turn a pilot into a genuine go or no-go signal that a technical team, an executive stakeholder, and an auditor can all independently verify against the same data. Observability and control evidence, the two dimensions most often left off an informal pilot review, are exactly the ones that determine whether the result holds up to scrutiny once the pilot is over.
Frequently Asked Questions
Which pilot scoring dimensions are most commonly skipped?
Observability and control evidence. Technical dimensions like compatibility and latency get measured almost by default, but negotiated-algorithm telemetry and audit-usable documentation are frequently left out, which limits how usable the pilot’s result is beyond internal technical confidence.
Why should a pilot deliberately trigger a rollback rather than just documenting a rollback plan?
Because a plan that has never been executed is unvalidated. Measuring actual rollback time, including detection and confirmation, surfaces gaps a paper plan cannot, and rollback is exactly the failure mode with the highest consequence if it doesn’t work as expected.
Should scoring thresholds be set before or after the pilot runs?
Before. Setting thresholds after results come in risks unconsciously calibrating them to match whatever the pilot happened to produce, which defeats the purpose of having measurable criteria at all.
Does every dimension need to pass for a pilot to authorize broader rollout?
Yes, ideally. Treating the scorecard as a checklist where a majority pass is sufficient risks overlooking a single critical failure, such as a broken rollback path or missing observability, that a partial pass rate would obscure.
How does this scorecard relate to the lab testing described in the PQC lab design guide?
The lab establishes the representative test environment and baseline measurements; this scorecard is what’s applied during the subsequent production-adjacent pilot stage, using the lab’s baseline data as the comparison point for the pilot’s real-world results.
