Silicon-Verified Scorecard
This is the complete format for every Chrysalis delivery. Accuracy, latency, and energy are measured on your actual chip. The scorecard is a compliance artifact, not a summary.
3-class detection. 91.4% cross-validation accuracy. Verified on customer silicon; rig-captured rows pending commissioning.
| Metric | Value | Methodology | ||
|---|---|---|---|---|
| 01 | Model Identity | |||
| Task class | Acoustic event detection | Customer-defined problem statement | ||
| Architecture | 1D-CNN, 3-layer | Auto-selected via rig search | ||
| Class labels | glass-break / impact / ambient | Predefined class taxonomy | ||
| 02 | Data Collection | |||
| Collection method | PENDING RIG COMMISSIONING | Closed-loop rig capture | ||
| Total labeled samples | PENDING RIG COMMISSIONING | Rig auto-label pipeline | ||
| Train / validation split | PENDING RIG COMMISSIONING | Stratified by class | ||
| 03 | Accuracy | |||
| Cross-validation accuracy | 91.4% | 5-fold stratified CV | ||
| False positive rate | 4.1% | Confusion matrix, ambient class | ||
| False negative rate | 6.2% | Confusion matrix, glass-break class | ||
| 04 | Latency (on-silicon) | |||
| Inference time (median) | PENDING RIG COMMISSIONING | 1,000-run on-chip timer | ||
| Inference time (p99) | PENDING RIG COMMISSIONING | 1,000-run on-chip timer | ||
| Frame pre-processing | PENDING RIG COMMISSIONING | Rig instrumentation | ||
| 05 | Energy (on-silicon) | |||
| Energy per inference | PENDING RIG COMMISSIONING | EnergyRunner methodology | ||
| Peak current draw | PENDING RIG COMMISSIONING | On-silicon shunt measurement | ||
| Sleep current | PENDING RIG COMMISSIONING | EnergyRunner methodology | ||
| 06 | Memory Footprint | |||
| Flash usage | 62 KB | Linker map analysis | ||
| RAM peak | 18 KB | Static analysis + runtime probe | ||
| Quantization | INT8 | Post-training quantization | ||
| 07 | MLPerf Tiny Alignment | |||
| Benchmark category | Keyword spotting (adapted) | MLPerf Tiny v1.1 reference | ||
| Scenario | Single-stream | MLPerf Tiny inference scenario | ||
| Closed-division conformance | In review | Self-certification path | ||
| 08 | Reproducibility | |||
| Rig identifier | CRY-RIG-001 | Physical rig serial number | ||
| Binary hash (SHA-256) | a3f7c1d9…b82e04 | Delivered binary checksum | ||
| Scorecard revision | 1.0.0 | Semantic versioning | ||
Latency, energy, power and rig-captured labeling rows pending rig commissioning. Accuracy and memory values measured on customer silicon via CRY-RIG-001.
View full scorecard methodology→Methodology: all 8 sections
| Metric | Value | Methodology |
|---|---|---|
| Model hash (SHA-256) | Computed at binary freeze | Deterministic build |
| Target chip | Customer-supplied silicon | Physical device |
| Framework | TensorFlow Lite Micro / ONNX Runtime | Chip-agnostic |
| Quantisation | INT8 post-training | Edge inference standard |
Every scorecard section references the same hash. The binary is the artifact. The hash is the chain of custody.
| Metric | Value | Methodology |
|---|---|---|
| Collection method | Auto-labeled via real sensor chain | Rig-captured |
| Labeling pipeline | Closed-loop rig stimulus + annotation | No synthetic augmentation |
| Class split (train / val / test) | 70 / 15 / 15 | Stratified random split |
| Held-out test set custody | Sealed before training starts | Blind evaluation |
Data is collected through the same sensor chain the deployed model will run on. Sim-to-real gap is zero by construction.
| Metric | Value | Methodology |
|---|---|---|
| Cross-validation accuracy | Reported per class | k-fold, k=5 |
| Held-out test accuracy | Reported per class | Single blind pass |
| Confusion matrix | Full NxN table attached | Per-class false positive/negative |
| Reference: glass-break result | 91.4% cross-validation | 3-class: glass-break / impact / ambient |
Accuracy numbers in the scorecard come from the sealed test set, not the training curve. Cross-validation is reported separately.
| Metric | Value | Methodology |
|---|---|---|
| Measurement point | Inference call on-chip | Hardware timer, not profiler estimate |
| Sample count | 1,000 consecutive inferences | Warm cache, production clock |
| Reported metrics | P50, P95, P99, max | Distribution, not mean alone |
| Clock frequency | Customer production setting | No synthetic boost |
Latency is measured at production clock. A simulator estimate is not accepted as a substitute. The rig captures the number from the chip.
| Metric | Value | Methodology |
|---|---|---|
| Measurement instrument | On-rig current probe | EnergyRunner methodology |
| Reported unit | µJ per inference | Integrated over inference window |
| Sleep state included | Yes, duty-cycle energy reported | System-level, not core-only |
| Battery life estimate | Derived from duty-cycle data | Stated assumptions attached |
Energy measurement follows EnergyRunner conventions. The number includes the wake-from-sleep cost, because that is what ships in production.
| Metric | Value | Methodology |
|---|---|---|
| Benchmark alignment | MLPerf Tiny v1.1 | Keyword spotting / anomaly detection tracks |
| Submission type | Open, reference implementation | Reproducing the closed submission format |
| Scoring script | Attached as run artifact | Auditable |
| Deviation notes | Any deviation documented in scorecard | Transparent delta reporting |
Chrysalis aligns to MLPerf Tiny format so procurement teams can compare across chips and suppliers using a common frame.
| Metric | Value | Methodology |
|---|---|---|
| Rig dimensions | 730 × 400 × 300 mm | T-slot aluminium frame |
| Active modalities | Audio, vision, vibration, gas | Gym-module hot-swap |
| Trainer interface | Common gym interface / DGX Spark Trainer | Standardised harness |
| Chip interface | Customer-supplied board + breakout | Chip-agnostic |
The rig configuration used for measurement is photographed and logged. Any substitution invalidates the scorecard hash.
| Metric | Value | Methodology |
|---|---|---|
| EU AI Act relevance | High-risk system documentation | Article 9 / Annex IV alignment |
| Scorecard versioning | Git-tagged per re-verification run | Immutable history |
| Re-verification cadence | Annual (subscription) or on-demand | Triggered by drift monitor |
| Deliverable format | PDF + machine-readable JSON | Both attached to release tag |
Regulated verticals need a living document history, not a one-time snapshot. The subscription tier keeps the audit trail current.
Scorecard questions
The scorecard is structured as a compliance artifact, aligned to EU AI Act Annex IV documentation requirements for high-risk AI systems. It is not legal advice. Your procurement and legal teams decide whether it satisfies a specific regulatory submission.
Yes. The MLPerf Tiny alignment section uses a common measurement frame, so results from two scorecard runs on different chips are directly comparable, provided the model task is equivalent.
The monitoring subscription triggers a retraining run when drift exceeds the threshold. The re-verification scorecard is issued under a new version tag. The old scorecard is not overwritten. Both remain in the audit history.
The SHA-256 hash is computed at binary freeze, before any measurement run. The hash is printed in Section 1 and referenced in every subsequent section. If the binary changes, the hash changes and the scorecard is void.
Ready to run your chip
Send the chip, the sensor modality, and the detection problem. Chrysalis runs it through the rig and returns a go/no-go with preliminary accuracy and latency on your actual silicon.
START AN ENGAGEMENT
We return a go/no-go within the feasibility window.