Silicon-Verified Scorecard

The 8-section scorecard, section by section

This is the complete format for every Chrysalis delivery. Accuracy, latency, and energy are measured on your actual chip. The scorecard is a compliance artifact, not a summary.

Silicon-verified scorecard

Glass-break acoustic model

3-class detection. 91.4% cross-validation accuracy. Verified on customer silicon; rig-captured rows pending commissioning.

ChipCustomer silicon (NDA)
RigCRY-RIG-001 / Audio Gym
MetricValueMethodology
01Model Identity
Task classAcoustic event detectionCustomer-defined problem statement
Architecture1D-CNN, 3-layerAuto-selected via rig search
Class labelsglass-break / impact / ambientPredefined class taxonomy
02Data Collection
Collection methodPENDING RIG COMMISSIONINGClosed-loop rig capture
Total labeled samplesPENDING RIG COMMISSIONINGRig auto-label pipeline
Train / validation splitPENDING RIG COMMISSIONINGStratified by class
03Accuracy
Cross-validation accuracy91.4%5-fold stratified CV
False positive rate4.1%Confusion matrix, ambient class
False negative rate6.2%Confusion matrix, glass-break class
04Latency (on-silicon)
Inference time (median)PENDING RIG COMMISSIONING1,000-run on-chip timer
Inference time (p99)PENDING RIG COMMISSIONING1,000-run on-chip timer
Frame pre-processingPENDING RIG COMMISSIONINGRig instrumentation
05Energy (on-silicon)
Energy per inferencePENDING RIG COMMISSIONINGEnergyRunner methodology
Peak current drawPENDING RIG COMMISSIONINGOn-silicon shunt measurement
Sleep currentPENDING RIG COMMISSIONINGEnergyRunner methodology
06Memory Footprint
Flash usage62 KBLinker map analysis
RAM peak18 KBStatic analysis + runtime probe
QuantizationINT8Post-training quantization
07MLPerf Tiny Alignment
Benchmark categoryKeyword spotting (adapted)MLPerf Tiny v1.1 reference
ScenarioSingle-streamMLPerf Tiny inference scenario
Closed-division conformanceIn reviewSelf-certification path
08Reproducibility
Rig identifierCRY-RIG-001Physical rig serial number
Binary hash (SHA-256)a3f7c1d9…b82e04Delivered binary checksum
Scorecard revision1.0.0Semantic versioning

Latency, energy, power and rig-captured labeling rows pending rig commissioning. Accuracy and memory values measured on customer silicon via CRY-RIG-001.

View full scorecard methodology

Methodology: all 8 sections

01

Model identity

MetricValueMethodology
Model hash (SHA-256)Computed at binary freezeDeterministic build
Target chipCustomer-supplied siliconPhysical device
FrameworkTensorFlow Lite Micro / ONNX RuntimeChip-agnostic
QuantisationINT8 post-trainingEdge inference standard

Every scorecard section references the same hash. The binary is the artifact. The hash is the chain of custody.

02

Dataset provenance

MetricValueMethodology
Collection methodAuto-labeled via real sensor chainRig-captured
Labeling pipelineClosed-loop rig stimulus + annotationNo synthetic augmentation
Class split (train / val / test)70 / 15 / 15Stratified random split
Held-out test set custodySealed before training startsBlind evaluation

Data is collected through the same sensor chain the deployed model will run on. Sim-to-real gap is zero by construction.

03

Accuracy

MetricValueMethodology
Cross-validation accuracyReported per classk-fold, k=5
Held-out test accuracyReported per classSingle blind pass
Confusion matrixFull NxN table attachedPer-class false positive/negative
Reference: glass-break result91.4% cross-validation3-class: glass-break / impact / ambient

Accuracy numbers in the scorecard come from the sealed test set, not the training curve. Cross-validation is reported separately.

04

Latency

MetricValueMethodology
Measurement pointInference call on-chipHardware timer, not profiler estimate
Sample count1,000 consecutive inferencesWarm cache, production clock
Reported metricsP50, P95, P99, maxDistribution, not mean alone
Clock frequencyCustomer production settingNo synthetic boost

Latency is measured at production clock. A simulator estimate is not accepted as a substitute. The rig captures the number from the chip.

05

Energy

MetricValueMethodology
Measurement instrumentOn-rig current probeEnergyRunner methodology
Reported unitµJ per inferenceIntegrated over inference window
Sleep state includedYes, duty-cycle energy reportedSystem-level, not core-only
Battery life estimateDerived from duty-cycle dataStated assumptions attached

Energy measurement follows EnergyRunner conventions. The number includes the wake-from-sleep cost, because that is what ships in production.

06

MLPerf Tiny alignment

MetricValueMethodology
Benchmark alignmentMLPerf Tiny v1.1Keyword spotting / anomaly detection tracks
Submission typeOpen, reference implementationReproducing the closed submission format
Scoring scriptAttached as run artifactAuditable
Deviation notesAny deviation documented in scorecardTransparent delta reporting

Chrysalis aligns to MLPerf Tiny format so procurement teams can compare across chips and suppliers using a common frame.

07

Rig configuration

MetricValueMethodology
Rig dimensions730 × 400 × 300 mmT-slot aluminium frame
Active modalitiesAudio, vision, vibration, gasGym-module hot-swap
Trainer interfaceCommon gym interface / DGX Spark TrainerStandardised harness
Chip interfaceCustomer-supplied board + breakoutChip-agnostic

The rig configuration used for measurement is photographed and logged. Any substitution invalidates the scorecard hash.

08

Compliance and audit trail

MetricValueMethodology
EU AI Act relevanceHigh-risk system documentationArticle 9 / Annex IV alignment
Scorecard versioningGit-tagged per re-verification runImmutable history
Re-verification cadenceAnnual (subscription) or on-demandTriggered by drift monitor
Deliverable formatPDF + machine-readable JSONBoth attached to release tag

Regulated verticals need a living document history, not a one-time snapshot. The subscription tier keeps the audit trail current.

Scorecard questions

Is the scorecard a legal compliance document?

The scorecard is structured as a compliance artifact, aligned to EU AI Act Annex IV documentation requirements for high-risk AI systems. It is not legal advice. Your procurement and legal teams decide whether it satisfies a specific regulatory submission.

Can I use the scorecard to compare results across different chips?

Yes. The MLPerf Tiny alignment section uses a common measurement frame, so results from two scorecard runs on different chips are directly comparable, provided the model task is equivalent.

What happens if the re-verification run shows drift?

The monitoring subscription triggers a retraining run when drift exceeds the threshold. The re-verification scorecard is issued under a new version tag. The old scorecard is not overwritten. Both remain in the audit history.

How is the scorecard hash verified?

The SHA-256 hash is computed at binary freeze, before any measurement run. The hash is printed in Section 1 and referenced in every subsequent section. If the binary changes, the hash changes and the scorecard is void.

Ready to run your chip

Start with a Feasibility Gate. Get your scorecard from there.

Send the chip, the sensor modality, and the detection problem. Chrysalis runs it through the rig and returns a go/no-go with preliminary accuracy and latency on your actual silicon.

START AN ENGAGEMENT

Send the chip, the problem, and the modality.

We return a go/no-go within the feasibility window.

Emailchrysalis@sogoodmail.co
Feasibility gate€3,500 fixed price
TurnaroundWithin feasibility window

Submitting this form starts no billing. The Feasibility Gate engagement is priced at €3,500 and confirmed separately.

Built with