A strong example is not the same as a complete benchmark
When a Cardiomic peak lines up clearly with a reference heart sound, the result can look convincing. But a few strong examples cannot show how consistently the same method works across thousands of recordings.
That is why we ran Cardiomic across almost the entire public CirCor DigiScope Phonocardiogram Dataset. The source set contained 3,163 recordings, totaling 19 hours and 57 minutes, from 942 subjects and multiple auscultation locations. The benchmark successfully processed 3,162 recordings, or 99.97% of the dataset.

The result was useful precisely because it was mixed. Among valid matched intervals, Cardiomic followed rhythm timing closely. Across the complete, unstratified dataset, however, many Cardiomic peaks did not match the selected S2 target and many reference S2 events remained unmatched.
These findings describe two different questions: whether Cardiomic matched the selected acoustic events, and whether the timing between valid matched events agreed with the reference.
This was an engineering benchmark against a public reference dataset. It was not a clinical validation, a diagnosis, or a validation of consumer smartphone microphones.
The short answer
There is no single percentage that describes the accuracy of the whole system.
For matching Cardiomic’s unclassified acoustic peaks to CirCor S2-start annotations within a 100 ms tolerance, the benchmark found:
| S2-target detection metric | Result |
|---|---|
| Precision | 29.86% |
| Recall | 44.63% |
| F1 score | 35.78% |
These pooled results do not support a claim of reliable S2 detection across the full, unstratified CirCor dataset.
The strongest rhythm result appeared among 21,330 eligible RR intervals formed by events that were matched and consecutive in both Cardiomic and the reference sequence:
| Rhythm metric | Result |
|---|---|
| RR mean absolute error | 21.66 ms |
| RR bias | -6.09 ms |
| RR Lin concordance correlation coefficient | 0.92 |
| Interval heart-rate mean absolute error | 4.41 BPM |
| Interval heart-rate Lin CCC | 0.84 |
The benchmark found useful rhythm information when valid consecutive matches were present. It also showed that event matching and valid-sequence coverage still need substantial improvement.
What the benchmark actually measured
CirCor provides phonocardiogram recordings with reference annotations. S2 start was selected as the target event, and a Cardiomic peak was counted as a match when it fell within 100 ms of that target.
Cardiomic currently exports unclassified acoustic peaks. It does not label each exported event as S1 or S2. A peak counted as “extra” against an S2 target may represent noise, an artifact, an S1-related sound, another acoustic event, or a genuine detection outside the matching window. The benchmark cannot establish the physiological identity of every unmatched peak.
The run used Cardiomic 2.0.3-debug, export schema 2, CirCor 1.0.3, validation policy 3.0, 4,000 Hz source audio, and the latest Cardiomic session for each recording. Rhythm metrics were calculated only from pairs consecutive in both sequences. No interpolation replaced missing events.
One recording was excluded because its reference was incomplete. No validation run failed, and no recording lacked a Cardiomic session. The 99.97% figure therefore describes processing coverage, not detection accuracy.
What worked
The benchmark processed 3,162 recordings consistently and found strong rhythm agreement after restricting the analysis to valid matched sequences.
Of the 21,330 eligible RR intervals, 82.61% were within 30 ms of the reference, 89.72% within 50 ms, and 94.97% within 100 ms. For interval heart rate, 78.95% were within 5 BPM and 89.33% within 10 BPM.
Recording-level results retained comparatively strong agreement for mean RR and mean heart rate, while variability metrics were weaker:
| Recording-level metric | Eligible recordings | MAE | Lin CCC |
|---|---|---|---|
| Mean RR | 2,222 | 25.65 ms | 0.88 |
| Median RR | 2,222 | 26.35 ms | 0.87 |
| Mean heart rate | 2,222 | 5.23 BPM | 0.81 |
| RMSSD | 1,778 | 27.14 ms | 0.50 |
| SDNN | 1,898 | 18.80 ms | 0.57 |
RMSSD and SDNN are particularly sensitive to extra events, missed events, short valid sequences, and local timing errors. No clinical interpretation of HRV was made in this benchmark.
What remains unresolved
Across the dataset, Cardiomic produced 95,307 detections for 63,763 reference S2 events. There were 28,457 matches, 66,850 unmatched Cardiomic events, and 35,306 unmatched reference S2 events.
In addition, 395 recordings had no matched event and 940 had no valid RR interval. Those 940 recordings were not included in the rhythm-error denominator. The strong RR results must be read together with this limitation.
Performance also varied by auscultation location:
| Location | Precision | Recall | F1 | RR MAE |
|---|---|---|---|---|
| Aortic (AV) | 27.15% | 47.48% | 34.55% | 27.10 ms |
| Mitral (MV) | 25.25% | 41.75% | 31.47% | 20.15 ms |
| Pulmonary (PV) | 39.07% | 51.95% | 44.60% | 20.34 ms |
| Tricuspid (TV) | 28.86% | 37.69% | 32.69% | 20.30 ms |
Recordings marked with a murmur had lower event-level F1 than those without a murmur, at 30.92% versus 37.78%. Recordings with an abnormal outcome also had lower F1 than those with a normal outcome, at 32.98% versus 38.46%. These are descriptive subgroup results, not evidence that Cardiomic detects murmurs or clinical outcomes.
Why recording conditions may matter
CirCor is valuable partly because its recordings are not uniformly clean laboratory signals. They may contain low-amplitude heart sounds, movement, handling or contact noise, clipping, changing sensor pressure, environmental sound, murmurs, and overlapping acoustic components.
Cardiomic uses a dynamic threshold and temporal guardrails. A strong noise transient may affect the detector’s short-term state and may be followed by delayed or suppressed detections. In principle, one noisy episode could create both an unmatched Cardiomic event and missed cardiac events immediately afterward.
This is an engineering hypothesis, not a conclusion of the current benchmark. The analysis did not include event-level artifact labels or measure detector-state transitions around each error. It cannot determine how much of the pooled result came from corrupted audio, post-noise recovery, target-definition mismatch, or other detector behavior.
Cardiomic classified 1,809 of 3,163 sessions as internally reliable, or 57.19%. This internal gate combines RR count, rhythm stability, an audio-signal-quality score, heart rate, and, when available, deviation from a recent baseline.
That figure is not event-detection precision, clinical reliability, or usable-recording time. Because the benchmark was not stratified by this flag, it must not be compared directly with precision, recall, or F1.
Why the first examples looked stronger
Early visual inspection found recordings in which Cardiomic peaks repeatedly aligned with CirCor S2 windows. Those examples were real. Two of them, 13918 AV and 13918 PV, appear among the recordings with 100% recall in the complete benchmark.
But clear examples demonstrate only that the method can sometimes work very well. They do not estimate its overall performance. The full benchmark therefore replaces the earlier exploratory 91.2% window-overlap figure as the primary public result.
What this means when using Cardiomic
The benchmark reinforces an important boundary. One recording can provide immediate access to a heart sound and rhythm pattern, but it should not be treated as a diagnosis or final conclusion.
Before interpreting a difference, inspect the recording itself. Background noise, movement, placement, contact pressure, and the phone’s audio pathway can affect what the app receives. A clear waveform and audible repeating pattern provide more useful material for observation than a session dominated by noise or unstable contact.
Repeated sessions create context that one recording cannot provide. Using the same phone, a similar body position, a familiar chest location, and a quiet moment makes later comparisons more grounded. Repetition does not remove the limitations revealed by this benchmark, but it helps distinguish an isolated recording from a recurring personal pattern.
The appropriate value is observation with context, not certainty from one result.
What happens next
The next validation cycle has three priorities.
First, performance must be separated by recording quality using reproducible noise measures, Cardiomic’s quality indicators, and, where feasible, blinded human quality grades. Results should also be stratified by the current reliable flag.
Second, detector state must become observable around matches and errors. This will allow the proposed relationship between noise, threshold changes, guardrails, and nearby missed events to be tested rather than assumed.
Third, algorithm changes must use a fixed development, validation, and locked-test protocol. Improvements should reduce unmatched detections, recover more valid sequences, and preserve or improve rhythm agreement. Smartphone capture requires a separate prospective study with synchronized references across phone models, environments, positions, and participants.
Each material change should be rerun under the same versioned policy. Improvements, regressions, and subgroup results should all be published.
What we can responsibly claim
Cardiomic demonstrated strong rhythm agreement among valid matched intervals in the CirCor dataset. Pooled agreement with S2-start targets was limited under mixed, unstratified recording conditions, and the analysis did not separate clean-signal performance from errors associated with noise or post-noise detector state.
The benchmark does not support claims that Cardiomic reliably detects S2 across the full dataset, provides clinically validated heart-rate or HRV measurements, diagnoses a condition, or has been validated across consumer smartphones.
A practical takeaway
This benchmark is a transparent baseline, not a final verdict. It identifies rhythm information worth preserving, event matching that must improve, and a measurement gap around recording quality and detector behavior.
For the person using Cardiomic, the implication is simple: treat each session as an observation, check whether the signal is clear enough to review, and avoid drawing conclusions from one recording.
Record again tomorrow under similar conditions. Over time, comparable sessions can build a more useful personal reference while validation continues to define what the system can and cannot reliably measure.
Data source and benchmark provenance
The reference data came from the CirCor DigiScope Phonocardiogram Dataset v1.0.3 on PhysioNet, cited by PhysioNet as Oliveira et al. (2022), DOI: 10.13026/tshs-mw03.
The benchmark report was generated on August 18, 2026, using Cardiomic 2.0.3-debug.
Cardiomic is intended for acoustic self-observation and development research. It is not a substitute for professional medical evaluation. If you have symptoms or concerns about your heart, seek advice from a qualified healthcare professional.
