How often does a peak detected by Cardiomic line up with a reference heart sound? Our Validation v1 tested this across 3,162 recordings from CirCor, a collection of recorded heart sounds with reference markings.
The comparison was specific: acoustic peaks detected by Cardiomic against the start of the second heart sound, called S2, marked in CirCor. We counted a match when a peak fell within 100 milliseconds—one tenth of a second—of that reference mark.
In the part of each recording covered by the reference markings, about 50 out of every 100 Cardiomic peaks matched S2. Looking from the other direction, Cardiomic matched about 42 out of every 100 reference S2 sounds, leaving about 58 unmatched.
Those two results answer different questions: how many detected peaks matched the target, and how many targets the app found. Neither is a single accuracy score for the whole app.
What matched, and what did not?
We evaluated the same recordings in two ways. Full recording includes all detected peaks. Annotated span only, or span only, includes peaks from the first reference-marked section to the last, leaving out the beginning and end beyond those boundaries.
Span only does not select cleaner recordings. The selected stretch can still contain noise and gaps in the markings. It also includes the sounds between S2 events, rather than keeping only the S2 sounds themselves.
| What we counted | Full recording | Annotated span only |
|---|---|---|
| Recordings compared | 3,162 | 3,162 |
| Reference S2 sounds | 63,763 | 63,763 |
| Cardiomic peaks evaluated | 95,307 | 53,547 |
| Peaks that matched S2 | 26,955 | 26,945 |
| Peaks with no S2 match | 68,352 | 26,602 |
| Reference S2 sounds left unmatched | 36,808 | 36,818 |
| Share of peaks that matched S2 (precision) | 28.28% | 50.32% |
| Share of reference S2 sounds found (recall) | 42.27% | 42.26% |
| Combined matching score (F1) | 33.89% | 45.94% |
Precision counts matches out of all evaluated peaks. Recall counts matches out of all reference S2 sounds. F1 combines these two percentages, so a high score requires both finding the reference sounds and limiting unmatched peaks. The table combines events across all recordings.
With span only, 26,945 peaks matched and 26,602 did not. Another 36,818 reference S2 sounds had no matching peak. These are the successes and misses for this particular comparison.
An unmatched peak is not necessarily a false heart sound. Cardiomic’s peaks do not carry an S1 or S2 label: a peak might correspond to the first heart sound, another sound, noise, or a sound outside the allowed timing distance. Outside the reference-marked stretch, there may be no reference available to judge it against.
Span only leaves out 41,760 peaks while losing just 10 matches. This explains why the share of matching peaks rises so much. The app has not detected more S2 sounds; we have changed which part of the recording counts in the comparison. In both views, more than half of the reference S2 sounds remain unmatched.
How close was the rhythm when peaks matched?
Finding a sound and measuring the time between sounds are separate tasks. For the rhythm comparison, we used only pairs of matched events that followed each other in both Cardiomic and CirCor, without a skipped event in either sequence. Missing matches were not filled in.
| Rhythm comparison | Full recording | Annotated span only |
|---|---|---|
| Time intervals compared | 20,556 | 20,553 |
| Average size of the timing error | 17.18 ms | 17.17 ms |
| Average signed timing difference | -2.05 ms | -2.06 ms |
| Timing agreement score (1 = perfect agreement) | 0.96 | 0.96 |
| Average heart-rate error per interval | 3.48 BPM | 3.48 BPM |
| Heart-rate agreement score (1 = perfect agreement) | 0.93 | 0.93 |
| Intervals with an error of 30 ms or less | 85.53% | 85.54% |
| Intervals with an error of 50 ms or less | 92.83% | 92.85% |
| Intervals with an error of 100 ms or less | 97.74% | 97.75% |
A millisecond (ms) is one thousandth of a second. The average timing error was about 17 milliseconds. The negative signed difference means Cardiomic’s intervals were, on average, slightly shorter than the reference intervals. The agreement scores summarize how closely the paired measurements agree; they are not percentages of correct detections.
These results were almost identical in both views because almost the same matched pairs remained. They describe the spacing between acoustic events, not a comparison with electrical heart signals from an ECG.
There is an important limit: 1,101 recordings had no qualifying pair of matches. Only 2,061 of the 3,162 recordings contributed to the rhythm results. A small error among matched pairs does not show that rhythm was measured equally well throughout every recording.
Measures of how much the intervals varied agreed less closely than average rhythm. This study therefore does not establish that Cardiomic can provide clinically reliable measurements of heart-rate variability.
Why results differ between recordings
Performance varied with the chest location where the sound was recorded. The table uses the same F1 score introduced above, which combines finding reference S2 sounds with limiting unmatched peaks.
| Recording location | Full-recording F1 | Span-only F1 | Average timing error, full / span |
|---|---|---|---|
| Aortic (AV) | 32.12% | 47.87% | 20.11 / 20.11 ms |
| Mitral (MV) | 29.93% | 43.51% | 16.78 / 16.75 ms |
| Pulmonary (PV) | 42.78% | 53.65% | 16.31 / 16.31 ms |
| Tricuspid (TV) | 30.84% | 38.77% | 16.29 / 16.26 ms |
These are the four main chest recording locations in CirCor. Four additional recordings labeled Phc are included in the overall totals but not in this table. The pulmonary location had the highest matching score among the four main locations.
Recordings marked as containing a murmur also had lower matching scores. In span only, F1 was 37.93% when a murmur was present and 48.35% when it was absent. This comparison does not mean Cardiomic can identify murmurs.
The study does not yet tell us how much of the error comes from noise, other heart sounds, or the way Cardiomic detects peaks. Answering that requires reviewing recording quality and individual errors, rather than assuming every unmatched peak has the same cause.
What Validation v1 means for a Cardiomic session
The useful finding is that the timing between consecutive matches was close to the reference. The equally important limitation is that many reference S2 sounds did not match, and many recordings could not contribute to that timing comparison.
Validation v1 used the CirCor DigiScope heart-sound dataset, recorded with an electronic stethoscope. We examined 3,163 recordings; one lacked usable S2 reference markings, leaving 3,162 for comparison. Both views used the same Cardiomic sessions and the same matching rule.
This is a test against recorded reference sounds. It does not establish the same performance when recording through a phone microphone, and it does not validate diagnosis.
For everyday use, listen back to a session and notice whether the repeating sound pattern is easy to review. A number becomes more useful when you can relate it to the recording itself. Comparing sessions under similar conditions can help you notice recurring patterns, but repetition does not remove the measurement limits shown here.
Next steps: testing the new markers
Validation v1 provides a starting point for a second benchmark article about Cardiomic’s new markers. The main question is straightforward: can the markers match more reference sounds and follow longer stretches of rhythm, while keeping timing errors small?
We will compare the existing acoustic peaks and the new markers on the same recordings, showing full-recording and span-only results side by side. Each marker needs a clear meaning before it can be judged against a reference. Appearing on the waveform is not enough to establish that a marker correctly identifies a particular heart sound.
The broader study should show how many markers matched, how many did not, how many reference sounds were missed, and how much of each recording could be followed. It should also examine different chest locations, sound quality, recordings with murmurs, and interrupted patterns.
Recordings used to adjust the method should be kept separate from those used for the final test, including recordings from the same person. Difficult cases should be reported alongside successful ones. A further study using phone recordings and simultaneous reference measurements will be needed to understand everyday capture performance.
For your next observation, listen back to a session in Cardiomic, then record again tomorrow under similar conditions. Compare the sound and repeating pattern while the next benchmark investigates what the new markers can add.
