Data & AI
How a heart recording becomes a decision: what a phonocardiogram looks like, what the AI sees at every step, how the model was trained and how well it works. Hover over, or tab to, any info button (the small i) or coloured part for a short explanation.
Reading a phonocardiogram
A phonocardiogram (PCG) is a recording of the sounds a heart makes. Every heartbeat has two main sounds: the first heart sound (S1) and the second heart sound (S2). The gap from S1 to S2 is systole, and the gap from S2 to the next S1 is diastole. A heart murmur is an extra sound that fills part of the beat. In this made-up example a murmur fills every systole.
Hover over a coloured band, or hover over or tab to a legend button, to see what it means.
What the AI sees, step by step
Pick a real recording and a start time. On this public snapshot no server runs anything: the windows were precomputed every 5 s on the snapshot date, with the same feature code the classifier uses, and this page shows what comes out of each stage. The slider shows the saved window at or before the start you pick. A recording from the test split was never used for training.
Recordings from the CirCor DigiScope Phonocardiogram Dataset v1.0.3 (PhysioNet, ODC-By 1.0); full citation under How the model was trained.
Loading recordings…
Scroll sideways to see the whole diagram.
Step 1 of 6: The raw window
The classifier never looks at a whole recording at once. It cuts the recording into 5 s windows and starts a new window every 2.5 s, so each window holds several heartbeats. This plot is the window you chose, as the file stores it, with its average level (a constant offset) subtracted.
Step 2 of 6: Keep 20–600 Hz, then fix the loudness
The analog front end is designed to pass 20–600 Hz and weaken the rest. The software removes everything outside that band exactly, so every condition sees the same band. Then the window is divided by its own RMS level (its typical size), so a louder and a quieter recording with the same shape give the same numbers. We band-limit first so the loudness is measured only in the band the features use.
RMS divisor
—
Waiting for a window
Step 3 of 6: MFCC: the spectrum, frame by frame
The window is cut into short frames, 128 ms long with a new frame every 32 ms. For each frame, 20 numbers called MFCCs (mel-frequency cepstral coefficients) describe the shape of the spectrum between 20 and 600 Hz, using 32 mel bands. The picture has one column per frame and one row per coefficient.
Why MFCCs? The mel scale bends frequency to match hearing, but it is almost a straight line below 1 kHz, so the bending does little in this band. MFCCs still help here because they compress the spectrum into a few numbers and make those numbers less correlated with each other (decorrelation).
Step 4 of 6: Summarise into 60 numbers
The model needs the same amount of input for every window, so each of the 20 coefficients is boiled down to three numbers over the whole window: its mean, its standard deviation (how much it varies) and the standard deviation of its delta (how much its frame-to-frame change varies). 20 × 3 = 60 numbers. This is everything the classifier is told about the window.
Show the 60 numbers as a table
| MFCC | Mean | Std | Std of delta |
|---|
Step 5 of 6: The classifier adds up the evidence
Logistic regression is small and fast, and its decision can be read as a sum. Each of the 60 numbers is first standardised: z = (x − mean) ÷ spread, using the average and spread of the training windows. Each z is multiplied by a weight the model learned. Those contributions, plus a fixed intercept, add up to one score: the logit. A positive contribution pushes towards murmur and a negative one towards normal. A logit of zero is an even lean (p = 0.5), but a window is labelled “Murmur present” only when p reaches the threshold t, so the decision line is logit = ln(t ÷ (1 − t)), about — for this model. The bars show the features that pushed hardest; all the others are combined in one bar.
The logit is squeezed into a probability between 0 and 1 with the sigmoid function, p = 1 ÷ (1 + e^(−logit)), and compared with the threshold.
Confidence is the probability the model gives to the label it chose. It is not a calibrated clinical probability: it only shows how strongly the model leans, on the data it was trained on. It can be below 50 % because the threshold is not 0.5.
Step 6 of 6: From windows to a recording decision
In Review, the Monitor's main result (Whole recording) is one decision per recording; its This window row, and Live mode, show single-window results instead. For the recording decision the model scores every 5 s window of the recording (—), averages the window probabilities and compares the mean with the same threshold.
How the model was trained
The model learns from recordings that doctors have already labelled. Training reads only the training half of the data; it never opens a test file.
The data
The CirCor DigiScope Phonocardiogram Dataset, version 1.0.3, from PhysioNet under the Open Data Commons Attribution License 1.0. It holds heart sounds from 942 paediatric patients: 3163 recordings at 4 kHz, each from one of up to four auscultation locations (aortic AV, pulmonary PV, tricuspid TV and mitral MV). A recording is labelled “murmur present” if a murmur was audible at that location, and “murmur absent” if the patient has no murmur. Recordings of patients labelled Unknown are left out, and so are recordings of a patient with a murmur at a location where it was not audible.
This public site includes the waveforms of a few demo recordings from this dataset (most or all of each recording), used under the Open Data Commons Attribution License 1.0. Cite the dataset when you use it:
- Oliveira J, Renna F, Costa P, et al. The CirCor DigiScope Phonocardiogram Dataset (version 1.0.3). PhysioNet; 2022. doi:10.13026/tshs-mw03
- Oliveira JH, Renna F, Costa P, et al. The CirCor DigiScope Dataset: From Murmur Detection to Murmur Classification. IEEE Journal of Biomedical and Health Informatics. 2021. doi:10.1109/JBHI.2021.3137048
- Pollard T, Moody BE, Lehman L, et al. PhysioNet as a global platform for biomedical research. Nature Health. 2026. doi:10.1038/s44360-026-00096-z
The recipe
- Split by child, 50/50 (seed 3641). Half of the children train the model and the other half are held out. No child has recordings on both sides.
- Make windows. Every training recording is cut into 5 s windows, a new one every 2.5 s, and each window becomes 60 numbers (steps 1–4 above). A window gets the label of its recording.
- Compare candidates (three strengths of logistic regression and a random forest) with 5-fold cross-validation, grouped by child.
- Pick the model: the best logistic regression by AUC, unless the random forest beats it by more than one fold standard deviation.
- Fix the threshold on the out-of-fold training predictions with Youden's rule. It is set once and never changed for any test condition.
- Fit the winner on all training windows and save it.
Training patient IDs
—
Training recordings
—
Training windows
—
Threshold
—
| Model | AUC, mean ± SD over folds | Balanced accuracy at 0.5 | Sensitivity at 0.5 | Specificity at 0.5 |
|---|
Why logistic regression? It is cheap to run on a Raspberry Pi and its decision can be explained feature by feature (step 5). A random forest is trained only as a comparison row.
How well it works
The experiment scores the same held-out files under different conditions, with the same model, features and threshold, and compares the results. The bench set has one recording each from a balanced subset of the held-out children: equal numbers of children with and without a murmur, so not every held-out child is in it (the file counts are in the table below). The test split has every recording of the held-out children.
Headline numbers
Accuracy
—
Sensitivity
—
Specificity
—
AUC
—
Balanced accuracy
—
Re-acquired
NOT MEASURED
Needs the hardware
| Condition, set and files | Accuracy | Sensitivity | Specificity | AUC | Balanced accuracy | Always-absent baseline |
|---|
Under each rate are its 95% Wilson confidence interval and the number of files behind it.
| Grade |
|---|
Confusion matrices
A confusion matrix counts every file by what it really is (rows) and what the model said (columns). Hover over, or tab to, a cell for what it means.
Simulated bandpass against injected Simulated
Each bench file was scored twice with the same model and threshold: as read from disk (injected) and after a software model of the 20–600 Hz bandpass filter (simulated). The change is simulated minus injected. This models the filter only; it is not a measurement through hardware.
| Metric | Injected | Simulated | Change | 95% interval of the change |
|---|
The intervals are descriptive. The McNemar tests below are the inferential result: with fewer than about 10 discordant files (files whose decision changed one way or the other) an interval can exclude 0 while the exact test does not.
| Files that are | Files | Changed | To “present” | To “absent” | p (exact McNemar) | p (Holm) |
|---|
Honest caveats
- Quiet murmurs are hard. On the injected bench set, sensitivity is — for grade I/VI murmurs against — for grades II/VI and III/VI.
- Test-split intervals are optimistic. The test split has several recordings per child, but the interval maths treats every file as independent, so its intervals come out too narrow. The bench set has one recording per child and does not have this problem.
- Accuracy alone flatters the test split. A model that always says “absent” already scores — there. The bench set is balanced, so its baseline is —. Read sensitivity and balanced accuracy too.
- Absolute accuracy may be a little optimistic. The design was refined while earlier experimental splits of the same labelled patients were in use, so it was not chosen blind to the final test half. This matters much less for the comparison between conditions, which scores the same files with the same model and threshold.
- Live shows one window at a time. A Live result is one 5 s window, so its error rates are the window-level ones, not the recording-level ones above. On the injected bench set: —. The two can differ noticeably, so read a single Live result with that in mind.
- Simulated is not hardware. The simulated bandpass models the filter only, not the loudspeaker, the silicone layer, the sensor or the noise of real electronics. The re-acquired condition will measure those and is not measured yet.
Where to look in the code
pcg/features.py— the one feature path: windows, band-limit, RMS, MFCC and the 60 numbers.pcg/explain.py— recomputes each stage of one window, and builds the example signal, for this page.pcg/labels.py— the label rule and the split by child.pcg/train.py— cross-validation, model choice and the threshold.pcg/model.py— the saved model and the recording-level rule (mean of the window probabilities).pcg/evaluate.py— scoring a condition and comparing two of them.hmi/app.py— the JSON routes/api/explainand/api/info/*this page reads.hmi/static/data-and-ai.js,hmi/static/data-and-ai-results.js,hmi/static/site.js— draw this page. Waveforms usehmi/static/plot.js, unchanged since the midterm.