Digital Phonocardiogram

Human–machine interface · final build

Static snapshot. Numbers and examples were computed on 2026-10-05 UTC from the committed model and results. The live monitor runs on the Raspberry Pi.

Data & AI

How a heart recording becomes a decision: what a phonocardiogram looks like, what the AI sees at every step, how the model was trained and how well it works. Hover over, or tab to, any info button (the small i) or coloured part for a short explanation.

Reading a phonocardiogram

A phonocardiogram (PCG) is a recording of the sounds a heart makes. Every heartbeat has two main sounds: the first heart sound (S1) and the second heart sound (S2). The gap from S1 to S2 is systole, and the gap from S2 to the next S1 is diastole. A heart murmur is an extra sound that fills part of the beat. In this made-up example a murmur fills every systole.

Hover over a coloured band, or hover over or tab to a legend button, to see what it means.

Waveform plot. This browser does not support the canvas element.
Time axis
Amplitude axis
Cursor readout

Figure 1. A phonocardiogram with its parts marked. Illustrative signal generated by our code; not a patient recording.

What the AI sees, step by step

Pick a real recording and a start time. On this public snapshot no server runs anything: the windows were precomputed every 5 s on the snapshot date, with the same feature code the classifier uses, and this page shows what comes out of each stage. The slider shows the saved window at or before the start you pick. A recording from the test split was never used for training.

0.0 s

Recordings from the CirCor DigiScope Phonocardiogram Dataset v1.0.3 (PhysioNet, ODC-By 1.0); full citation under How the model was trained.

Loading recordings…

Raspberry Pi: one 5 s window 1 Raw window 5 s of samples 2 Band-limit, ÷ RMS 20–600 Hz, same loudness 3 MFCC frames 20 numbers per 32 ms 4 Summary numbers 60 per window 5 Logistic regression weights, sum, probability 6 Recording decision average vs threshold

Scroll sideways to see the whole diagram.

Figure 2. The six stages every 5 s window goes through, in the same order as the steps below. Hover over, or tab to, a stage to read what it does.

Step 1 of 6: The raw window

The classifier never looks at a whole recording at once. It cuts the recording into 5 s windows and starts a new window every 2.5 s, so each window holds several heartbeats. This plot is the window you chose, as the file stores it, with its average level (a constant offset) subtracted.

Waveform plot.

Figure 3. The raw 5 s window. The vertical scale fits this plot, so it can differ from the next one.

Step 2 of 6: Keep 20–600 Hz, then fix the loudness

The analog front end is designed to pass 20–600 Hz and weaken the rest. The software removes everything outside that band exactly, so every condition sees the same band. Then the window is divided by its own RMS level (its typical size), so a louder and a quieter recording with the same shape give the same numbers. We band-limit first so the loudness is measured only in the band the features use.

Waveform plot.

Figure 4. The same window after band-limiting to 20–600 Hz, before the division by RMS.

RMS divisor

—

Waiting for a window

Step 3 of 6: MFCC: the spectrum, frame by frame

The window is cut into short frames, 128 ms long with a new frame every 32 ms. For each frame, 20 numbers called MFCCs (mel-frequency cepstral coefficients) describe the shape of the spectrum between 20 and 600 Hz, using 32 mel bands. The picture has one column per frame and one row per coefficient.

Why MFCCs? The mel scale bends frequency to match hearing, but it is almost a straight line below 1 kHz, so the bending does little in this band. MFCCs still help here because they compress the spectrum into a few numbers and make those numbers less correlated with each other (decorrelation).

MFCC heatmap.

Time axis
Coefficient axis
Colour scale
Figure 5. The MFCC matrix of the window. Move the pointer over it, or focus it and press the arrow keys, to read a value.

Step 4 of 6: Summarise into 60 numbers

The model needs the same amount of input for every window, so each of the 20 coefficients is boiled down to three numbers over the whole window: its mean, its standard deviation (how much it varies) and the standard deviation of its delta (how much its frame-to-frame change varies). 20 × 3 = 60 numbers. This is everything the classifier is told about the window.

Bar chart of the 60 summary numbers.

Figure 6. Three groups of 20 bars, one bar per MFCC: the mean of each coefficient (left), its standard deviation (middle) and the standard deviation of its delta (right). Each group has its own vertical scale, printed under it.
Show the 60 numbers as a table
The 60 summary numbers of this window
MFCCMeanStdStd of delta

Step 5 of 6: The classifier adds up the evidence

Logistic regression is small and fast, and its decision can be read as a sum. Each of the 60 numbers is first standardised: z = (x − mean) ÷ spread, using the average and spread of the training windows. Each z is multiplied by a weight the model learned. Those contributions, plus a fixed intercept, add up to one score: the logit. A positive contribution pushes towards murmur and a negative one towards normal. A logit of zero is an even lean (p = 0.5), but a window is labelled “Murmur present” only when p reaches the threshold t, so the decision line is logit = ln(t ÷ (1 − t)), about — for this model. The bars show the features that pushed hardest; all the others are combined in one bar.

← towards normaltowards murmur →
    Figure 7. How much each feature pushed the logit for this window. The dashed mark on the Logit row is the decision line.

    The logit is squeezed into a probability between 0 and 1 with the sigmoid function, p = 1 ÷ (1 + e^(−logit)), and compared with the threshold.

    —
    Probability
    Threshold
    Confidence

    Confidence is the probability the model gives to the label it chose. It is not a calibrated clinical probability: it only shows how strongly the model leans, on the data it was trained on. It can be below 50 % because the threshold is not 0.5.

    Step 6 of 6: From windows to a recording decision

    In Review, the Monitor's main result (Whole recording) is one decision per recording; its This window row, and Live mode, show single-window results instead. For the recording decision the model scores every 5 s window of the recording (—), averages the window probabilities and compares the mean with the same threshold.

    Open the Monitor and pick the same recording

    How the model was trained

    The model learns from recordings that doctors have already labelled. Training reads only the training half of the data; it never opens a test file.

    The data

    The CirCor DigiScope Phonocardiogram Dataset, version 1.0.3, from PhysioNet under the Open Data Commons Attribution License 1.0. It holds heart sounds from 942 paediatric patients: 3163 recordings at 4 kHz, each from one of up to four auscultation locations (aortic AV, pulmonary PV, tricuspid TV and mitral MV). A recording is labelled “murmur present” if a murmur was audible at that location, and “murmur absent” if the patient has no murmur. Recordings of patients labelled Unknown are left out, and so are recordings of a patient with a murmur at a location where it was not audible.

    This public site includes the waveforms of a few demo recordings from this dataset (most or all of each recording), used under the Open Data Commons Attribution License 1.0. Cite the dataset when you use it:

    • Oliveira J, Renna F, Costa P, et al. The CirCor DigiScope Phonocardiogram Dataset (version 1.0.3). PhysioNet; 2022. doi:10.13026/tshs-mw03
    • Oliveira JH, Renna F, Costa P, et al. The CirCor DigiScope Dataset: From Murmur Detection to Murmur Classification. IEEE Journal of Biomedical and Health Informatics. 2021. doi:10.1109/JBHI.2021.3137048
    • Pollard T, Moody BE, Lehman L, et al. PhysioNet as a global platform for biomedical research. Nature Health. 2026. doi:10.1038/s44360-026-00096-z

    The recipe

    1. Split by child, 50/50 (seed 3641). Half of the children train the model and the other half are held out. No child has recordings on both sides.
    2. Make windows. Every training recording is cut into 5 s windows, a new one every 2.5 s, and each window becomes 60 numbers (steps 1–4 above). A window gets the label of its recording.
    3. Compare candidates (three strengths of logistic regression and a random forest) with 5-fold cross-validation, grouped by child.
    4. Pick the model: the best logistic regression by AUC, unless the random forest beats it by more than one fold standard deviation.
    5. Fix the threshold on the out-of-fold training predictions with Youden's rule. It is set once and never changed for any test condition.
    6. Fit the winner on all training windows and save it.

    How well it works

    The experiment scores the same held-out files under different conditions, with the same model, features and threshold, and compares the results. The bench set has one recording each from a balanced subset of the held-out children: equal numbers of children with and without a murmur, so not every held-out child is in it (the file counts are in the table below). The test split has every recording of the held-out children.

    Where to look in the code