How it works
What happens between a heart sound and the answer shown on the Monitor page, and the experiment this project was built for.
What this is
A phonocardiogram (PCG) is a recording of the sounds a heart makes. This instrument draws that sound as a waveform and asks a small machine-learning model whether a murmur (an extra whooshing sound between the main heart sounds) is present. It is an EECS 3641 course project and a research instrument for bench testing: not a medical device, no diagnostic claims, and no patient testing.
The question
How much murmur-detection accuracy is lost when the sound arrives through a real acquisition chain instead of straight from a dataset file, and where in the chain does the loss come from? The project does not try to beat published accuracy.
The system, end to end
Follow the arrows from left to right, lane by lane. Two routes lead to the same classifier. In the injected (control) route the dataset file goes straight into the software. In the re-acquired route the same file is played as sound, picked up again by a sensor, passed through the analog electronics and digitised. Hover, focus or tap any block to read what it does.
Scroll sideways to see the whole diagram.
Dataset file A heart-sound recording from the CirCor DigiScope dataset, stored as a .wav file at 4 kHz. Both test conditions start from the same file.
Loudspeaker Plays the recording as sound inside the coupler, so the signal has to leave the computer and be picked up again.
3D-printed coupler Holds the loudspeaker, the silicone layer and the sensor in a fixed arrangement, so every recording is played the same way.
Silicone skin layer A cast silicone layer that stands in for skin and tissue between the loudspeaker and the sensor.
Microphone The sensor on top of the silicone. It turns the sound back into a small electrical signal. It replaced the original piezoelectric contact sensor, and its exact type is not final.
Signal conditioning The analog electronics between the sensor and the converter: input protection, a high-impedance voltage amplifier and a 20–600 Hz bandpass filter, as designed. The Hardware page shows them.
ADC The analog-to-digital converter turns the voltage into numbers the Raspberry Pi can read. Which converter interface is used, and so the sample rate and bit depth, is not decided yet.
Acquisition and 5 s windows Collects the samples and cuts them into 5 s analysis windows. Today a replay of .wav files stands in for the converter; a hardware source will plug into the same interface.
Band-limit and normalise Keeps only 20–600 Hz of each window, then divides by the window's loudness (its RMS), so the result does not depend on volume.
MFCC features Turns each window into 20 MFCC curves (a compact description of the sound's spectrum over time), then summarises them as 60 numbers.
Logistic regression Combines the 60 numbers into one probability that the window holds a murmur. A recording's decision is the average over its windows compared with a fixed threshold.
Web HMI The Monitor page: the waveform, and the classification with its confidence.
Injected (control) The dashed route. The dataset file goes straight from disk into the feature code and skips every physical stage. It is the baseline the other route is compared with.
Re-acquired The solid route. The same file is played through the loudspeaker, the coupler and the silicone, picked up by the microphone, conditioned and digitised. It is not measured yet because the hardware is not finished.
The experiment
The same recordings, the same model, the same features and the same decision threshold are scored in each condition. Only the way the sound reaches the classifier changes, so a difference in the results is caused by what the sound passed through. The threshold is fixed from the training data and is never adjusted for a condition.
Every condition is scored on the bench set: one recording each from a balanced subset of the held-out children (children the model never trained on), with equal numbers of children with and without a murmur. Not every held-out child is in it; the counts are shown below.
Injected (control)
Checking resultsThe dataset file is read straight from disk into the feature code. Nothing can change it on the way, so this is the baseline. Any difference from it is caused by what the sound passed through.
Re-acquired
NOT MEASUREDThe same file played through the loudspeaker, coupler and silicone, picked up by the microphone and digitised. This is the project's headline result: how much of the baseline accuracy survives the real chain. It needs the hardware.
Simulated bandpass
Checking resultsA software model of the 20–600 Hz bandpass filter applied to the same files. It shows how much of any loss the filter's response alone could explain. It is a model, not a measurement, and never a re-acquired result.
What is measured
- Accuracy
- The share of recordings the classifier got right.
- Sensitivity
- Of the recordings with a murmur, the share it flagged as murmur present.
- Specificity
- Of the recordings without a murmur, the share it cleared as absent.
- AUC (area under the ROC curve)
- How well the model's probability ranks murmur recordings above normal ones, whatever threshold is used. 1 is perfect and 0.5 is chance.
To tell whether a change between two conditions is real and not luck, the files are paired and each class is tested separately with a per-class McNemar test, which looks only at the files where the two conditions disagree.
Injected (control) on the bench set
Accuracy
—
—
Sensitivity
—
—
Specificity
—
—
AUC
—
1 is perfect, 0.5 is chance
Threshold
—
fixed from training data
Loading the results…
The full results, with confusion matrices and the simulated comparison, are on the Data & AI page.
Accuracy The share of recordings classified correctly. Always answering "absent" would score the share shown as the baseline below it.
Sensitivity Of the recordings that have a murmur, the share the classifier flagged. The interval is a 95 % Wilson confidence interval.
Specificity Of the recordings without a murmur, the share the classifier cleared. The interval is a 95 % Wilson confidence interval.
Decision threshold The probability a recording's average window probability must reach to be called "murmur present". It was fixed from the training data and is never changed for a condition.
AUC Area under the ROC curve: how well the model's probability ranks murmur recordings above normal ones, whatever threshold is used.
The software layers
The code is split into layers that talk through small, fixed interfaces. That is what lets the data source change from a file to live hardware without touching the rest. Hover, focus or tap a block to see which files it is.
Scroll sideways to see the whole diagram.
A swappable source
Review mode reads a file. Live mode reads blocks of samples from a BlockSource. Today FileReplaySource replays .wav files in real time and no hardware is attached. A hardware capture source will implement the same interface, so the server, the page and the classifier do not change.
One feature path
The same feature code serves training, both test conditions and the live view. That is what keeps the comparison fair: the only thing that differs between conditions is the audio that goes in.
A renderer that did not change
plot.js draws the waveform and has not changed since the midterm. The display built for the example waveform now shows real recordings and the live replay.
Sources and stream Where samples come from. Review reads a .wav file (WavFileSource). Live reads blocks through a BlockSource; today FileReplaySource replays files in real time, and a hardware capture source will implement the same interface.
Features pcg/features.py is the only feature path: band-limit to 20–600 Hz, divide by the loudness, compute MFCCs, then 60 summary numbers per 5 s window.
Model pcg/model.py loads the trained logistic regression and pcg/classify.py defines what a classification looks like. Without a trained model a placeholder answers instead.
JSON API hmi/app.py is a small Flask server. It calls pcg and sends the results as JSON; it does no signal maths of its own.
Controller hmi.js runs in the browser. It asks the server for a window of samples, keeps the page state (mode, recording, position) and fills in the readouts. It never draws on the canvas itself.
Renderer plot.js draws one window of samples on a canvas with labelled axes and a cursor. It never makes network requests, and it has not changed since the midterm.
Where to look in the code
pcg/sources.py,pcg/stream.py— where samples come from:WavFileSource, theBlockSourceinterface,FileReplaySourceand theLiveMonitor.pcg/features.py— the one feature path: windows, band-limit, normalise, MFCCs, the 60 numbers.pcg/model.py,pcg/classify.py,pcg/train.py— the classifier, what its answer looks like, and how it is trained.pcg/evaluate.py— scores each condition on the bench set and compares two conditions.pcg/simulate_chain.py— the simulated bandpass condition.hmi/app.py— the Flask server and its JSON routes.hmi/static/hmi.js,hmi/static/plot.js— the Monitor's controller and renderer.hmi/static/how-it-works.js— fills this page's result tiles from/api/info/results.