Sudden noise during mobile phone calls, occasional pop noises during headphone playback, and smart speaker distortion within specific frequency ranges may occur only briefly, but they can significantly impact user experience. For audio product manufacturers, rapidly detecting such anomalies during mass production testing and accurately locating their occurrence time has become a critical aspect of audio quality control.
Traditionally, audio quality inspection has relied heavily on manual listening. Quality inspectors repeatedly play audio samples and evaluate issues such as noise, distortion, popping sounds, and audio dropouts based on subjective experience. However, as products enter mass production, the limitations of manual inspection become increasingly apparent: limited efficiency, inconsistent evaluation standards among inspectors, and potential auditory fatigue caused by prolonged listening. More importantly, transient anomalies lasting only tens or hundreds of milliseconds can easily be missed during inspection.
To address this challenge, the Liankang Information technology team developed a dedicated audio quality assessment algorithm. This innovation has been granted a national invention patent titled "An Automatic Audio Quality Scoring Method and System Combining MFCC and Time-Domain Statistical Features" (Patent No.: ZL 2025 1 1516445.X). The method enables automatic audio quality scoring and anomaly localization through audio feature extraction, segmentation analysis, and machine learning models.

How Can Machines "Understand" Audio Quality Differences?
The human ear relies on multidimensional perception of loudness, timbre, noise, distortion, and transient changes to evaluate sound quality. Although machines cannot perceive audio in the same way humans do, they can convert audio signals into computable and comparable feature sets.
This patented method combines two types of audio features. The first type is MFCCs (Mel-frequency cepstral coefficients), which describe the energy distribution and timbral characteristics of audio signals across different frequency ranges. By analyzing MFCC features, the system can characterize spectral differences between reference audio and test audio.
The second type consists of time-domain statistical features, including root mean square (RMS) energy, zero crossing rate, maximum amplitude, skewness, and kurtosis. These features represent sound energy, waveform transition characteristics, amplitude peaks, and statistical distribution characteristics, respectively.
The proposed system combines MFCC features with time-domain statistical features, enabling detection results to capture not only timbral and spectral variations but also sudden waveform anomalies.
How Can Audio Quality Detection Become More Accurate, Intuitive, and Reliable?

Flow diagram of automatic audio quality scoring method combining MFCC and time-domain statistical features
01 Combining Average and Peak Values for Improved Audio Anomaly Detection
After audio preprocessing, the audio signal is divided into 1-second segments, and each segment is further divided into ten 100-millisecond sub-segments.
The average value represents overall audio quality within each second, while the peak value captures short-term anomalies such as sudden pop noises and clipping events. Combining these two features prevents transient anomalies from being masked by average values, improving detection accuracy.
The system then compares feature differences between reference audio and test audio and generates an evaluation score using a Support Vector Machine (SVM) model.
02 0–5 Scoring System Enables More Intuitive Audio Quality Evaluation
The system outputs a score from 0 to 5 according to the degree of audio distortion.
A score of 0 indicates almost no difference from the reference audio, while higher scores represent more severe distortion. A score of 5 indicates severe signal loss or poor audio quality.
Compared with simple pass/fail judgment, the scoring system provides a more intuitive evaluation of audio quality degradation. R&D engineers can compare different design solutions based on the score, while production lines can define thresholds for automatic retesting and defective product screening.
03 Lightweight Model for Stable Production-Line Deployment
A production-line inspection system must not only provide accurate results but also operate reliably, respond quickly, and remain easy to maintain.
The proposed method uses a Support Vector Machine (SVM) model instead of complex deep learning networks, reducing requirements for training samples and computational resources.
The SVM model requires only several tens of megabytes of storage and can run on standard industrial computers without GPU acceleration. Through small-sample learning, the system achieves lightweight deployment and stable model performance, improving detection accuracy and reliability.

The figure is a schematic diagram of the test results of sample b on the production line shown in the present invention
Application Scenarios
This patented technology is applicable to mobile phones, wireless headphones, smart speakers, wearable devices, automotive audio terminals, and other electronic products equipped with speakers, microphones, or audio playback functions.
In speaker testing, the system compares reference audio with the audio signals played and captured by the device to identify noise, pop sounds, clipping, distortion, and signal loss during playback.
In microphone testing, the system evaluates differences between the device-recorded signal and the reference signal, helping identify sensitivity abnormalities, excessive background noise, frequency response deviations, and signal acquisition path failures.
During the R&D validation phase, the system serves as an objective evaluation tool, helping engineers compare audio performance across different design solutions. During production testing, it can be integrated with automation equipment, test software, and data management systems to achieve automatic data collection, scoring, anomaly localization, and result traceability.
The value of this patented technology lies not only in the proposed scoring method, but also in its capability to be successfully deployed in real-world customer testing environments. In the future, Liankang Information will continue focusing on practical customer testing requirements, integrating more algorithms and innovative technologies into R&D verification and mass production processes, and transforming audio inspection from subjective judgment to data-driven evaluation.