Embedded Scream Detection: Offline Scream Recognition SDK for SafetyNEW
The Value of Scream Detection: Scenarios and Workflow
A scream is a human distress signal with a distinctive acoustic shape — high energy, sustained vocal pitch. It is also easily drowned in noise and easily confused with singing, cheering and children playing. Detection is therefore a problem of evidence and discrimination, not just sensitivity. Unlike cameras, microphones work in darkness and behind doors — and respect privacy, which matters for bedrooms and bathrooms-adjacent spaces.
Scenario families: home security cameras, solo-living care (the contact chain), and night-time public spaces (a safety net over the streets).
The Signature: Scream vs Speech vs Singing
Below is a real scream recording — waveform (top) and Mel spectrogram (bottom). The sustained, high-energy human vocalization spans energy across the whole band:

Audio samples (real recordings — press play):
🔊 Woman scream (real sample)
🔊 Another scream recording
Energy — loudness well above conversational baselines in the same environment
Pitch — held high and stable — a "sustained pull-up", whereas speech fluctuates and singing glides across a wide range
Duration — screams last longer than a surprised yelp; duration thresholds separate momentary exclamations from sustained distress
How Recognition Works: Preprocessing → Features → Algorithms
The engine combines energy + pitch evidence, discrimination against speech / singing / children playing via multi-class negative training, and relative-energy judgment — so detection does not depend on absolute loudness in noisy environments. The main confusion source is singing, handled by the wide-glide-vs-sustained-pitch distinction; single exclamations fall below the duration threshold.
Detection and Linkage Flow
A single surprised yelp is recorded with a normal reminder; a sustained scream escalates with push plus audio clip (call escalation configurable). Linkage differs by scene: home (notify family + camera), solo monitoring (emergency contact chain), public spaces (duty desk + nearest patrol).
Positioning note: this is a safety-assistance tool to speed up human response — for real emergencies, dial emergency services.
Solo Living: The Contact Chain
1. Scream detection — sustained scream/shout identified on-device
2. Grading — strength, duration, time-of-day
3. Instant notification — emergency contact with audio clip
4. Contact two-way talk — confirm status through the device
5. Escalation when unanswered — neighbor, property management, community
6. Record and configuration refinement — false-alarm review, contact chain tuning
Public Spaces: The Night Safety Net
On night-time streets and around transit hubs, sensing nodes on lamp posts feed a duty desk: event with location, auto-retrieved camera, nearest patrol dispatched. Verify before dispatch keeps resources focused on real incidents.
Events cluster in the evening and night hours; deployment focus differs by scene — quiet streets favor wider spacing and lighting-gap coverage; transit hubs rely on relative-energy stability; campuses combine with guard workflows and can reuse gunshot-detection deployment practices.
Accuracy and Performance
Relative-energy judgment keeps detection stable (92%+ illustrative) from quiet rooms to noisy streets. The key trade-off: sensitivity vs false alarms is a per-scene configuration — solo-living care leans toward "notify more"; public scenes lean toward "verify first".
Note: performance figures are based on internal test environments; actual results depend on hardware and deployment scenarios.
Platform and Hardware Requirements
From camera and care-device SoCs upward; CPU under 50 MHz for the standard model; low-power always-listening supported for battery devices; no camera required in privacy-sensitive deployments.
C API and Embedded Integration
The scream detection library exposes a concise streaming C API: the caller just keeps feeding 16 kHz mono PCM; framing, Mel preprocessing and model inference run internally, and frame-level probabilities are aggregated by the alarm strategy into event callbacks.
Full interface declaration (scream_detect.h):
/**
* scream_detect.h — 尖叫声识别统一接口
*
* 封装 Mel 预处理 + 推理引擎 + 报警策略, 内部模型消费线程处理音频。
* 与录音模块 (audio_capture.h) 相互独立: 调用者自行决定音频来源
* (录音回调 / wav 文件 / 网络流), 通过 scream_detect_feed 送入, 数据任意大小。
*
* 用法 (实时录音模式):
* scream_detect_t *d = scream_detect_create(mgk_path, NULL, NULL);
* scream_detect_set_listener(d, on_frame, on_onset, on_offset, NULL);
* scream_detect_start(d); // 启动内部模型消费线程
* audio_capture_start(rec, capture_cb, d); // 录音回调里调 scream_detect_feed
* ...
* scream_detect_stop(d); // 排空缓冲, 停止线程
* scream_detect_destroy(d);
*
* 用法 (wav 文件模式):
* scream_detect_t *d = scream_detect_create(mgk_path, NULL, NULL);
* scream_detect_set_listener(d, on_frame, on_onset, on_offset, NULL);
* scream_detect_start(d);
* 循环读文件: scream_detect_feed(d, pcm, n); // 任意数据大小
* scream_detect_stop(d);
* scream_detect_destroy(d);
*/
#ifndef SCREAM_DETECT_H
#define SCREAM_DETECT_H
#include <stdint.h>
#ifdef __cplusplus
extern "C" {
#endif
/* 识别事件 (报警策略输出, 用于事件结束回调) */
typedef struct {
float start_time; /* 事件开始时间 (秒) */
float end_time; /* 事件结束时间 (秒) */
float confidence; /* 事件置信度 */
float max_confidence; /* 事件内最大帧置信度 */
int frame_count; /* 事件持续帧数 */
} scream_detect_event_t;
/* 帧级回调: 每帧识别结果 (模型线程内执行) */
typedef void (*scream_detect_frame_cb_t)(float scream_prob, float timestamp,
void *user_data);
/* 事件开始回调: 策略判定尖叫事件开始, 只有开始时间 */
typedef void (*scream_detect_onset_cb_t)(float start_time, void *user_data);
/* 事件结束回调: 策略判定尖叫事件结束 (或停止识别时未结束的事件), 完整事件信息 */
typedef void (*scream_detect_offset_cb_t)(const scream_detect_event_t *event,
void *user_data);
typedef struct scream_detect_s scream_detect_t;
/* 创建/销毁; alarm_name/alarm_params 可传 NULL (用默认策略及参数) */
scream_detect_t *scream_detect_create(const char *mgk_path, /* 模型文件路径 (必填) */
const char *alarm_name, /* 报警策略名, NULL=默认 */
const char *alarm_params); /* 策略参数 key=val,key=val, NULL=默认 */
void scream_detect_destroy(scream_detect_t *det);
/* 设置事件回调 (create 后调用, 也可在运行中调整); 不需要的回调传 NULL */
void scream_detect_set_listener(scream_detect_t *det,
scream_detect_frame_cb_t on_frame,
scream_detect_onset_cb_t on_onset,
scream_detect_offset_cb_t on_offset,
void *user_data);
/* 启动/停止识别: 启动内部模型消费线程 / 排空缓冲后停止线程 */
int scream_detect_start(scream_detect_t *det);
void scream_detect_stop(scream_detect_t *det);
int scream_detect_is_running(scream_detect_t *det);
/* 设置事件识别策略 (可在运行中调整) */
int scream_detect_set_alarm(scream_detect_t *det, const char *alarm_name,
const char *alarm_params);
/* 送入 PCM 数据 (16bit 单声道 16kHz), 线程安全, 任意数据大小 */
int scream_detect_feed(scream_detect_t *det, const int16_t *pcm, int num_samples);
#ifdef __cplusplus
}
#endif
#endif /* SCREAM_DETECT_H */
A minimal WAV-inference demo (excerpt; the full file ships at src/screamDetect/c/scream_demo.c):
#include <stdio.h>
#include <stdlib.h>
#include <stdint.h>
#include "scream_detect.h"
#define DEFAULT_MGK "scream_detect_v7.mgk" /* 模型文件 */
#define DEFAULT_WAV "test_scream.wav" /* 16kHz 单声道 16bit PCM */
#define FEED_CHUNK 16000 /* 每次送入 1 秒音频 */
/* 帧级回调: 每帧输出尖叫声概率 (约 1 秒一帧) */
static void on_frame(float scream_prob, float timestamp, void *user_data)
{
(void)user_data;
printf("%8.2fs scream=%.4f\n", timestamp, scream_prob);
}
/* 事件开始回调: 策略判定尖叫事件开始 */
static void on_onset(float start_time, void *user_data)
{
(void)user_data;
printf("[EVENT] scream start at %.2fs\n", start_time);
}
/* 事件结束回调: 应用层可据此做后续统计或分级响应 */
static void on_offset(const scream_detect_event_t *ev, void *user_data)
{
(void)user_data;
printf("[EVENT] scream end at %.2fs (dur=%.2fs, conf=%.3f, frames=%d)\n",
ev->end_time, ev->end_time - ev->start_time,
ev->confidence, ev->frame_count);
}
static int read_wav_pcm(const char *path, int16_t **pcm, int *n, int *sr); /* 完整实现见源文件 */
int main(void)
{
scream_detect_t *det;
int16_t *pcm = NULL;
int num_samples = 0, sample_rate = 0;
int pos;
if (read_wav_pcm(DEFAULT_WAV, &pcm, &num_samples, &sample_rate) != 0)
return 1;
/* 1. 创建识别器: 模型文件 + 默认报警策略 (NULL) */
det = scream_detect_create(DEFAULT_MGK, NULL, NULL);
if (!det) return 1;
/* 2. 注册回调 (均为可选) */
scream_detect_set_listener(det, on_frame, on_onset, on_offset, NULL);
/* 3. 启动内部模型消费线程 */
scream_detect_start(det);
/* 4. 分块送入 PCM; 实时录音时改在录音回调里 feed */
for (pos = 0; pos < num_samples; pos += FEED_CHUNK) {
int n = num_samples - pos;
if (n > FEED_CHUNK) n = FEED_CHUNK;
scream_detect_feed(det, pcm + pos, n);
}
/* 5. 停止并销毁 */
scream_detect_stop(det);
scream_detect_destroy(det);
free(pcm);
return 0;
}
Build and run:
$(CC) scream_demo.c -I. -L. -lscreamdetect -lpthread -lm -o scream_demo
./scream_demoConclusion
Whether it's one apartment or one kilometer of street, the goal is the same: make a call for help impossible to miss. An online trial with scenario deployment support is available.
Want to try the detection yourself?